EDBT 2026 Demo / reviewers in the wild / expert
Trung Le 0001
dblp:88/8728
· DBLP profile ↗
134ranked-venue papers
28as first author
71since 2021 · last 2026
0000-0003-0414-9067ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 113 · 24 first-author · 60 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 2 first-author · 17 since 2021Databases, data management, data science and information retrieval · 16 · 5 first-author · 3 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DIET: Machine Unlearning on a Data-DietabstractMachine Unlearning (MU) aims to remove the influence of specific knowledge from a pretrained model. Existing methods often rely on retained training data to preserve utility; such dependence is impractical due to privacy and scalability constraints. A further complication arises when unlearning is applied to vision-language models (VLMs), where entangled multimodal representations make targeted forgetting especially challenging. We propose DIET, a principled retain-data-free unlearning method for VLMs that addresses these challenges by leveraging the geometry of hyperbolic space. The core idea is to push forget embeddings toward class-mismatched prototypes located at the boundary of the hyperbolic space. In hyperbolic geometry, points near the boundary become infinitely distant from interior points. As a result, moving forget embeddings to the boundary makes their influence on the model asymptotically negligible. To formalize this, we guide the forgetting process using the Busemann function, which quantifies directional distance to the boundary. We further develop an adaptive scheme based on optimal transport that selects mismatched prototypes for each forget embedding, enabling flexible unlearning dynamics. Extensive experiments on fine-grained datasets such as Flowers102, OxfordPets, and StanfordCars show that DIET achieves an average forget accuracy of 8.06%, while preserving 69.04% utility using only 16 samples per concept, significantly outperforming the best retain-free baselines with a 117.5% improvement in model utility, and showing competitive performance to retain-data baselines with only a 3.79% drop Nilakshan Kunananthaseelan, Jing Wu 0021, Trung Le 0001, Gholamreza Haffari, Mehrtash Harandi |
AAAI | 3 |
| 2026 | CTPD: Cross Tokenizer Preference DistillationabstractWhile knowledge distillation has seen widespread use in pre-training and instruction tuning, its application to aligning language models with human preferences remains underexplored, particularly in the more realistic cross-tokenizer setting. The incompatibility of tokenization schemes between teacher and student models has largely prevented fine-grained, white-box distillation of preference information. To address this gap, we propose Cross-Tokenizer Preference Distillation (CTPD), the first unified framework for transferring human-aligned behavior between models with heterogeneous tokenizers. CTPD introduces three key innovations: (1) Aligned Span Projection, which maps teacher and student tokens to shared character-level spans for precise supervision transfer; (2) a cross-tokenizer adaptation of Token-level Importance Sampling (TIS-DPO) for improved credit assignment; and (3) a Teacher-Anchored Reference, allowing the student to directly leverage the teacher’s preferences in a DPO-style objective. Our theoretical analysis grounds CTPD in importance sampling, and experiments across multiple benchmarks confirm its effectiveness, with significant performance gains over existing methods. These results establish CTPD as a practical and general solution for preference distillation across diverse tokenization schemes, opening the door to more accessible and efficient alignment of language models. Phi Van Dat, Ngan Nguyen, Ngo Van Linh 0001, Trung Le 0001, Thanh Hong Nguyen |
AAAI | 5 |
| 2026 | MCW-KD: Multi-Cost Wasserstein Knowledge Distillation for Large Language Models
Hoang Tran Vuong, Tue Le, Quyen Tran, Ngo Van Linh 0001, Trung Le 0001 |
AAAI | 5 |
| 2026 | MTA: Multi-Granular Trajectory Alignment for Large Language Model DistillationabstractPham Khanh Chi, Quoc Phong Dao, Thuat Nguyen, Linh Ngo Van, Trung Le, Thanh Hong Nguyen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pham Khanh Chi, Quoc Phong Dao, Thuat Nguyen, Ngo Van Linh 0001, Trung Le 0001, Thanh Hong Nguyen |
ACL (1) | 5 |
| 2026 | SRA: Span Representation Alignment for Large Language Model DistillationabstractQuoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Tung Nguyen, Linh Ngo Van, Nguyen Thi Ngoc Diep, Trung Le. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Trung Le 0001 |
ACL (1) | 7 |
| 2026 | TALAS: Teacher-Anchored Layer Alignment with Adaptive Sharpness-Aware Minimization for Embedding DistillationabstractQuoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Linh Ngo Van, Nguyen Thi Ngoc Diep, Thien Huu Nguyen, Trung Le. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Thien Huu Nguyen, Trung Le 0001 |
ACL (1) | 7 |
| 2026 | Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language ModelsabstractCuong Pham, Anh Dung Hoang, Cuong C. Nguyen, Trung Le, Gustavo Carneiro, Thanh-Toan Do. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Cuong Pham 0007, Dung Anh Hoang, Cuong Nguyen 0006, Trung Le 0001, Gustavo Carneiro 0001, Thanh-Toan Do |
ACL (1) | 4 |
| 2026 | LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language ModelsabstractMinh Chu Xuan, Tien-Phat Nguyen, Linh Ngo Van, Dinh Viet Sang, Nguyen Thi Ngoc Diep, Trung Le. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Minh Chu Xuan, Tien-Phat Nguyen, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Trung Le 0001 |
ACL (1) | 6 |
| 2025 | Erasing Undesirable Influence in Diffusion ModelsabstractDiffusion models are highly effective at generating high-quality images but pose risks, such as the unintentional generation of NSFW (not safe for work) content. Although various techniques have been proposed to mitigate unwanted influences in diffusion models while preserving overall performance, achieving a balance between these goals remains challenging. In this work, we introduce EraseDiff, an algorithm designed to preserve the utility of the diffusion model on retained data while removing the unwanted information associated with the data to be forgotten. Our approach formulates this task as a constrained optimization problem using the value function, resulting in a natural first-order algorithm for solving the optimization problem. By altering the generative process to deviate away from the groundtruth denoising trajectory, we update parameters for preservation while controlling constraint reduction to ensure effective erasure, striking an optimal trade-off. Extensive experiments and thorough comparisons with state-of-the-art algorithms demonstrate that EraseDiff effectively preserves the model’s utility, efficacy, and efficiency.$\color{Red} {\text {WARNING}}$: This paper contains sexually explicit imagery that may be offensive in nature. Jing Wu 0021, Trung Le 0001, Munawar Hayat, Mehrtash Harandi |
CVPR | 2 |
| 2025 | Enhancing Dataset Distillation via Non-Critical Region RefinementabstractDataset distillation has become a popular method for compressing large datasets into smaller, more efficient representations while preserving critical information for model training. Data features are broadly categorized into two types: instance-specific features, which capture unique, fine-grained details of individual examples, and class-general features, which represent shared, broad patterns across a class. However, previous approaches often struggle to balance these features—some focus solely on class-general patterns, neglecting finer instance details, while others prioritize instance-specific features, overlooking the shared characteristics essential for class-level understanding. In this paper, we introduce the Non-Critical Region Refinement Dataset Distillation (NRR-DD) method, which preserves instance-specific details and fine-grained regions in synthetic data while enriching non-critical regions with class-general information. This approach enables models to leverage all pixel information, capturing both feature types and enhancing overall performance. Additionally, we present Distance-Based Representative (DBR) knowledge transfer, which eliminates the need for soft labels in training by relying on the distance between synthetic data predictions and one-hot encoded labels. Experimental results show that NRR-DD achieves state-of-the-art performance on both small- and large-scale datasets. Furthermore, by storing only two distances per instance, our method delivers comparable results across various settings. The code is available at https://github.com/tmtuan1307/NRR-DD. Minh-Tuan Tran, Trung Le 0001, Xuan-May Le, Thanh-Toan Do, Dinh Q. Phung |
CVPR | 2 |
| 2025 | Preserving Clusters in Prompt Learning for Unsupervised Domain AdaptationabstractRecent approaches leveraging multi-modal pre-trained models like CLIP for Unsupervised Domain Adaptation (UDA) have shown significant promise in bridging domain gaps and improving generalization by utilizing rich semantic knowledge and robust visual representations learned through extensive pre-training on diverse image-text datasets. While these methods achieve state-of-the-art performance across benchmarks, much of the improvement stems from base pseudolabels (CLIP zero-shot predictions) and self-training mechanisms. Thus, the training mechanism exhibits a key limitation wherein the visual embedding distribution in target domains can deviate from the visual embedding distribution in the pre-trained model, leading to misguided signals from class descriptions. This work introduces a fresh solution to reinforce these pseudo-labels and facilitate target-prompt learning, by exploiting the geometry of visual and text embeddings - an aspect that is overlooked by existing methods. We first propose to directly leverage the reference predictions (from source prompts) based on the relationship between source and target visual embeddings. We later show that there is a strong clustering behavior observed between visual and text embeddings in pre-trained multi-modal models. Building on optimal transport theory, we transform this insight into a novel strategy to enforce the clustering property in text embeddings, further enhancing the alignment in the target domain. Our experiments and ablation studies validate the effectiveness of the proposed approach, demonstrating superior performance and improved quality of target prompts in terms of representation. Tung Long Vuong, Hoang Phan, Vy Vo, Anh Bui, Thanh-Toan Do, Trung Le 0001, Dinh Q. Phung |
CVPR | 6 |
| 2025 | MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic CorporaabstractContinually updating model-based indexes in generative retrieval with new documents remains challenging, as full retraining is computationally expensive and impractical under resource constraints.We propose MixLoRA-DSI, a novel framework that combines an expandable mixture of Low-Rank Adaptation experts with a layer-wise out-of-distribution (OOD)driven expansion strategy.Instead of allocating new experts for each new corpus, our proposed expansion strategy enables sublinear parameter growth by selectively introducing new experts only when significant number of OOD documents are detected.Experiments on NQ320k and MS MARCO Passage demonstrate that MixLoRA-DSI outperforms full-model update baselines, with minimal parameter overhead and substantially lower training costs. 1 Tuan-Luc Huynh, Thuy-Trang Vu, Weiqing Wang 0001, Trung Le 0001, Dragan Gasevic, Yuan-Fang Li, Thanh-Toan Do |
EMNLP | 4 |
| 2025 | EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsabstractKnowledge distillation (KD) is crucial for compressing large text embedding models, but faces challenges when teacher and student models use different tokenizers (Cross-Tokenizer KD -CTKD).Vocabulary mismatches impede the transfer of relational knowledge encoded in deep representations, such as hidden states and attention matrices, which are vital for producing high-quality embeddings.Existing CTKD methods often focus on direct output alignment, neglecting this crucial structural information.We propose a novel framework tailored for CTKD embedding model distillation.We first map tokens one-to-one via Minimum Edit Distance (MinED).Then, we distill intra-model relational knowledge by aligning attention matrix patterns using Centered Kernel Alignment, focusing on the top-m most important tokens of the directly mapped tokens.Simultaneously, we align final hidden states via Optimal Transport with Importance-Scored Mass Assignment, which emphasizes semantically important token representations, based on importance scores derived from attention weights.We evaluate distillation from state-of-the-art embedding models (e.g., LLM2Vec, BGE) to a Bert-base-uncased model on embedding-reliant tasks such as text classification, sentence pair classification, and semantic textual similarity.Our proposed framework significantly outperforms existing CTKD baselines.By preserving attention structure and prioritizing key representations, our approach yields smaller, highfidelity embedding models despite tokenizer differences. Minh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep, Ngo Van Linh 0001, Thien Huu Nguyen, Trung Le 0001 |
EMNLP | 7 |
| 2025 | Beyond Losses Reweighting: Empowering Multi-Task Learning via the Generalization Perspective
Hoang Phan, Lam Tran, Quyen Tran, Ngoc N. Tran, Tuan Truong, Nhat Ho, Dinh Q. Phung, Trung Le 0001 |
ICCV | 9 |
| 2025 | A Good Teacher Adapts Their Knowledge for Distillation
Chengyao Qian, Trung Le 0001, Mehrtash Harandi |
ICCV | 2 |
| 2025 | Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find ThemabstractConcept erasure has emerged as a promising technique for mitigating the risk of harmful content generation in diffusion models by selectively unlearning undesirable concepts. The common principle of previous works to remove a specific concept is to map it to a fixed generic concept, such as a neutral concept or just an empty text prompt. In this paper, we demonstrate that this fixed-target strategy is suboptimal, as it fails to account for the impact of erasing one concept on the others. To address this limitation, we model the concept space as a graph and empirically analyze the effects of erasing one concept on the remaining concepts. Our analysis uncovers intriguing geometric properties of the concept space, where the influence of erasing a concept is confined to a local region. Building on this insight, we propose the Adaptive Guided Erasure (AGE) method, which dynamically selects optimal target concepts tailored to each undesirable concept, minimizing unintended side effects. Experimental results show that AGE significantly outperforms state-of-the-art erasure methods on preserving unrelated concepts while maintaining effective erasure performance. Our code is published at {https://github.com/tuananhbui89/Adaptive-Guided-Erasure}. Anh Tuan Bui, Thuy-Trang Vu, Long Tung Vuong, Trung Le 0001, Paul Montague, Tamas Abraham, Junae Kim, Dinh Q. Phung |
ICLR | 4 |
| 2025 | Improved Training Technique for Latent Consistency ModelsabstractConsistency models are a new family of generative models capable of producing high-quality samples in either a single step or multiple steps. Recently, consistency models have demonstrated impressive performance, achieving results on par with diffusion models in the pixel space. However, the success of scaling consistency training to large-scale datasets, particularly for text-to-image and video generation tasks, is determined by performance in the latent space. In this work, we analyze the statistical differences between pixel and latent spaces, discovering that latent data often contains highly impulsive outliers, which significantly degrade the performance of iCT in the latent space. To address this, we replace Pseudo-Huber losses with Cauchy losses, effectively mitigating the impact of outliers. Additionally, we introduce a diffusion loss at early timesteps and employ optimal transport (OT) coupling to further enhance performance. Lastly, we introduce the adaptive scaling-$c$ scheduler to manage the robust training process and adopt Non-scaling LayerNorm in the architecture to better capture the statistics of the features and reduce outlier impact. With these strategies, we successfully train latent consistency models capable of high-quality sampling with one or two steps, significantly narrowing the performance gap between latent consistency and diffusion models. The implementation is released here: \url{https://github.com/quandao10/sLCT/} Quan Dao, Khanh Doan, Di Liu 0003, Trung Le 0001, Dimitris N. Metaxas |
ICLR | 4 |
| 2025 | Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among PromptsabstractPrompt-based techniques, such as prompt-tuning and prefix-tuning, have gained prominence for their efficiency in fine-tuning large pre-trained models. Despite their widespread adoption, the theoretical foundations of these methods remain limited. For instance, in prefix-tuning, we observe that a key factor in achieving performance parity with full fine-tuning lies in the reparameterization strategy. However, the theoretical principles underpinning the effectiveness of this approach have yet to be thoroughly examined. Our study demonstrates that reparameterization is not merely an engineering trick but is grounded in deep theoretical foundations. Specifically, we show that the reparameterization strategy implicitly encodes a shared structure between prefix key and value vectors. Building on recent insights into the connection between prefix-tuning and mixture of experts models, we further illustrate that this shared structure significantly improves sample efficiency in parameter estimation compared to non-shared alternatives. The effectiveness of prefix-tuning across diverse tasks is empirically confirmed to be enhanced by the shared structure, through extensive experiments in both visual and language domains. Additionally, we uncover similar structural benefits in prompt-tuning, offering new perspectives on its success. Our findings provide theoretical and empirical contributions, advancing the understanding of prompt-based methods and their underlying mechanisms. Minh Le, Quyen Tran, Trung Le 0001, Nhat Ho |
ICLR | 5 |
| 2025 | Boosting Multiple Views for pretrained-based Continual LearningabstractRecent research has shown that Random Projection (RP) can effectively improve the performance of pre-trained models in Continual learning (CL). The authors hypothesized that using RP to map features onto a higher-dimensional space can make them more linearly separable. In this work, we theoretically analyze the role of RP and present its benefits for improving the model’s generalization ability
in each task and facilitating CL overall. Additionally, we take this result to the next level by proposing a Multi-View Random Projection scheme for a stronger ensemble classifier. In particular, we train a set of linear experts, among which diversity is encouraged based on the principle of AdaBoost, which was initially very challenging to apply to CL. Moreover, we employ a task-based adaptive backbone
with distinct prompts dedicated to each task for better representation learning. To properly select these task-specific components and mitigate potential feature shifts caused by misprediction, we introduce a simple yet effective technique called the self-improvement process. Experimentally, our method consistently outperforms state-of-the-art baselines across a wide range of datasets. Quyen Tran, Tung Lam Tran, Khanh Doan, Toan Tran 0003, Dinh Q. Phung, Khoat Than, Trung Le 0001 |
ICLR | 7 |
| 2025 | Promoting Ensemble Diversity with Interactive Bayesian Distributional Robustness for Fine-tuning Foundation ModelsabstractWe introduce Interactive Bayesian Distributional Robustness (IBDR), a novel Bayesian inference framework that allows modeling the interactions between particles, thereby enhancing ensemble quality through increased particle diversity. IBDR is grounded in a generalized theoretical framework that connects the distributional population loss with the approximate posterior, motivating a practical dual optimization procedure that enforces distributional robustness while fostering particle diversity. We evaluate IBDR's performance against various baseline methods using the VTAB-1K benchmark and the common reasoning language task. The results consistently show that IBDR outperforms these baselines, underscoring its effectiveness in real-world applications. Ngoc-Quan Pham, Tuan Truong, Quyen Tran, Tan M. Nguyen, Dinh Q. Phung, Trung Le 0001 |
ICML | 6 |
| 2025 | RepLoRA: Reparameterizing Low-rank Adaptation via the Perspective of Mixture of ExpertsabstractLow-rank Adaptation (LoRA) has emerged as a powerful and efficient method for fine-tuning large-scale foundation models. Despite its popularity, the theoretical understanding of LoRA has remained underexplored. In this paper, we present a theoretical analysis of LoRA by examining its connection to the Mixture of Experts models. Under this framework, we show that a simple technique, reparameterizing LoRA matrices, can notably accelerate the low-rank matrix estimation process. In particular, we prove that reparameterization can reduce the data needed to achieve a desired estimation error from an exponential to a polynomial scale. Motivated by this insight, we propose Reparameterized Low-Rank Adaptation (RepLoRA), incorporating a lightweight MLP to reparameterize the LoRA matrices. Extensive experiments across multiple domains demonstrate that RepLoRA consistently outperforms vanilla LoRA. With limited data, RepLoRA surpasses LoRA by a substantial margin of up to 40.0% and achieves LoRA’s performance using only 30.0% of the training data, highlighting the theoretical and empirical robustness of our PEFT method. Tuan Truong, Minh Le, Trung Le 0001, Nhat Ho |
ICML | 5 |
| 2025 | Improving Generalization with Flat Hilbert Bayesian InferenceabstractWe introduce Flat Hilbert Bayesian Inference (FHBI), an algorithm designed to enhance generalization in Bayesian inference. Our approach involves an iterative two-step procedure with an adversarial functional perturbation step and a functional descent step within the reproducing kernel Hilbert spaces. This methodology is supported by a theoretical analysis that extends previous findings on generalization ability from finite-dimensional Euclidean spaces to infinite-dimensional functional spaces. To evaluate the effectiveness of FHBI, we conduct comprehensive comparisons against nine baseline methods on the VTAB-1K benchmark, which encompasses 19 diverse datasets across various domains with diverse semantics. Empirical results demonstrate that FHBI consistently outperforms the baselines by notable margins, highlighting its practical efficacy. Tuan Truong, Quyen Tran, Ngoc-Quan Pham, Nhat Ho, Dinh Q. Phung, Trung Le 0001 |
ICML | 6 |
| 2025 | Mutual-pairing Data Augmentation for Fewshot Continual Relation ExtractionabstractNguyen Hoang Anh, Quyen Tran, Thanh Xuan Nguyen, Nguyen Thi Ngoc Diep, Linh Ngo Van, Thien Huu Nguyen, Trung Le. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Nguyen Hoang Anh, Quyen Tran, Nguyen Thi Ngoc Diep, Ngo Van Linh 0001, Thien Huu Nguyen, Trung Le 0001 |
NAACL (Long Papers) | 7 |
| 2025 | Token-Level Self-Play with Importance-Aware Guidance for Large Language ModelsabstractLeveraging the power of Large Language Models (LLMs) through preference optimization is crucial for aligning model outputs with human values. Direct Preference Optimization (DPO) has recently emerged as a simple yet effective method by directly optimizing on preference data without the need for explicit reward models. However, DPO typically relies on human-labeled preference data, which can limit its scalability. Self-Play Fine-Tuning (SPIN) addresses this by allowing models to generate their own rejected samples, reducing the dependence on human annotations. Nevertheless, SPIN uniformly applies learning signals across all tokens, ignoring the fine-grained quality variations within responses. As the model improves, rejected samples increasingly contain high-quality tokens, making the uniform treatment of tokens suboptimal. In this paper, we propose SWIFT (Self-Play Weighted Fine-Tuning), a fine-grained self-refinement method that assigns token-level importance weights estimated from a stronger teacher model. Beyond alignment, we also demonstrate that SWIFT serves as an effective knowledge distillation strategy by using the teacher not for logits matching, but for reward-guided token weighting. Extensive experiments on diverse benchmarks and settings demonstrate that SWIFT consistently surpasses both existing alignment approaches and conventional knowledge distillation methods. Tue Le, Hoang Tran Vuong, Quyen Tran, Ngo Van Linh 0001, Mehrtash Harandi, Trung Le 0001 |
NeurIPS | 6 |
| 2025 | Unveiling m-Sharpness Through the Structure of Stochastic Gradient NoiseabstractSharpness-aware minimization (SAM) has emerged as a highly effective technique to improve model generalization, but its underlying principles are not fully understood. We investigate m-sharpness, where SAM performance improves monotonically as the micro-batch size for computing perturbations decreases, a phenomenon critical for distributed training yet lacking rigorous explanation. We leverage an extended Stochastic Differential Equation (SDE) framework and analyze stochastic gradient noise (SGN) to characterize the dynamics of SAM variants, including n-SAM and m-SAM. Our analysis reveals that stochastic perturbations induce an implicit variance-based sharpness regularization whose strength increases as m decreases. Motivated by this insight, we propose Reweighted SAM (RW-SAM), which employs sharpness-weighted sampling to mimic the generalization benefits of m-SAM while remaining parallelizable. Comprehensive experiments validate our theory and method. Haocheng Luo, Mehrtash Harandi, Dinh Q. Phung, Trung Le 0001 |
NeurIPS | 4 |
| 2025 | Geometry-Aware Collaborative Multi-Solutions Optimizer for Model Fine-Tuning with Parameter EfficiencyabstractWe propose a framework grounded in gradient flow theory and informed by geometric structure that provides multiple diverse solutions for a given task, ensuring collaborative results that enhance performance and adaptability across different tasks. This framework enables flexibility, allowing for efficient task-specific fine-tuning while preserving the knowledge of the pre-trained foundation models. Extensive experiments across transfer learning, few-shot learning, and domain generalization show that our proposed approach consistently outperforms existing Bayesian methods, delivering strong performance with affordable computational overhead and offering a practical solution by updating only a small subset of parameters. Van-Anh Nguyen, Trung Le 0001, Mehrtash Harandi, Ehsan Abbasnejad, Thanh-Toan Do, Dinh Q. Phung |
NeurIPS | 2 |
| 2025 | PromptDSI: Prompt-Based Rehearsal-Free Continual Learning for Document Retrieval
Tuan-Luc Huynh, Thuy-Trang Vu, Weiqing Wang 0001, Yinwei Wei, Trung Le 0001, Dragan Gasevic, Yuan-Fang Li, Thanh-Toan Do |
ECML/PKDD (7) | 5 |
| 2025 | HVQ-VAE: Variational auto-encoder with hyperbolic vector quantization
Shangyu Chen, Pengfei Fang, Mehrtash Harandi, Trung Le 0001, Jianfei Cai 0001, Dinh Q. Phung |
Comput. Vis. Image Underst. | 4 |
| 2025 | DeepVulMatch: Learning and Matching Latent Vulnerability Representations for Dual-Granularity Vulnerability DetectionabstractDeep learning (DL) models are widely used to detect software vulnerabilities, but identifying vulnerabilities at the line level remains challenging due to varied coding styles and the spread of vulnerabilities across multiple lines. We observe that vulnerable line embeddings tend to form clusters in the feature space, which can help models capture hidden patterns more effectively. In this article, we propose a novel approach that leverages vector quantization (VQ) and optimal transport (OT) to exploit the clustering characteristics of vulnerable line embeddings and enhance detection performance. Specifically, we extract vulnerable line embeddings from the training data to form a vulnerability collection, which we condense into a compact vulnerability codebook using VQ and OT. Inspired by static analysis tools that rely on pattern matching, our model uses this codebook to match latent vulnerability representations during inference. Our approach also introduces dual-granularity detection, predicting both vulnerable functions and, when a function is predicted vulnerable, identifying the specific vulnerable lines within it. We evaluate our approach against 12 baselines on two large-scale datasets of real-world open-source vulnerabilities. Our method achieves the highest F1 scores at both the function and line levels. Trung Le 0001, Van Nguyen 0002, Chakkrit Tantithamthavorn, Dinh Q. Phung |
IEEE Trans. Reliab. | 2 |
| 2024 | Text-Enhanced Data-Free Approach for Federated Class-Incremental LearningabstractFederated Class-Incremental Learning (FCIL) is an underexplored yet pivotal issue, involving the dynamic addition of new classes in the context of federated learning. In this field, Data-Free Knowledge Transfer (DFKT) plays a crucial role in addressing catastrophic forgetting and data privacy problems. However, prior approaches lack the crucial synergy between DFKT and the model training phases, causing DFKT to encounter difficulties in generating high-quality data from a non-anchored latent space of the old task model. In this paper, we introduce LANDER (Label Text Centered Data-Free Knowledge Transfer) to address this issue by utilizing label text embeddings (LTE) produced by pretrained language models. Specifically, during the model training phase, our approach treats LTE as anchor points and constrains the feature embeddings of corresponding training samples around them, enriching the surrounding area with more meaningful information. In the DFKT phase, by using these LTE anchors, LANDER can synthesize more meaningful samples, thereby effectively addressing the forgetting problem. Additionally, instead of tightly constraining embeddings toward the anchor, the Bounding Loss is introduced to encourage sample embeddings to remain flexible within a defined radius. This approach preserves the natural differences in sample embeddings and mitigates the embedding overlap caused by heterogeneous federated settings. Extensive experiments conducted on CIFAR100, Tiny-ImageNet, and ImageNet demonstrate that LANDER significantly outperforms previous methods and achieves state-of-the-art performance in FCIL. The code is available at https://github.com/tmtuan1307/lander. Minh-Tuan Tran, Trung Le 0001, Xuan-May Le, Mehrtash Harandi, Dinh Q. Phung |
CVPR | 2 |
| 2024 | NAYER: Noisy Layer Data Generation for Efficient and Effective Data-free Knowledge DistillationabstractData-Free Knowledge Distillation (DFKD) has made significant recent strides by transferring knowledge from a teacher neural network to a student neural network without accessing the original data. Nonetheless, existing approaches encounter a significant challenge when attempting to generate samples from random noise inputs, which inherently lack meaningful information. Consequently, these models struggle to effectively map this noise to the ground-truth sample distribution, resulting in prolonging training times and low-quality outputs. In this paper, we propose a novel Noisy Layer Generation method (NAYER) which re-locates the random source from the input to a noisy layer and utilizes the meaningful constant label-text embedding (LTE) as the input. LTE is generated by using the language model once, and then it is stored in memory for all subsequent training processes. The significance of LTE lies in its ability to contain substantial meaningful inter-class information, enabling the generation of high-quality samples with only a few training steps. Simultaneously, the noisy layer plays a key role in addressing the issue of diversity in sample generation by preventing the model from overemphasizing the constrained label information. By reinitializing the noisy layer in each iteration, we aim to facilitate the generation of diverse samples while still retaining the method's efficiency, thanks to the ease of learning provided by LTE. Experiments carried out on multiple datasets demonstrate that our NAYER not only outperforms the state-of-the-art methods but also achieves speeds 5 to 15 times faster than previous approaches. The code is available at https://github.com/tmtuan1307/nayer. Minh-Tuan Tran, Trung Le 0001, Xuan-May Le, Mehrtash Harandi, Quan Hung Tran, Dinh Q. Phung |
CVPR | 2 |
| 2024 | MetaAug: Meta-data Augmentation for Post-training Quantization
Cuong Pham 0007, Hoang Anh Dung, Cuong Nguyen 0006, Trung Le 0001, Dinh Q. Phung, Gustavo Carneiro 0001, Thanh-Toan Do |
ECCV (27) | 4 |
| 2024 | Preserving Generalization of Language models in Few-shot Continual Relation ExtractionabstractFew-shot Continual Relations Extraction (FCRE) is an emerging and dynamic area of study where models can sequentially integrate knowledge from new relations with limited labeled data while circumventing catastrophic forgetting and preserving prior knowledge from pre-trained backbones.In this work, we introduce a novel method that leverages oftendiscarded language model heads.By employing these components via a mutual information maximization strategy, our approach helps maintain prior knowledge from the pre-trained backbone and strategically aligns the primary classification head, thereby enhancing model performance.Furthermore, we explore the potential of Large Language Models (LLMs), renowned for their wealth of knowledge, in addressing FCRE challenges.Our comprehensive experimental results underscore the efficacy of the proposed method and offer valuable insights for future work. Quyen Tran, Nguyen Hoang Anh, Trung Le 0001, Ngo Van Linh 0001, Thien Huu Nguyen |
EMNLP | 5 |
| 2024 | Sharpness-Aware Data Generation for Zero-shot QuantizationabstractZero-shot quantization aims to learn a quantized model from a pre-trained full-precision model with no access to original real training data. The common idea in zero-shot quantization approaches is to generate synthetic data for quantizing the full-precision model. While it is well-known that deep neural networks with low sharpness have better generalization ability, none of the previous zero-shot quantization works considers the sharpness of the quantized model as a criterion for generating training data. This paper introduces a novel methodology that takes into account quantized model sharpness in synthetic data generation to enhance generalization. Specifically, we first demonstrate that sharpness minimization can be attained by maximizing gradient matching between the reconstruction loss gradients computed on synthetic and real validation data, under certain assumptions. We then circumvent the problem of the gradient matching without real validation set by approximating it with the gradient matching between each generated sample and its neighbors. Experimental evaluations on CIFAR-100 and ImageNet datasets demonstrate the superiority of the proposed method over the state-of-the-art techniques in low-bit quantization settings. Hoang Anh Dung, Cuong Pham 0007, Trung Le 0001, Jianfei Cai 0001, Thanh-Toan Do |
ICML | 3 |
| 2024 | Optimal Transport for Structure Learning Under Missing DataabstractCausal discovery in the presence of missing data introduces a chicken-and-egg dilemma. While the goal is to recover the true causal structure, robust imputation requires considering the dependencies or, preferably, causal relations among variables. Merely filling in missing values with existing imputation methods and subsequently applying structure learning on the complete data is empirically shown to be sub-optimal. To address this problem, we propose a score-based algorithm for learning causal structures from missing data based on optimal transport. This optimal transport viewpoint diverges from existing score-based approaches that are dominantly based on expectation maximization. We formulate structure learning as a density fitting problem, where the goal is to find the causal model that induces a distribution of minimum Wasserstein distance with the observed data distribution. Our framework is shown to recover the true causal graphs more effectively than competing methods in most simulations and real-data settings. Empirical evidence also shows the superior scalability of our approach, along with the flexibility to incorporate any off-the-shelf causal discovery methods for complete data. Vy Vo, He Zhao 0001, Trung Le 0001, Edwin V. Bonilla, Dinh Q. Phung |
ICML | 3 |
| 2024 | Parameter Estimation in DAGs from Incomplete Data via Optimal TransportabstractEstimating the parameters of a probabilistic directed graphical model from incomplete data is a long-standing challenge. This is because, in the presence of latent variables, both the likelihood function and posterior distribution are intractable without assumptions about structural dependencies or model classes. While existing learning methods are fundamentally based on likelihood maximization, here we offer a new view of the parameter learning problem through the lens of optimal transport. This perspective licenses a general framework that operates on any directed graphs without making unrealistic assumptions on the posterior over the latent variables or resorting to variational approximations. We develop a theoretical framework and support it with extensive empirical evidence demonstrating the versatility and robustness of our approach. Across experiments, we show that not only can our method effectively recover the ground-truth parameters but it also performs comparably or better than competing baselines on downstream applications. Vy Vo, Trung Le 0001, Long Tung Vuong, He Zhao 0001, Edwin V. Bonilla, Dinh Q. Phung |
ICML | 2 |
| 2024 | Erasing Undesirable Concepts in Diffusion Models with Adversarial PreservationabstractDiffusion models excel at generating visually striking content from text but can inadvertently produce undesirable or harmful content when trained on unfiltered internet data. A practical solution is to selectively removing target concepts from the model, but this may impact the remaining concepts. Prior approaches have tried to balance this by introducing a loss term to preserve neutral content or a regularization term to minimize changes in the model parameters, yet resolving this trade-off remains challenging. In this work, we propose to identify and preserving concepts most affected by parameter changes, termed as *adversarial concepts*. This approach ensures stable erasure with minimal impact on the other concepts. We demonstrate the effectiveness of our method using the Stable Diffusion model, showing that it outperforms state-of-the-art erasure methods in eliminating unwanted content while maintaining the integrity of other unrelated elements. Our code is available at \url{https://github.com/tuananhbui89/Erasing-Adversarial-Preservation}. Anh Bui, Tung Long Vuong, Khanh Doan, Trung Le 0001, Paul Montague, Tamas Abraham, Dinh Q. Phung |
NeurIPS | 4 |
| 2024 | Explicit Eigenvalue Regularization Improves Sharpness-Aware MinimizationabstractSharpness-Aware Minimization (SAM) has attracted significant attention for its effectiveness in improving generalization across various tasks. However, its underlying principles remain poorly understood. In this work, we analyze SAM’s training dynamics using the maximum eigenvalue of the Hessian as a measure of sharpness and propose a third-order stochastic differential equation (SDE), which reveals that the dynamics are driven by a complex mixture of second- and third-order terms. We show that alignment between the perturbation vector and the top eigenvector is crucial for SAM’s effectiveness in regularizing sharpness, but find that this alignment is often inadequate in practice, which limits SAM's efficiency. Building on these insights, we introduce Eigen-SAM, an algorithm that explicitly aims to regularize the top Hessian eigenvalue by aligning the perturbation vector with the leading eigenvector. We validate the effectiveness of our theory and the practical advantages of our proposed approach through comprehensive experiments. Code is available at https://github.com/RitianLuo/EigenSAM. Haocheng Luo, Tuan Truong, Tung Pham 0001, Mehrtash Harandi, Dinh Q. Phung, Trung Le 0001 |
NeurIPS | 6 |
| 2024 | Enhancing Domain Adaptation through Prompt Gradient AlignmentabstractPrior Unsupervised Domain Adaptation (UDA) methods often aim to train a domain-invariant feature extractor, which may hinder the model from learning sufficiently discriminative features. To tackle this, a line of works based on prompt learning leverages the power of large-scale pre-trained vision-language models to learn both domain-invariant and specific features through a set of domain-agnostic and domain-specific learnable prompts. Those studies typically enforce invariant constraints on representation, output, or prompt space to learn such prompts. Differently, we cast UDA as a multiple-objective optimization problem in which each objective is represented by a domain loss. Under this new framework, we propose aligning per-objective gradients to foster consensus between them. Additionally, to prevent potential overfitting when fine-tuning this deep learning architecture, we penalize the norm of these gradients. To achieve these goals, we devise a practical gradient update procedure that can work under both single-source and multi-source UDA. Empirically, our method consistently surpasses other vision language model adaptation methods by a large margin on a wide range of benchmarks. The implementation is available at https://github.com/VietHoang1512/PGA. Viet Hoang Phan, Tung Lam Tran, Quyen Tran, Trung Le 0001 |
NeurIPS | 4 |
| 2024 | Frequency Attention for Knowledge DistillationabstractKnowledge distillation is an attractive approach for learning compact deep neural networks, which learns a lightweight student model by distilling knowledge from a complex teacher model. Attention-based knowledge distillation is a specific form of intermediate feature-based knowledge distillation that uses attention mechanisms to encourage the student to better mimic the teacher. However, most of the previous attention-based distillation approaches perform attention in the spatial domain, which primarily affects local regions in the input image. This may not be sufficient when we need to capture the broader context or global information necessary for effective knowledge transfer. In frequency domain, since each frequency is determined from all pixels of the image in spatial domain, it can contain global information about the image. Inspired by the benefits of the frequency domain, we propose a novel module that functions as an attention mechanism in the frequency domain. The module consists of a learnable global filter that can adjust the frequencies of student’s features under the guidance of the teacher’s features, which encourages the student’s features to have patterns similar to the teacher’s features. We then propose an enhanced knowledge review-based distillation model by leveraging the proposed frequency attention module. The extensive experiments with various teacher and student architectures on image classification and object detection benchmark datasets show that the proposed approach outperforms other knowledge distillation methods. Cuong Pham 0007, Van-Anh Nguyen, Trung Le 0001, Dinh Q. Phung, Gustavo Carneiro 0001, Thanh-Toan Do |
WACV | 3 |
| 2024 | AIBugHunter: A Practical tool for predicting, classifying and repairing software vulnerabilitiesabstractAbstract Many Machine Learning(ML)-based approaches have been proposed to automatically detect, localize, and repair software vulnerabilities. While ML-based methods are more effective than program analysis-based vulnerability analysis tools, few have been integrated into modern Integrated Development Environments (IDEs), hindering practical adoption. To bridge this critical gap, we propose in this article AIBugHunter , a novel Machine Learning-based software vulnerability analysis tool for C/C++ languages that is integrated into the Visual Studio Code (VS Code) IDE. AIBugHunter helps software developers to achieve real-time vulnerability detection, explanation, and repairs during programming. In particular, AIBugHunter scans through developers’ source code to (1) locate vulnerabilities, (2) identify vulnerability types, (3) estimate vulnerability severity, and (4) suggest vulnerability repairs. We integrate our previous works (i.e., LineVul and VulRepair) to achieve vulnerability localization and repairs. In this article, we propose a novel multi-objective optimization (MOO)-based vulnerability classification approach and a transformer-based estimation approach to help AIBugHunter accurately identify vulnerability types and estimate severity. Our empirical experiments on a large dataset consisting of 188K+ C/C++ functions confirm that our proposed approaches are more accurate than other state-of-the-art baseline methods for vulnerability classification and estimation. Furthermore, we conduct qualitative evaluations including a survey study and a user study to obtain software practitioners’ perceptions of our AIBugHunter tool and assess the impact that AIBugHunter may have on developers’ productivity in security aspects. Our survey study shows that our AIBugHunter is perceived as useful where 90% of the participants consider adopting our AIBugHunter during their software development. Last but not least, our user study shows that our AIBugHunter can enhance developers’ productivity in combating cybersecurity issues during software development. AIBugHunter is now publicly available in the Visual Studio Code marketplace. Chakkrit Tantithamthavorn, Trung Le 0001, Yuki Kume, Van Nguyen 0002, Dinh Q. Phung, John C. Grundy |
Empir. Softw. Eng. | 3 |
| 2024 | Vision Transformer Inspired Automated Vulnerability RepairabstractRecently, automated vulnerability repair approaches have been widely adopted to combat increasing software security issues. In particular, transformer-based encoder-decoder models achieve competitive results. Whereas vulnerable programs may only consist of a few vulnerable code areas that need repair, existing AVR approaches lack a mechanism guiding their model to pay more attention to vulnerable code areas during repair generation. In this article, we propose a novel vulnerability repair framework inspired by the Vision Transformer based approaches for object detection in the computer vision domain. Similar to the object queries used to locate objects in object detection in computer vision, we introduce and leverage vulnerability queries (VQs) to locate vulnerable code areas and then suggest their repairs. In particular, we leverage the cross-attention mechanism to achieve the cross-match between VQs and their corresponding vulnerable code areas. To strengthen our cross-match and generate more accurate vulnerability repairs, we propose to learn a novel vulnerability mask (VM) and integrate it into decoders’ cross-attention, which makes our VQs pay more attention to vulnerable code areas during repair generation. In addition, we incorporate our VM into encoders’ self-attention to learn embeddings that emphasize the vulnerable areas of a program. Through an extensive evaluation using the real-world 5,417 vulnerabilities, our approach outperforms all of the automated vulnerability repair baseline methods by 2.68% to 32.33%. Additionally, our analysis of the cross-attention map of our approach confirms the design rationale of our VM and its effectiveness. Finally, our survey study with 71 software practitioners highlights the significance and usefulness of AI-generated vulnerability repairs in the realm of software security. The training code and pre-trained models are available at https://github.com/awsm-research/VQM. Van Nguyen 0002, Chakkrit Tantithamthavorn, Dinh Q. Phung, Trung Le 0001 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2024 | Deep Domain Adaptation With Max-Margin Principle for Cross-Project Imbalanced Software Vulnerability DetectionabstractSoftware vulnerabilities (SVs) have become a common, serious, and crucial concern due to the ubiquity of computer software. Many AI-based approaches have been proposed to solve the software vulnerability detection (SVD) problem to ensure the security and integrity of software applications (in both the development and testing phases). However, there are still two open and significant issues for SVD in terms of (i) learning automatic representations to improve the predictive performance of SVD, and (ii) tackling the scarcity of labeled vulnerability datasets that conventionally need laborious labeling effort by experts. In this paper, we propose a novel approach to tackle these two crucial issues. We first exploit the automatic representation learning with deep domain adaptation for SVD. We then propose a novel cross-domain kernel classifier leveraging the max-margin principle to significantly improve the transfer learning process of SVs from imbalanced labeled into imbalanced unlabeled projects. Our approach is the first work that leverages solid body theories of the max-margin principle, kernel methods, and bridging the gap between source and target domains for imbalanced domain adaptation (DA) applied in cross-project SVD . The experimental results on real-world software datasets show the superiority of our proposed method over state-of-the-art baselines. In short, our method obtains a higher performance on F1-measure, one of the most important measures in SVD, from 1.83% to 6.25% compared to the second highest method in the used datasets. Van Nguyen 0002, Trung Le 0001, Chakkrit Tantithamthavorn, John C. Grundy, Dinh Q. Phung |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | Global-Local Regularization Via Distributional RobustnessabstractDespite superior performance in many situations, deep neural networks are often vulnerable to adversarial examples and distribution shifts, limiting model generalization ability in real-world applications. To alleviate these problems, recent approaches leverage distributional robustness optimization (DRO) to find the most challenging distribution, and then minimize loss function over this most challenging distribution. Regardless of having achieved some improvements, these DRO approaches have some obvious limitations. First, they purely focus on local regularization to strengthen model robustness, missing a global regularization effect that is useful in many real-world applications (e.g., domain adaptation, domain generalization, and adversarial machine learning). Second, the loss functions in the existing DRO approaches operate in only the most challenging distribution, hence decouple with the original distribution, leading to a restrictive modeling capability. In this paper, we propose a novel regularization technique, following the veins of Wasserstein-based DRO framework. Specifically, we define a particular joint distribution and Wasserstein-based uncertainty, allowing us to couple the original and most challenging distributions for enhancing modeling capability and applying both local and global regularizations. Empirical studies on different learning problems demonstrate that our proposed approach significantly outperforms the existing regularization approaches in various domains. Hoang Phan, Trung Le 0001, Anh Tuan Bui, Nhat Ho, Dinh Q. Phung |
AISTATS | 2 |
| 2023 | ChatGPT for Vulnerability Detection, Classification, and Repair: How Far Are We?abstractLarge language models (LLMs) like ChatGPT (i.e., gpt-3.5-turbo and gpt-4) exhibited remarkable advancement in a range of software engineering tasks associated with source code such as code review and code generation. In this paper, we undertake a comprehensive study by instructing ChatGPT for four prevalent vulnerability tasks: function and line-level vulnerability prediction, vulnerability classification, severity estimation, and vulnerability repair. We compare ChatGPT with state-of-the-art language models designed for software vulnerability purposes. Through an empirical assessment employing extensive real-world datasets featuring over 190,000 C/C++ functions, we found that ChatGPT achieves limited performance, trailing behind other language models in vulnerability contexts by a significant margin. The experimental outcomes highlight the challenging nature of vulnerability prediction tasks, requiring domain-specific expertise. Despite ChatGPT's substantial model scale, exceeding that of source code-pre-trained language models (e.g., CodeBERT) by a factor of 14,000, the process of fine-tuning remains imperative for ChatGPT to generalize for vulnerability prediction tasks. We publish the studied dataset, experimental prompts for ChatGPT, and experimental results at https://github.com/awsm-research/ChatGPT4Vul. Chakkrit Tantithamthavorn, Van Nguyen 0002, Trung Le 0001 |
APSEC | 4 |
| 2023 | An Additive Instance-Wise Approach to Multi-class Model Interpretation
Vy Vo, Van Nguyen 0002, Trung Le 0001, Quan Hung Tran, Gholamreza Haffari, Seyit Ahmet Çamtepe, Dinh Q. Phung |
ICLR | 3 |
| 2023 | Vector Quantized Wasserstein Auto-EncoderabstractLearning deep discrete latent presentations offers a promise of better symbolic and summarized abstractions that are more useful to subsequent downstream tasks. Inspired by the seminal Vector Quantized Variational Auto-Encoder (VQ-VAE), most of work in learning deep discrete representations has mainly focused on improving the original VQ-VAE form and none of them has studied learning deep discrete representations from the generative viewpoint. In this work, we study learning deep discrete representations from the generative viewpoint. Specifically, we endow discrete distributions over sequences of codewords and learn a deterministic decoder that transports the distribution over the sequences of codewords to the data distribution via minimizing a WS distance between them. We develop further theories to connect it with the clustering viewpoint of WS distance, allowing us to have a better and more controllable clustering solution. Finally, we empirically evaluate our method on several well-known benchmarks, where it achieves better qualitative and quantitative performances than the other VQ-VAE variants in terms of the codebook utilization and image reconstruction/generation. Long Tung Vuong, Trung Le 0001, He Zhao 0001, Chuanxia Zheng, Mehrtash Harandi, Jianfei Cai 0001, Dinh Q. Phung |
ICML | 2 |
| 2023 | Feature-based Learning for Diverse and Privacy-Preserving Counterfactual ExplanationsabstractInterpretable machine learning seeks to understand the reasoning process of complex black-box systems that are long notorious for lack of explainability. One flourishing approach is through counterfactual explanations, which provide suggestions on what a user can do to alter an outcome. Not only must a counterfactual example counter the original prediction from the black-box classifier but it should also satisfy various constraints for practical applications. Diversity is one of the critical constraints that however remains less discussed. While diverse counterfactuals are ideal, it is computationally challenging to simultaneously address some other constraints. Furthermore, there is a growing privacy concern over the released counterfactual data. To this end, we propose a feature-based learning framework that effectively handles the counterfactual constraints and contributes itself to the limited pool of private explanation models. We demonstrate the flexibility and effectiveness of our method in generating diverse counterfactuals of actionability and plausibility. Our counterfactual engine is more efficient than counterparts of the same capacity while yielding the lowest re-identification risks. Vy Vo, Trung Le 0001, Van Nguyen 0002, He Zhao 0001, Edwin V. Bonilla, Gholamreza Haffari, Dinh Q. Phung |
KDD | 2 |
| 2023 | Cross-Adversarial Local Distribution Regularization for Semi-supervised Medical Image Segmentation
Thanh Nguyen-Duc, Trung Le 0001, Roland Bammer, He Zhao 0001, Jianfei Cai 0001, Dinh Q. Phung |
MICCAI (1) | 2 |
| 2023 | Optimal Transport Model Distributional RobustnessabstractDistributional robustness is a promising framework for training deep learning models that are less vulnerable to adversarial examples and data distribution shifts. Previous works have mainly focused on exploiting distributional robustness in the data space. In this work, we explore an optimal transport-based distributional robustness framework in model spaces. Specifically, we examine a model distribution within a Wasserstein ball centered on a given model distribution that maximizes the loss. We have developed theories that enable us to learn the optimal robust center model distribution. Interestingly, our developed theories allow us to flexibly incorporate the concept of sharpness awareness into training, whether it's a single model, ensemble models, or Bayesian Neural Networks, by considering specific forms of the center model distribution. These forms include a Dirac delta distribution over a single model, a uniform distribution over several models, and a general Bayesian Neural Network. Furthermore, we demonstrate that Sharpness-Aware Minimization (SAM) is a specific case of our framework when using a Dirac delta distribution over a single model, while our framework can be seen as a probabilistic extension of SAM. To validate the effectiveness of our framework in the aforementioned settings, we conducted extensive experiments, and the results reveal remarkable improvements compared to the baselines. Van-Anh Nguyen, Trung Le 0001, Anh Tuan Bui, Thanh-Toan Do, Dinh Q. Phung |
NeurIPS | 2 |
| 2023 | Flat Seeking Bayesian Neural NetworksabstractBayesian Neural Networks (BNNs) provide a probabilistic interpretation for deep learning models by imposing a prior distribution over model parameters and inferring a posterior distribution based on observed data. The model sampled from the posterior distribution can be used for providing ensemble predictions and quantifying prediction uncertainty. It is well-known that deep learning models with lower sharpness have better generalization ability. However, existing posterior inferences are not aware of sharpness/flatness in terms of formulation, possibly leading to high sharpness for the models sampled from them. In this paper, we develop theories, the Bayesian setting, and the variational inference approach for the sharpness-aware posterior. Specifically, the models sampled from our sharpness-aware posterior, and the optimal approximate posterior estimating this sharpness-aware posterior, have better flatness, hence possibly possessing higher generalization ability. We conduct experiments by leveraging the sharpness-aware posterior with state-of-the-art Bayesian Neural Networks, showing that the flat-seeking counterparts outperform their baselines in all metrics of interest. Van-Anh Nguyen, Tung Long Vuong, Hoang Phan, Thanh-Toan Do, Dinh Q. Phung, Trung Le 0001 |
NeurIPS | 6 |
| 2023 | Model and Feature Diversity for Bayesian Neural Networks in Mutual LearningabstractBayesian Neural Networks (BNNs) offer probability distributions for model parameters, enabling uncertainty quantification in predictions. However, they often underperform compared to deterministic neural networks. Utilizing mutual learning can effectively enhance the performance of peer BNNs. In this paper, we propose a novel approach to improve BNNs performance through deep mutual learning. The proposed approaches aim to increase diversity in both network parameter distributions and feature distributions, promoting peer networks to acquire distinct features that capture different characteristics of the input, which enhances the effectiveness of mutual learning. Experimental results demonstrate significant improvements in the classification accuracy, negative log-likelihood, and expected calibration error when compared to traditional mutual learning for BNNs. Van Cuong Pham, Cuong Nguyen 0006, Trung Le 0001, Dinh Q. Phung, Gustavo Carneiro 0001, Thanh-Toan Do |
NeurIPS | 3 |
| 2023 | Adversarial local distribution regularization for knowledge distillationabstractKnowledge distillation is a process of distilling information from a large model with significant knowledge capacity (teacher) to enhance a smaller model (student). Therefore, exploring the properties of the teacher is the key to improving student performance (e.g., teacher decision boundaries). One decision boundary exploring technique is to leverage adversarial attack methods, which add crafted perturbations within a ball constraint to clean inputs to create attack examples of the teacher called adversarial examples. These adversarial examples are informative examples because they are near decision boundaries. In this paper, we formulate a teacher adversarial local distribution, a set of all adversarial examples within the ball constraint given an input. This distribution is used to sufficiently explore the decision boundaries of the teacher by covering the full spectrum of possible teacher model perturbations. The student model is then regularized by matching the loss between teacher and student using these adversarial example inputs. We conducted a number of experiments on CIFAR-100 and Imagenet datasets to illustrate this teacher adversarial local distribution regularization (TALD) can be applied to improve performance of many existing knowledge distillation methods (e.g., KD, FitNet, CRD, VID, FT, etc.). Thanh Nguyen-Duc, Trung Le 0001, He Zhao 0001, Jianfei Cai 0001, Dinh Q. Phung |
WACV | 2 |
| 2023 | VulExplainer: A Transformer-Based Hierarchical Distillation for Explaining Vulnerability TypesabstractDeep learning-based vulnerability prediction approaches are proposed to help under-resourced security practitioners to detect vulnerable functions. However, security practitioners still do not know what type of vulnerabilities correspond to a given prediction (aka CWE-ID). Thus, a novel approach to explain the type of vulnerabilities for a given prediction is imperative. In this paper, we proposeVulExplainer, an approach to explain the type of vulnerabilities. We representVulExplaineras a vulnerability classification task. However, vulnerabilities have diverse characteristics (i.e., CWE-IDs) and the number of labeled samples in each CWE-ID is highly imbalanced (known as a highly imbalanced multi-class classification problem), which often lead to inaccurate predictions. Thus, we introduce a Transformer-based hierarchical distillation for software vulnerability classification in order to address the highly imbalanced types of software vulnerabilities. Specifically, we split a complex label distribution into sub-distributions based on CWE abstract types (i.e., categorizations that group similar CWE-IDs). Thus, similar CWE-IDs can be grouped and each group will have a more balanced label distribution. We learn TextCNN teachers on each of the simplified distributions respectively, however, they only perform well in their group. Thus, we build a transformer student model to generalize the performance of TextCNN teachers through our hierarchical knowledge distillation framework. Through an extensive evaluation using the real-world 8,636 vulnerabilities, our approach outperforms all of the baselines by 5%–29%. The results also demonstrate that our approach can be applied to Transformer-based architectures such as CodeBERT, GraphCodeBERT, and CodeGPT. Moreover, our method maintains compatibility with any Transformer-based model without requiring any architectural modifications but only adds a special distillation token to the input. These results highlight our significant contributions towards the fundamental and practical problem of explaining software vulnerability. Van Nguyen 0002, Chakkrit Tantithamthavorn, Trung Le 0001, Dinh Q. Phung |
IEEE Trans. Software Eng. | 4 |
| 2022 | On Global-view Based Defense via Adversarial Attack and Defense Risk Guaranteed BoundsabstractIt is well-known that deep neural networks (DNNs) are susceptible to adversarial attacks, which presents the most severe fragility of the deep learning system. Despite achieving impressive performance, most of the current state-of-the-art classifiers remain highly vulnerable to carefully crafted imperceptible, adversarial perturbations. Recent research attempts to understand neural network attack and defense have become increasingly urgent and important. While rapid progress has been made on this front, there is still an important theoretical gap in achieving guaranteed bounds on attack/defense models, leaving uncertainty in the quality and certified guarantees of these models. To this end, we systematically address this problem in this paper. More specifically, we formulate attack and defense in a generic setting where there exists a family of adversaries (i.e., attackers) for attacking a family of classifiers (i.e., defenders). We develop a novel class of f-divergences suitable for measuring divergence among multiple distributions. This equips us to study the interactions between attackers and defenders in a countervailing game where we formulate a joint risk on attack and defense schemes. This is followed by our key results on guaranteed upper and lower bounds on this risk that can provide a better understanding of the behaviors of those parties from the attack and defense perspectives, thereby having important implications to both attack and defense sides. Finally, benefited from our theory, we propose an empirical approach that bases on a global view to defend against adversarial attacks. The experimental results conducted on benchmark datasets show that the global view for attack/defense if exploited appropriately can help to improve adversarial robustness. Trung Le 0001, Anh Tuan Bui, Le Minh Tri Tue, He Zhao 0001, Paul Montague, Quan Hung Tran, Dinh Q. Phung |
AISTATS | 1 |
| 2022 | Particle-based Adversarial Local Distribution RegularizationabstractAdversarial training defense (ATD) and virtual adversarial training (VAT) are the two most effective methods to improve model robustness against attacks and model generalization. While ATD is usually applied in robust machine learning, VAT is used in semi-supervised learning and domain adaption. In this paper, we introduce a novel adversarial local distribution regularization. The adversarial local distribution is defined by a set of all adversarial examples within a ball constraint given a natural input. We illustrate this regularization is a general form of previous methods (e.g., PGD, TRADES, VAT and VADA). We conduct comprehensive experiments on MNIST, SVHN and CIFAR10 to illustrate that our method outperforms well-known methods such as PGD, TRADES and ADT in robust machine learning, VAT in semi-supervised learning and VADA in domain adaption. Our implementation is on Github: https://github.com/PotatoThanh/ALD-Regularization. Thanh Nguyen-Duc, Trung Le 0001, He Zhao 0001, Jianfei Cai 0001, Dinh Q. Phung |
AISTATS | 2 |
| 2022 | A Unified Wasserstein Distributional Robustness Framework for Adversarial Training
Anh Tuan Bui, Trung Le 0001, Quan Hung Tran, He Zhao 0001, Dinh Q. Phung |
ICLR | 2 |
| 2022 | On Transportation of Mini-batches: A Hierarchical ApproachabstractMini-batch optimal transport (m-OT) has been successfully used in practical applications that involve probability measures with a very high number of supports. The m-OT solves several smaller optimal transport problems and then returns the average of their costs and transportation plans. Despite its scalability advantage, the m-OT does not consider the relationship between mini-batches which leads to undesirable estimation. Moreover, the m-OT does not approximate a proper metric between probability measures since the identity property is not satisfied. To address these problems, we propose a novel mini-batch scheme for optimal transport, named Batch of Mini-batches Optimal Transport (BoMb-OT), that finds the optimal coupling between mini-batches and it can be seen as an approximation to a well-defined distance on the space of probability measures. Furthermore, we show that the m-OT is a limit of the entropic regularized version of the BoMb-OT when the regularized parameter goes to infinity. Finally, we carry out experiments on various applications including deep generative models, deep domain adaptation, approximate Bayesian computation, color transfer, and gradient flow to show that the BoMb-OT can be widely applied and performs well in various applications. Dang Nguyen 0002, Quoc Dinh Nguyen, Tung Pham 0001, Hung Hai Bui, Dinh Q. Phung, Trung Le 0001, Nhat Ho |
ICML | 7 |
| 2022 | Stochastic Multiple Target Sampling Gradient DescentabstractSampling from an unnormalized target distribution is an essential problem with many applications in probabilistic inference. Stein Variational Gradient Descent (SVGD) has been shown to be a powerful method that iteratively updates a set of particles to approximate the distribution of interest. Furthermore, when analysing its asymptotic properties, SVGD reduces exactly to a single-objective optimization problem and can be viewed as a probabilistic version of this single-objective optimization problem. A natural question then arises: ``Can we derive a probabilistic version of the multi-objective optimization?''. To answer this question, we propose Stochastic Multiple Target Sampling Gradient Descent (MT-SGD), enabling us to sample from multiple unnormalized target distributions. Specifically, our MT-SGD conducts a flow of intermediate distributions gradually orienting to multiple target distributions, which allows the sampled particles to move to the joint high-likelihood region of the target distributions. Interestingly, the asymptotic analysis shows that our approach reduces exactly to the multiple-gradient descent algorithm for multi-objective optimization, as expected. Finally, we conduct comprehensive experiments to demonstrate the merit of our approach to multi-task learning. Hoang Phan, Ngoc Tran, Trung Le 0001, Toan Tran 0003, Nhat Ho, Dinh Q. Phung |
NeurIPS | 3 |
| 2022 | VulRepair: a T5-based automated software vulnerability repairabstractAs software vulnerabilities grow in volume and complexity, researchers proposed various Artificial Intelligence (AI)-based approaches to help under-resourced security analysts to find, detect, and localize vulnerabilities. However, security analysts still have to spend a huge amount of effort to manually fix or repair such vulnerable functions. Recent work proposed an NMT-based Automated Vulnerability Repair, but it is still far from perfect due to various limitations. In this paper, we propose VulRepair, a T5-based automated software vulnerability repair approach that leverages the pre-training and BPE components to address various technical limitations of prior work. Through an extensive experiment with over 8,482 vulnerability fixes from 1,754 real-world software projects, we find that our VulRepair achieves a Perfect Prediction of 44%, which is 13%-21% more accurate than competitive baseline approaches. These results lead us to conclude that our VulRepair is considerably more accurate than two baseline approaches, highlighting the substantial advancement of NMT-based Automated Vulnerability Repairs. Our additional investigation also shows that our VulRepair can accurately repair as many as 745 out of 1,706 real-world well-known vulnerabilities (e.g., Use After Free, Improper Input Validation, OS Command Injection), demonstrating the practicality and significance of our VulRepair for generating vulnerability repairs, helping under-resourced security analysts on fixing vulnerabilities. Chakkrit Tantithamthavorn, Trung Le 0001, Van Nguyen 0002, Dinh Q. Phung |
ESEC/SIGSOFT FSE | 3 |
| 2022 | Cycle class consistency with distributional optimal transport and knowledge distillation for unsupervised domain adaptationabstractUnsupervised domain adaptation (UDA) aims to transfer knowledge from a model trained on a labeled source domain to an unlabeled target domain. To this end, we propose in this paper a novel cycle class-consistent model based on optimal transport (OT) and knowledge distillation. The model consists of two agents, a teacher and a student cooperatively working in a cycle process under the guidance of the distributional optimal transport and distillation manner. The OT distance is designed to bridge the gap between the distribution of the target data and a distribution over the source class-conditional distributions. The optimal probability matrix then provides pseudo labels to learn a teacher that achieves a good classification performance on the target domain. Knowledge distillation is performed in the next step in which the teacher distills and transfers its knowledge to the student. And finally, the student produces its prediction for the optimal transport step. This process forms a closed cycle in which the teacher and student networks are simultaneously trained to conduct transfer learning from the source to the target domain. Extensive experiments show that our proposed method outperforms existing methods, especially the class-aware and OT-based ones on benchmark datasets including Office-31, Office-Home, and ImageCLEF-DA. Tuan Nguyen 0004, Van Nguyen 0002, Trung Le 0001, He Zhao 0001, Quan Hung Tran, Dinh Q. Phung |
UAI | 3 |
| 2022 | Improving kernel online learning with a snapshot memory
Trung Le 0001, Dinh Q. Phung |
Mach. Learn. | 1 |
| 2022 | Robust Variational Learning for Multiclass Kernel Models With Stein RefinementabstractKernel-based models have a strong generalization ability, but most, including SVM, are vulnerable to the curse of kernelization. Moreover, their predictive performance is sensitive to hyperparameter tuning, which demands high computational resources. These problems render kernel methods problematic when dealing with large-scale datasets. To this end, we first formulate the optimization problem in a kernel-based learning setting as a posterior inference problem, and then develop a rich family of Recurrent Neural Network-based variational inference techniques. Unlike existing literature, which stops at the variational distribution and uses it as the surrogate for the true posterior distribution, here we further leverage Stein Variational Gradient Descent to further bring the variational distribution closer to the true posterior, we refer to this step asStein Refinement. Putting these altogether, we arrive at a robust and efficient variational learning method for multiclass kernel machines with extremely accurate approximation. Moreover, our formulation enables efficient learning of kernel parameters and hyperparameters which robustifies the proposed method against data uncertainties. The extensive experiments show that without tuning any parameter on modest quantities of data our method obtains comparable accuracy to LIBSVM, a well-known implementation of SVM, and outperforms other baselines, while being able to seamlessly scale with large-scale datasets. Trung Le 0001, Tu Dinh Nguyen, Geoffrey I. Webb, Dinh Q. Phung |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Improving Ensemble Robustness by Collaboratively Promoting and Demoting Adversarial RobustnessabstractEnsemble-based Adversarial Training is a principled approach to achieve robustness against adversarial attacks. An important technicality of this approach is to control the transferability of adversarial examples between ensemble members. We propose in this work a simple, but effective strategy to collaborate among committee models of an ensemble model. This is achieved via the secure and insecure sets defined for each model member on a given sample, hence help us to quantify and regularize the transferability. Consequently, our proposed framework provides the flexibility to reduce the adversarial transferability as well as promote the diversity of ensemble members, which are two crucial factors for better robustness in our ensemble approach. We conduct extensive and comprehensive experiments to demonstrate that our proposed method outperforms the state-of-the-art ensemble baselines, at the same time can detect a wide range of adversarial examples with a near perfect accuracy. Tuan-Anh Bui, Trung Le 0001, He Zhao 0001, Paul Montague, Olivier Y. de Vel, Tamas Abraham, Dinh Q. Phung |
AAAI | 2 |
| 2021 | STEM: An approach to Multi-source Domain Adaptation with GuaranteesabstractMulti-source Domain Adaptation (MSDA) is more practical but challenging than the conventional unsupervised domain adaptation due to the involvement of diverse multiple data sources. Two fundamental challenges of MSDA are: (i) how to deal with the diversity in the multiple source domains and (ii) how to cope with the data shift between the target domain and the source domains. In this paper, to address the first challenge, we propose a theoretical-guaranteed approach to combine domain experts locally trained on its own source domain to achieve a combined multi-source teacher that globally predicts well on the mixture of source domains. To address the second challenge, we propose to bridge the gap between the target domain and the mixture of source domains in the latent space via a generator or feature extractor. Together with bridging the gap in the latent space, we train a student to mimic the predictions of the teacher expert on both source and target examples. In addition, our approach is guaranteed with rigorous theory offered insightful justifications of how each component influences the transferring performance. Extensive experiments conducted on three benchmark datasets show that our proposed method achieves state-of-the-art performances to the best of our knowledge. Van-Anh Nguyen, Tuan Nguyen 0004, Trung Le 0001, Quan Hung Tran, Dinh Q. Phung |
ICCV | 3 |
| 2021 | Neural Topic Model via Optimal Transport
He Zhao 0001, Dinh Q. Phung, Viet Huynh, Trung Le 0001, Wray L. Buntine |
ICLR | 4 |
| 2021 | LAMDA: Label Matching Deep Domain AdaptationabstractDeep domain adaptation (DDA) approaches have recently been shown to perform better than their shallow rivals with better modeling capacity on complex domains (e.g., image, structural data, and sequential data). The underlying idea is to learn domain invariant representations on a latent space that can bridge the gap between source and target domains. Several theoretical studies have established insightful understanding and the benefit of learning domain invariant features; however, they are usually limited to the case where there is no label shift, hence hindering its applicability. In this paper, we propose and study a new challenging setting that allows us to use a Wasserstein distance (WS) to not only quantify the data shift but also to define the label shift directly. We further develop a theory to demonstrate that minimizing the WS of the data shift leads to closing the gap between the source and target data distributions on the latent space (e.g., an intermediate layer of a deep net), while still being able to quantify the label shift with respect to this latent space. Interestingly, our theory can consequently explain certain drawbacks of learning domain invariant features on the latent space. Finally, grounded on the results and guidance of our developed theory, we propose the Label Matching Deep Domain Adaptation (LAMDA) approach that outperforms baselines on real-world datasets for DA problems. Trung Le 0001, Tuan Nguyen 0004, Nhat Ho, Hung Hai Bui, Dinh Q. Phung |
ICML | 1 |
| 2021 | TIDOT: A Teacher Imitation Learning Approach for Domain Adaptation with Optimal TransportabstractUsing the principle of imitation learning and the theory of optimal transport we propose in this paper a novel model for unsupervised domain adaptation named Teacher Imitation Domain Adaptation with Optimal Transport (TIDOT). Our model includes two cooperative agents: a teacher and a student. The former agent is trained to be an expert on labeled data in the source domain, whilst the latter one aims to work with unlabeled data in the target domain. More specifically, optimal transport is applied to quantify the total of the distance between embedded distributions of the source and target data in the joint space, and the distance between predictive distributions of both agents, thus by minimizing this quantity TIDOT could mitigate not only the data shift but also the label shift. Comprehensive empirical studies show that TIDOT outperforms existing state-of-the-art performance on benchmark datasets. Tuan Nguyen 0004, Trung Le 0001, Nhan Dam, Quan Hung Tran, Truyen Nguyen, Dinh Q. Phung |
IJCAI | 2 |
| 2021 | Information-theoretic Source Code Vulnerability HighlightingabstractSoftware vulnerabilities are a crucial and serious concern in the software industry and computer security. A variety of methods have been proposed to detect vulnerabilities in real-world software. Recent methods based on deep learning approaches for automatic feature extraction have improved software vulnerability identification compared with machine learning approaches based on hand-crafted feature extraction. However, these methods can usually only detect software vulnerabilities at a function or program level, which is much less informative because, out of hundreds (thousands) of code statements in a program or function, only a few core statements contribute to a software vulnerability. This requires us to find a way to detect software vulnerabilities at a fine-grained level. In this paper, we propose a novel method based on the concept of mutual information that can help us to detect and isolate software vulnerabilities at a fine-grained level (i.e., several statements that are highly relevant to a software vulnerability that include the core vulnerable statements) in both unsupervised and semi-supervised contexts. We conduct comprehensive experiments on real-world software projects to demonstrate that our proposed method can detect vulnerabilities at a fine-grained level by identifying several statements that mostly contribute to the vulnerability detection decision. Van Nguyen 0002, Trung Le 0001, Olivier Y. de Vel, Paul Montague, John C. Grundy, Dinh Q. Phung |
IJCNN | 2 |
| 2021 | On Learning Domain-Invariant Representations for Transfer Learning with Multiple SourcesabstractDomain adaptation (DA) benefits from the rigorous theoretical works that study its insightful characteristics and various aspects, e.g., learning domain-invariant representations and its trade-off. However, it seems not the case for the multiple source DA and domain generalization (DG) settings which are remarkably more complicated and sophisticated due to the involvement of multiple source domains and potential unavailability of target domain during training. In this paper, we develop novel upper-bounds for the target general loss which appeal us to define two kinds of domain-invariant representations. We further study the pros and cons as well as the trade-offs of enforcing learning each domain-invariant representation. Finally, we conduct experiments to inspect the trade-off of these representations for offering practical hints regarding how to use them in practice and explore other interesting properties of our developed theory. Trung Le 0001, Long Vuong, Toan Tran 0003, Anh Tuan Tran 0001, Hung Hai Bui, Dinh Q. Phung |
NeurIPS | 2 |
| 2021 | Most: multi-source domain adaptation via optimal transport for student-teacher learningabstractMulti-source domain adaptation (DA) is more challenging than conventional DA because the knowledge is transferred from several source domains to a target domain. To this end, we propose in this paper a novel model for multi-source DA using the theory of optimal transport and imitation learning. More specifically, our approach consists of two cooperative agents: a teacher classifier and a student classifier. The teacher classifier is a combined expert that leverages knowledge of domain experts that can be theoretically guaranteed to handle perfectly source examples, while the student classifier acting on the target domain tries to imitate the teacher classifier acting on the source domains. Our rigorous theory developed based on optimal transport makes this cross-domain imitation possible and also helps to mitigate not only the data shift but also the label shift, which are inherently thorny issues in DA research. We conduct comprehensive experiments on real-world datasets to demonstrate the merit of our approach and its optimal transport based imitation learning viewpoint. Experimental results show that our proposed method achieves state-of-the-art performance on benchmark datasets for multi-source domain adaptation including Digits-five, Office-Caltech10, and Office-31 to the best of our knowledge. Tuan Nguyen 0004, Trung Le 0001, He Zhao 0001, Quan Hung Tran, Truyen Nguyen, Dinh Q. Phung |
UAI | 2 |
| 2020 | Explain by Evidence: An Explainable Memory-based Neural Network for Question AnsweringabstractInterpretability and explainability of deep neural networks are challenging due to their scale, complexity, and the agreeable notions on which the explaining process rests.Previous work, in particular, has focused on representing internal components of neural networks through humanfriendly visuals and concepts.On the other hand, in real life, when making a decision, human tends to rely on similar situations and/or associations in the past.Hence arguably, a promising approach to make the model transparent is to design it in a way such that the model explicitly connects the current sample with the seen ones, and bases its decision on these samples.Grounded on that principle, we propose in this paper an explainable, evidence-based memory network architecture, which learns to summarize the dataset and extract supporting evidences to make its decision.Our model achieves state-of-the-art performance on two popular question answering datasets (i.e.TrecQA and WikiQA).Via further analysis, we show that this model can reliably trace the errors it has made in the validation step to the training instances that might have caused these errors.We believe that this error-tracing capability provides significant benefit in improving dataset quality in many applications. Quan Hung Tran, Nhan Dam, Tuan Manh Lai, Franck Dernoncourt, Trung Le 0001, Nham Le, Dinh Q. Phung |
COLING | 5 |
| 2020 | Improving Adversarial Robustness by Enforcing Local and Global Compactness
Tuan-Anh Bui, Trung Le 0001, He Zhao 0001, Paul Montague, Olivier Y. de Vel, Tamas Abraham, Dinh Q. Phung |
ECCV (27) | 2 |
| 2020 | Parameterized Rate-Distortion Stochastic EncoderabstractWe propose a novel gradient-based tractable approach for the Blahut-Arimoto (BA) algorithm to compute the rate-distortion function where the BA algorithm is fully parameterized. This results in a rich and flexible framework to learn a new class of stochastic encoders, termed PArameterized RAte-DIstortion Stochastic Encoder (PARADISE). The framework can be applied to a wide range of settings from semi-supervised, multi-task to supervised and robust learning. We show that the training objective of PARADISE can be seen as a form of regularization that helps improve generalization. With an emphasis on robust learning we further develop a novel posterior matching objective to encourage smoothness on the loss function and show that PARADISE can significantly improve interpretability as well as robustness to adversarial attacks on the CIFAR-10 and ImageNet datasets. In particular, on the CIFAR-10 dataset, our model reduces standard and adversarial error rates in comparison to the state-of-the-art by 50% and 41%, respectively without the expensive computational cost of adversarial training. Quan Hoang, Trung Le 0001, Dinh Q. Phung |
ICML | 2 |
| 2020 | Explain2Attack: Text Adversarial Attacks via Cross-Domain InterpretabilityabstractTraining robust deep learning models for downstream tasks is a critical challenge. Research has shown that down-stream models can be easily fooled with adversarial inputs that look like the training data, but slightly perturbed, in a way imperceptible to humans. Understanding the behavior of natural language models under these attacks is crucial to better defend these models against such attacks. In the black-box attack setting, where no access to model parameters is available, the attacker can only query the output information from the targeted model to craft a successful attack. Current black-box state-of-the-art models are costly in both computational complexity and number of queries needed to craft successful adversarial examples. For real world scenarios, the number of queries is critical, where less queries are desired to avoid suspicion towards an attacking agent. In this paper, we propose Explain2Attack, a black-box adversarial attack on text classification task. Instead of searching for important words to be perturbed by querying the target model, Explain2Attack employs an interpretable substitute model from a similar domain to learn word importance scores. We show that our framework either achieves or out-performs attack rates of the state-of-the-art models, yet with lower queries cost and higher efficiency. Mahmoud Hossam, Trung Le 0001, He Zhao 0001, Dinh Q. Phung |
ICPR | 2 |
| 2020 | Stein Variational Gradient Descent with Variance ReductionabstractProbabilistic inference is a common and important task in statistical machine learning. The recently proposed Stein variational gradient descent (SVGD) is a generic Bayesian inference method that has been shown to be successfully applied in a wide range of contexts, especially in dealing with large datasets, where existing probabilistic inference methods have been known to be ineffective. In a large-scale data setting, SVGD employs the mini-batch strategy but its mini-batch estimator has large variance, hence compromising its estimation quality in practice. To this end, we propose in this paper a generic SVGD-based inference method that can significantly reduce the variance of mini-batch estimator when working with large datasets. Our experiments on 14 datasets show that the proposed method enjoys substantial and consistent improvements compared with baseline methods in binary classification task and its pseudo-online learning setting, and regression task. Furthermore, our framework is generic and applicable to a wide range of probabilistic inference problems such as in Bayesian neural networks and Markov random fields. Nhan Dam, Trung Le 0001, Viet Huynh, Dinh Q. Phung |
IJCNN | 2 |
| 2020 | OptiGAN: Generative Adversarial Networks for Goal Optimized Sequence GenerationabstractOne of the challenging problems in sequence generation tasks is the optimized generation of sequences with specific desired goals. Current sequential generative models mainly generate sequences to closely mimic the training data, without direct optimization of desired goals or properties specific to the task. We introduce OptiGAN, a generative model that incorporates both Generative Adversarial Networks (GAN) and Reinforcement Learning (RL) to optimize desired goal scores using policy gradients. We apply our model to text and real-valued sequence generation, where our model is able to achieve higher desired scores out-performing GAN and RL baselines, while not sacrificing output sample diversity. Mahmoud Hossam, Trung Le 0001, Viet Huynh, Michael Papasimeon, Dinh Q. Phung |
IJCNN | 2 |
| 2020 | Code Pointer Network for Binary Function Scope IdentificationabstractFunction identification is a preliminary step in binary analysis for many extensive applications from malware detection, common vulnerability detection and binary instrumentation to name a few. In this paper, we propose the Code Pointer Network that leverages the underlying idea of a pointer network to efficiently and effectively tackle function scope identification - the hardest and most crucial task in function identification. We establish extensive experiments to compare our proposed method with the deep learning based baseline. Experimental results demonstrate that our proposed method significantly outperforms the state-of-the-art baseline in terms of both predictive performance and running time. Van Nguyen 0002, Trung Le 0001, Tue Le, Olivier Y. de Vel, Paul Montague, Dinh Q. Phung |
IJCNN | 2 |
| 2020 | Code Action Network for Binary Function Scope Identification
Van Nguyen 0002, Trung Le 0001, Tue Le, Olivier Y. de Vel, Paul Montague, John C. Grundy, Dinh Q. Phung |
PAKDD (1) | 2 |
| 2020 | Deep Cost-Sensitive Kernel Machine for Binary Software Vulnerability Detection
Tuan Nguyen 0004, Trung Le 0001, Olivier Y. de Vel, Paul Montague, John C. Grundy, Dinh Q. Phung |
PAKDD (2) | 2 |
| 2020 | Dual-Component Deep Domain Adaptation: A New Approach for Cross Project Software Vulnerability Detection
Van Nguyen 0002, Trung Le 0001, Olivier Y. de Vel, Paul Montague, John C. Grundy, Dinh Q. Phung |
PAKDD (1) | 2 |
| 2019 | Robust Anomaly Detection in Videos Using Multilevel RepresentationsabstractDetecting anomalies in surveillance videos has long been an important but unsolved problem. In particular, many existing solutions are overly sensitive to (often ephemeral) visual artifacts in the raw video data, resulting in false positives and fragmented detection regions. To overcome such sensitivity and to capture true anomalies with semantic significance, one natural idea is to seek validation from abstract representations of the videos. This paper introduces a framework of robust anomaly detection using multilevel representations of both intensity and motion data. The framework consists of three main components: 1) representation learning using Denoising Autoencoders, 2) level-wise representation generation using Conditional Generative Adversarial Networks, and 3) consolidating anomalous regions detected at each representation level. Our proposed multilevel detector shows a significant improvement in pixel-level Equal Error Rate, namely 11.35%, 12.32% and 4.31% improvement in UCSD Ped 1, UCSD Ped 2 and Avenue datasets respectively. In addition, the model allowed us to detect mislabeled anomalies in the UCDS Ped 1. Hung Vu, Tu Dinh Nguyen, Trung Le 0001, Wei Luo 0001, Dinh Q. Phung |
AAAI | 3 |
| 2019 | Maximal Divergence Sequential Autoencoder for Binary Software Vulnerability Detection
Tue Le, Tuan Nguyen 0004, Trung Le 0001, Dinh Q. Phung, Paul Montague, Olivier Y. de Vel, Lizhen Qu |
ICLR (Poster) | 3 |
| 2019 | Three-Player Wasserstein GAN via Amortised DualityabstractWe propose a new formulation for learning generative adversarial networks (GANs) using optimal transport cost (the general form of Wasserstein distance) as the objective criterion to measure the dissimilarity between target distribution and learned distribution. Our formulation is based on the general form of the Kantorovich duality which is applicable to optimal transport with a wide range of cost functions that are not necessarily metric. To make optimising this duality form amenable to gradient-based methods, we employ a function that acts as an amortised optimiser for the innermost optimisation problem. Interestingly, the amortised optimiser can be viewed as a mover since it strategically shifts around data points. The resulting formulation is a sequential min-max-min game with 3 players: the generator, the critic, and the mover where the new player, the mover, attempts to fool the critic by shifting the data around. Despite involving three players, we demonstrate that our proposed formulation can be trained reasonably effectively via a simple alternative gradient learning strategy. Compared with the existing Lipschitz-constrained formulations of Wasserstein GAN on CIFAR-10, our model yields significantly better diversity scores than weight clipping and comparable performance to gradient penalty method. Nhan Dam, Quan Hoang, Trung Le 0001, Tu Dinh Nguyen, Hung Hai Bui, Dinh Q. Phung |
IJCAI | 3 |
| 2019 | Learning Generative Adversarial Networks from Multiple Data SourcesabstractGenerative Adversarial Networks (GANs) are a powerful class of deep generative models. In this paper, we extend GAN to the problem of generating data that are not only close to a primary data source but also required to be different from auxiliary data sources. For this problem, we enrich both GANs' formulations and applications by introducing pushing forces that thrust generated samples away from given auxiliary data sources. We term our method Push-and-Pull GAN (P2GAN). We conduct extensive experiments to demonstrate the merit of P2GAN in two applications: generating data with constraints and addressing the mode collapsing problem. We use CIFAR-10, STL-10, and ImageNet datasets and compute Fréchet Inception Distance to evaluate P2GAN's effectiveness in addressing the mode collapsing problem. The results show that P2GAN outperforms the state-of-the-art baselines. For the problem of generating data with constraints, we show that P2GAN can successfully avoid generating specific features such as black hair. Trung Le 0001, Quan Hoang, Hung Vu, Tu Dinh Nguyen, Hung Hai Bui, Dinh Q. Phung |
IJCAI | 1 |
| 2019 | Deep Domain Adaptation for Vulnerable Code Function IdentificationabstractDue to the ubiquity of computer software, software vulnerability detection (SVD) has become crucial in the software industry and in the field of computer security. Two significant issues in SVD arise when using machine learning, namely: i) how to learn automatic features that can help improve the predictive performance of vulnerability detection and ii) how to overcome the scarcity of labeled vulnerabilities in projects that require the laborious labeling of code by software security experts. In this paper, we address these two crucial concerns by proposing a novel architecture which leverages deep domain adaptation with automatic feature learning for software vulnerability identification. Based on this architecture, we keep the principles and reapply the state-of-the-art deep domain adaptation methods to indicate that deep domain adaptation for SVD is plausible and promising. Moreover, we further propose a novel method named Semi-supervised Code Domain Adaptation Network (SCDAN) that can efficiently utilize and exploit information carried in unlabeled target data by considering them as the unlabeled portion in a semi-supervised learning context. The proposed SCDAN method enforces the clustering assumption, which is a key principle in semi-supervised learning. The experimental results using six real-world software project datasets show that our SCDAN method and the baselines using our architecture have better predictive performance by a wide margin compared with the Deep Code Network (VulDeePecker) method without domain adaptation. Also, the proposed SCDAN significantly outperforms the DIRT-T which to the best of our knowledge is currently the-state-of-the-art method in deep domain adaptation and other baselines. Van Nguyen 0002, Trung Le 0001, Tue Le, Olivier Y. de Vel, Paul Montague, Lizhen Qu, Dinh Q. Phung |
IJCNN | 2 |
| 2019 | GoGP: scalable geometric-based Gaussian process for online regression
Trung Le 0001, Vu Nguyen 0001, Tu Dinh Nguyen, Dinh Q. Phung |
Knowl. Inf. Syst. | 1 |
| 2018 | Clustering Induced Kernel LearningabstractLearning rich and expressive kernel functions is a challenging task in kernel-based supervised learning. Multiple kernel learning (MKL) approach addresses this problem by combining a mixed variety of kernels and letting the optimization solver choose the most appropriate combination. However, most of existing methods are parametric in the sense that they require a predefined list of kernels. Hence, there appears a substantial trade-off between computation and the modeling risk of not being able to explore more expressive and suitable kernel functions. Moreover, current existing approaches to combine kernels cannot exploit clustering structure carried in data, especially when data are heterogeneous. In this work, we present a new framework that leverages Bayesian nonparametric models (i.e, automatically grow kernel functions) with multiple kernel learning to develop a new framework that enjoys the nonparametric flavor in the context of multiple kernel learning. In particular, we propose Clustering Induced Kernel Learning (CIK) method that can automatically discover clustering structure from the data and train a single kernel machine to fit data in each discovered cluster simultaneously. The outcome of our proposed method includes both clustering analysis and multiple kernel classifier for a given dataset. We conduct extensive experiments on several benchmark datasets. The experimental results show that our method can improve classification and clustering performance when datasets have complex clustering structure with different preferred kernels. Nhan Dam, Trung Le 0001, Tu Dinh Nguyen, Dinh Q. Phung |
ACML | 3 |
| 2018 | Batch Normalized Deep Boltzmann MachinesabstractTraining Deep Boltzmann Machines (DBMs) is a challenging task in deep generative model studies. The careless training usually leads to a divergence or a useless model. We discover that this phenomenon is due to the change of DBM layers’ input signals during model parameter updates, similar to other deterministic deep networks such as Convolutional Neural Networks (CNNs) or Recurrent Neural Networks (RNNs). The change of layers’ input distributions not only complicates the learning process but also causes redundant neurons that simply imitate the others’ behaviors. Although this phenomenon can be coped using batch normalization in deep learning, integrating this technique into the probabilistic network of DBMs is a challenging problem since it has to satisfy two conditions of energy function and conditional probabilities. In this paper, we introduce Batch Normalized Deep Boltzmann Machines (BNDBMs) that meet both aforementioned conditions and successfully combine batch normalization and DBMs into the same framework. However, unlike CNNs, due to the probabilistic nature of DBMs, training DBMs with batch normalization has some differences: i) fixing shift parameters $\bnshift$ but learning scale parameters $\bnscale$; ii) avoiding normalizing the first hidden layer and iii) maintaining multiple pairs of population means and variances per neuron rather than one pair in CNNs. We observe that our proposed BNDBMs can stabilize the input signals of network layers and facilitate the training process as well as improve the model quality. More interestingly, BNDBMs can be trained successfully without pretraining, which is usually a mandatory step in most existing DBMs. The experimental results in MNIST, Fashion-MNIST and Caltech 101 Silhouette datasets show that our BNDBMs outperform DBMs and centered DBMs in terms of feature representation and classification accuracy ($3.98%$ and $5.84%$ average improvement for pretraining and no pretraining respectively). Hung Vu, Tu Dinh Nguyen, Trung Le 0001, Wei Luo 0001, Dinh Q. Phung |
ACML | 3 |
| 2018 | MGAN: Training Generative Adversarial Nets with Multiple Generators
Quan Hoang, Tu Dinh Nguyen, Trung Le 0001, Dinh Q. Phung |
ICLR (Poster) | 3 |
| 2018 | Bayesian Multi-Hyperplane Machine for Pattern RecognitionabstractCurrent existing multi-hyperplane machine approach deals with high-dimensional and complex datasets by approximating the input data region using a parametric mixture of hyperplanes. Consequently, this approach requires an excessively time-consuming parameter search to find the set of optimal hyper-parameters. Another serious drawback of this approach is that it is often suboptimal since the optimal choice for the hyper-parameter is likely to lie outside the searching space due to the space discretization step required in grid search. To address these challenges, we propose in this paper BAyesian Multi-hyperplane Machine (BAMM). Our approach departs from a Bayesian perspective, and aims to construct an alternative probabilistic view in such a way that its maximum-a-posteriori (MAP) estimation reduces exactly to the original optimization problem of a multi-hyperplane machine. This view allows us to endow prior distributions over hyper-parameters and augment auxiliary variables to efficiently infer model parameters and hyper-parameters via Markov chain Monte Carlo (MCMC) method. We then employ a Stochastic Gradient Descent (SGD) framework to scale our model up with ever-growing large datasets. Extensive experiments demonstrate the capability of our proposed method in learning the optimal model without using any parameter tuning, and in achieving comparable accuracies compared with the state-of-art baselines; in the meantime our model can seamlessly handle with large-scale datasets. Trung Le 0001, Tu Dinh Nguyen, Dinh Q. Phung |
ICPR | 2 |
| 2018 | Geometric Enclosing NetworksabstractTraining model to generate data has increasingly attracted research attention and become important in modern world applications. We propose in this paper a new geometry-based optimization approach to address this problem. Orthogonal to current state-of-the-art density-based approaches, most notably VAE and GAN, we present a fresh new idea that borrows the principle of minimal enclosing ball to train a generator G\left(\bz\right) in such a way that both training and generated data, after being mapped to the feature space, are enclosed in the same sphere. We develop theory to guarantee that the mapping is bijective so that its inverse from feature space to data space results in expressive nonlinear contours to describe the data manifold, hence ensuring data generated are also lying on the data manifold learned from training data. Our model enjoys a nice geometric interpretation, hence termed Geometric Enclosing Networks (GEN), and possesses some key advantages over its rivals, namely simple and easy-to-control optimization formulation, avoidance of mode collapsing and efficiently learn data manifold representation in a completely unsupervised manner. We conducted extensive experiments on synthesis and real-world datasets to illustrate the behaviors, strength and weakness of our proposed GEN, in particular its ability to handle multi-modal data and quality of generated data. Trung Le 0001, Hung Vu, Tu Dinh Nguyen, Dinh Q. Phung |
IJCAI | 1 |
| 2018 | Robust Bayesian Kernel Machine via Stein Variational Gradient Descent for Big DataabstractKernel methods are powerful supervised machine learning models for their strong generalization ability, especially on limited data to effectively generalize on unseen data. However, most kernel methods, including the state-of-the-art LIBSVM, are vulnerable to the curse of kernelization, making them infeasible to apply to large-scale datasets. This issue is exacerbated when kernel methods are used in conjunction with a grid search to tune their kernel parameters and hyperparameters which brings in the question of model robustness when applied to real datasets. In this paper, we propose a robust Bayesian Kernel Machine (BKM) - a Bayesian kernel machine that exploits the strengths of both the Bayesian modelling and kernel methods. A key challenge for such a formulation is the need for an efficient learning algorithm. To this end, we successfully extended the recent Stein variational theory for Bayesian inference for our proposed model, resulting in fast and efficient learning and prediction algorithms. Importantly our proposed BKM is resilient to the curse of kernelization, hence making it applicable to large-scale datasets and robust to parameter tuning, avoiding the associated expense and potential pitfalls with current practice of parameter tuning. Our extensive experimental results on 12 benchmark datasets show that our BKM without tuning any parameter can achieve comparable predictive performance with the state-of-the-art LIBSVM and significantly outperforms other baselines, while obtaining significantly speedup in terms of the total training time compared with its rivals Trung Le 0001, Tu Dinh Nguyen, Dinh Q. Phung, Geoffrey I. Webb |
KDD | 2 |
| 2018 | Jointly Predicting Affective and Mental Health Scores Using Deep Neural Networks of Visual Cues on the Web
Van Nguyen 0002, Thin Nguyen, Mark E. Larsen, Bridianne O'Dea, Duc Thanh Nguyen, Trung Le 0001, Dinh Q. Phung, Svetha Venkatesh, Helen Christensen |
WISE (2) | 7 |
| 2017 | GoGP: Fast Online Regression with Gaussian ProcessesabstractOne of the most current challenging problems in Gaussian process regression (GPR) is to handle large-scale datasets and to accommodate an online learning setting where data arrive irregularly on the fly. In this paper, we introduce a novel online Gaussian process model that could scale with massive datasets. Our approach is formulated based on alternative representation of the Gaussian process under geometric and optimization views, hence termed geometric-based online GP (GoGP). We developed theory to guarantee that with a good convergence rate our proposed algorithm always produces a (sparse) solution which is close to the true optima to any arbitrary level of approximation accuracy specified a priori. Furthermore, our method is proven to scale seamlessly not only with large-scale datasets, but also to adapt accurately with streaming data. We extensively evaluated our proposed model against state-of-the-art baselines using several large-scale datasets for online regression task. The experimental results show that our GoGP delivered comparable, or slightly better, predictive performance while achieving a magnitude of computational speedup compared with its rivals under online setting. More importantly, its convergence behavior is guaranteed through our theoretical analysis, which is rapid and stable while achieving lower errors. Trung Le 0001, Vu Nguyen 0001, Tu Dinh Nguyen, Dinh Q. Phung |
ICDM | 1 |
| 2017 | Large-scale Online Kernel Learning with Random Feature ReparameterizationabstractA typical online kernel learning method faces two fundamental issues: the complexity in dealing with a huge number of observed data points (a.k.a the curse of kernelization) and the difficulty in learning kernel parameters, which often assumed to be fixed. Random Fourier feature is a recent and effective approach to address the former by approximating the shift-invariant kernel function via Bocher's theorem, and allows the model to be maintained directly in the random feature space with a fixed dimension, hence the model size remains constant w.r.t. data size. We further introduce in this paper the reparameterized random feature (RRF), a random feature framework for large-scale online kernel learning to address both aforementioned challenges. Our initial intuition comes from the so-called "reparameterization trick" [Kingma et al., 2014] to lift the source of randomness of Fourier components to another space which can be independently sampled, so that stochastic gradient of the kernel parameters can be analytically derived. We develop a well-founded underlying theory for our method, including a general way to reparameterize the kernel, and a new tighter error bound on the approximation quality. This view further inspires a direct application of stochastic gradient descent for updating our model under an online learning setting. We then conducted extensive experiments on several large-scale datasets where we demonstrate that our work achieves state-of-the-art performance in both learning efficacy and efficiency. Tu Dinh Nguyen, Trung Le 0001, Hung Hai Bui, Dinh Q. Phung |
IJCAI | 2 |
| 2017 | Discriminative Bayesian Nonparametric ClusteringabstractWe propose a general framework for discriminative Bayesian nonparametric clustering to promote the inter-discrimination among the learned clusters in a fully Bayesian nonparametric (BNP) manner. Our method combines existing BNP clustering and discriminative models by enforcing latent cluster indices to be consistent with the predicted labels resulted from probabilistic discriminative model. This formulation results in a well-defined generative process wherein we can use either logistic regression or SVM for discrimination. Using the proposed framework, we develop two novel discriminative BNP variants: the discriminative Dirichlet process mixtures, and the discriminative-state infinite HMMs for sequential data. We develop efficient data-augmentation Gibbs samplers for posterior inference. Extensive experiments in image clustering and dynamic location clustering demonstrate that by encouraging discrimination between induced clusters, our model enhances the quality of clustering in comparison with the traditional generative BNP models. Vu Nguyen 0001, Dinh Q. Phung, Trung Le 0001, Hung Hai Bui |
IJCAI | 3 |
| 2017 | Dual Discriminator Generative Adversarial NetsabstractWe propose in this paper a novel approach to tackle the problem of mode collapse encountered in generative adversarial network (GAN). Our idea is intuitive but proven to be very effective, especially in addressing some key limitations of GAN. In essence, it combines the Kullback-Leibler (KL) and reverse KL divergences into a unified objective function, thus it exploits the complementary statistical properties from these divergences to effectively diversify the estimated density in capturing multi-modes. We term our method dual discriminator generative adversarial nets (D2GAN) which, unlike GAN, has two discriminators; and together with a generator, it also has the analogy of a minimax game, wherein a discriminator rewards high scores for samples from data distribution whilst another discriminator, conversely, favoring data from the generator, and the generator produces data to fool both two discriminators. We develop theoretical analysis to show that, given the maximal discriminators, optimizing the generator of D2GAN reduces to minimizing both KL and reverse KL divergences between data distribution and the distribution induced from the data generated by the generator, hence effectively avoiding the mode collapsing problem. We conduct extensive experiments on synthetic and real-world large-scale datasets (MNIST, CIFAR-10, STL-10, ImageNet), where we have made our best effort to compare our D2GAN with the latest state-of-the-art GAN's variants in comprehensive qualitative and quantitative evaluations. The experimental results demonstrate the competitive and superior performance of our approach in generating good quality and diverse samples over baselines, and the capability of our method to scale up to ImageNet database. Tu Dinh Nguyen, Trung Le 0001, Hung Vu, Dinh Q. Phung |
NIPS | 2 |
| 2017 | Supervised Restricted Boltzmann Machines
Tu Dinh Nguyen, Dinh Q. Phung, Viet Huynh, Trung Le 0001 |
UAI | 4 |
| 2017 | Approximation Vector Machines for Large-scale Online LearningabstractOne of the most challenging problems in kernel online learning is to bound the model size and to promote model sparsity. Sparse models not only improve computation and memory usage, but also enhance the generalization capacity -- a principle that concurs with the law of parsimony. However, inappropriate sparsity modeling may also significantly degrade the performance. In this paper, we propose Approximation Vector Machine (AVM), a model that can simultaneously encourage sparsity and safeguard its risk in compromising the performance. In an online setting context, when an incoming instance arrives, we approximate this instance by one of its neighbors whose distance to it is less than a predefined threshold. Our key intuition is that since the newly seen instance is expressed by its nearby neighbor the optimal performance can be analytically formulated and maintained. We develop theoretical foundations to support this intuition and further establish an analysis for the common loss functions including Hinge, smooth Hinge, and Logistic (i.e., for the classification task) and $\ell_{1}$, $\ell_{2}$, and $\varepsilon$-insensitive (i.e., for the regression task) to characterize the gap between the approximation and optimal solutions. This gap crucially depends on two key factors including the frequency of approximation (i.e., how frequent the approximation operation takes place) and the predefined threshold. We conducted extensive experiments for classification and regression tasks in batch and online modes using several benchmark datasets. The quantitative results show that our proposed AVM obtained comparable predictive performances with current state-of-the-art methods while simultaneously achieving significant computational speed-up due to the ability of the proposed AVM in maintaining the model size. Trung Le 0001, Tu Dinh Nguyen, Vu Nguyen 0001, Dinh Q. Phung |
J. Mach. Learn. Res. | 1 |
| 2016 | Multiple Kernel Learning with Data AugmentationabstractThe motivations of multiple kernel learning (MKL) approach are to increase kernel expressiveness capacity and to avoid the expensive grid search over a wide spectrum of kernels. A large amount of work has been proposed to improve the MKL in terms of the computational cost and the sparsity of the solution. However, these studies still either require an expensive grid search on the model parameters or scale unsatisfactorily with the numbers of kernels and training samples. In this paper, we address these issues by conjoining MKL, Stochastic Gradient Descent (SGD) framework, and data augmentation technique. The pathway of our proposed method is developed as follows. We first develop a maximum-a-posteriori (MAP) view for MKL under a probabilistic setting and described in a graphical model. This view allows us to develop data augmentation technique to make the inference for finding the optimal parameters feasible, as opposed to traditional approach of training MKL via convex optimization techniques. As a result, we can use the standard SGD framework to learn weight matrix and extend the model to support online learning. We validate our method on several benchmark datasets in both batch and online settings. The experimental results show that our proposed method can learn the parameters in a principled way to eliminate the expensive grid search while gaining a significant computational speedup comparing with the state-of-the-art baselines. Trung Le 0001, Vu Nguyen 0001, Tu Dinh Nguyen, Dinh Q. Phung |
ACML | 2 |
| 2016 | Nonparametric Budgeted Stochastic Gradient DescentabstractOne of the most challenging problems in kernel online learning is to bound the model size. Budgeted kernel online learning addresses this issue by bounding the model size to a predefined budget. However, determining an appropriate value for such predefined budget is arduous. In this paper, we propose the Nonparametric Budgeted Stochastic Gradient Descent that allows the model size to automatically grow with data in a principled way. We provide theoretical analysis to show that our framework is guaranteed to converge for a large collection of loss functions (e.g. Hinge, Logistic, L2, L1, and \varepsilon-insensitive) which enables the proposed algorithm to perform both classification and regression tasks without hurting the ideal convergence rate O\left(\frac1T\right) of the standard Stochastic Gradient Descent. We validate our algorithm on the real-world datasets to consolidate the theoretical claims. Trung Le 0001, Vu Nguyen 0001, Tu Dinh Nguyen, Dinh Q. Phung |
AISTATS | 1 |
| 2016 | Fast Support Vector Clustering
Tung Pham 0001, Trung Le 0001, Dat Tran 0001 |
ESANN | 2 |
| 2016 | One-Pass Logistic Regression for Label-Drift and Large-Scale Classification on Distributed SystemsabstractLogistic regression (LR) for classification is the workhorse in industry, where a set of predefined classes is required. The model, however, fails to work in the case where the class labels are not known in advance, a problem we term label-drift classification. Label-drift classification problem naturally occurs in many applications, especially in the context of streaming settings where the incoming data may contain samples categorized with new classes that have not been previously seen. Additionally, in the wave of big data, traditional LR methods may fail due to their expense of running time. In this paper, we introduce a novel variant of LR, namely one-pass logistic regression (OLR) to offer a principled treatment for label-drift and large-scale classifications. To handle largescale classification for big data, we further extend our OLR to a distributed setting for parallelization, termed sparkling OLR (Spark-OLR). We demonstrate the scalability of our proposed methods on large-scale datasets with more than one hundred million data points. The experimental results show that the predictive performances of our methods are comparable orbetter than those of state-of-the-art baselines whilst the executiontime is much faster at an order of magnitude. In addition, the OLR and Spark-OLR are invariant to data shuffling and have no hyperparameter to tune that significantly benefits data practitioners and overcomes the curse of big data cross-validationto select optimal hyperparameters. Vu Nguyen 0001, Tu Dinh Nguyen, Trung Le 0001, Svetha Venkatesh, Dinh Q. Phung |
ICDM | 3 |
| 2016 | Distributed data augmented support vector machine on SparkabstractSupport vector machines (SVMs) are widely-used for classification in machine learning and data mining tasks. However, they traditionally have been applied to small to medium datasets. Recent need to scale up with data size has attracted research attention to develop new methods and implementation for SVM to perform tasks at scale. Distributed SVMs are relatively new and studied recently, but the distributed implementation for SVM with data augmentation has not been developed. This paper introduces a distributed data augmentation implementation for SVM on Apache Spark, a recent advanced and popular platform for distributed computing that has been employed widely in research as well as in industry. We term our implementation sparkling vector machine (SkVM) which supports both classification and regression tasks by scanning through the data exactly once. In addition, we further develop a framework to handle the data with new classes arriving under an online classification setting where new data points can have labels that have not previously seen - a problem we term label-drift classification. We demonstrate the scalability of our proposed method on large-scale datasets with more than one hundred million data points. The experimental results show that the predictive performances of our method are comparable or better than those of baselines whilst the execution time is much faster at an order of magnitude. Tu Dinh Nguyen, Vu Nguyen 0001, Trung Le 0001, Dinh Q. Phung |
ICPR | 3 |
| 2016 | Fast Kernel-based method for anomaly detectionabstractAnomaly detection (AD) involves detecting abnormality from normality and has a wide spectrum of applications in reality. Kernel-based methods for AD have been proven robust with diverse data distributions and offering good generalization ability. Stochastic gradient descent (SGD) method has recently emerged as a promising framework to devise ultra-fast learning methods. In this paper, we conjoin the advantages of Kernel-based method and SGD-based method to propose fast learning methods for anomaly detection. We validate the proposed methods on 8 benchmark datasets in UCI repository and KDD cup 1999 dataset. The experimental results show that the proposed methods offer a comparable one-class classification accuracy while simultaneously achieving a significantly computational speed-up. Trung Le 0001, Van Nguyen 0002, Dat Tran 0001 |
IJCNN | 2 |
| 2016 | Fuzzy Kernel Stochastic Gradient Descent machinesabstractStochastic Gradient Descent (SGD) based method offers a viable solution to training large-scale dataset. However, the traditional SGD-based methods cannot get benefit from the distribution or geometry information carried in data. The reason is that these methods make use of the uniform distribution over the entire training set so as to sample the next data point for updating the model. We address this issue by incorporating the distribution or geometry information carried in the data into the sampling procedure. In particular, we utilize the fuzzy-membership evaluation methods which allow transferring the distribution or geometry information carried in the data to the fuzzy memberships. The fuzzy memberships is then normalized to a discrete distribution from which the next data point is sampled. This allows the training staying more focused on the important data points and tending to ignore the less impact data points, e.g., the noises and outliers. We validate the proposed methods on 8 benchmark datasets. The experimental results show that the proposed methods are comparable with the standard SGD-based method in training time while offering a significant improvement in classification accuracy. Tuan Nguyen 0004, Phuong Duong, Trung Le 0001, Viet Ngo, Dat Tran 0001, Wanli Ma 0003 |
IJCNN | 3 |
| 2016 | Dual Space Gradient Descent for Online LearningabstractOne crucial goal in kernel online learning is to bound the model size. Common approaches employ budget maintenance procedures to restrict the model sizes using removal, projection, or merging strategies. Although projection and merging, in the literature, are known to be the most effective strategies, they demand extensive computation whilst removal strategy fails to retain information of the removed vectors. An alternative way to address the model size problem is to apply random features to approximate the kernel function. This allows the model to be maintained directly in the random feature space, hence effectively resolve the curse of kernelization. However, this approach still suffers from a serious shortcoming as it needs to use a high dimensional random feature space to achieve a sufficiently accurate kernel approximation. Consequently, it leads to a significant increase in the computational cost. To address all of these aforementioned challenges, we present in this paper the Dual Space Gradient Descent (DualSGD), a novel framework that utilizes random features as an auxiliary space to maintain information from data points removed during budget maintenance. Consequently, our approach permits the budget to be maintained in a simple, direct and elegant way while simultaneously mitigating the impact of the dimensionality issue on learning performance. We further provide convergence analysis and extensively conduct experiments on five real-world datasets to demonstrate the predictive performance and scalability of our proposed method in comparison with the state-of-the-art baselines. Trung Le 0001, Tu Dinh Nguyen, Vu Nguyen 0001, Dinh Q. Phung |
NIPS | 1 |
| 2016 | Sparse Adaptive Multi-hyperplane Machine
Trung Le 0001, Vu Nguyen 0001, Dinh Q. Phung |
PAKDD (1) | 2 |
| 2016 | Budgeted Semi-supervised Support Vector Machine
Trung Le 0001, Phuong Duong, Mi Dinh, Tu Dinh Nguyen, Vu Nguyen 0001, Dinh Q. Phung |
UAI | 1 |
| 2015 | Graph-based semi-supervised Support Vector Data Description for novelty detectionabstractSupport Vector Data Description (SVDD) is a well-known supervised learning method for novelty detection purpose. For its classification task, SVDD requires a fully-labeled dataset. Nonetheless, contemporary datasets always consist of a collection of labeled data samples jointly a much larger collection of unlabeled ones. This fact impedes the usage of SVDD in the real-world problems. In this paper, we propose to utilize the information implicated in a spectral graph to leverage SVDD in the context of semi-supervised learning. The theory and experiment evidence that the proposed method is able to efficiently employ the information carried in the spectral graph to not only enhance the generalization ability of SVDD but also enforce the cluster assumption which is crucial for a semi-supervised learning method. Phuong Duong, Van Nguyen 0002, Mi Dinh, Trung Le 0001, Dat Tran 0001, Wanli Ma 0003 |
IJCNN | 4 |
| 2015 | Least square Support Vector Machine for large-scale datasetabstractSupport Vector Machine (SVM) is a very well-known tool for classification and regression problems. Many applications require SVMs with non-linear kernels for accurate classification. Training time complexity for SVMs with non-linear kernels is typically quadratic in the size of the training dataset. In this paper, we depart from the very well-known variation of SVM, the so-called Least Square Support Vector Machine, and apply Steepest Sub-gradient Descent method to propose Steepest Sub-gradient Descent Least Square Support Vector Machine (SGDLSSVM). It is theoretically proven that the convergent rate of the proposed method to gain ε - precision solution is O (log (1/ε)). The experiments established on the large-scale datasets indicate that the proposed method offers the comparable classification accuracies while being faster than the baselines. Trung Le 0001, Vinh Lai, Dat Tran 0001, Wanli Ma 0003 |
IJCNN | 2 |
| 2015 | Fast One-Class Support Vector Machine for Novelty Detection
Trung Le 0001, Dinh Q. Phung, Svetha Venkatesh |
PAKDD (2) | 1 |
| 2014 | Robust Support Vector MachineabstractSupport Vector Machine (SVM) is a well-known kernel-based method for binary classification problem. SVM aims at constructing the optimal middle hyperplane which induces the largest margin. It is proven that in a linearly separable case, this middle hyperplane offers the high accuracy on universal datasets. However, real world datasets often contain overlapping regions and therefore, the decision hyperplane should be adjusted according to the profiles of the datasets. In this paper, we propose Robust Support Vector Machine (RSVM), where the hyperplanes can be properly adjusted to accommodate the real world datasets. By setting the value of the adjustment factor properly, RSVM can handle well the datasets with any possible profiles. Our experiments on the benchmark datasets demonstrate the superiority of the RSVM for both binary and one-class classification problems. Trung Le 0001, Dat Tran 0001, Wanli Ma 0003, Thien Pham, Phuong Duong, Minh Nguyen 0004 |
IJCNN | 1 |
| 2014 | Using EEG artifacts for BCI applicationsabstractBrain computer interface (BCI) is about the communication channel between the brain of a human subject and a computerized device. Electroencephalography (EEG) signals are the primary choice as the sources of interpreting the intention of the human subject. EEG signals have a long history of being used in human health for the purposes of studying brain activities and medical diagnosis. EEG signals are very weak and are subject to the contamination from many artifact signals. For the applications in human health, true EEG signals, without the contamination, is highly desirable. However, for the purposes of BCI, where stable patterns from the source signals are critical, the origins of the signals are of less concern. In this paper, we propose a BCI, which is simple to implement and easy to use, by taking the advantage of EEG artifacts, generated by a number of purposely designed voluntary facial muscle movements. Wanli Ma 0003, Dat Tran 0001, Trung Le 0001, Shang-Ming Zhou |
IJCNN | 3 |
| 2014 | Kernel-based semi-supervised learning for novelty detectionabstractOne-class Support Vector Machine (OCSVM) is a well-known method for novelty detection. However, OCSVM regards all negative data samples as a common symbol and thereby not being able to utilize the information carried by them. Furthermore, OCSVM requires a fully labeled data set and cannot work efficiently with data set with both labeled and unlabeled data samples which is very popular nowadays. In this paper, we first extend the model of OCSVM to enable efficiently using the negative data samples. We then propose two methods to integrate the semi-supervised learning paradigm to the extended model for novelty detection purpose. Van Nguyen 0002, Trung Le 0001, Thien Pham, Mi Dinh |
IJCNN | 2 |
| 2013 | Maximal margin learning vector quantisationabstractKernel Generalised Learning Vector Quantisation (KGLVQ) was proposed to extend Generalised Learning Vector Quantisation into the kernel feature space to deal with complex class boundaries and thus yielded promising performance for complex classification tasks in pattern recognition. However KGLVQ does not follow the maximal margin principle, which is crucial for kernel-based learning methods. In this paper we propose a maximal margin approach (MLVQ) to the KGLVQ algorithm. MLVQ inherits the merits of KGLVQ and also follows the maximal margin principle to improve the generalisation capability. Experiments performed on the well-known data sets available in UCI repository show promising classification results for the proposed method. Trung Le 0001, Dat Tran 0001, Van Nguyen 0002, Wanli Ma 0003 |
IJCNN | 1 |
| 2013 | Fuzzy entropy semi-supervised support vector data descriptionabstractSupport Vector Data Description (SVDD) is known as one of the best kernel-based methods for one-class classification problems. SVDD requires fully labelled data sets. However, in reality, an abundant amount of data can be easily collected, while the labelling process is often expensive, time-consuming, and error-prone. Therefore, partially labelled data sets are popular and easy to obtain. In this paper, we propose a semi-supervised learning method, Fuzzy Entropy Semi-supervised SVDD (FS3VDD), to extend SVDD to cope with partially labelled data sets. The learning model employs fuzzy membership and fuzzy entropy to help the labelling of the unlabeled data. Trung Le 0001, Dat Tran 0001, Tien Tran, Wanli Ma 0003 |
IJCNN | 1 |
| 2013 | Fuzzy Multi-Sphere Support Vector Data Description
Trung Le 0001, Dat Tran 0001, Wanli Ma 0003 |
PAKDD (2) | 1 |
| 2013 | EEG-Based Person Verification Using Multi-Sphere SVDD and UBM
Phuoc Nguyen, Dat Tran 0001, Trung Le 0001, Xu Huang 0001, Wanli Ma 0003 |
PAKDD (1) | 3 |
| 2013 | Proximity multi-sphere support vector clustering
Trung Le 0001, Dat Tran 0001, Phuoc Nguyen, Wanli Ma 0003, Dharmendra Sharma 0001 |
Neural Comput. Appl. | 1 |
| 2012 | Fuzzy Multi-sphere Support Vector Data DescriptionabstractMulti-sphere Support Vector Data Description (MS-SVDD) has been proposed in our previous work. MS-SVDD aims to build a set of spherically shaped boundaries that provide a better data description to the normal dataset and an iterative learning algorithm that determines the set of spherically shaped boundaries. MS-SVDD could improve classification rate for one-class classification problems comparing with SVDD. However MS-SVDD requires a small abnormal data set to build the spherically shaped boundaries for the normal data set. In this paper, we propose a new fuzzy MS-SVDD that can be used when only the normal data set is available. Experimental results on 14 well-known datasets and a comparison between fuzzy MS-SVDD and SVDD are also presented. Trung Le 0001, Dat Tran 0001, Wanli Ma 0003, Dharmendra Sharma 0001 |
FUZZ-IEEE | 1 |
| 2012 | Time Domain Parameters for Online Feedback fNIRS-Based Brain-Computer Interface Systems
Tuan Hoang, Dat Tran 0001, Khoa Truong, Trung Le 0001, Xu Huang 0001, Dharmendra Sharma 0001, Toi Vo |
ICONIP (2) | 4 |
| 2012 | Maximal Margin Approach to Kernel Generalised Learning Vector Quantisation for Brain-Computer Interface
Trung Le 0001, Dat Tran 0001, Tuan Hoang, Dharmendra Sharma 0001 |
ICONIP (3) | 1 |
| 2012 | Deterministic Annealing Multi-Sphere Support Vector Data Description
Trung Le 0001, Dat Tran 0001, Wanli Ma 0003, Dharmendra Sharma 0001 |
ICONIP (3) | 1 |
| 2012 | A unified model for support vector machine and support vector data descriptionabstractSupport vector machine (SVM) and support vector data description (SVDD) are the well-known kernel-based methods for pattern classification. SVM constructs an optimal hyperplane whereas SVDD constructs an optimal hypersphere to separate data between two classes. SVM and SVDD have been compared in pattern classification experiments however there is no theoretical work on comparison between these methods. This paper presents a new theoretical model to unify SVM and SVDD. The proposed model constructs two optimal points to generate a general decision boundary which can be transformed to hyperplane for SVM or hypersphere for SVDD. Trung Le 0001, Dat Tran 0001, Wanli Ma 0003, Dharmendra Sharma 0001 |
IJCNN | 1 |
| 2011 | Generalised Support Vector Machine for Brain-Computer Interface
Trung Le 0001, Dat Tran 0001, Tuan Hoang, Wanli Ma 0003, Dharmendra Sharma 0001 |
ICONIP (1) | 1 |
| 2011 | A Novel Parameter Refinement Approach to One Class Support Vector Machine
Trung Le 0001, Dat Tran 0001, Wanli Ma 0003, Dharmendra Sharma 0001 |
ICONIP (2) | 1 |
| 2011 | Multi-Sphere Support Vector Clustering
Trung Le 0001, Dat Tran 0001, Phuoc Nguyen, Wanli Ma 0003, Dharmendra Sharma 0001 |
ICONIP (2) | 1 |
| 2011 | Multiple distribution data description learning method for novelty detectionabstractCurrent data description learning methods for novelty detection such as support vector data description and small sphere with large margin construct a spherically shaped boundary around a normal data set to separate this set from abnormal data. The volume of this sphere is minimized to reduce the chance of accepting abnormal data. However those learning methods do not guarantee that the single spherically shaped boundary can best describe the normal data set if there exist some distinctive data distributions in this set. We propose in this paper a new data description learning method that constructs a set of spherically shaped boundaries to provide a better data description to the normal data set. An optimisation problem is proposed and solving this problem results in an iterative learning algorithm to determine the set of spherically shaped boundaries. We prove that the classification error will be reduced after each iteration in our learning method. Experimental results on 23 well-known data sets show that the proposed method provides lower classification error rates. Trung Le 0001, Dat Tran 0001, Phuoc Nguyen, Wanli Ma 0003, Dharmendra Sharma 0001 |
IJCNN | 1 |
| 2011 | Multiple Distribution Data Description Learning Algorithm for Novelty Detection
Trung Le 0001, Dat Tran 0001, Wanli Ma 0003, Dharmendra Sharma 0001 |
PAKDD (2) | 1 |
| 2010 | A Theoretical Framework for Multi-sphere Support Vector Data Description
Trung Le 0001, Dat Tran 0001, Wanli Ma 0003, Dharmendra Sharma 0001 |
ICONIP (2) | 1 |
| 2010 | An optimal sphere and two large margins approach for novelty detectionabstractWe introduce a new model to deal with imbalanced data sets for novelty detection problems where the normal class of training data set can be majority or minority class. The key idea is to construct an optimal hypersphere such that the inside margin between the surface of this sphere and the normal data and the outside margin between that surface and the abnormal data are as large as possible. Depending on a specific real application of novelty detection, the two margins can be adjusted to achieve the best true positive and false positive rates. Experimental results on a number of data sets showed that the proposed model can provide better performance comparing with current models for novelty detection. Trung Le 0001, Dat Tran 0001, Wanli Ma 0003, Dharmendra Sharma 0001 |
IJCNN | 1 |
| 2010 | Fuzzy support vector machines for age and gender classification
Phuoc Nguyen, Trung Le 0001, Dat Tran 0001, Xu Huang 0001, Dharmendra Sharma 0001 |
INTERSPEECH | 2 |