EDBT 2026 Demo / reviewers in the wild / expert
Thang Bui
dblp:171/2650
· DBLP profile ↗
17ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 9 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mining Attribute-Based Access Control Policies with Large Language Models [Work in Progress Paper]abstractAttribute-Based Access Control (ABAC) supports flexible, expressive authorization policies but is often difficult to author from scratch. Mining ABAC policies from existing permissions and attribute data can reduce this effort, yet conventional approaches rely on specialized algorithms. This paper explores whether large language models (LLMs) can serve as zero-shot ABAC policy miners, producing usable policies directly from attribute data and observed permissions without any task-specific training. Thang Bui, Esteban Martinez Mota |
SACMAT | 1 |
| 2025 | Enhancing Link Prediction with Attentive-HINGE: A Vietnamese Case Study in Educational Knowledge Graphs
Nhan Vo, Nhat Truong, Tuan Bui, Khang Huynh, Thang Bui |
IEEE Big Data | 7 |
| 2025 | Text-JEPA: A Joint Embedding Predictive Architecture for the Conversion of Natural Language into First-Order Logic
Trong Le, Phat Thai, Sang Nguyen, Minh Hua, Ngan Pham, Thang Bui, Tuan Bui |
ICCCI (1) | 6 |
| 2025 | Speaking in Words, Thinking in Logic: A Dual-Process Framework in QA SystemsabstractRecent advances in large language models (LLMs) have significantly enhanced question-answering (QA) capabilities, particularly in open-domain contexts. However, in closed-domain scenarios such as education, healthcare, and law, users demand not only accurate answers but also transparent reasoning and explainable decision-making processes. While neural-symbolic (NeSy) frameworks have emerged as a promising solution—leveraging LLMs for natural language understanding and symbolic systems for formal reasoning—existing approaches often rely on large-scale models and exhibit inefficiencies in translating natural language into formal logic representations.To address these limitations, we introduce Text-JEPA (Text-based Joint-Embedding Predictive Architecture), a lightweight yet effective framework for converting natural language into first-order logic (NL2FOL). Drawing inspiration from dual-system cognitive theory, Text-JEPA emulates System 1 by efficiently generating logic representations, while the Z3 solver operates as System 2, enabling robust logical inference. To rigorously evaluate the NL2FOL-to-reasoning pipeline, we propose a comprehensive evaluation framework comprising three custom metrics: conversion_score, reasoning_score, and Spearman_rho_score, which collectively capture the quality of logical translation and its downstream impact on reasoning accuracy.Empirical results on domain-specific datasets demonstrate that Text-JEPA achieves competitive performance with significantly lower computational overhead compared to larger LLM-based systems. Our findings highlight the potential of structured, interpretable reasoning frameworks for building efficient and explainable QA systems in specialized domains. Tuan Bui, Trong Le, Phat Thai, Sang Nguyen, Minh Hua, Ngan Pham, Thang Bui |
IJCNN | 7 |
| 2025 | Rao-Blackwellised Reparameterisation GradientsabstractLatent Gaussian variables have been popularised in probabilistic machine learning. In turn, gradient estimators are the machinery that facilitates gradient-based optimisation for models with latent Gaussian variables. The reparameterisation trick is often used as the default estimator as it is simple to implement and yields low-variance gradients for variational inference. In this work, we propose the R2-G2 estimator as the Rao-Blackwellisation of the reparameterisation gradient estimator. Interestingly, we show that the local reparameterisation gradient estimator for Bayesian MLPs is an instance of the R2-G2 estimator and Rao-Blackwellisation. This lets us extend benefits of Rao-Blackwellised gradients to a suite of probabilistic models. We show that initial training with R2-G2 consistently yields better performance in models with multiple applications of the reparameterisation trick. Kevin H. Lam, Thang Bui, George Deligiannidis, Yee Whye Teh |
NeurIPS | 2 |
| 2025 | ABAC Lab: An Interactive Platform for Attribute-based Access Control Policy Analysis, Tools, and Datasets [Dataset/Tool Paper]abstractAttribute-Based Access Control (ABAC) provides expressiveness and flexibility, making it a compelling model for enforcing fine-grained access control policies. To facilitate the transition to ABAC, extensive research has been conducted to develop methodologies, frameworks, and tools that assist policy administrators in adapting the model. Despite these efforts, challenges remain in the availability and benchmarking of ABAC datasets. Specifically, there is a lack of clarity on how datasets can be systematically acquired, no standardized benchmarking practices to evaluate existing methodologies and their effectiveness, and limited access to real-world datasets suitable for policy analysis and testing. Thang Bui, Anthony Matricia, Emily Contreras, Ryan Mauvais, Luis Medina 0005, Israel Serrano |
SACMAT | 1 |
| 2024 | Toward Context-Aware Federated Learning Assessment: A Reality CheckabstractFederated learning (FL) enabled creating models that are competitive to centralized machine learning models, without compromising user privacy. Participating FL clients train local models on their data and only share model weights. An FL server aggregates these weights into global weights that are pushed to clients for the next training round. Despite FL research growth, most of this work is conceived in experimental simulated environments that do not reflect its applicability to real-world scenarios. Also, existing open-source FL testbeds/frameworks have drawbacks that prohibit convenient deployment over a large spectrum of heterogeneous clients in realistic environments. These drawbacks include simulations, unrealistic data sets, not supporting heterogeneity, and not having realistic environment control in terms of network and client churn, for example. In this article, we introduce (RealFL) a novel, realistic, open-source, and extendable platform for FL that supports a large scale of heterogeneous clients. It enables a realistic assessment of FL solutions by controlling various environmental parameters, e.g., network, client churn, data distribution, training complexity, and client heterogeneity. Using these parameters, we assess RealFL performance through an extensive evaluation. Preliminary evaluation shows a performance gap of up to 72% in training time and 27% in accuracy between FL-simulated environments and RealFL. Moreover, extensive evaluation reveals that realistic environmental parameters could affect accuracy by up to 52.7%, training time by up to 77.5%, and communication overhead by up to 98%. Hend Gedawy, Khaled A. Harras, Thang Bui, Temoor Tanveer |
IEEE Internet Things J. | 3 |
| 2023 | Bridging the Chasm Between Ideal and Realistic Federated Learning: A Measurements StudyabstractFederated Learning is being hailed as a privacy-preserving machine learning alternative, by allowing models to be distributively trained on source devices owning their data. Most FL solutions, and their assessments, however, assume superior environmental reliability, despite the more realistic variances in environmental factors such as device and network capacity, data distribution, and device churn. As such, we argue in this paper, that there is a growing chasm between current FL assessment setups and the evolving FL assessment needs. Motivated by this chasm, we conduct, to the best of our knowledge, the first empirical measurement study of FL performance given realistic environmental factors. Our study quantifies the impact of these environmental factors on FL performance in terms of training time, accuracy, and communication overhead. Our findings have broad implications for the future development of FL including client admission control and scheduling optimizations. Hend Gedawy, Khaled A. Harras, Temoor Tanveer, Thang Bui |
CloudCom | 4 |
| 2023 | RealFL: A Realistic Platform for Federated LearningabstractFederated Learning (FL) enabled creating models that are competitive to centralized Machine Learning models while preserving privacy by allowing clients to train data locally. Despite FL research growth, most of the work assessment and existing open-source FL testbeds/frameworks have drawbacks that prohibit convenient deployment over a large spectrum of heterogeneous clients in realistic environments. These drawbacks include simulations, unrealistic datasets, not supporting heterogeneity, and not having a realistic environment control in terms of network and client churn, for example. In this paper, we introduce (RealFL) a novel, realistic, open-source, and extendable platform for FL that supports a large scale of heterogeneous clients. It enables realistic assessment of FL solutions by controlling various environmental parameters; e.g. network, client churn, data distribution, training complexity, and client heterogeneity. Using these parameters, we assess RealFL performance through an extensive evaluation. The results show a performance gap of up to 77.5% in training time and 23.9% in accuracy between FL unrealistic environments and RealFL. Hend Gedawy, Khaled A. Harras, Thang Bui, Temoor Tanveer |
MSWiM | 3 |
| 2021 | q-Paths: Generalizing the geometric annealing path using power meansabstractMany common machine learning methods involve the geometric annealing path, a sequence of intermediate densities between two distributions of interest constructed using the geometric average. While alternatives such as the moment-averaging path have demonstrated performance gains in some settings, their practical applicability remains limited by exponential family endpoint assumptions and a lack of closed form energy function. In this work, we introduce $q$-paths, a family of paths which is derived from a generalized notion of the mean, includes the geometric and arithmetic mixtures as special cases, and admits a simple closed form involving the deformed logarithm function from nonextensive thermodynamics. Following previous analysis of the geometric path, we interpret our $q$-paths as corresponding to a $q$-exponential family of distributions, and provide a variational representation of intermediate densities as minimizing a mixture of $\alpha$-divergences to the endpoints. We show that small deviations away from the geometric path yield empirical gains for Bayesian inference using Sequential Monte Carlo and generative model evaluation using Annealed Importance Sampling. Vaden Masrani, Rob Brekelmans, Thang Bui, Frank Nielsen, Aram Galstyan, Greg Ver Steeg, Frank D. Wood |
UAI | 3 |
| 2020 | A Decision Tree Learning Approach for Mining Relationship-Based Access Control PoliciesabstractRelationship-based access control (ReBAC) provides a high level of expressiveness and flexibility that promotes security and information sharing, by allowing policies to be expressed in terms of chains of relationships between entities. ReBAC policy mining algorithms have the potential to significantly reduce the cost of migration from legacy access control systems to ReBAC, by partially automating the development of a ReBAC policy. This paper presents new algorithms, called DTRM (Decision Tree ReBAC Miner) and DTRM-, based on decision trees, for mining ReBAC policies from access control lists (ACLs) and information about entities. Compared to state-of-the-art ReBAC mining algorithms, our algorithms are significantly faster, achieve comparable policy quality, and can mine policies in a richer language. Thang Bui, Scott D. Stoller |
SACMAT | 1 |
| 2019 | Efficient and Extensible Policy Mining for Relationship-Based Access ControlabstractRelationship-based access control (ReBAC) is a flexible and expressive framework that allows policies to be expressed in terms of chains of relationship between entities as well as attributes of entities. ReBAC policy mining algorithms have a potential to significantly reduce the cost of migration from legacy access control systems to ReBAC, by partially automating the development of a ReBAC policy. Existing ReBAC policy mining algorithms support a policy language with a limited set of operators; this limits their applicability. Thang Bui, Scott D. Stoller, Hieu Le 0001 |
SACMAT | 1 |
| 2019 | Greedy and evolutionary algorithms for mining relationship-based access control policies
Thang Bui, Scott D. Stoller |
Comput. Secur. | 1 |
| 2018 | Mining hierarchical temporal roles with multiple metricsabstractTemporal role-based access control (TRBAC) extends role-based access control to limit the times at which roles are enabled. This paper presents a new algorithm for mining high-quality TRBAC policies from timed ACLs (i.e., ACLs with time limits in the entries) and optionally user attribute information. Such algorithms have potential to significantly reduce the cost of migration from timed ACLs to TRBAC. The algorithm is parameterized by the policy quality metric. We consider multiple quality metrics, including number of roles, weighted structural complexity (a generalization of policy size), and (when user attribute information is available) interpretability, i.e., how well role membership can be characterized in terms of user attributes. Ours is the first TRBAC policy mining algorithm that produces hierarchical policies, and the first that optimizes weighted structural complexity or interpretability. In experiments with datasets based on real-world ACL policies, our algorithm is more effective than previous algorithms at optimizing policy quality. Scott D. Stoller, Thang Bui |
J. Comput. Secur. | 2 |
| 2017 | Fast Distributed Evaluation of Stateful Attribute-Based Access Control Policies
Thang Bui, Scott D. Stoller |
DBSec | 1 |
| 2017 | Mining Relationship-Based Access Control PoliciesabstractRelationship-based access control (ReBAC) provides a high level of expressiveness and flexibility that promotes security and information sharing. We formulate ReBAC as an object-oriented extension of attribute-based access control (ABAC) in which relationships are expressed using fields that refer to other objects, and path expressions are used to follow chains of relationships between objects. Thang Bui, Scott D. Stoller |
SACMAT | 1 |
| 2016 | Mining Hierarchical Temporal Roles with Multiple Metrics
Scott D. Stoller, Thang Bui |
DBSec | 2 |