EDBT 2026 Demo / reviewers in the wild / expert
Tho Tran
dblp:337/2038 · also Tho Tran Huu
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 47% Generative modeling · 22% Learning theory · 12% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › neural collapse
class-imbalanced neural collapse |
1.4 | 2 | 2024 | Neural Collapse for Cross-entropy Class-Imbalanced Learning with Unconstrained ReLU Features Model · ICML 2024 Neural Collapse in Deep Linear Networks: From Balanced to Imbalanced Data · ICML 2023 |
Machine learning › Deep learning architectures and training
neural collapse |
1.4 | 2 | 2024 | Neural Collapse for Cross-entropy Class-Imbalanced Learning with Unconstrained ReLU Features Model · ICML 2024 Neural Collapse in Deep Linear Networks: From Balanced to Imbalanced Data · ICML 2023 |
Machine learning › Learning theory
neural network theory |
1.4 | 2 | 2024 | Neural Collapse for Cross-entropy Class-Imbalanced Learning with Unconstrained ReLU Features Model · ICML 2024 Neural Collapse in Deep Linear Networks: From Balanced to Imbalanced Data · ICML 2023 |
Machine learning › Deep learning architectures and training
equivariant neural network |
0.9 | 1 | 2025 | Equivariant Polynomial Functional Networks · ICML 2025 |
Machine learning › Graph learning › graph neural network
message passing |
0.9 | 1 | 2025 | Equivariant Polynomial Functional Networks · ICML 2025 |
Machine learning › Deep learning architectures and training › neural operator
neural functional network |
0.9 | 1 | 2025 | Equivariant Neural Functional Networks for Transformers · ICLR 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | Equivariant Neural Functional Networks for Transformers · ICLR 2025 |
Mathematical optimization
optimal transport |
0.9 | 1 | 2025 | Tree-Sliced Wasserstein Distance: A Geometric Perspective · ICML 2025 |
Mathematical optimization › optimal transport › wasserstein distance
sliced wasserstein distance |
0.9 | 1 | 2025 | Tree-Sliced Wasserstein Distance: A Geometric Perspective · ICML 2025 |
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder |
0.8 | 1 | 2024 | Beyond Vanilla Variational Autoencoders: Detecting Posterior Collapse in Conditional and Hierarchical Variational Autoencoders · ICLR 2024 |
Machine learning › Representation and self-supervised learning › representation learning
feature geometry |
0.8 | 1 | 2024 | Neural Collapse for Cross-entropy Class-Imbalanced Learning with Unconstrained ReLU Features Model · ICML 2024 |
Machine learning › Generative modeling › variational autoencoder
posterior collapse |
0.8 | 1 | 2024 | Beyond Vanilla Variational Autoencoders: Detecting Posterior Collapse in Conditional and Hierarchical Variational Autoencoders · ICLR 2024 |
Machine learning › Generative modeling
variational autoencoder |
0.8 | 1 | 2024 | Beyond Vanilla Variational Autoencoders: Detecting Posterior Collapse in Conditional and Hierarchical Variational Autoencoders · ICLR 2024 |
Machine learning › Representation and self-supervised learning
equivariance |
0.3 | 1 | 2025 | Equivariant Neural Functional Networks for Transformers · ICLR 2025 |
Machine learning › Generative modeling
generative model |
0.3 | 1 | 2025 | Tree-Sliced Wasserstein Distance: A Geometric Perspective · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.2 | 1 | 2024 | Beyond Vanilla Variational Autoencoders: Detecting Posterior Collapse in Conditional and Hierarchical Variational Autoencoders · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
tree systems · 1.7radon transform · 1.7optimal transport · 1.7cross-entropy loss analysis · 1.4polynomial equivariant layer · 0.9parameter sharing · 0.9group action · 0.9equivariant networks · 0.9variational inference · 0.8theoretical analysis · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Equivariant Neural Functional Networks for TransformersabstractThis paper systematically explores neural functional networks (NFN) for transformer architectures. NFN are specialized neural networks that treat the weights, gradients, or sparsity patterns of a deep neural network (DNN) as input data and have proven valuable for tasks such as learnable optimizers, implicit data representations, and weight editing. While NFN have been extensively developed for MLP and CNN, no prior work has addressed their design for transformers, despite the importance of transformers in modern deep learning. This paper aims to address this gap by providing a systematic study of NFN for transformers. We first determine the maximal symmetric group of the weights in a multi-head attention module as well as a necessary and sufficient condition under which two sets of hyperparameters of the multi-head attention module define the same function. We then define the weight space of transformer architectures and its associated group action, which leads to the design principles for NFN in transformers. Based on these, we introduce Transformer-NFN, an NFN that is equivariant under this group action. Additionally, we release a dataset of more than 125,000 Transformers model checkpoints trained on two datasets with two different tasks, providing a benchmark for evaluating Transformer-NFN and encouraging further research on transformer training and performance. Hoang V. Tran, Thieu Vo, An Nguyen The, Tho Tran, Minh-Khoi Nguyen-Nhat, Duy-Tung Pham, Tan M. Nguyen |
ICLR | 4 |
| 2025 | Tree-Sliced Wasserstein Distance: A Geometric PerspectiveabstractMany variants of Optimal Transport (OT) have been developed to address its heavy computation. Among them, notably, Sliced Wasserstein (SW) is widely used for application domains by projecting the OT problem onto one-dimensional lines, and leveraging the closed-form expression of the univariate OT to reduce the computational burden. However, projecting measures onto low-dimensional spaces can lead to a loss of topological information. To mitigate this issue, in this work, we propose to replace one-dimensional lines with a more intricate structure, called tree systems. This structure is metrizable by a tree metric, which yields a closed-form expression for OT problems on tree systems. We provide an extensive theoretical analysis to formally define tree systems with their topological properties, introduce the concept of splitting maps, which operate as the projection mechanism onto these structures, then finally propose a novel variant of Radon transform for tree systems and verify its injectivity. This framework leads to an efficient metric between measures, termed Tree-Sliced Wasserstein distance on Systems of Lines (TSW-SL). By conducting a variety of experiments on gradient flows, image style transfer, and generative models, we illustrate that our proposed approach performs favorably compared to SW and its variants. Hoang V. Tran, Huyen Trang Pham, Tho Tran, Minh-Khoi Nguyen-Nhat, Thanh T. Chu, Tam Le, Tan M. Nguyen |
ICML | 3 |
| 2025 | Equivariant Polynomial Functional NetworksabstractA neural functional network (NFN) is a specialized type of neural network designed to process and learn from entire neural networks as input data. Recent NFNs have been proposed with permutation and scaling equivariance based on either graph-based message-passing mechanisms or parameter-sharing mechanisms. Compared to graph-based models, parameter-sharing-based NFNs built upon equivariant linear layers exhibit lower memory consumption and faster running time. However, their expressivity is limited due to the large size of the symmetric group of the input neural networks. The challenge of designing a permutation and scaling equivariant NFN that maintains low memory consumption and running time while preserving expressivity remains unresolved. In this paper, we propose a novel solution with the development of MAGEP-NFN (**M**onomial m**A**trix **G**roup **E**quivariant **P**olynomial **NFN**). Our approach follows the parameter-sharing mechanism but differs from previous works by constructing a nonlinear equivariant layer represented as a polynomial in the input weights. This polynomial formulation enables us to incorporate additional relationships between weights from different input hidden layers, enhancing the model's expressivity while keeping memory consumption and running time low, thereby addressing the aforementioned challenge. We provide empirical evidence demonstrating that MAGEP-NFN achieves competitive performance and efficiency compared to existing baselines. Thieu Vo, Hoang V. Tran, Tho Tran, An Nguyen The, Minh-Khoi Nguyen-Nhat, Duy-Tung Pham, Tan M. Nguyen |
ICML | 3 |
| 2024 | Beyond Vanilla Variational Autoencoders: Detecting Posterior Collapse in Conditional and Hierarchical Variational AutoencodersabstractThe posterior collapse phenomenon in variational autoencoder (VAE), where the variational posterior distribution closely matches the prior distribution, can hinder the quality of the learned latent variables. As a consequence of posterior collapse, the latent variables extracted by the encoder in VAE preserve less information from the input data and thus fail to produce meaningful representations as input to the reconstruction process in the decoder. While this phenomenon has been an actively addressed topic related to VAE performance, the theory for posterior collapse remains underdeveloped, especially beyond the standard VAE. In this work, we advance the theoretical understanding of posterior collapse to two important and prevalent yet less studied classes of VAE: conditional VAE and hierarchical VAE. Specifically, via a non-trivial theoretical analysis of linear conditional VAE and hierarchical VAE with two levels of latent, we prove that the cause of posterior collapses in these models includes the correlation between the input and output of the conditional VAE and the effect of learnable encoder variance in the hierarchical VAE. We empirically validate our theoretical findings for linear conditional and hierarchical VAE and demonstrate that these results are also predictive for non-linear cases with extensive experiments. Hien Dang 0003, Tho Tran, Tan M. Nguyen, Nhat Ho |
ICLR | 2 |
| 2024 | Neural Collapse for Cross-entropy Class-Imbalanced Learning with Unconstrained ReLU Features ModelabstractThe current paradigm of training deep neural networks for classification tasks includes minimizing the empirical risk, pushing the training loss value towards zero even after the training classification error has vanished. In this terminal phase of training, it has been observed that the last-layer features collapse to their class-means and these class-means converge to the vertices of a simplex Equiangular Tight Frame (ETF). This phenomenon is termed as Neural Collapse ($\mathcal{NC}$). However, this characterization only holds in class-balanced datasets where every class has the same number of training samples. When the training dataset is class-imbalanced, some $\mathcal{NC}$ properties will no longer hold true, for example, the geometry of class-means will skew away from the simplex ETF. In this paper, we generalize $\mathcal{NC}$ to imbalanced regime for cross-entropy loss under the unconstrained ReLU features model. We demonstrate that while the within-class features collapse property still holds in this setting, the class-means will converge to a structure consisting of orthogonal vectors with lengths dependent on the number of training samples. Furthermore, we find that the classifier weights (i.e., the last-layer linear classifier) are aligned to the scaled and centered class-means, with scaling factors dependent on the number of training samples of each class. This generalizes $\mathcal{NC}$ in the class-balanced setting. We empirically validate our results through experiments on practical architectures and dataset. Hien Dang 0003, Tho Tran, Tan M. Nguyen, Nhat Ho |
ICML | 2 |
| 2024 | Revisiting Kernel Attention with Correlated Gaussian Process RepresentationabstractTransformers have increasingly become the de facto method to model sequential data with state-of-the-art performance. Due to its widespread use, being able to estimate and calibrate its modeling uncertainty is important to understand and design robust transformer models. To achieve this, previous works have used Gaussian processes (GPs) to perform uncertainty calibration for the attention units of transformers and attained notable successes. However, such approaches have to confine the transformers to the space of symmetric attention to ensure the necessary symmetric requirement of their GP’s kernel specification, which reduces the representation capacity of the model. To mitigate this restriction, we propose the Correlated Gaussian Process Transformer (CGPT), a new class of transformers whose self-attention units are modeled as cross-covariance between two correlated GPs (CGPs). This allows asymmetries in attention and can enhance the representation capacity of GP-based transformers. We also derive a sparse approximation for CGP to make it scale better. Our empirical studies show that both CGP-based and sparse CGP-based transformers achieve better performance than state-of-the-art GP-based transformers on a variety of benchmark tasks. Long Minh Bui, Tho Tran, Duy Dinh, Tan M. Nguyen, Trong Nghia Hoang |
UAI | 2 |
| 2023 | Neural Collapse in Deep Linear Networks: From Balanced to Imbalanced DataabstractModern deep neural networks have achieved impressive performance on tasks from image classification to natural language processing. Surprisingly, these complex systems with massive amounts of parameters exhibit the same structural properties in their last-layer features and classifiers across canonical datasets when training until convergence. In particular, it has been observed that the last-layer features collapse to their class-means, and those class-means are the vertices of a simplex Equiangular Tight Frame (ETF). This phenomenon is known as Neural Collapse (NC). Recent papers have theoretically shown that NC emerges in the global minimizers of training problems with the simplified ``unconstrained feature model''. In this context, we take a step further and prove the NC occurrences in deep linear networks for the popular mean squared error (MSE) and cross entropy (CE) losses, showing that global solutions exhibit NC properties across the linear layers. Furthermore, we extend our study to imbalanced data for MSE loss and present the first geometric analysis of NC under bias-free setting. Our results demonstrate the convergence of the last-layer features and classifiers to a geometry consisting of orthogonal vectors, whose lengths depend on the amount of data in their corresponding classes. Finally, we empirically validate our theoretical analyses on synthetic and practical network architectures with both balanced and imbalanced scenarios. Hien Dang 0003, Tho Tran, Stanley J. Osher, Hung Tran-The, Nhat Ho, Tan M. Nguyen |
ICML | 2 |