Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hoang V. Tran

dblp:387/4324 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Deep learning architectures and training · 55% Learning theory · 19% Optimization for machine learning · 9%
Theoretical computer science
4 papers
Mathematical optimization · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization
optimal transport
3.542025
Tree-Sliced Wasserstein Distance with Nonlinear Projection · ICML 2025
Tree-Sliced Wasserstein Distance: A Geometric Perspective · ICML 2025
Distance-Based Tree-Sliced Wasserstein Distance · ICLR 2025
Machine learning › Learning theory
probability metric
1.722025
Distance-Based Tree-Sliced Wasserstein Distance · ICLR 2025
Spherical Tree-Sliced Wasserstein Distance · ICLR 2025
Mathematical optimization › optimal transport › wasserstein distance
sliced wasserstein distance
1.722025
Tree-Sliced Wasserstein Distance: A Geometric Perspective · ICML 2025
Spherical Tree-Sliced Wasserstein Distance · ICLR 2025
Machine learning › Deep learning architectures and training
equivariant neural network
1.622025
Equivariant Polynomial Functional Networks · ICML 2025
Monomial Matrix Group Equivariant Neural Functional Networks · NeurIPS 2024
Machine learning › Deep learning architectures and training › neural operator
neural functional network
1.622025
Equivariant Neural Functional Networks for Transformers · ICLR 2025
Monomial Matrix Group Equivariant Neural Functional Networks · NeurIPS 2024
Machine learning › Graph learning › graph neural network
message passing
0.912025
Equivariant Polynomial Functional Networks · ICML 2025
Machine learning › Optimization for machine learning › optimal transport
sliced wasserstein distance
0.912025
Tree-Sliced Wasserstein Distance with Nonlinear Projection · ICML 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Equivariant Neural Functional Networks for Transformers · ICLR 2025
Mathematical optimization › optimal transport
wasserstein distance
0.912025
Tree-Sliced Wasserstein Distance with Nonlinear Projection · ICML 2025
Machine learning › Deep learning architectures and training › training dynamics
weight symmetry
0.812024
Monomial Matrix Group Equivariant Neural Functional Networks · NeurIPS 2024
Machine learning › Generative modeling
generative model
0.522025
Tree-Sliced Wasserstein Distance: A Geometric Perspective · ICML 2025
Distance-Based Tree-Sliced Wasserstein Distance · ICLR 2025
Machine learning › Representation and self-supervised learning
equivariance
0.312025
Equivariant Neural Functional Networks for Transformers · ICLR 2025
Machine learning › Deep learning architectures and training › symmetry-aware learning
permutation invariance
0.212024
Monomial Matrix Group Equivariant Neural Functional Networks · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

radon transform · 5.2tree-sliced optimal transport · 1.7tree systems · 1.7tree sampling · 1.7spherical radon transform · 1.7optimal transport · 1.7nonlinear projection · 1.7gradient flow · 1.7group action · 0.9equivariant networks · 0.9
YearPublicationVenuePosition
2025 Spherical Tree-Sliced Wasserstein Distance
abstract
Sliced Optimal Transport (OT) simplifies the OT problem in high-dimensional spaces by projecting supports of input measures onto one-dimensional lines, then exploiting the closed-form expression of the univariate OT to reduce the computational burden of OT. Recently, the Tree-Sliced method has been introduced to replace these lines with more intricate structures, known as tree systems. This approach enhances the ability to capture topological information of integration domains in Sliced OT while maintaining low computational cost. Inspired by this approach, in this paper, we present an adaptation of tree systems on OT problem for measures supported on a sphere. As counterpart to the Radon transform variant on tree systems, we propose a novel spherical Radon transform, with a new integration domain called spherical trees. By leveraging this transform and exploiting the spherical tree structures, we derive closed-form expressions for OT problems on the sphere. Consequently, we obtain an efficient metric for measures on the sphere, named Spherical Tree-Sliced Wasserstein (STSW) distance. We provide an extensive theoretical analysis to demonstrate the topology of spherical trees, the well-definedness and injectivity of our Radon transform variant, which leads to an orthogonally invariant distance between spherical measures. Finally, we conduct a wide range of numerical experiments, including gradient flows and self-supervised learning, to assess the performance of our proposed metric, comparing it to recent benchmarks.
Hoang V. Tran, Thanh T. Chu, Minh-Khoi Nguyen-Nhat, Huyen Trang Pham, Tam Le, Tan M. Nguyen
ICLR1
2025 Distance-Based Tree-Sliced Wasserstein Distance
abstract
To overcome computational challenges of Optimal Transport (OT), several variants of Sliced Wasserstein (SW) has been developed in the literature. These approaches exploit the closed-form expression of the univariate OT by projecting measures onto one-dimensional lines. However, projecting measures onto low-dimensional spaces can lead to a loss of topological information. Tree-Sliced Wasserstein distance on Systems of Lines (TSW-SL) has emerged as a promising alternative that replaces these lines with a more intricate structure called tree systems. The tree structures enhance the ability to capture topological information of the metric while preserving computational efficiency. However, at the core of TSW-SL, the splitting maps, which serve as the mechanism for pushing forward measures onto tree systems, focus solely on the position of the measure supports while disregarding the projecting domains. Moreover, the specific splitting map used in TSW-SL leads to a metric that is not invariant under Euclidean transformations, a typically expected property for OT on Euclidean space. In this work, we propose a novel class of splitting maps that generalizes the existing one studied in TSW-SL enabling the use of all positional information from input measures, resulting in a novel Distance-based Tree-Sliced Wasserstein (*Db-TSW*) distance. In addition, we introduce a simple tree sampling process better suited for Db-TSW, leading to an efficient GPU-friendly implementation for tree systems, similar to the original SW. We also provide a comprehensive theoretical analysis of proposed class of splitting maps to verify the injectivity of the corresponding Radon Transform, and demonstrate that Db-TSW is an Euclidean invariant metric. We empirically show that Db-TSW significantly improves accuracy compared to recent SW variants while maintaining low computational cost via a wide range of experiments on gradient flows, image style transfer, and generative models.
Hoang V. Tran, Minh-Khoi Nguyen-Nhat, Huyen Trang Pham, Thanh T. Chu, Tam Le, Tan M. Nguyen
ICLR1
2025 Equivariant Neural Functional Networks for Transformers
abstract
This paper systematically explores neural functional networks (NFN) for transformer architectures. NFN are specialized neural networks that treat the weights, gradients, or sparsity patterns of a deep neural network (DNN) as input data and have proven valuable for tasks such as learnable optimizers, implicit data representations, and weight editing. While NFN have been extensively developed for MLP and CNN, no prior work has addressed their design for transformers, despite the importance of transformers in modern deep learning. This paper aims to address this gap by providing a systematic study of NFN for transformers. We first determine the maximal symmetric group of the weights in a multi-head attention module as well as a necessary and sufficient condition under which two sets of hyperparameters of the multi-head attention module define the same function. We then define the weight space of transformer architectures and its associated group action, which leads to the design principles for NFN in transformers. Based on these, we introduce Transformer-NFN, an NFN that is equivariant under this group action. Additionally, we release a dataset of more than 125,000 Transformers model checkpoints trained on two datasets with two different tasks, providing a benchmark for evaluating Transformer-NFN and encouraging further research on transformer training and performance.
Hoang V. Tran, Thieu Vo, An Nguyen The, Tho Tran, Minh-Khoi Nguyen-Nhat, Duy-Tung Pham, Tan M. Nguyen
ICLR1
2025 Tree-Sliced Wasserstein Distance: A Geometric Perspective
abstract
Many variants of Optimal Transport (OT) have been developed to address its heavy computation. Among them, notably, Sliced Wasserstein (SW) is widely used for application domains by projecting the OT problem onto one-dimensional lines, and leveraging the closed-form expression of the univariate OT to reduce the computational burden. However, projecting measures onto low-dimensional spaces can lead to a loss of topological information. To mitigate this issue, in this work, we propose to replace one-dimensional lines with a more intricate structure, called tree systems. This structure is metrizable by a tree metric, which yields a closed-form expression for OT problems on tree systems. We provide an extensive theoretical analysis to formally define tree systems with their topological properties, introduce the concept of splitting maps, which operate as the projection mechanism onto these structures, then finally propose a novel variant of Radon transform for tree systems and verify its injectivity. This framework leads to an efficient metric between measures, termed Tree-Sliced Wasserstein distance on Systems of Lines (TSW-SL). By conducting a variety of experiments on gradient flows, image style transfer, and generative models, we illustrate that our proposed approach performs favorably compared to SW and its variants.
Hoang V. Tran, Huyen Trang Pham, Tho Tran, Minh-Khoi Nguyen-Nhat, Thanh T. Chu, Tam Le, Tan M. Nguyen
ICML1
2025 Tree-Sliced Wasserstein Distance with Nonlinear Projection
abstract
Tree-Sliced methods have recently emerged as an alternative to the traditional Sliced Wasserstein (SW) distance, replacing one-dimensional lines with tree-based metric spaces and incorporating a splitting mechanism for projecting measures. This approach enhances the ability to capture the topological structures of integration domains in Sliced Optimal Transport while maintaining low computational costs. Building on this foundation, we propose a novel nonlinear projectional framework for the Tree-Sliced Wasserstein (TSW) distance, substituting the linear projections in earlier versions with general projections, while ensuring the injectivity of the associated Radon Transform and preserving the well-definedness of the resulting metric. By designing appropriate projections, we construct efficient metrics for measures on both Euclidean spaces and spheres. Finally, we validate our proposed metric through extensive numerical experiments for Euclidean and spherical datasets. Applications include gradient flows, self-supervised learning, and generative models, where our methods demonstrate significant improvements over recent SW and TSW variants.
Hoang V. Tran, Thanh T. Chu, Huyen Trang Pham, Laurent El Ghaoui, Tam Le, Tan M. Nguyen
ICML2
2025 Equivariant Polynomial Functional Networks
abstract
A neural functional network (NFN) is a specialized type of neural network designed to process and learn from entire neural networks as input data. Recent NFNs have been proposed with permutation and scaling equivariance based on either graph-based message-passing mechanisms or parameter-sharing mechanisms. Compared to graph-based models, parameter-sharing-based NFNs built upon equivariant linear layers exhibit lower memory consumption and faster running time. However, their expressivity is limited due to the large size of the symmetric group of the input neural networks. The challenge of designing a permutation and scaling equivariant NFN that maintains low memory consumption and running time while preserving expressivity remains unresolved. In this paper, we propose a novel solution with the development of MAGEP-NFN (**M**onomial m**A**trix **G**roup **E**quivariant **P**olynomial **NFN**). Our approach follows the parameter-sharing mechanism but differs from previous works by constructing a nonlinear equivariant layer represented as a polynomial in the input weights. This polynomial formulation enables us to incorporate additional relationships between weights from different input hidden layers, enhancing the model's expressivity while keeping memory consumption and running time low, thereby addressing the aforementioned challenge. We provide empirical evidence demonstrating that MAGEP-NFN achieves competitive performance and efficiency compared to existing baselines.
Thieu Vo, Hoang V. Tran, Tho Tran, An Nguyen The, Minh-Khoi Nguyen-Nhat, Duy-Tung Pham, Tan M. Nguyen
ICML2
2024 Monomial Matrix Group Equivariant Neural Functional Networks
abstract
Neural functional networks (NFNs) have recently gained significant attention due to their diverse applications, ranging from predicting network generalization and network editing to classifying implicit neural representation. Previous NFN designs often depend on permutation symmetries in neural networks' weights, which traditionally arise from the unordered arrangement of neurons in hidden layers. However, these designs do not take into account the weight scaling symmetries of $\operatorname{ReLU}$ networks, and the weight sign flipping symmetries of $\operatorname{sin}$ or $\operatorname{Tanh}$ networks. In this paper, we extend the study of the group action on the network weights from the group of permutation matrices to the group of monomial matrices by incorporating scaling/sign-flipping symmetries. Particularly, we encode these scaling/sign-flipping symmetries by designing our corresponding equivariant and invariant layers. We name our new family of NFNs the Monomial Matrix Group Equivariant Neural Functional Networks (Monomial-NFN). Because of the expansion of the symmetries, Monomial-NFN has much fewer independent trainable parameters compared to the baseline NFNs in the literature, thus enhancing the model's efficiency. Moreover, for fully connected and convolutional neural networks, we theoretically prove that all groups that leave these networks invariant while acting on their weight spaces are some subgroups of the monomial matrix group. We provide empirical evidences to demonstrate the advantages of our model over existing baselines, achieving competitive performance and efficiency. The code is publicly available at https://github.com/MathematicalAI-NUS/Monomial-NFN.
Hoang V. Tran, Ngoc Thieu Vo, Tho Huu, An Nguyen The, Tan M. Nguyen
NeurIPS1