Huyen Trang Pham

dblp:404/7359 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
4 papers
Mathematical optimization · 100%
Artificial intelligence
5 papers
Learning theory · 44% Deep learning architectures and training · 22% Optimization for machine learning · 22%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization
optimal transport
3.542025
Tree-Sliced Wasserstein Distance with Nonlinear Projection · ICML 2025
Tree-Sliced Wasserstein Distance: A Geometric Perspective · ICML 2025
Distance-Based Tree-Sliced Wasserstein Distance · ICLR 2025
Machine learning › Learning theory
probability metric
1.722025
Distance-Based Tree-Sliced Wasserstein Distance · ICLR 2025
Spherical Tree-Sliced Wasserstein Distance · ICLR 2025
Mathematical optimization › optimal transport › wasserstein distance
sliced wasserstein distance
1.722025
Tree-Sliced Wasserstein Distance: A Geometric Perspective · ICML 2025
Spherical Tree-Sliced Wasserstein Distance · ICLR 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts · ICLR 2025
Machine learning › Optimization for machine learning › optimal transport
sliced wasserstein distance
0.912025
Tree-Sliced Wasserstein Distance with Nonlinear Projection · ICML 2025
Mathematical optimization › optimal transport
wasserstein distance
0.912025
Tree-Sliced Wasserstein Distance with Nonlinear Projection · ICML 2025
Machine learning › Generative modeling
generative model
0.522025
Tree-Sliced Wasserstein Distance: A Geometric Perspective · ICML 2025
Distance-Based Tree-Sliced Wasserstein Distance · ICLR 2025

Methods — techniques the papers use, named apart from their topics

radon transform · 5.2tree-sliced optimal transport · 1.7tree systems · 1.7tree sampling · 1.7spherical radon transform · 1.7optimal transport · 1.7nonlinear projection · 1.7gradient flow · 1.7perturbation analysis · 0.9least-squares estimation · 0.9
YearPublicationVenuePosition
2025 Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
abstract
The cosine router in Mixture of Experts (MoE) has recently emerged as an attractive alternative to the conventional linear router. Indeed, the cosine router demonstrates favorable performance in image and language tasks and exhibits better ability to mitigate the representation collapse issue, which often leads to parameter redundancy and limited representation potentials. Despite its empirical success, a comprehensive analysis of the cosine router in MoE has been lacking. Considering the least square estimation of the cosine routing MoE, we demonstrate that due to the intrinsic interaction of the model parameters in the cosine router via some partial differential equations, regardless of the structures of the experts, the estimation rates of experts and model parameters can be as slow as $\mathcal{O}(1/\log^{\tau}(n))$ where $\tau > 0$ is some constant and $n$ is the sample size. Surprisingly, these pessimistic non-polynomial convergence rates can be circumvented by the widely used technique in practice to stabilize the cosine router --- simply adding noises to the $\ell^2$-norms in the cosine router, which we refer to as *perturbed cosine router*. Under the strongly identifiable settings of the expert functions, we prove that the estimation rates for both the experts and model parameters under the perturbed cosine routing MoE are significantly improved to polynomial rates. Finally, we conduct extensive simulation studies in both synthetic and real data settings to empirically validate our theoretical results.
Pedram Akbarian, Huyen Trang Pham, Thien Trang Nguyen Vu, Shujian Zhang, Nhat Ho
ICLR3
2025 Spherical Tree-Sliced Wasserstein Distance
abstract
Sliced Optimal Transport (OT) simplifies the OT problem in high-dimensional spaces by projecting supports of input measures onto one-dimensional lines, then exploiting the closed-form expression of the univariate OT to reduce the computational burden of OT. Recently, the Tree-Sliced method has been introduced to replace these lines with more intricate structures, known as tree systems. This approach enhances the ability to capture topological information of integration domains in Sliced OT while maintaining low computational cost. Inspired by this approach, in this paper, we present an adaptation of tree systems on OT problem for measures supported on a sphere. As counterpart to the Radon transform variant on tree systems, we propose a novel spherical Radon transform, with a new integration domain called spherical trees. By leveraging this transform and exploiting the spherical tree structures, we derive closed-form expressions for OT problems on the sphere. Consequently, we obtain an efficient metric for measures on the sphere, named Spherical Tree-Sliced Wasserstein (STSW) distance. We provide an extensive theoretical analysis to demonstrate the topology of spherical trees, the well-definedness and injectivity of our Radon transform variant, which leads to an orthogonally invariant distance between spherical measures. Finally, we conduct a wide range of numerical experiments, including gradient flows and self-supervised learning, to assess the performance of our proposed metric, comparing it to recent benchmarks.
Hoang V. Tran, Thanh T. Chu, Minh-Khoi Nguyen-Nhat, Huyen Trang Pham, Tam Le, Tan M. Nguyen
ICLR4
2025 Distance-Based Tree-Sliced Wasserstein Distance
abstract
To overcome computational challenges of Optimal Transport (OT), several variants of Sliced Wasserstein (SW) has been developed in the literature. These approaches exploit the closed-form expression of the univariate OT by projecting measures onto one-dimensional lines. However, projecting measures onto low-dimensional spaces can lead to a loss of topological information. Tree-Sliced Wasserstein distance on Systems of Lines (TSW-SL) has emerged as a promising alternative that replaces these lines with a more intricate structure called tree systems. The tree structures enhance the ability to capture topological information of the metric while preserving computational efficiency. However, at the core of TSW-SL, the splitting maps, which serve as the mechanism for pushing forward measures onto tree systems, focus solely on the position of the measure supports while disregarding the projecting domains. Moreover, the specific splitting map used in TSW-SL leads to a metric that is not invariant under Euclidean transformations, a typically expected property for OT on Euclidean space. In this work, we propose a novel class of splitting maps that generalizes the existing one studied in TSW-SL enabling the use of all positional information from input measures, resulting in a novel Distance-based Tree-Sliced Wasserstein (*Db-TSW*) distance. In addition, we introduce a simple tree sampling process better suited for Db-TSW, leading to an efficient GPU-friendly implementation for tree systems, similar to the original SW. We also provide a comprehensive theoretical analysis of proposed class of splitting maps to verify the injectivity of the corresponding Radon Transform, and demonstrate that Db-TSW is an Euclidean invariant metric. We empirically show that Db-TSW significantly improves accuracy compared to recent SW variants while maintaining low computational cost via a wide range of experiments on gradient flows, image style transfer, and generative models.
Hoang V. Tran, Minh-Khoi Nguyen-Nhat, Huyen Trang Pham, Thanh T. Chu, Tam Le, Tan M. Nguyen
ICLR3
2025 Tree-Sliced Wasserstein Distance: A Geometric Perspective
abstract
Many variants of Optimal Transport (OT) have been developed to address its heavy computation. Among them, notably, Sliced Wasserstein (SW) is widely used for application domains by projecting the OT problem onto one-dimensional lines, and leveraging the closed-form expression of the univariate OT to reduce the computational burden. However, projecting measures onto low-dimensional spaces can lead to a loss of topological information. To mitigate this issue, in this work, we propose to replace one-dimensional lines with a more intricate structure, called tree systems. This structure is metrizable by a tree metric, which yields a closed-form expression for OT problems on tree systems. We provide an extensive theoretical analysis to formally define tree systems with their topological properties, introduce the concept of splitting maps, which operate as the projection mechanism onto these structures, then finally propose a novel variant of Radon transform for tree systems and verify its injectivity. This framework leads to an efficient metric between measures, termed Tree-Sliced Wasserstein distance on Systems of Lines (TSW-SL). By conducting a variety of experiments on gradient flows, image style transfer, and generative models, we illustrate that our proposed approach performs favorably compared to SW and its variants.
Hoang V. Tran, Huyen Trang Pham, Tho Tran, Minh-Khoi Nguyen-Nhat, Thanh T. Chu, Tam Le, Tan M. Nguyen
ICML2
2025 Tree-Sliced Wasserstein Distance with Nonlinear Projection
abstract
Tree-Sliced methods have recently emerged as an alternative to the traditional Sliced Wasserstein (SW) distance, replacing one-dimensional lines with tree-based metric spaces and incorporating a splitting mechanism for projecting measures. This approach enhances the ability to capture the topological structures of integration domains in Sliced Optimal Transport while maintaining low computational costs. Building on this foundation, we propose a novel nonlinear projectional framework for the Tree-Sliced Wasserstein (TSW) distance, substituting the linear projections in earlier versions with general projections, while ensuring the injectivity of the associated Radon Transform and preserving the well-definedness of the resulting metric. By designing appropriate projections, we construct efficient metrics for measures on both Euclidean spaces and spheres. Finally, we validate our proposed metric through extensive numerical experiments for Euclidean and spherical datasets. Applications include gradient flows, self-supervised learning, and generative models, where our methods demonstrate significant improvements over recent SW and TSW variants.
Hoang V. Tran, Thanh T. Chu, Huyen Trang Pham, Laurent El Ghaoui, Tam Le, Tan M. Nguyen
ICML4