Nam H. Nguyen

dblp:76/2975 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0002-1254-0069ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Theory of computation · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 61% Efficient and distributed learning · 22% Time series and sequential data · 17%
Theoretical computer science
4 papers
Mathematical optimization · 42% Information theory · 41% Algorithms and data structures · 8%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
foundation model
0.812024
AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability Data · AAAI 2024
Machine learning › Deep learning architectures and training › foundation model
time series foundation model
0.812024
AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability Data · AAAI 2024
Data mining › time series analysis › time series forecasting
multivariate time series forecasting
0.812024
AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability Data · AAAI 2024
Data mining › time series analysis
time series forecasting
0.812024
AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability Data · AAAI 2024
Machine learning › Time series and sequential data › time series analysis › time series forecasting
long-term time series forecasting
0.712023
A Time Series is Worth 64 Words: Long-term Forecasting with Transformers · ICLR 2023
Machine learning › Deep learning architectures and training
transformer
0.712023
A Time Series is Worth 64 Words: Long-term Forecasting with Transformers · ICLR 2023
Information theory › signal processing › compressed sensing
sparse recovery
0.532013
Robust Lasso With Missing and Grossly Corrupted Observations · IEEE Trans. Inf. Theory 2013
Exact Recoverability From Dense Corrupted Observations via ℓ1-Minimization · IEEE Trans. Inf. Theory 2013
Robust Lasso with missing and grossly corrupted observations · NIPS 2011
Machine learning › Efficient and distributed learning
model compression
0.412020
Pruning Deep Neural Networks with $\ell_{0}$-constrained Optimization · ICDM 2020
Machine learning › Efficient and distributed learning › model compression
pruning
0.412020
Pruning Deep Neural Networks with $\ell_{0}$-constrained Optimization · ICDM 2020
Information theory › signal processing
compressed sensing
0.322013
Robust Lasso With Missing and Grossly Corrupted Observations · IEEE Trans. Inf. Theory 2013
Exact Recoverability From Dense Corrupted Observations via ℓ1-Minimization · IEEE Trans. Inf. Theory 2013
Mathematical optimization › continuous optimization
convex optimization
0.322013
Robust Lasso With Missing and Grossly Corrupted Observations · IEEE Trans. Inf. Theory 2013
Exact Recoverability From Dense Corrupted Observations via ℓ1-Minimization · IEEE Trans. Inf. Theory 2013
Mathematical optimization › continuous optimization › convex optimization › norm optimization
l1 minimization
0.212013
Exact Recoverability From Dense Corrupted Observations via ℓ1-Minimization · IEEE Trans. Inf. Theory 2013
Mathematical optimization › statistical estimation › regression › sparse regression
lasso
0.212013
Robust Lasso With Missing and Grossly Corrupted Observations · IEEE Trans. Inf. Theory 2013
Approximation and online algorithms
robust recovery
0.212013
Exact Recoverability From Dense Corrupted Observations via ℓ1-Minimization · IEEE Trans. Inf. Theory 2013
Mathematical optimization
statistical estimation
0.212013
Robust Lasso With Missing and Grossly Corrupted Observations · IEEE Trans. Inf. Theory 2013
Information theory › signal processing › compressed sensing
support recovery
0.212013
Robust Lasso With Missing and Grossly Corrupted Observations · IEEE Trans. Inf. Theory 2013
Mathematical optimization › statistical estimation › regression
robust regression
0.112011
Robust Lasso with missing and grossly corrupted observations · NIPS 2011
Algorithms and data structures › matrix approximation
low-rank approximation
0.112009
A fast and efficient algorithm for low-rank approximation of a matrix · STOC 2009
Algorithms and data structures
numerical linear algebra
0.112009
A fast and efficient algorithm for low-rank approximation of a matrix · STOC 2009
Computational complexity › learning theory
sample complexity
0.012011
Robust Lasso with missing and grossly corrupted observations · NIPS 2011
Mathematical optimization
least squares
0.012009
A fast and efficient algorithm for low-rank approximation of a matrix · STOC 2009

Methods — techniques the papers use, named apart from their topics

autoencoder · 1.5TSMixer · 1.5transformer · 0.7patching · 0.7channel independence · 0.7l0-constrained optimization · 0.4cutting plane algorithm · 0.4restricted eigenvalue · 0.2l1 minimization · 0.2extended lasso · 0.2restricted eigenvalue analysis · 0.1gaussian design analysis · 0.1randomized sampling · 0.1preprocessing · 0.1orthonormal transformation · 0.1
YearPublicationVenuePosition
2026 Channel-Independence for Traffic Forecasting: A Cascaded Spatio-Temporal MLP Framework
abstract
The criticality of efficient traffic forecasting in Intelligent Transportation System (ITS) has garnered significant academic attention. This study addresses the prevalent issue of distribution shift in real-world datasets, which often degrades performance, and explores the effectiveness of the channel-independence (CI), a technique recently proposed to mitigate this issue. While Spatio-Temporal Graph Neural Networks (STGNNs) are noted for their flexibility to represent road structures, their designs typically lack the capability to integrate CI without disrupting the spatial relationships, potentially limiting the performance. We present a novel approach that successfully integrates CI into spatial-temporal forecasting by incorporating distinct temporal, spatial, and predefined graph structure information within each channel. Moreover, STGNNs frequently emphasize intricate designs, which result in increased computational demands while offering only marginal improvements in accuracy. This paper presents ST-MLP, a streamlined spatio-temporal model constructed exclusively from cascaded Multi-Layer Perceptron (MLP) modules and linear layers. Experimental results indicate that ST-MLP outperforms numerous existing STGNNs in both accuracy and computational efficiency. Our findings advocate for further investigation into more streamlined and effective neural network architectures within spatial-temporal forecasting research.
Zepu Wang, Yuqi Nie, Yang Liu 0246, John M. Mulvey, H. Vincent Poor, Azzedine Boukerche, Nam H. Nguyen, Peng Sun 0007
IEEE Trans. Intell. Transp. Syst.7
2025 VITRO: Vocabulary Inversion for Time-series Representation Optimization
abstract
Although LLMs have demonstrated remarkable capabilities in processing and generating textual data, their pretrained vocabularies are ill-suited for capturing the nuanced temporal dynamics and patterns inherent in time series. The discrete, symbolic nature of natural language tokens, which these vocabularies are designed to represent, does not align well with the continuous, numerical nature of time series data. To address this fundamental limitation, we propose VITRO. Our method adapts textual inversion optimization from the vision-language domain in order to learn a new time series per-dataset vocabulary that bridges the gap between the discrete, semantic nature of natural language and the continuous, numerical nature of time series data. We show that learnable time series-specific pseudo-word embeddings represent time series data better than existing general language model vocabularies, with VITRO-enhanced methods achieving state-of-the-art performance in long-term forecasting across most datasets.
Filippos Bellos, Nam H. Nguyen, Jason J. Corso
ICASSP2
2024 AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability Data
abstract
The efficiency of business processes relies on business key performance indicators (Biz-KPIs), that can be negatively impacted by IT failures. Business and IT Observability (BizITObs) data fuses both Biz-KPIs and IT event channels together as multivariate time series data. Forecasting Biz-KPIs in advance can enhance efficiency and revenue through proactive corrective measures. However, BizITObs data generally exhibit both useful and noisy inter-channel interactions between Biz-KPIs and IT events that need to be effectively decoupled. This leads to suboptimal forecasting performance when existing multivariate forecasting models are employed. To address this, we introduce AutoMixer, a time-series Foundation Model (FM) approach, grounded on the novel technique of channel-compressed pretrain and finetune workflows. AutoMixer leverages an AutoEncoder for channel-compressed pretraining and integrates it with the advanced TSMixer model for multivariate time series forecasting. This fusion greatly enhances the potency of TSMixer for accurate forecasts and also generalizes well across several downstream tasks. Through detailed experiments and dashboard analytics, we show AutoMixer's capability to consistently improve the Biz-KPI's forecasting accuracy (by 11-15%) which directly translates to actionable business insights.
Santosh Palaskar, Vijay Ekambaram, Arindam Jati, Neelamadhav Gantayat, Avirup Saha, Seema Nagar, Nam H. Nguyen, Pankaj Dayama 0001, Renuka Sindhgatta, Prateeti Mohapatra, Jayant Kalagnanam, Nandyala Hemachandra, Narayan Rangaraj
AAAI7
2023 A Time Series is Worth 64 Words: Long-term Forecasting with Transformers
Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam
ICLR2
2021 A Scale Invariant Measure of Flatness for Deep Network Minima
abstract
It has been empirically observed that the flatness of minima obtained from training deep networks seems to correlate with better generalization. However, for deep networks with positively homogeneous activations, most measures of flatness are not invariant to rescaling of the network parameters. This means that the measure of flatness can be made as small or as large as possible through rescaling, rendering the quantitative measures meaningless. In this paper we show that for deep networks with positively homogenous activations, these rescalings constitute equivalence relations, and that these equivalence relations induce a quotient manifold structure in the parameter space. Using an appropriate Riemannian metric, we propose a Hessian-based measure for flatness that is invariant to rescaling and perform simulations to empirically verify our claim. Finally we perform experiments to verify that our flatness measure correlates with generalization by using minibatch stochastic gradient descent with different batch sizes to find deep network minima with different generalization properties.
Akshay Rangamani, Nam H. Nguyen, Dzung T. Phan, Sang (Peter) Chin, Trac D. Tran
ICASSP2
2021 Quantum circuit representation of Bayesian networks
Sima Esfandiarpour Borujeni, Saideep Nannapaneni, Nam H. Nguyen, Elizabeth C. Behrman, James Edward Steck
Expert Syst. Appl.3
2020 Pruning Deep Neural Networks with $\ell_{0}$-constrained Optimization
abstract
Deep neural networks (DNNs) give state-of-the-art accuracy in many tasks, but they can require large amounts of memory storage, energy consumption, and long inference times. Modern DNNs can have hundreds of million parameters, which make it difficult for DNNs to be deployed in some applications with low-resource environments. Pruning redundant connections without sacrificing accuracy is one of popular approaches to overcome these limitations. We propose two l0-constrained optimization models for pruning deep neural networks layer-by-layer. The first model is devoted to a general activation function, while the second one is specifically for a ReLU. We introduce an efficient cutting plane algorithm to solve the latter to optimality. Our experiments show that the proposed approach achieves competitive compression rates over several state-of-the-art baseline methods.
Dzung T. Phan, Lam M. Nguyen, Nam H. Nguyen, Jayant Kalagnanam
ICDM3
2020 Benchmarking Neural Networks For Quantum Computations
abstract
The power of quantum computers is still somewhat speculative. Although they are certainly faster than classical ones at some tasks, the class of problems they can efficiently solve has not been mapped definitively onto known classical complexity theory. This means that we do not know for which calculations there will be a "quantum advantage," once an algorithm is found. One way to answer the question is to find those algorithms, but finding truly quantum algorithms turns out to be very difficult. In previous work, over the past three decades, we have pursued the idea of using techniques of machine learning to develop algorithms for quantum computing. Here, we compare the performance of standard real- and complex-valued classical neural networks with that of one of our models for a quantum neural network, on both classical problems and on an archetypal quantum problem: the computation of an entanglement witness. The quantum network is shown to need far fewer epochs and a much smaller network to achieve comparable or better results.
Nam H. Nguyen, Elizabeth C. Behrman, Mohamed A. Moustafa, James Edward Steck
IEEE Trans. Neural Networks Learn. Syst.1
2013 Exact Recoverability From Dense Corrupted Observations via ℓ1-Minimization
abstract
This paper confirms a surprising phenomenon first observed by Wright under a different setting: givenmhighly corrupted measurementsy=AΩ·x*+e*, whereAΩ·is a submatrix whose rows are selected uniformly at random from rows of an orthogonal matrixAande*is an unknown sparse error vector whose nonzero entries may be unbounded, we show that with high probability, ℓ1-minimization can recover the sparse signal of interestx*exactly from onlym=Cμ2k(logn)2, wherekis the number of nonzero components ofx*and μ =nmaxij Aij2, even if a significant fraction of the measurements are corrupted. We further guarantee that stable recovery is possible when measurements are polluted by both gross sparse and small dense errors:y=AΩ·x*+e*+ ν, where ν is the small dense noise with bounded energy. Numerous simulation results under various settings are also presented to verify the validity of the theory as well as to illustrate the promising potential of the proposed framework.
Nam H. Nguyen, Trac D. Tran
IEEE Trans. Inf. Theory1
2013 Robust Lasso With Missing and Grossly Corrupted Observations
abstract
This paper studies the problem of accurately recovering ak-sparse vector β*∈ \BBRpfrom highly corrupted linear measurementsy=Xβ*+e*+w, wheree*∈ \BBRnis a sparse error vector whose nonzero entries may be unbounded andwis a stochastic noise term. We propose a so-called extended Lasso optimization which takes into consideration sparse prior information of both β*ande*. Our first result shows that the extended Lasso can faithfully recover both the regression as well as the corruption vector. Our analysis relies on the notion of extended restricted eigenvalue for the design matrixX. Our second set of results applies to a general class of Gaussian design matrixXwith i.i.d. rowsN(0,Σ), for which we can establish a surprising result: the extended Lasso can recover exact signed supports of both β*ande*from only Ω(klogplogn) observations, even when a linear fraction of observations is grossly corrupted. Our analysis also shows that this amount of observations required to achieve exact signed support is indeed optimal.
Nam H. Nguyen, Trac D. Tran
IEEE Trans. Inf. Theory1
2011 Robust multi-sensor classification via joint sparse representation
Nam H. Nguyen, Nasser M. Nasrabadi, Trac D. Tran
FUSION1
2011 Robust Lasso with missing and grossly corrupted observations
abstract
This paper studies the problem of accurately recovering a sparse vector $\beta^{\star}$ from highly corrupted linear measurements $y = X \beta^{\star} + e^{\star} + w$ where $e^{\star}$ is a sparse error vector whose nonzero entries may be unbounded and $w$ is a bounded noise. We propose a so-called extended Lasso optimization which takes into consideration sparse prior information of both $\beta^{\star}$ and $e^{\star}$. Our first result shows that the extended Lasso can faithfully recover both the regression and the corruption vectors. Our analysis is relied on a notion of extended restricted eigenvalue for the design matrix $X$. Our second set of results applies to a general class of Gaussian design matrix $X$ with i.i.d rows $\oper N(0, \Sigma)$, for which we provide a surprising phenomenon: the extended Lasso can recover exact signed supports of both $\beta^{\star}$ and $e^{\star}$ from only $\Omega(k \log p \log n)$ observations, even the fraction of corruption is arbitrarily close to one. Our analysis also shows that this amount of observations required to achieve exact signed support is optimal.
Nam H. Nguyen, Nasser M. Nasrabadi, Trac D. Tran
NIPS1
2009 A fast and efficient heuristic nuclear-norm algorithm for affine rank minimization
abstract
The problem of affine rank minimization seeks to find the minimum rank matrix that satisfies a set of linear equality constraints. Generally, since affine rank minimization is NP-hard, a popular heuristic method is to minimize the nuclear norm that is a sum of singular values of the matrix variable. A recent intriguing paper shows that if the linear transform that defines the set of equality constraints is nearly isometrically distributed and the number of constraints is at least O(r(m + n) logmn), where r and m times n are the rank and size of the minimum rank matrix, minimizing the nuclear norm yields exactly the minimum rank matrix solution. Unfortunately, it takes a large amount of computational complexity and memory buffering to solve the nuclear norm minimization problem with known nearly isometric transforms. This paper presents a fast and efficient algorithm for nuclear norm minimization that employs structurally random matrices for its linear transform and a projected subgradient method that exploits the unique features of structurally random matrices to substantially speed up the optimization process. Theoretically, we show that nuclear norm minimization using structurally random linear constraints guarantees the minimum rank matrix solution if the number of linear constraints is at least O(r(m+n) log3mn). Extensive simulations verify that structurally random transforms still retain optimal performance while their implementation complexity is just a fraction of that of completely random transforms, making them promising candidates for large scale applications.
Thong T. Do, Yi Chen 0014, Nam H. Nguyen, Lu Gan 0002, Trac D. Tran
ICASSP3
2009 A fast and efficient algorithm for low-rank approximation of a matrix
abstract
The low-rank matrix approximation problem involves finding of a rank k version of a m x n matrix A, labeled Ak, such that Ak is as "close" as possible to the best SVD approximation version of A at the same rank level. Previous approaches approximate matrix A by non-uniformly adaptive sampling some columns (or rows) of A, hoping that this subset of columns contain enough information about A. The sub-matrix is then used for the approximation process. However, these approaches are often computationally intensive due to the complexity in the adaptive sampling. In this paper, we propose a fast and efficient algorithm which at first pre-processes matrix A in order to spread out information (energy) of every columns (or rows) of A, then randomly selects some of its columns (or rows). Finally, a rank-k approximation is generated from the row space of these selected sets. The preprocessing step is performed by uniformly randomizing signs of entries of A and transforming all columns of A by an orthonormal matrix F with existing fast implementation (e.g. Hadamard, FFT, DCT...). Our main contribution is summarized as follows. 1) We show that by uniformly selecting at random d rows of the preprocessed matrix with d = ( 1/η k max {log k, log 1/β} ), we guarantee the relative Frobenius norm error approximation: (1 + η) norm{A - Ak}F with probability at least 1 - 5β. 2) With d above, we establish a spectral norm error approximation: (2 + √2m/d) norm{A - Ak}2 with probability at least 1 - 2β. 3) The algorithm requires 2 passes over the data and runs in time (mn log d + (m+n) d2) which, as far as the best of our knowledge, is the fastest algorithm when the matrix A is dense. 4) As a bonus, applying this framework to the well-known least square approximation problem min norm{A x - b} where A ∈ Rm x r, we show that by randomly choosing d = (1/η γ r log m), the approximation solution is proportional to the optimal one with a factor of η and with extremely high probability, (1 - 6 m-γ), say.
Nam H. Nguyen, Thong T. Do, Trac D. Tran
STOC1