Duc Hoang

dblp:169/3005 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 84% 3D vision · 7% Multi-agent systems · 3%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Reconfigurable computing and FPGAs · 69% Performance modeling and evaluation · 31%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA architecture
1.012026
KANELÉ: Kolmogorov-Arnold Networks for Efficient LUT-based Evaluation · FPGA 2026
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
hardware-aware neural architecture search
0.812024
Zero-Shot Neural Architecture Search: Challenges, Solutions, and Opportunities · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.812024
Zero-Shot Neural Architecture Search: Challenges, Solutions, and Opportunities · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
training-free neural architecture search
0.812024
Zero-Shot Neural Architecture Search: Challenges, Solutions, and Opportunities · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Efficient and distributed learning
model compression
0.712023
Don't just prune by magnitude! Your mask topology is a secret weapon · NeurIPS 2023
Machine learning › Efficient and distributed learning › model compression
pruning
0.712023
Don't just prune by magnitude! Your mask topology is a secret weapon · NeurIPS 2023
Machine learning › Efficient and distributed learning › model compression › pruning › DNN pruning
pruning at initialization
0.712023
Don't just prune by magnitude! Your mask topology is a secret weapon · NeurIPS 2023
Machine learning › Efficient and distributed learning › model compression
sparse neural network
0.712023
Don't just prune by magnitude! Your mask topology is a secret weapon · NeurIPS 2023
Computer vision › 3D vision › pose estimation
3d hand pose estimation
0.412020
MM-Hand: 3D-Aware Multi-Modal Guided Hand Generation for 3D Hand Pose Synthesis · ACM Multimedia 2020
Performance modeling and evaluation
benchmarking
0.212024
Zero-Shot Neural Architecture Search: Challenges, Solutions, and Opportunities · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Performance modeling and evaluation › benchmarking › machine learning benchmarking
neural architecture search benchmarks
0.212024
Zero-Shot Neural Architecture Search: Challenges, Solutions, and Opportunities · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Knowledge, reasoning and agents › Multi-agent systems
graph connectivity
0.212023
Don't just prune by magnitude! Your mask topology is a secret weapon · NeurIPS 2023
Machine learning › Graph learning
spectral graph theory
0.212023
Don't just prune by magnitude! Your mask topology is a secret weapon · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

zero-shot proxy · 1.5accuracy prediction · 1.5kolmogorov-arnold network · 1.0weighted spectral gap · 0.7ramanujan graph · 0.7multimodal guidance · 0.4generative network · 0.4curriculum learning · 0.4
YearPublicationVenuePosition
2026 KANELÉ: Kolmogorov-Arnold Networks for Efficient LUT-based Evaluation
abstract
FPGA ’26, Seaside, CA, USA
Duc Hoang, Aarush Gupta, Philip C. Harris
FPGA1
2026 hls4ml: A Flexible, Open Source Platform for Deep Learning Acceleration on Reconfigurable Hardware
abstract
We present hls4ml , a free and open source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this article, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.
Jan-Frederik Schulte, Benjamin Ramhorst, Jovan Mitrevski, Nicolò Ghielmetti, Enrico Lupi, Dimitrios Danopoulos, Vladimir Loncar, Javier M. Duarte, David Burnette, Lauri Laatu, Stylianos Tzelepis, Konstantinos Axiotis, Quentin Berthet, Haoyan Wang, Suleyman Demirsoy, Marco Colombo, Thea Aarrestad, Sioni Summers, Maurizio Pierini, Giuseppe Di Guglielmo, Jennifer Ngadiuba, Javier Campos, Benjamin Hawks, Abhijith Gandrakota, Farah Fahim, George A. Constantinides, Zhiqiang Que, Wayne Luk, Alexander D. Tapper, Duc Hoang, Noah Paladino, Philip C. Harris, Bo-Cheng Lai, Manuel Valentin, Ryan Forelli, Seda Ogrenci Memik, Lino Gerlach, Rian Brooks Flynn, Mia Liu, Daniel Diaz 0003, Elham E Khoda, Melissa Quinnan, Russell Solares, Santosh Parajuli, Mark S. Neubauer, Christian Herwig, Ho Fung Tsoi, Dylan S. Rankin, Shih-Chieh Hsu, Scott Hauck
ACM Trans. Reconfigurable Technol. Syst.33
2024 Zero-Shot Neural Architecture Search: Challenges, Solutions, and Opportunities
abstract
Recently, zero-shot (or training-free) Neural Architecture Search (NAS) approaches have been proposed to liberate NAS from the expensive training process. The key idea behind zero-shot NAS approaches is to design proxies that can predict the accuracy of some given networks without training the network parameters. The proxies proposed so far are usually inspired by recent progress in theoretical understanding of deep learning and have shown great potential on several datasets and NAS benchmarks. This paper aims to comprehensively review and compare the state-of-the-art (SOTA) zero-shot NAS approaches, with an emphasis on their hardware awareness. To this end, we first review the mainstream zero-shot proxies and discuss their theoretical underpinnings. We then compare these zero-shot proxies through large-scale experiments and demonstrate their effectiveness in both hardware-aware and hardware-oblivious NAS scenarios. Finally, we point out several promising ideas to design better proxies.
Guihong Li, Duc Hoang, Kartikeya Bhardwaj, Ming Lin 0002, Zhangyang Wang, Radu Marculescu
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Don't just prune by magnitude! Your mask topology is a secret weapon
abstract
Recent years have witnessed significant progress in understanding the relationship between the connectivity of a deep network's architecture as a graph, and the network's performance. A few prior arts connected deep architectures to expander graphs or Ramanujan graphs, and particularly,[7] demonstrated the use of such graph connectivity measures with ranking and relative performance of various obtained sparse sub-networks (i.e. models with prune masks) without the need for training. However, no prior work explicitly explores the role of parameters in the graph's connectivity, making the graph-based understanding of prune masks and the magnitude/gradient-based pruning practice isolated from one another. This paper strives to fill in this gap, by analyzing the Weighted Spectral Gap of Ramanujan structures in sparse neural networks and investigates its correlation with final performance. We specifically examine the evolution of sparse structures under a popular dynamic sparse-to-sparse network training scheme, and intriguingly find that the generated random topologies inherently maximize Ramanujan graphs. We also identify a strong correlation between masks, performance, and the weighted spectral gap. Leveraging this observation, we propose to construct a new "full-spectrum coordinate'' aiming to comprehensively characterize a sparse neural network's promise. Concretely, it consists of the classical Ramanujan's gap (structure), our proposed weighted spectral gap (parameters), and the constituent nested regular graphs within. In this new coordinate system, a sparse subnetwork's L2-distance from its original initialization is found to have nearly linear correlated with its performance. Eventually, we apply this unified perspective to develop a new actionable pruning method, by sampling sparse masks to maximize the L2-coordinate distance. Our method can be augmented with the "pruning at initialization" (PaI) method, and significantly outperforms existing PaI methods. With only a few iterations of training (e.g 500 iterations), we can get LTH-comparable performance as that yielded via "pruning after training", significantly saving pre-training costs. Codes can be found at: https://github.com/VITA-Group/FullSpectrum-PAI.
Duc Hoang, Souvik Kundu 0009, Shiwei Liu 0003, Zhangyang Wang
NeurIPS1
2022 AutoMARS: Searching to Compress Multi-Modality Recommendation Systems
abstract
Web applications utilize Recommendation Systems (RS) to address the problem of consumer over-choices. Recent works have taken advantage of multi-modality or multi-view, input information (such as user interaction, images, texts, rating scores) to boost recommendation system performance compared with using single-modality information. However, the use of multi-modality input demands much higher computational cost and storage capacity. On the other hand, the real-world RS services usually have strict budgets on both time and space for a good customer experience. As a result, the model efficiency of multi-modality recommendation systems has gained increasing importance. While unfortunately, to the best of our knowledge, there is no existing study of a generic compression framework for multi-modality RS. In this paper, we investigate, for the first time, how to compress a multi-modality recommendation system with a fixed budget. Assuming that input information from different modalities are of unequal importance, a good compression algorithm should learn to automatically allocate different resource budgets to each input, based on their importance in maximally preserving recommendation efficacy. To this end, we leverage the tools of neural architecture search (NAS) and distillation and propose Auto Multi-modAlity Recommendation System (AutoMARS), a unified modality-aware model compression framework dedicated to multi-modality recommendation systems. We demonstrate the effectiveness and generality of AutoMARS by testing it on three different Amazon datasets of various sparsity. AutoMARS demonstrates superior multi-modality compression performance than previous state-of-the-art compression methods. For example on the Amazon Beauty dataset, we achieve on average a 20% higher accuracy over previous state-of-the-art methods, while enjoying 65% reduction over baselines. Codes are available at: https://github.com/VITA-Group/AutoMARS.
Duc Hoang, Haotao Wang, Handong Zhao, Ryan Rossi, Sungchul Kim, Kanak Mahadik, Zhangyang Wang
CIKM1
2020 MM-Hand: 3D-Aware Multi-Modal Guided Hand Generation for 3D Hand Pose Synthesis
abstract
Estimating the 3D hand pose from a monocular RGB image is important but challenging. A solution is training on large-scale RGB hand images with accurate 3D hand keypoint annotations. However, it is too expensive in practice. Instead, we develop a learning-based approach to synthesize realistic, diverse, and 3D pose-preserving hand images under the guidance of 3D pose information. We propose a 3D-aware multi-modal guided hand generative network (MM-Hand), together with a novel geometry-based curriculum learning strategy. Our extensive experimental results demonstrate that the 3D-annotated images generated by MM-Hand qualitatively and quantitatively outperform existing options. Moreover, the augmented data can consistently improve the quantitative performance of the state-of-the-art 3D hand pose estimators on two benchmark datasets. The code will be available at https://github.com/ScottHoang/mm-hand.
Zhenyu Wu 0002, Duc Hoang, Shih-Yao Lin 0001, Yusheng Xie, Liangjian Chen, Yen-Yu Lin, Zhangyang Wang, Wei Fan 0001
ACM Multimedia2
2019 Inferring Convolutional Neural Networks' Accuracies from Their Architectural Characterizations
abstract
The challenge of choosing an appropriate convolutional neural network (CNN) architecture for specific applications and different data sets is still poorly understood in the literature. This is problematic, since CNNs have shown strong promise for analyzing scientific data from many domains including particle imaging detectors. In this paper, we proposed a systematic language that is useful for comparison between different CNN's architectures before training time. This helps us predict whether a network can perform better than a certain threshold accuracy before training up to 70% accuracy using simple machine learning models. Additionally, we found a coefficient of determination of 0.966 for an Ordinary Least Squares model in a regression task to predict accuracy of a large population of networks.
Duc Hoang, Jesse Hamer, Gabriel N. Perdue, Steven R. Young, Jonathan A. Miller, Anushree Ghosh
ICMLA1
2018 Development of an ENVISAT Altimetry Processor Providing Sea Level Continuity Between Open Ocean and Arctic Leads
abstract
Over the Arctic regions, current conventional altimetry products suffer from a lack of coverage or from degraded performance due to the inadequacy of the standard processing applied in the ground segments. This paper presents a set of dedicated algorithms able to process consistently returns from open ocean and from sea-ice leads in the Arctic Ocean (detection of water surfaces and derivation of water levels using returns from these surfaces). This processing extends the area over which a precise sea level can be computed. In the frame of the European Space Agency Sea Level Climate Change Initiative (http://cci.esa.int), we have first developed a new surface identification method combining two complementary solutions, one using a multiple-criteria approach (in particular the backscattering coefficient and the peakiness coefficient of the waveforms) and one based on a supervised neural network approach. Then, a new physical model has been developed (modified from the Brown model to include anisotropy in the scattering from calm protected water surfaces) and has been implemented in a maximum likelihood estimation retracker. This allows us to process both sea-ice lead waveforms (characterized by their peaky shapes) and ocean waveforms (more diffuse returns), guaranteeing, by construction, continuity between open ocean and ice-covered regions. This new processing has been used to produce maps of Arctic sea level anomaly from 18-Hz ENVIronment SATellite/RA-2 data.
Jean-Christophe Poisson, Graham D. Quartly, Andrey A. Kurekin, Pierre Thibaut, Duc Hoang, Francesco Nencioli
IEEE Trans. Geosci. Remote. Sens.5
2015 SPARK 2014 and GNATprove - A competition report from builders of an industrial-strength verifying compiler
Duc Hoang, Yannick Moy, Angela Wallenburg, Roderick Chapman
Int. J. Softw. Tools Technol. Transf.1