Truong Dinh Do

dblp:375/8381 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 81% Language models and text generation · 19%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference acceleration
1.922026
UniSpec: Training-Free Speculative Decoding for Robust LLM Acceleration Across Languages and Hardware · ACL (1) 2026
SPECTRA: Faster Large Language Model Inference with Optimized Internal and External Speculation · ACL (1) 2025
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
1.922026
UniSpec: Training-Free Speculative Decoding for Robust LLM Acceleration Across Languages and Hardware · ACL (1) 2026
SPECTRA: Faster Large Language Model Inference with Optimized Internal and External Speculation · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model inference
0.912025
SPECTRA: Faster Large Language Model Inference with Optimized Internal and External Speculation · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

tree expansion · 1.0draft-and-verify · 1.0confidence score estimation · 1.0internal speculation · 0.9external speculation · 0.9draft model · 0.9
YearPublicationVenuePosition
2026 UniSpec: Training-Free Speculative Decoding for Robust LLM Acceleration Across Languages and Hardware
abstract
Speculative decoding accelerates large language model (LLM) inference through a draftand-verify paradigm, yet existing methods face three key limitations: reliance on fixed draft templates that ignore device-specific verification costs, lack of mechanisms to assess draft token quality, and suboptimal tree expansion strategies.We introduce UNISPEC, a trainingfree, lossless speculative decoding framework that enables robust, plug-and-play LLM acceleration across diverse hardware configurations and languages.UNISPEC incorporates three novel components: (1) a device-aware calibration mechanism that determines the optimal draft size by measuring the acceptancetime trade-off on each target device; (2) a confidence score estimation module that assigns quality scores to n-grams based on the verifier's token probabilities, enabling selective retention of high-quality draft candidates; and (3) an improved tree expansion strategy that broadens first-level exploration and applies threshold-based filtering to prune lowconfidence nodes.To comprehensively evaluate multilingual performance, we create a comprehensive benchmark, covering seven languages across seven generation tasks.Experiments with various LLM architectures, hardware environments, and languages demonstrate that UNISPEC consistently outperforms existing training-free methods, achieving speedups of up to 2.6× while maintaining output quality identical to standard autoregressive decoding.Our code and benchmark are publicly available.
Truong Dinh Do, Nguyen-Khang Le, Minh Le Nguyen 0001
ACL (1)1
2025 SPECTRA: Faster Large Language Model Inference with Optimized Internal and External Speculation
abstract
Inference with modern Large Language Models (LLMs) is both computationally expensive and time-consuming.Speculative decoding has emerged as a promising solution, but existing approaches face key limitations: training-based methods require a draft model that is challenging to obtain and lacks generalizability, while training-free methods offer limited speedup gains.In this work, we present SPECTRA, a novel framework for accelerating LLM inference without the need for additional training or modification to the original LLM.SPECTRA introduces two new techniques for efficiently utilizing internal and external speculation, each outperforming corresponding state-of-the-art (SOTA) methods independently.When combined, these techniques achieve up to a 4.08x speedup across various benchmarks and LLM architectures, significantly surpassing existing training-free approaches.The implementation of SPECTRA is publicly available.
Nguyen-Khang Le, Truong Dinh Do, Minh Le Nguyen 0001
ACL (1)2
2025 Improving hierarchical semantic parsing with LLMs: Demonstration selection and chain-of-thought prompting via semantic fragment decoding
Phuong Minh Nguyen 0001, Truong Dinh Do, Minh Le Nguyen 0001
Knowl. Based Syst.2
2024 ZeLa: Advancing Zero-Shot Multilingual Semantic Parsing with Large Language Models and Chain-of-Thought Strategies
abstract
In recent years, there have been significant advancements in semantic parsing tasks, thanks to the introduction of pre-trained language models. However, a substantial gap persists between English and other languages due to the scarcity of annotated data. One promising strategy to bridge this gap involves augmenting multilingual datasets using labeled English data and subsequently leveraging this augmented dataset for training semantic parsers (known as zero-shot multilingual semantic parsing). In our study, we propose a novel framework to effectively perform zero-shot multilingual semantic parsing under the support of large language models (LLMs). Given data annotated pairs (sentence, semantic representation) in English, our proposed framework automatically augments data in other languages via multilingual chain-of-thought (CoT) prompting techniques that progressively construct the semantic form in these languages. By breaking down the entire semantic representation into sub-semantic fragments, our CoT prompting technique simplifies the intricate semantic structure at each step, thereby facilitating the LLMs in generating accurate outputs more efficiently. Notably, this entire augmentation process is achieved without the need for any demonstration samples in the target languages (zero-shot learning). In our experiments, we demonstrate the effectiveness of our method by evaluating it on two well-known multilingual semantic parsing datasets: MTOP and MASSIVE.
Truong Dinh Do, Phuong Minh Nguyen 0001
LREC/COLING1