Khanh Le

dblp:01/5906 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
abstract
RNN-Transducer (RNN-T) is a widely adopted architecture in speech recognition, integrating acoustic and language modeling in an end-to-end framework. However, the RNN-T predictor tends to over-rely on consecutive word dependencies in training data, leading to high deletion error rates, particularly with less common or out-of-domain phrases. Existing solutions, such as regularization and data augmentation, often compromise other aspects of performance. We propose SegAug, an alignment-based augmentation technique that generates contextually varied audio-text pairs with low sentence-level semantics. This method encourages the model to focus more on acoustic features while diversifying the learned textual patterns of its internal language model, thereby reducing deletion errors and enhancing overall performance. Evaluations on the LibriSpeech and Tedlium-v3 datasets demonstrate a relative WER reduction of up to 12.5% on small-scale and 6.9% on large-scale settings. Notably, most of the improvement stems from reduced deletion errors, with relative reductions of 45.4% and 18.5%, respectively. These results highlight SegAug’s effectiveness in improving RNN-T’s robustness, offering a promising solution for enhancing speech recognition performance across diverse and challenging scenarios.
Khanh Le, Tuan Vu Ho, Dung Tran, Duc Thanh Chau
ICASSP1
2025 ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
abstract
Deploying ASR models at an industrial scale poses significant challenges in hardware resource management, especially for long-form transcription tasks where audio may last for hours. Large Conformer models, despite their capabilities, are limited to processing only 15 minutes of audio on an 80GB GPU. Furthermore, variable input lengths worsen inefficiencies, as standard batching leads to excessive padding, increasing resource consumption and execution time. To address this, we introduce ChunkFormer, an efficient ASR model that uses chunk-wise processing with relative right context, enabling long audio transcriptions on low-memory GPUs. ChunkFormer handles up to 16 hours of audio on an 80GB GPU, 1.5x longer than the current state-of-the-art FastConformer, while also boosting long-form transcription performance with up to 7.7% absolute reduction on word error rate and maintaining accuracy on shorter tasks compared to Conformer. By eliminating the need for padding in standard batching, ChunkFormer’s masked batching technique reduces execution time and memory usage by more than 3x in batch processing, substantially reducing costs for a wide range of ASR systems, particularly regarding GPU resources for models serving in real-world applications.
Khanh Le, Tuan Vu Ho, Dung Tran, Duc Thanh Chau
ICASSP1
2025 Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization
Luong Ho, Khanh Le, Vinh Pham, Bao Nguyen, Tan Tran, Duc Chau
INTERSPEECH2
2024 Algorithmic Aspects of Left-Orderings of Solvable Baumslag-Solitar Groups via its Dynamical Realization
Meng-Che, Khanh Le, Dino Rossegger
CiE2
2024 Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
Khanh Le, Duc Chau
INTERSPEECH1
2024 VN-SLU: A Vietnamese Spoken Language Understanding Dataset
Tuyen Tran, Khanh Le, Ngoc Dang Nguyen, Huyen Ngo, Woomyoung Park, Thi Thu Trang Nguyen
INTERSPEECH2
2019 Exploiting Q-Learning in Extending the Network Lifetime of Wireless Sensor Networks with Holes
abstract
Geographic routing is one of the most popular routing protocols in wireless sensor networks (WSNs) due to its simplicity and efficiency. However, with the occurrence of holes, geographic routing incurs with the so-called local minimum problem that may lead to a long hole detour path as well as the traffic concentration around the hole boundary. In consequence, the network lifetime is shortened. In this paper, we aim at proposing a lightweight distributed geographic routing protocol, which can prolong the lifetime of WSNs under the hole occurrence. Our main idea is to exploiting Q-learning technique to estimate the distance from a node to the holes. The routing decision is then determined based on the residual energy of the nodes, their estimated distance to the holes, and their distance to the destination. The simulation experiments show that our protocol strongly outperforms state-of-the-art protocols in terms of the network lifetime, packet latency and energy consumption. Specifically, our proposed protocol extends the network lifetime by more than 12% compared to the existing protocols.
Khanh Le, Thanh-Hung Nguyen, Kien Nguyen 0002, Phi-Le Nguyen
ICPADS1
2018 Load balanced and constant stretch routing in the vicinity of holes in WSNs
abstract
Because of its simplicity and scalability, geographic routing is a popular approach in wireless sensor networks, which can achieve a near-optimal routing path in the networks without holes (i.e., regions without working sensors). With the occurrence of holes, however, geographic routing faces the problems of load imbalance and routing path enlargement. In the literature, several proposals have attempted to fix these issues, but the majority of them considers only the cases when both the source and the destination stay fairly far from the holes. Recently, a few work has been proposed to tackle the problem of routing in the vicinity of routing holes. However, none of them addresses the two problems (i.e., load imbalance and routing path enlargement) concurrently, and none of them can solve the problem of load imbalance thoroughly. In this paper, we introduce a novel approach in dealing with routing in the vicinity of holes, that is the first to target and solve both the load imbalance and path enlargement problems. The theoretical analysis proves that the routing path stretch of our proposed protocol can be controlled to be as small as 1 + ε (for any predefined ε> 0) and the simulation experiments show that our protocol strongly outperforms the existing protocols in terms of load balancing.
Phi-Le Nguyen, Yusheng Ji, Khanh Le, Thanh-Hung Nguyen
CCNC3
2017 Probabilistic modelling for congestion detection on wireless sensor networks
abstract
Recently, Wireless Sensor Networks (WSNs) attract many researches due to their real applications. WSN is actually a network whose main components are sensors and channels. Based on applications, these components can be worked independently or separately with each others to capture information, process and send it to sink. However, in congestion-based aspect, most researches are assumed that environmental working of components are perfect, i.e. they omit packet-loss aspect due to failed sensors or broken links. This causes a limitation to rationally represent a WSN. Thus, in this proposal, using the reliable probability property, we define a Discrete Time Stochastic Petri Net Model for congestion detection on WSN in order to represent all working scenarios for components on the one hand, and calculate the congestion probability in the network on the other hand. After that, we also present a new algorithm to analyse on that model. Our straight example through this paper emphasises the idea of our model.
Khanh Le, Giang V. Trinh, Thang H. Bui, Thanh Tho Quan
CoDIT1
2016 The Flat Finger: Exploring Area Touches on Smartwatches
abstract
Smartwatches are emerging device category that feature highly limited input and display surfaces. We explore how touch contact areas, such as lines generated by flat fingers, can be used to increase input expressivity in these diminutive systems in three ways. Firstly, we present four design themes that emerged from an ideation workshop in which five designers proposed concepts for smartwatch touch area interaction. Secondly, we describe a sensor unit and study that captured user performance with 31 area touches and contrasted this against standard targeting performance. Finally, we describe three demonstration applications that instantiate ideas from the workshop and deploy the most reliably and rapidly produced area touches. We report generally positive user reactions to these demonstrators: the area touch interactions were perceived as quick, convenient and easy to learn and remember. Together this work characterizes how designers can use area touches in watch UIs, which area touches are most appropriate and how users respond to this interaction style.
Ian Oakley, Carina Lindahl, Khanh Le, Doyoung Lee, Md. Rasel Islam
CHI3
2008 The WINLAB Network Centric Cognitive Radio Hardware Platform - WiNC2R
Zoran Miljanic, Ivan Seskar, Khanh Le, Dipankar Raychaudhuri
Mob. Networks Appl.3
2008 Cognitive radio technology: From distributed spectrum coordination to adaptive network collaboration
Dipankar Raychaudhuri, Xiangpeng Jing, Ivan Seskar, Khanh Le, Joseph B. Evans
Pervasive Mob. Comput.4