Thong Le

dblp:223/4258 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 77% Optimization for machine learning · 23%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
adversarial attack
0.412019
Query-Efficient Hard-label Black-box Attack: An Optimization-based Approach · ICLR (Poster) 2019
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.412019
Query-Efficient Hard-label Black-box Attack: An Optimization-based Approach · ICLR (Poster) 2019
Machine learning › Trustworthy machine learning › robustness › adversarial attack
hard-label black-box attack
0.412019
Query-Efficient Hard-label Black-box Attack: An Optimization-based Approach · ICLR (Poster) 2019
Machine learning › Optimization for machine learning › black-box optimization
zeroth-order optimization
0.412019
Query-Efficient Hard-label Black-box Attack: An Optimization-based Approach · ICLR (Poster) 2019
Machine learning › Trustworthy machine learning › robustness
adversarial examples
0.112019
Query-Efficient Hard-label Black-box Attack: An Optimization-based Approach · ICLR (Poster) 2019

Methods — techniques the papers use, named apart from their topics

zeroth-order optimization · 0.4random gradient estimation · 0.4
YearPublicationVenuePosition
2021 Alignment Restricted Streaming Recurrent Neural Network Transducer
abstract
There is a growing interest in the speech community in developing Recurrent Neural Network Transducer (RNN-T) models for automatic speech recognition (ASR) applications. RNN-T is trained with a loss function that does not enforce temporal alignment of the training transcripts and audio. As a result, RNN-T models built with uni-directional long short term memory (LSTM) encoders tend to wait for longer spans of input audio, before streaming already decoded ASR tokens. In this work, we propose a modification to the RNN-T loss function and develop Alignment Restricted RNN-T (Ar-RNN-T) models, which utilize audio-text alignment in-formation to guide the loss computation. We compare the proposed method with existing works, such as monotonic RNN-T, on LibriSpeech and in-house datasets. We show that the Ar-RNN-T loss provides a refined control to navigate the trade-offs between the token emission delays and the Word Error Rate (WER). The Ar-RNN-T models also improve downstream applications such as the ASR End-pointing by guaranteeing token emissions within any given range of latency. Moreover, the Ar-RNN-T loss allows for bigger batch sizes and 4 times higher throughput for our LSTM model architecture, enabling faster training and convergence on GPUs.
Jay Mahadeokar, Yuan Shangguan, Gil Keren, Thong Le, Ching-Feng Yeh, Christian Fügen, Michael L. Seltzer
SLT6
2019 Query-Efficient Hard-label Black-box Attack: An Optimization-based Approach
Minhao Cheng, Thong Le, Huan Zhang 0001, Jinfeng Yi, Cho-Jui Hsieh
ICLR (Poster)2