Toan H. Vu

dblp:188/1190 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
2since 2021 · last 2023
0000-0001-8775-120XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Deep learning architectures and training · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Human-computer interaction and pervasive computing
1 paper
Ubiquitous computing and smart environments · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
convolutional neural network
0.412020
Learning to Remember Beauty Products · ACM Multimedia 2020
Machine learning › Deep learning architectures and training
memory-augmented neural networks
0.412020
Learning to Remember Beauty Products · ACM Multimedia 2020
Machine learning › Deep learning architectures and training
multi-scale feature fusion
0.412020
Learning to Remember Beauty Products · ACM Multimedia 2020
Information retrieval › image retrieval › product image retrieval
beauty product image retrieval
0.412020
Learning to Remember Beauty Products · ACM Multimedia 2020
Information retrieval
image retrieval
0.412020
Learning to Remember Beauty Products · ACM Multimedia 2020
Machine learning › Deep learning architectures and training
recurrent neural network
0.212016
Transportation Mode Detection on Mobile Devices Using Recurrent Nets · ACM Multimedia 2016
Ubiquitous computing and smart environments
mobile sensing
0.212016
Transportation Mode Detection on Mobile Devices Using Recurrent Nets · ACM Multimedia 2016
Ubiquitous computing and smart environments › context recognition › activity recognition
transportation mode detection
0.212016
Transportation Mode Detection on Mobile Devices Using Recurrent Nets · ACM Multimedia 2016

Methods — techniques the papers use, named apart from their topics

distance loss · 0.9data augmentation · 0.9attention mechanism · 0.9recurrent neural network · 0.5accelerometer signal processing · 0.5
YearPublicationVenuePosition
2023 EMIX: A Data Augmentation Method for Speech Emotion Recognition
abstract
In the last few years, many deep learning (DL) models have been developed to improve the accuracy of speech emotion recognition (SER). However, as SER datasets are generally small and insufficient due to their difficult and expensive collection, the DL models are prone to overfitting, so their performance is limited. In this paper, we introduce a novel data augmentation (DA) method for the SER problem, namely EMix, which is simple but effective. The method creates new data by mixing pairs of selective samples from the original data. The generated mixtures will be noisier or less ambiguous than their constructive ones. To verify the effectiveness of the proposed DA, we develop a transformer-based network for the SER task, and experiment with the two public datasets including IEMOCAP and Crema-D. The experimental results demonstrate the superiority of EMix over other DA methods. In comparison with state-of-the-art methods, our approach shows competitive performance.
An Dang, Toan H. Vu, Le Dinh Nguyen, Jia-Ching Wang
ICASSP2
2023 Anti-aliasing convolution neural network of finger vein recognition for virtual reality (VR) human-robot equipment of metaverse
Nghi C. Tran, Toan H. Vu, Tzu-Chiang Tai, Jia-Ching Wang
J. Supercomput.3
2020 Encoder-Recurrent Decoder Network for Single Image Dehazing
abstract
This paper develops a deep learning model, called Encoder-Recurrent Decoder Network (ERDN), which recovers the clear image from a degrade hazy image without using the atmospheric scattering model. The proposed model consists of two key components- an encoder and a decoder. The encoder is constructed by a residual efficient spatial pyramid (rESP) module such that it can effectively process hazy images at any resolution to extract relevant features at multiple contextual levels. The decoder has a recurrent module which sequentially aggregates encoded features from high levels to low levels to generate haze-free images. The network is trained end-to-end given pairs of hazy-clear images. Experimental results on the RESIDE-Standard dataset demonstrate that the proposed model achieves a competitive dehazing performance compared to the state-of-the-art methods in term of PSNR and SSIM.
An Dang, Toan H. Vu, Jia-Ching Wang
ICASSP2
2020 Learning to Remember Beauty Products
abstract
This paper develops a deep learning model for the beauty product image retrieval problem. The proposed model has two main components- an encoder and a memory. The encoder extracts and aggregates features from a deep convolutional neural network at multiple scales to get feature embeddings. With the use of an attention mechanism and a data augmentation method, it learns to focus on foreground objects and neglect background on images, so can it extract more relevant features. The memory consists of representative states of all database images as its stacks, and it can be updated during training process. Based on the memory, we introduce a distance loss to regularize embedding vectors from the encoder to be more discriminative. Our model is fully end-to-end, requires no manual feature aggregation and post-processing. Experimental results on the Perfect-500K dataset demonstrate the effectiveness of the proposed model with a significant retrieval accuracy.
Toan H. Vu, An Dang, Jia-Ching Wang
ACM Multimedia1
2016 Transportation Mode Detection on Mobile Devices Using Recurrent Nets
abstract
We present an approach to the use of Recurrent Neural Networks (RNN) for transportation mode detection (TMD) on mobile devices. The proposed model, called Control Gate-based Recurrent Neural Network (CGRNN), is an end-to-end model that works directly with raw signals from an embedded accelerometer. As mobile devices have limited computational resources, we evaluate the model in terms of accuracy, computational cost, and memory usage. Experiments on the HTC transportation mode dataset demonstrate that our proposed model not only exhibits remarkable accuracy, but also is efficient with low resource consumption.
Toan H. Vu, Le Dung, Jia-Ching Wang
ACM Multimedia1