Hongda Mao

dblp:01/8845 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-0514-5028ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Head2Body: Body Pose Generation from Multi-Sensory Head-Mounted Inputs
Minh Tran 0004, Hongda Mao, Qingshuang Chen, Yelin Kim
ICCV2
2025 Efficient Visual Place Recognition Through Multimodal Semantic Knowledge Integration
Sitao Zhang, Hongda Mao, Qingshuang Chen, Yelin Kim
ICCV2
2023 Augmentation Robust Self-Supervised Learning for Human Activity Recognition
abstract
Human Activity Recognition (HAR) is widely applied on wearable devices in our daily lives. However, acquiring high-quality wearable sensor data set with ground-truths is challenging due to the high cost in collecting data and necessity of domain experts. In order to achieve generalization from limited data, we study augmentation-based Self-Supervised Learning (SSL) for data from wearable devices. However, there is an issue in one of the most popular SSL approaches, contrastive learning: it is sensitive to the choice of data augmentations. To resolve this, we first propose to combine contrastive learning with generative learning, which is robust to augmentations. Second, we propose an automatic augmentation policy search method to discover the most promising augmentation policy. We empirically verify our approaches on three public HAR datasets. Experimental results show that our proposed SSL approach is robust to augmentations, and delivers higher accuracy than contrastive learning. Additionally, with the searched augmentation policy we are able to further improve the accuracy of HAR task.
Yuhang Li 0001, Dae Lee, Dae Hoon Park, Hongda Mao, Huyen Do, Jonathan Chung 0001, Dinesh Nair
ICASSP5
2022 Dynamically Pruning Segformer for Efficient Semantic Segmentation
abstract
As one of the successful Transformer-based models in computer vision tasks, SegFormer demonstrates superior performance in semantic segmentation. Nevertheless, the high computational cost greatly challenges the deployment of SegFormer on edge devices. In this paper, we seek to design a lightweight SegFormer for efficient semantic segmentation. Based on the observation that neurons in SegFormer layers exhibit large variances across different images, we propose a dynamic gated linear layer, which prunes the most uninformative set of neurons based on the input instance. To improve the dynamically pruned SegFormer, we also introduce two-stage knowledge distillation to transfer the knowledge within the original teacher to the pruned student network. Experimental results show that our method can significantly reduce the computation overhead of SegFormer without an apparent performance drop. For instance, we can achieve 36.9% mIoU with only 3.3G FLOPs on ADE20K, saving more than 60% computation with the drop of only 0.5% in mIoU.
Haoli Bai, Hongda Mao, Dinesh Nair
ICASSP2
2021 Fusion of Embeddings Networks for Robust Combination of Text Dependent and Independent Speaker Recognition
abstract
By implicitly recognizing a user based on his/her speech input, speaker identification enables many downstream applications, such as personalized system behavior and expedited shopping checkouts.Based on whether the speech content is constrained or not, both text-dependent (TD) and text-independent (TI) speaker recognition models may be used.We wish to combine the advantages of both types of models through an ensemble system to make more reliable predictions.However, any such combined approach has to be robust to incomplete inputs, i.e., when either TD or TI input is missing.As a solution we propose a fusion of embeddings network (FOEnet) architecture, combining joint learning with neural attention.We compare FOEnet with four competitive baseline methods on a dataset of voice assistant inputs, and show that it achieves higher accuracy than the baseline and score fusion methods, especially in the presence of incomplete inputs.
Ruirui Li 0002, Chelsea J.-T. Ju, Zeya Chen, Hongda Mao, Oguz Elibol, Andreas Stolcke
Interspeech4
2020 Bridging Mixture Density Networks with Meta-Learning for Automatic Speaker Identification
abstract
Speaker identification answers the fundamental question "Who is speaking" The identification technology enables various downstream applications to provide a personalized experience. Both the prevalent i-vector based solutions and the state-of-the-art deep learning solutions usually treat all users equally, with no distinctions between new users and existing users, during the training process. We notice that a good many new users start with limited labeled training data, which often results in inferior predicting performance of recognizing users' voices. To alleviate the disadvantage caused by training data deficiency, we propose a Mixture Density Network- based Meta-Learning method (MDNML) for speaker identification. MDNML emphasizes the expeditious process of learning to recognize new users where each has only a few seconds of labeled data. We conduct experiments on the LibriSpeech dataset and compare MDNML with four state-of-the-art baseline methods. The results conclude that MDNML achieves higher accuracy in recognizing new users with limited labeled utterances than all baseline methods. Our proposed solution significantly expedites the learning by transferring the knowledge learned from the existing user base through gradient-based meta-learning. We consider our work to be a steppingstone for more sophisticated meta-learning frameworks for accelerating voice recognition. Furthermore, we discuss a strategy for enhancing the accuracy by incorporating the notion of household-based acoustic profiles with MDNML.
Ruirui Li 0002, Jyun-Yu Jiang, Xian Wu 0001, Hongda Mao, Chu-Cheng Hsieh, Wei Wang 0010
ICASSP4
2019 ThunderNet: A Turbo Unified Network for Real-Time Semantic Segmentation
abstract
Recent research in pixel-wise semantic segmentation has increasingly focused on the development of very complicated deep neural networks, which require a large amount of computational resources. The ability to perform dense predictions in real-time, therefore, becomes tantamount to achieving high accuracies. This real-time demand turns out to be fundamental particularly on the mobile platform and other GPU-powered embedded systems like NVIDIA Jetson TX series. In this paper, we present a fast and efficient lightweight network called Turbo Unified Network (ThunderNet). With a minimum backbone truncated from ResNet18, ThunderNet unifies the pyramid pooling module with our customized decoder. Our experimental results show that ThunderNet can achieve 64.0% mIoU on CityScapes, with real-time performance of 96.2 fps on a Titan XP GPU (512x1024), and 20.9 fps on Jetson TX2 (256x512).
Hongda Mao, Vassilis Athitsos
WACV2
2010 A convex neighbor-constrained active contour model for image segmentation
abstract
A large number of real images possess the property of intensity non-homogeneity, which hinders them from being segmented by many image segmentation approaches. Recently, region-based active contour models utilizing local information have been introduced to segment images with intensity non-homogeneity. However, all these models are not convex, thus a good initial guess is required, which limits their practical application. In this paper, we propose a convex neighbor-constrained active contour model to segment images with intensity non-homogeneity. With different shapes and sizes of the neighborhood for each point, our model can accurately capture the region information of a given image. Our model is convex, and therefore it is independent of the initial condition and allows for automatic segmentation. To minimize energy functional of the model, we choose the efficient and fast Split Bregman method. Experimental results on synthetic and real images demonstrate the superior performance of our model.
Hongda Mao, Huafeng Liu 0003
ICIP1