Insoo Oh

dblp:211/1247 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-3938-4932ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021
YearPublicationVenuePosition
2026 MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation
abstract
In this work, we present MUSE (Model-based Uncertainty-aware Similarity Estimation), a training-free framework for model-based zero-shot 2D object detection and segmentation. First, MUSE incorporates 2D multi-view templates from 3D unseen objects and 2D object proposals from the input query image, respectively. In the embedding stage, we propose a new feature embedding scheme which integrates class and patch embeddings. Specifically, the patch embeddings are normalized using the generalized mean pooling (GeM). In the matching stage, a joint similarity score is introduced, which integrates an absolute score and a relative score. Finally, we update the similarity score using an uncertainty-aware object prior. MUSE achieves state-of-the-art performance on the BOP Challenge 2025, ranking first in the Classic Core, H3, and Industrial tracks—without any additional training or fine-tuning. Therefore, we believe that MUSE is a promising framework for zero-shot 2D object detection and segmentation.
Sungmin Cho, Sungbum Park, Insoo Oh
WACV3
2024 X-Singer: Code-Mixed Singing Voice Synthesis via Cross-Lingual Learning
Ji-Sang Hwang, HyeongRae Noh, Yoonseok Hong, Insoo Oh
INTERSPEECH4
2023 RWEN-TTS: Relation-Aware Word Encoding Network for Natural Text-to-Speech Synthesis
abstract
With the advent of deep learning, a huge number of text-to-speech (TTS) models which produce human-like speech have emerged. Recently, by introducing syntactic and semantic information w.r.t the input text, various approaches have been proposed to enrich the naturalness and expressiveness of TTS models. Although these strategies showed impressive results, they still have some limitations in utilizing language information. First, most approaches only use graph networks to utilize syntactic and semantic information without considering linguistic features. Second, most previous works do not explicitly consider adjacent words when encoding syntactic and semantic information, even though it is obvious that adjacent words are usually meaningful when encoding the current word. To address these issues, we propose Relation-aware Word Encoding Network (RWEN), which effectively allows syntactic and semantic information based on two modules (i.e., Semantic-level Relation Encoding and Adjacent Word Relation Encoding). Experimental results show substantial improvements compared to previous works.
Shinhyeok Oh, HyeongRae Noh, Yoonseok Hong, Insoo Oh
AAAI4
2022 UCSM-DNN: User and Card Style Modeling with Deep Neural Networks for Personalized Game AI
abstract
This paper tries to resolve long waiting time to find a matching person in player versus player mode of online sports games, such as baseball, soccer and basketball. In player versus player mode, game playing AI which is instead of player needs to be not just smart as human but also show variety to improve user experience against AI. Therefore a need to design game playing AI agents with diverse personalized styles rises. To this end, we propose a personalized game AI which encodes user style vectors and card style vectors with a general DNN, named UCSM-DNN. Extensive experiments show that UCSM-DNN shows improved performance in terms of personalized styles, which enrich user experiences. UCSM-DNN has already been integrated into popular mobile baseball game: MaguMagu 2021 as personalized game AI.
Daegeun Choe, Youngbak Jo, Shindong Kang, Shounan An, Insoo Oh
AAAI5
2022 MONICA2: Mobile Neural Voice Command Assistants towards Smaller and Smarter
abstract
In this paper, we propose on-device voice command assistants for mobile games to increase user experiences even in hands-busy situations such as driving and cooking. Since most of the current mobile games cost large memory (e.g. more than 1GB memory), so it is necessary to reduce memory usage further to integrate voice commands systems on mobile clients. Therefore a need to design an on-device automatic speech recognition system that costs minimal memory and CPU resources rises. To this end, we apply cross layer parameter sharing to Conformer, named MONICA2 which results in lower memory usage for on-device speech recognition. MONICA2 reduces the number of parameters of deep neural network by 58%, with minimal recognition accuracy degradation measured in word error rate on Librispeech benchmark. As an on-device voice command user interface, MONICA2 costs only 12.8MB mobile memory and the average inference time for 3-seconds voice command is about 30ms, which is profiled in Samsung Galaxy S9. As far as we know, MONICA2 is the most memory efficient yet accurate on-device speech recognition which could be applied to various applications such as mobile games, IoT devices, etc.
Yoonseok Hong, Shounan An, Sunwoo Im, Jaegeon Jo, Insoo Oh
AAAI5
2019 Robust Keyword Spotting via Recycle-Pooling for Mobile Game
Shounan An, Myungwoo Lee, Insoo Oh
INTERSPEECH6