VLDB 2026 Research / reviewers in the wild / expert
Sungho Shin
dblp:131/9168
· DBLP profile ↗
21ranked-venue papers
8as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High-Quality Unknown Object Instance Segmentation via Quadruple Boundary Error RefinementabstractAccurate and efficient segmentation of unknown objects in unstructured environments is essential for robotic manipulation. Unknown Object Instance Segmentation (UOIS), which aims to identify all objects in unknown categories and backgrounds, has become a key capability for various robotic tasks. However, existing methods struggle with over-segmentation and under-segmentation, leading to failures in manipulation tasks such as grasping. To address these challenges, we propose QuBER (Quadruple Boundary Error Refinement), a novel error-informed refinement approach for high-quality UOIS. QuBER first estimates quadruple boundary errors-true positive, true negative, false positive, and false negative pixels-at the instance boundaries of the initial segmentation. It then refines the segmentation using an error-guided fusion mechanism, effectively correcting both fine-grained and instance-level segmentation errors. Extensive evaluations on three public benchmarks demonstrate that QuBER outperforms state-of-the-art methods and consistently improves various UOIS methods while maintaining a fast inference time of less than 0.1 seconds. Furthermore, we show that QuBER improves the success rate of grasping target objects in cluttered environments. Code and supplementary materials are available at https://sites.google.com/view/uois-quber. Seunghyeok Back, Sangbeom Lee, Kangmin Kim, Joosoon Lee, Sungho Shin, Jemo Maeng, Kyoobin Lee |
ICRA | 5 |
| 2024 | Domain-Specific Block Selection and Paired-View Pseudo-Labeling for Online Test-Time AdaptationabstractTest-time adaptation (TTA) aims to adapt a pre-trained model to a new test domain without access to source data after deployment. Existing approaches typically rely on self-training with pseudo-labels since ground-truth cannot be obtained from test data. Although the quality of pseudo labels is important for stable and accurate long-term adaptation, it has not been previously addressed. In this work, we propose DPLOT, a simple yet effective TTA framework that consists of two components: (1) domain-specific block selection and (2) pseudo-label generation using paired-view images. Specifically, we select blocks that involve domain-specific feature extraction and train these blocks by entropy minimization. After blocks are adjusted for current test domain, we generate pseudo-labels by averaging given test images and corresponding flipped counterparts. By simply using flip augmentation, we prevent a decrease in the quality of the pseudo-labels, which can be caused by the domain gap resulting from strong augmentation. Our experimental results demonstrate that DPLOT outperforms previous TTA methods in CIFAR10-C, CIFAR100-C, and ImageNet-C benchmarks, reducing error by up to 5.4%, 9.1%, and 2.9%, respectively. Also, we provide an extensive analysis to demonstrate effectiveness of our framework. Code is available at https://github.com/gist-ailab/domain-specific-block-selection-and-paired-view-pseudo-labeling-for-online-TTA. Yeonguk Yu, Sungho Shin, Seunghyeok Back, Minhwan Ko, Sangjun Noh, Kyoobin Lee |
CVPR | 2 |
| 2024 | Curriculum Fine-tuning of Vision Foundation Model for Medical Image Classification Under Label NoiseabstractDeep neural networks have demonstrated remarkable performance in various vision tasks, but their success heavily depends on the quality of the training data. Noisy labels are a critical issue in medical datasets and can significantly degrade model performance. Previous clean sample selection methods have not utilized the well pre-trained features of vision foundation models (VFMs) and assumed that training begins from scratch. In this paper, we propose CUFIT, a curriculum fine-tuning paradigm of VFMs for medical image classification under label noise. Our method is motivated by the fact that linear probing of VFMs is relatively unaffected by noisy samples, as it does not update the feature extractor of the VFM, thus robustly classifying the training samples. Subsequently, curriculum fine-tuning of two adapters is conducted, starting with clean sample selection from the linear probing phase. Our experimental results demonstrate that CUFIT outperforms previous methods across various medical image benchmarks. Specifically, our method surpasses previous baselines by 5.0\%, 2.1\%, 4.6\%, and 5.8\% at a 40\% noise rate on the HAM10000, APTOS-2019, BloodMnist, and OrgancMnist datasets, respectively. Furthermore, we provide extensive analyses to demonstrate the impact of our method on noisy label detection. For instance, our method shows higher label precision and recall compared to previous approaches. Our work highlights the potential of leveraging VFMs in medical image classification under challenging conditions of noisy labels. Yeonguk Yu, Minhwan Ko, Sungho Shin, Kangmin Kim, Kyoobin Lee |
NeurIPS | 3 |
| 2024 | Exploring using jigsaw puzzles for out-of-distribution detectionabstractOut-of-distribution (OOD) detection involves binary classification whether the given data is from outside the training data or not. Previous studies proposed outlier exposure (OE) that trains the model on an outlier dataset designed to represent potential future OOD data, thereby enhancing OOD detection performance. However, obtaining an outlier dataset representing all possible future OOD data can be challenging, and such dataset may be unavailable in some cases. This study proposes a novel approach to expose the model to jigsaw puzzles generated from training images as the outlier data. Specifically, the model is trained to have a low LogitNorm for given jigsaw puzzles. We argue that jigsaw puzzles can effectively represent future OOD data because they contain similar background information as the in-distribution data but with their semantic information destroyed. Our experimental results demonstrate that our approach outperforms previous competitive OOD detection methods and effectively detects semantically shifted OOD examples. Our code is available at https://github.com/gist-ailab/jigsaw-training-OOD. Yeonguk Yu, Sungho Shin, Minhwan Ko, Kyoobin Lee |
Comput. Vis. Image Underst. | 2 |
| 2023 | Block Selection Method for Using Feature Norm in Out-of-Distribution DetectionabstractDetecting out-of-distribution (OOD) inputs during the inference stage is crucial for deploying neural networks in the real world. Previous methods typically relied on the highly activated feature map outputted by the network. In this study, we revealed that the norm of the feature map obtained from a block other than the last block can serve as a better indicator for OOD detection. To leverage this insight, we propose a simple framework that comprises two metrics: FeatureNorm, which computes the norm of the feature map, and NormRatio, which calculates the ratio of FeatureNorm for ID and OOD samples to evaluate the OOD detection performance of each block. To identify the block that provides the largest difference between FeatureNorm of ID and FeatureNorm of OOD, we create jigsaw puzzles as pseudo OOD from ID training samples and compute NormRatio, selecting the block with the highest value. After identifying the suitable block, OOD detection using FeatureNorm outperforms other methods by reducing FPR95 by up to 52.77% on CIFAR10 benchmark and up to 48.53% on ImageNet benchmark. We demonstrate that our framework can generalize to various architectures and highlight the significance of block selection, which can also improve previous OOD detection methods. Our code is available at https://github.com/gistailab/block-selection-for-OOD-detection. Yeonguk Yu, Sungho Shin, Seongju Lee, Changhyun Jun 0002, Kyoobin Lee |
CVPR | 2 |
| 2023 | Probability propagation for faster and efficient point cloud segmentation using a neural networkabstractNeural networks (NN) have shown promising performance in point cloud segmentation (PCS). However, the measured points are too numerous to be used as model input at once. It results in a long inference time and high computational cost due to iterative sampling and inference. This study proposes Probability Propagation (PP) as a stochastic upsampling method. PP propagates the predicted probability of a sampled part of a point cloud into the other unpredicted points by considering proximity. By replacing the iterative inference of NN with PP, large point clouds can be dealt with quickly and efficiently. We investigated the effectiveness of PP using the ShapeNet benchmark on various settings: sampling methods (random, farthest point, and Poisson disk sampling) with sampling ratios (5%, 10%, 20%, 39%, and 78%) for NN and the stochastic mapping conditions (uniform, linear, cosine, Gaussian, and exponential distributions) for PP. Using NN with PP achieved higher performance and faster inference speed than when using NN alone. For the farthest point sampling method of 5% sampling ratio, NN+PP improved the instance mIoU by 2.457%p with 102 times faster speed compared to that when using NN alone. The result indicates that PP can significantly contribute to the improvement of performance and efficiency in PCS when used in edge AI systems. Hogeon Seo, Sangjun Noh, Sungho Shin, Kyoobin Lee |
Pattern Recognit. Lett. | 3 |
| 2022 | Teaching Where to Look: Attention Similarity Knowledge Distillation for Low Resolution Face Recognition
Sungho Shin, Joosoon Lee, Junseok Lee 0003, Yeonguk Yu, Kyoobin Lee |
ECCV (12) | 1 |
| 2022 | LightTrader : World's first AI-enabled High-Frequency Trading Solution with 16 TFLOPS / 64 TOPS Deep Learning Inference AcceleratorsabstractWe present the world’s first AI-enabled high-frequency trading (HFT) system, LightTrader , which integrates the custom AI accelerators and the FPGA-based conventional HFT pipeline for the low-latency-high-throughput trading solutions with a reduced query miss rate. For better utilization, adaptive job scheduling methods are also proposed to further improve the performance, where layer-wise workload scaling and dynamic voltage-frequency scaling (DVFS) techniques progressively adjust the workloads of AI accelerators, in conjunction with the architecture support. LightTrader integrating TSMC 7nm tape-out accelerators solely achieves 6x speed-up of DNN processing and 30-50x reduction of query miss rate without the scheduling method while the scheduling scheme further improves the energy efficiency by 25% and reduces the query miss rate by 2.4x . Hyunsung Kim 0003, Sungyeob Yoo, Jaewan Bae, Kyeongryeol Bong, Yoonho Boo, Karim Charfi, Hyo-Eun Kim, Hyun Suk Kim, Jinseok Kim 0006, Byungjae Lee, Myeongbo Shim, Sungho Shin, Jeong Seok Woo, Joo-Young Kim 0001, Sunghyun Park 0006, Jinwook Oh |
HCS | 13 |
| 2021 | Stochastic Precision Ensemble: Self-Knowledge Distillation for Quantized Deep Neural NetworksabstractThe quantization of deep neural networks (QDNNs) has been actively studied for deployment in edge devices. Recent studies employ the knowledge distillation (KD) method to improve the performance of quantized networks. In this study, we propose stochastic precision ensemble training for QDNNs (SPEQ). SPEQ is a knowledge distillation training scheme; however, the teacher is formed by sharing the model parameters of the student network. We obtain the soft labels of the teacher by randomly changing the bit precision of the activation stochastically at each layer of the forward-pass computation. The student model is trained with these soft labels to reduce the activation quantization noise. The cosine similarity loss is employed, instead of the KL-divergence, for KD training. As the teacher model changes continuously by random bit-precision assignment, it exploits the effect of stochastic ensemble KD. SPEQ outperforms the existing quantization training methods in various tasks, such as image classification, question-answering, and transfer learning without the need for cumbersome teacher networks. Yoonho Boo, Sungho Shin, Jungwook Choi, Wonyong Sung |
AAAI | 2 |
| 2021 | SQWA: Stochastic Quantized Weight Averaging For Improving The Generalization Capability Of Low-Precision Deep Neural NetworksabstractLow-precision deep neural networks (DNNs) are very needed for efficient implementations, but severe quantization of weights often sacrifices the generalization capability and lowers the test accuracy. We present a new quantized neural network optimization approach, stochastic quantized weight averaging (SQWA), to design low-precision DNNs with good generalization capability using model averaging. The proposed approach includes (1) floating-point model training, (2) direct quantization of weights, (3) capturing multiple low precision models during retraining with cyclical learning rates, (4) averaging the captured models, and (5) re-quantizing the averaged model and fine-tuning it with low-learning rates. With SQWA training, we could develop the best performing QDNNs for image classification on ImageNet datasets and also for semantic segmentation on Pascal VOC 2012 dataset. Sungho Shin, Yoonho Boo, Wonyong Sung |
ICASSP | 1 |
| 2020 | HLHLp: Quantized Neural Networks Training for Reaching Flat Minima in Loss SurfaceabstractQuantization of deep neural networks is extremely essential for efficient implementations. Low-precision networks are typically designed to represent original floating-point counterparts with high fidelity, and several elaborate quantization algorithms have been developed. We propose a novel training scheme for quantized neural networks to reach flat minima in the loss surface with the aid of quantization noise. The proposed training scheme employs high-low-high-low precision in an alternating manner for network training. The learning rate is also abruptly changed at each stage for coarse- or fine-tuning. With the proposed training technique, we show quite good performance improvements for convolutional neural networks when compared to the previous fine-tuning based quantization scheme. We achieve the state-of-the-art results for recurrent neural network based language modeling with 2-bit weight and activation. Sungho Shin, Jinhwan Park, Yoonho Boo, Wonyong Sung |
AAAI | 1 |
| 2020 | Automatic Detection and Identification of Fasteners with Simple Visual Calibration using Synthetic DataabstractIn this paper, we present a deep learning-based approach to detect and identify multiple fasteners from various camera poses. To distinguish fasteners of similar size and shape from each other, we propose a part identifier network and simple visual calibration method using a reference image. Though the camera poses changes, the model can infer the actual scale of detected parts by just capturing a reference object at once. Also, we present a synthetic data generation pipeline that adopts domain randomization and can automatically generate a training set for various fastener identification. In the experiment, we evaluated the real-world performance of the fully synthetically trained model and showed that it could be directly applied to real-world part identification. This indicates that our approach has the potential to accelerate the model retraining procedure for various part identification tasks since data acquisition requires almost no cost. Sangjun Noh, Seunghyeok Back, Raeyoung Kang, Sungho Shin, Kyoobin Lee |
ETFA | 4 |
| 2019 | Memorization Capacity of Deep Neural Networks under Parameter QuantizationabstractMost deep neural networks (DNNs) require complex models to achieve high performance. Parameter quantization is widely used for reducing the implementation complexities. Previous studies on quantization were mostly based on extensive simulation using training data on a specific model. We choose a different approach and attempt to measure the per-parameter capacity of DNN models and interpret the results to obtain insights on optimum quantization of parameters. This research uses artificially generated data and generic forms of fully connected DNNs, convolutional neural networks, and recurrent neural networks. We conduct memorization and classification tests to study the effects of the number and precision of the parameters on the performance. The model and the per-parameter capacities are assessed by measuring the mutual information between the input and the classified output. To get insight for parameter quantization when performing real tasks, the training and test performances are compared. Yoonho Boo, Sungho Shin, Wonyong Sung |
ICASSP | 2 |
| 2019 | Workload-aware Automatic Parallelization for Multi-GPU DNN TrainingabstractDeep neural networks (DNNs) have emerged as successful solutions for variety of artificial intelligence applications, but their very large and deep models impose high computational requirements during training. Multi-GPU parallelization is a popular option to accelerate demanding computations in DNN training, but most state-of-the-art multi-GPU deep learning frameworks not only require users to have an in-depth understanding of the implementation of the frameworks themselves, but also apply parallelization in a straight-forward way without optimizing GPU utilization. In this work, we propose a workload-aware auto-parallelization framework (WAP) for DNN training, where the work is automatically distributed to multiple GPUs based on the workload characteristics. We evaluate WAP using TensorFlow with popular DNN benchmarks (AlexNet and VGG-16), and show competitive training throughput compared with the state-of-the-art frameworks, and also demonstrate that WAP automatically optimizes GPU assignment based on the workload's compute requirements, thereby improving energy efficiency. Sungho Shin, Youngmin Jo, Jungwook Choi, Swagath Venkataramani, Vijayalakshmi Srinivasan, Wonyong Sung |
ICASSP | 1 |
| 2019 | Scalable nonlinear programming framework for parameter estimation in dynamic biological system modelsabstractWe present a nonlinear programming (NLP) framework for the scalable solution of parameter estimation problems that arise in dynamic modeling of biological systems. Such problems are computationally challenging because they often involve highly nonlinear and stiff differential equations as well as many experimental data sets and parameters. The proposed framework uses cutting-edge modeling and solution tools which are computationally efficient, robust, and easy-to-use. Specifically, our framework uses a time discretization approach that: i) avoids repetitive simulations of the dynamic model, ii) enables fully algebraic model implementations and computation of derivatives, and iii) enables the use of computationally efficient nonlinear interior point solvers that exploit sparse and structured linear algebra techniques. We demonstrate these capabilities by solving estimation problems for synthetic human gut microbiome community models. We show that an instance with 156 parameters, 144 differential equations, and 1,704 experimental data points can be solved in less than 3 minutes using our proposed framework (while an off-the-shelf simulation-based solution framework requires over 7 hours). We also create large instances to show that the proposed framework is scalable and can solve problems with up to 2,352 parameters, 2,304 differential equations, and 20,352 data points in less than 15 minutes. The proposed framework is flexible and easy-to-use, can be broadly applied to dynamic models of biological systems, and enables the implementation of sophisticated estimation techniques to quantify parameter uncertainty, to diagnose observability/uniqueness issues, to perform model selection, and to handle outliers. Sungho Shin, Ophelia S. Venturelli, Victor M. Zavala |
PLoS Comput. Biol. | 1 |
| 2018 | Fully Neural Network Based Speech Recognition on Mobile and Embedded DevicesabstractReal-time automatic speech recognition (ASR) on mobile and embedded devices has been of great interests for many years. We present real-time speech recognition on smartphones or embedded systems by employing recurrent neural network (RNN) based acoustic models, RNN based language models, and beam-search decoding. The acoustic model is end-to-end trained with connectionist temporal classification (CTC) loss. The RNN implementation on embedded devices can suffer from excessive DRAM accesses because the parameter size of a neural network usually exceeds that of the cache memory and the parameters are used only once for each time step. To remedy this problem, we employ a multi-time step parallelization approach that computes multiple output samples at a time with the parameters fetched from the DRAM. Since the number of DRAM accesses can be reduced in proportion to the number of parallelization steps, we can achieve a high processing speed. However, conventional RNNs, such as long short-term memory (LSTM) or gated recurrent unit (GRU), do not permit multi-time step parallelization. We construct an acoustic model by combining simple recurrent units (SRUs) and depth-wise 1-dimensional convolution layers for multi-time step parallelization. Both the character and word piece models are developed for acoustic modeling, and the corresponding RNN based language models are used for beam search decoding. We achieve a competitive WER for WSJ corpus using the entire model size of around 15MB and achieve real-time speed using only a single core ARM without GPU or special hardware. Jinhwan Park, Yoonho Boo, Iksoo Choi, Sungho Shin, Wonyong Sung |
NeurIPS | 4 |
| 2017 | Fixed-point optimization of deep neural networks with adaptive step size retrainingabstractFixed-point optimization of deep neural networks plays an important role in hardware based design and low-power implementations. Many deep neural networks show fairly good performance even with 2- or 3-bit precision when quantized weights are fine-tuned by retraining. We propose an improved fixed-point optimization algorithm that estimates the quantization step size dynamically during the retraining. In addition, a gradual quantization scheme is also tested, which sequentially applies fixed-point optimizations from high- to low-precision. The experiments are conducted for feed-forward deep neural networks (FFDNNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs). Sungho Shin, Yoonho Boo, Wonyong Sung |
ICASSP | 1 |
| 2016 | Fixed-point performance analysis of recurrent neural networksabstractRecurrent neural networks have shown excellent performance in many applications; however they require increased complexity in hardware or software based implementations. The hardware complexity can be much lowered by minimizing the word-length of weights and signals. This work analyzes the fixed-point performance of recurrent neural networks using a retrain based quantization method. The quantization sensitivity of each layer in RNNs is studied, and the overall fixed-point optimization results minimizing the capacity of weights while not sacrificing the performance are presented. A language model and a phoneme recognition examples are used. Sungho Shin, Kyuyeon Hwang, Wonyong Sung |
ICASSP | 1 |
| 2016 | Dynamic hand gesture recognition for wearable devices with low complexity recurrent neural networksabstractGesture recognition is a very essential technology for many wearable devices. While previous algorithms are mostly based on statistical methods including the hidden Markov model, we develop two dynamic hand gesture recognition techniques using low complexity recurrent neural network (RNN) algorithms. One is based on video signal and employs a combined structure of a convolutional neural network (CNN) and an RNN. The other uses accelerometer data and only requires an RNN. Fixed-point optimization that quantizes most of the weights into two bits is conducted to optimize the amount of memory size for weight storage and reduce the power consumption in hardware and software based implementations. Sungho Shin, Wonyong Sung |
ISCAS | 1 |
| 2015 | Finding hidden relevant documents buried in scientific documents by terminological paraphrases
Sung-Pil Choi, Sungho Shin, Hanmin Jung, Daesung Lee 0001 |
Multim. Tools Appl. | 2 |
| 2014 | Contextual keyword extraction by building sentences with crowdsourcing
Soon Gill Hong, Sungho Shin, Mun Yong Yi |
Multim. Tools Appl. | 2 |