Sung-Ho Bae

dblp:76/2068 · DBLP profile ↗
← Back
40ranked-venue papers
9as first author
26since 2021 · last 2027
0000-0003-2677-3186ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 20 · 20 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Pseudo-image spatiotemporal augmentation for cellular network traffic prediction
Maryam Qamar, Tahir Khalil, Md. Atikuzzaman, Chaoning Zhang, Sung-Ho Bae
Expert Syst. Appl.5
2026 I-INR: Iterative Implicit Neural Representations
abstract
Implicit Neural Representations (INRs) have revolutionized signal processing and computer vision by modeling signals as continuous, differentiable functions parameterized by neural networks. However, INRs are prone to the spectral bias problem, limiting their ability to retain high-frequency information, and often struggle with noise robustness. Motivated by recent trends in iterative refinement processes, we propose Iterative Implicit Neural Representations (I-INRs). This novel plug-and-play framework iteratively refines signal reconstructions to restore high-frequency details, improve noise robustness, and enhance generalization, ultimately delivering superior reconstruction quality. I-INRs integrate seamlessly into existing INR architectures with only a 0.5–2% increase in parameters. During reconstruction, the iterative refinement adds just 0.8–1.6% additional FLOPs over the baseline while delivering a substantial performance boost of up to +2.0 PSNR. Extensive experiments demonstrate that I-INRs consistently outperform WIRE, SIREN, and Gauss across various computer vision tasks, including image fitting, image denoising, and object occupancy prediction.
Ali Haider, Muhammad Salman Ali, Maryam Qamar, Tahir Khalil, Soo Ye Kim, Jihyong Oh, Enzo Tartaglione, Sung-Ho Bae
AAAI8
2026 Post Training Quantization for Efficient Dataset Condensation
abstract
Dataset Condensation (DC) distills knowledge from large datasets into smaller ones, accelerating training and reducing storage requirements. However, despite notable progress, prior methods have largely overlooked the potential of quantization for further reducing storage costs. In this paper, we take the first step to explore post-training quantization in dataset condensation, demonstrating its effectiveness in reducing storage size while maintaining representation quality without requiring expensive training cost. However, we find that at extremely low bit-widths (e.g., 2-bit), conventional quantization leads to substantial degradation in representation quality, negatively impacting the networks trained on these data. To address this, we propose a novel patch-based post-training quantization approach that ensures localized quantization with minimal loss of information. To reduce the overhead of quantization parameters, especially for small patch sizes, we employ quantization-aware clustering to identify similar patches and subsequently aggregate them for efficient quantization. Furthermore, we introduce a refinement module to align the distribution between original images and their dequantized counterparts, compensating for quantization errors. Our method is a plug-and-play framework that can be applied to synthetic images generated by various DC methods. Extensive experiments across diverse benchmarks including CIFAR-10/100, Tiny ImageNet, and ImageNet subsets demonstrate that our method consistently outperforms prior works under the same storage constraints. Notably, our method doubles the test accuracy of existing methods at extreme compression regimes (e.g., from 26.0% to 54.1% for DM at IPC=1), while operating directly on 2-bit images without additional distillation.
Linh-Tam Tran, Sung-Ho Bae
AAAI2
2026 Lightweight LLM Agent Memory with Small Language Models
abstract
Jiaquan Zhang, Chaoning Zhang, Shuxu Chen, Zhenzhen Huang, Pengcheng Zheng, Zhicheng Wang, Ping Guo, Fan Mo, Sung-Ho Bae, Jie Zou, Jiwei Wei, Yang Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiaquan Zhang, Chaoning Zhang, Shuxu Chen, Zhenzhen Huang, Sung-Ho Bae, Jie Zou 0001, Jiwei Wei, Yang Yang 0002
ACL (1)9
2026 Efficient dataset condensation with learnable color representation and subset matching
Linh-Tam Tran, Quang Hieu Vo, Maryam Qamar, Chaoning Zhang, Hui Yong Kim, Sung-Ho Bae
Knowl. Based Syst.6
2026 Compression in 3D Gaussian Splatting: A Survey of Methods, Trends, and Future Directions
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a pioneering approach in explicit scene rendering and computer graphics. Unlike traditional neural radiance field (NeRF) methods, which typically rely on implicit, coordinate-based models to map spatial coordinates to pixel values, 3DGS utilizes millions of learnable 3D Gaussians. Its differentiable rendering technique and inherent capability for explicit scene representation and manipulation positions 3DGS as a potential game-changer for the next generation of 3D reconstruction and representation technologies. This enables 3DGS to deliver real-time rendering speeds while offering unparalleled editability levels. However, despite its advantages, 3DGS suffers from substantial memory and storage requirements, posing challenges for deployment on resource-constrained devices. In this survey, we provide a comprehensive overview focusing on the scalability and compression of 3DGS. We begin with a detailed background overview of 3DGS, followed by a structured taxonomy of existing compression methods. Additionally, we analyze and compare current methods from the topological perspective, evaluating their strengths and limitations in terms of fidelity, compression ratios, and computational efficiency. Furthermore, we explore how advancements in efficient NeRF representations can inspire future developments in 3DGS optimization. Finally, we conclude with current research challenges and highlight key directions for future exploration.
Muhammad Salman Ali, Chaoning Zhang, Marco Cagnazzo, Giuseppe Valenzise, Enzo Tartaglione, Sung-Ho Bae
IEEE Trans. Circuits Syst. Video Technol.6
2025 ELMGS: Enhancing Memory and Computation Scalability Through coMpression for 3D Gaussian Splatting
abstract
3D models have recently been popularized by the potentiality of end-to-end training offered first by Neural Radiance Fields and most recently by 3D Gaussian Splatting models. The latter has the big advantage of naturally providing fast training convergence and high editability. However, as the research around these is still in its infancy, there is still a gap in the literature regarding the model's scalability. In this work, we propose an approach enabling both memory and computation scalability of such models. More specifically, we propose an iterative pruning strategy that removes redundant information encoded in the model. We also enhance compressibility for the model by including a differentiable quantization and entropy coding estimator in the optimization strategy. Our results on popular benchmarks showcase the effectiveness of the proposed approach and open the road to the broad deployability of such a solution even on resource-constrained devices.
Muhammad Salman Ali, Sung-Ho Bae, Enzo Tartaglione
WACV2
2025 Unified link prediction modeling for enhanced knowledge graph completion task
Tri D. T. Nguyen, Ubaid Ur Rehman 0002, Musarrat Hussain, Rao Faizan, Jamil Hussain, Sung-Ho Bae, Jung Uk Kim, Seong Tae Kim 0001, Sungyoung Lee 0001
Expert Syst. Appl.6
2025 Effects of mixed sample data augmentation on interpretability of neural networks
Soyoun Won, Sung-Ho Bae, Seong Tae Kim 0001
Neural Networks2
2024 Descanning: From Scanned to the Original Images with a Color Correction Diffusion Model
abstract
A significant volume of analog information, i.e., documents and images, have been digitized in the form of scanned copies for storing, sharing, and/or analyzing in the digital world. However, the quality of such contents is severely degraded by various distortions caused by printing, storing, and scanning processes in the physical world. Although restoring high-quality content from scanned copies has become an indispensable task for many products, it has not been systematically explored, and to the best of our knowledge, no public datasets are available. In this paper, we define this problem as Descanning and introduce a new high-quality and large-scale dataset named DESCAN-18K. It contains 18K pairs of original and scanned images collected in the wild containing multiple complex degradations. In order to eliminate such complex degradations, we propose a new image restoration model called DescanDiffusion consisting of a color encoder that corrects the global color degradation and a conditional denoising diffusion probabilistic model (DDPM) that removes local degradations. To further improve the generalization ability of DescanDiffusion, we also design a synthetic data generation scheme by reproducing prominent degradations in scanned images. We demonstrate that our DescanDiffusion outperforms other baselines including commercial restoration products, objectively and subjectively, via comprehensive experiments and analyses.
Junghun Cha, Ali Haider, Seoyun Yang, Hoeyeong Jin, Subin Yang, A. F. M. Shahab Uddin, Jaehyoung Kim, Soo Ye Kim, Sung-Ho Bae
AAAI9
2024 Trimming the Fat: Efficient Compression of 3D Gaussian Splats through Pruning
Muhammad Salman Ali, Maryam Qamar, Sung-Ho Bae, Enzo Tartaglione
BMVC3
2024 G-SHARP: Globally Shared Kernel with Pruning for Efficient CNNs
abstract
Filter Decomposition (FD) methods have gained traction in compressing large neural networks by dividing weights into basis and coefficients. Recent advancements have focused on reducing weight redundancy by sharing either basis or coefficients stage-wise. However, traditional sharing approaches have overlooked the potential of sharing basis on a network-wide scale. In this study, we introduce an FD technique called G-SharP that elevates performance by using globally shared kernels throughout the network. To bolster the efficacy of G-SharP, we unveil a novel batch normalization-based co-efficient pruning strategy aiming to boost computational efficiency. Comprehensive evaluations show that our method notably diminishes computational demands and model size while incurring only a slight decline in performance. On benchmarks like CIFAR-10, ImageNet, PASCAL-VOC, and MS-COCO, G-SharP achieves significant reductions in model dimensions and FLOPs yet maintains accuracy levels akin to the original uncompressed models. Notably, G-SharP surpasses numerous leading lightweight models, striking a commendable balance between precision and efficiency.
Eunseop Shin, Incheon Cho, A. F. M. Shahab Uddin, Younho Jang, Sung-Ho Bae
ICASSP6
2024 Chain-of-Factors: A Zero-Shot Prompting Methodology Enabling Factor-Centric Reasoning in Large Language Models
abstract
Large language models (LLMs) have significantly improved numerous natural language processing tasks. However, their performance relies heavily on the provided instructions or prompts. Recently, several prompting methodologies have been developed to enhance the reasoning abilities of LLMs. Notably, the Chain-of-Thought (CoT) approach provides examples that help break down tasks into sub-steps, resulting in more accurate solutions. However, the process of generating detailed examples may not be user-friendly, as end users prefer providing task descriptions rather than a set of examples. In this study, we introduce Chain-of-Factors (CoF), an innovative zero-shot prompting methodology that incorporates task-specific instructions as a chain of factors into the prompt, aimed at enhancing the factor-centric reasoning abilities of LLMs. Experiments on three LLMs, including ChatGPT-3.5, Gemini, and GPT-4, show performance improvements ranging from 0.01% to 40.2% in accuracy on various symbolic reasoning and logical reasoning tasks compared with zero-shot and few-shot CoT. In summary, CoF enhances LLMs' reasoning abilities by including task-specific steps and instructions, while also decreasing the necessity for fine-tuning specific to each task.
Musarrat Hussain, Ubaid Ur Rehman 0002, Tri D. T. Nguyen, Sungyoung Lee 0001, Seong Tae Kim 0001, Sung-Ho Bae, Jung Uk Kim
ICMLA6
2024 ALICE: Adapt your Learnable Image Compression modEl for variable bitrates
abstract
When training a Learned Image Compression model, the loss function is minimized such that the encoder and the decoder attain a target Rate-Distorsion trade-off. Therefore, a distinct model shall be trained and stored at the transmitter and receiver for each target rate, fostering the quest for efficient variable bitrate compression schemes. This paper proposes plugging Low-Rank Adapters into a transformer-based pre-trained LIC model and training them to meet different target rates. With our method, encoding an image at a variable rate is as simple as training the corresponding adapters and plugging them into the frozen pre-trained model. Our experiments show performance comparable with state-of-the-art fixed-rate LIC models at a fraction of the training and deployment cost. We publicly released the code at https://github.com/EIDOSLAB/ALICE.
Gabriele Spadaro, Muhammad Salman Ali, Alberto Presta, Giommaria Pilo, Sung-Ho Bae, Jhony-Heriberto Giraldo-Zuluaga, Attilio Fiandrotti, Marco Grangetto, Enzo Tartaglione
VCIP5
2023 MST-compression: Compressing and Accelerating Binary Neural Networks with Minimum Spanning Tree
abstract
Binary neural networks (BNNs) have been widely adopted to reduce the computational cost and memory storage on edge-computing devices by using one-bit representation for activations and weights. However, as neural networks become wider/deeper to improve accuracy and meet practical requirements, the computational burden remains a significant challenge even on the binary version. To address these issues, this paper proposes a novel method called Minimum Spanning Tree (MST) compression that learns to compress and accelerate BNNs. The proposed architecture leverages an observation from previous works that an output channel in a binary convolution can be computed using another output channel and XNOR operations with weights that differ from the weights of the reused channel. We first construct a fully connected graph with vertices corresponding to output channels, where the distance between two vertices is the number of different values between the weight sets used for these outputs. Then, the MST of the graph with the minimum depth is proposed to reorder output calculations, aiming to reduce computational cost and latency. Moreover, we propose a new learning algorithm to reduce the total MST distance during training. Experimental results on benchmark models demonstrate that our method achieves significant compression ratios with negligible accuracy drops, making it a promising approach for resource-constrained edge-computing devices.
Quang Hieu Vo, Linh-Tam Tran, Sung-Ho Bae, Choong Seon Hong
ICCV3
2023 Towards Efficient Image Compression Without Autoregressive Models
abstract
Recently, learned image compression (LIC) has garnered increasing interest with its rapidly improving performance surpassing conventional codecs. A key ingredient of LIC is a hyperprior-based entropy model, where the underlying joint probability of the latent image features is modeled as a product of Gaussian distributions from each latent element. Since latents from the actual images are not spatially independent, autoregressive (AR) context based entropy models were proposed to handle the discrepancy between the assumed distribution and the actual distribution. Though the AR-based models have proven effective, the computational complexity is significantly increased due to the inherent sequential nature of the algorithm. In this paper, we present a novel alternative to the AR-based approach that can provide a significantly better trade-off between performance and complexity. To minimize the discrepancy, we introduce a correlation loss that forces the latents to be spatially decorrelated and better fitted to the independent probability model. Our correlation loss is proved to act as a general plug-in for the hyperprior (HP) based learned image compression methods. The performance gain from our correlation loss is ‘free’ in terms of computation complexity for both inference time and decoding time. To our knowledge, our method gives the best trade-off between the complexity and performance: combined with the Checkerboard-CM, it attains **90%** and when combined with ChARM-CM, it attains **98%** of the AR-based BD-Rate gains yet is around **50 times** and **30 times** faster than AR-based methods respectively
Muhammad Salman Ali, Yeongwoong Kim, Maryam Qamar, Sung-Chang Lim, Donghyun Kim 0017, Chaoning Zhang, Sung-Ho Bae, Hui Yong Kim
NeurIPS7
2023 Exploring the Optimal Bit Pair for a Quantized Generator and Discriminator
abstract
Generative Adversarial Networks (GANs) are hindered from real-world applications due to their high computational cost and memory requirements. Model compression techniques, such as quantization, pruning, and knowledge distillation, can compress neural networks, lower memory requirements, and model size. However, quantizing generators often leads to a suboptimal solution. In this paper, we propose a novel method to stabilize GAN quantization by quantizing the generator and discriminator with different bit precision. Our method maximizes the quantization efficiency by jointly quantizing the generator and discriminator, which we found to be dependent on each other’s quantization. Specifically, quantizing the discriminator enhances the performance of the quantized generator, while the discriminator’s optimal quantization bit depends on the generator’s quantization bit and architectural type. We conducted extensive experiments on various GAN models, including BigGAN, SAGAN, and SNGAN, using different quantization methods, such as LSQ, PACT, and DoReFa, on benchmark dataset (CIFAR10). The experimental results demonstrate that our joint quantization method achieves higher compression rates while offering better performance in Frechet Inception Distance (FID) and Inception Score (IS).
Subin Yang, Muhammad Salman Ali, A. F. M. Shahab Uddin, Sung-Ho Bae
VCIP4
2022 GLAMD: Global and Local Attention Mask Distillation for Object Detectors
Younho Jang, Wheemyung Shin, Jinbeom Kim, Simon S. Woo, Sung-Ho Bae
ECCV (10)5
2022 ZooD: Exploiting Model Zoo for Out-of-Distribution Generalization
abstract
Recent advances on large-scale pre-training have shown great potentials of leveraging a large set of Pre-Trained Models (PTMs) for improving Out-of-Distribution (OoD) generalization, for which the goal is to perform well on possible unseen domains after fine-tuning on multiple training domains. However, maximally exploiting a zoo of PTMs is challenging since fine-tuning all possible combinations of PTMs is computationally prohibitive while accurate selection of PTMs requires tackling the possible data distribution shift for OoD tasks. In this work, we propose ZooD, a paradigm for PTMs ranking and ensemble with feature selection. Our proposed metric ranks PTMs by quantifying inter-class discriminability and inter-domain stability of the features extracted by the PTMs in a leave-one-domain-out cross-validation manner. The top-K ranked models are then aggregated for the target OoD task. To avoid accumulating noise induced by model ensemble, we propose an efficient variational EM algorithm to select informative features. We evaluate our paradigm on a diverse model zoo consisting of 35 models for various OoD tasks and demonstrate: (i) model ranking is better correlated with fine-tuning ranking than previous methods and up to 9859x faster than brute-force fine-tuning; (ii) OoD generalization after model ensemble with feature selection outperforms the state-of-the-art methods and the accuracy on most challenging task DomainNet is improved from 46.5\% to 50.6\%. Furthermore, we provide the fine-tuning results of 35 PTMs on 7 OoD datasets, hoping to help the research of model zoo and OoD generalization. Code will be available at \href{https://gitee.com/mindspore/models/tree/master/research/cv/zood}{https://gitee.com/mindspore/models/tree/master/research/cv/zood}.
Qishi Dong, Fengwei Zhou, Chuanlong Xie, Tianyang Hu 0001, Yongxin Yang, Sung-Ho Bae, Zhenguo Li
NeurIPS7
2022 Facial Expression Recognition with Active Local Shape Pattern and Learned-Size Block Representations
abstract
Facial expression recognition has been studied broadly, and several works using local micro-pattern descriptors have obtained significant results. There are, however, open questions: how to design a discriminative and robust feature descriptor?, how to select expression-related most influential features?, and how to represent the face descriptor exploiting the most salient parts of the face? In this article, we address these three issues to achieve better performance in recognizing facial expressions. First, we propose a new feature descriptor, namely Local Shape Pattern (LSP), that describes the local shape structure of a pixel’s neighborhood based on the prominent directional information by analyzing the statistics of the neighborhood gradient, which allows it to be robust against subtle local noise and distortion. Furthermore, we propose a selection strategy for learning the influential codes being active in the expression affiliated changes by selecting them exhibiting statistical dominance and high spatial variance. Lastly, we learn the size of the salient facial blocks to represent the facial description with the notion that changes in expressions vary in size and location. We conduct person-independent experiments in existing datasets after combining above three proposals, and obtain an improved performance for the facial expression recognition task.
Md. Tauhid Bin Iqbal, Byungyong Ryu, Adín Ramírez Rivera, Farkhod Makhmudkhujaev, Oksam Chae, Sung-Ho Bae
IEEE Trans. Affect. Comput.6
2021 Adversarial Robustness for Unsupervised Domain Adaptation
abstract
Extensive Unsupervised Domain Adaptation (UDA) studies have shown great success in practice by learning transferable representations across a labeled source domain and an unlabeled target domain with deep models. However, current work focuses on improving the generalization ability of UDA models on clean examples without considering the adversarial robustness, which is crucial in real-world applications. Conventional adversarial training methods are not suitable for the adversarial robustness on the unlabeled target domain of UDA since they train models with adversarial examples generated by the supervised loss function. In this work, we propose to leverage intermediate representations learned by robust ImageNet models to improve the robustness of UDA models. Our method works by aligning the features of the UDA model with the robust features learned by ImageNet pre-trained models along with domain adaptation training. It utilizes both labeled and unlabeled domains and instills robustness without any adversarial intervention or label requirement during domain adaptation training. Our experimental results show that our method significantly improves adversarial robustness compared to the baseline while keeping clean accuracy on various UDA benchmarks.
Fengwei Zhou, Hang Xu 0004, Lanqing Hong, Ping Luo 0002, Sung-Ho Bae, Zhenguo Li
ICCV6
2021 Distilling Global and Local Logits with Densely Connected Relations
abstract
In prevalent knowledge distillation, logits in most image recognition models are computed by global average pooling, then used to learn to encode the high-level and task-relevant knowledge. In this work, we solve the limitation of this global logit transfer in this distillation context. We point out that it prevents the transfer of informative spatial information, which provides localized knowledge as well as rich relational information across contexts of an input scene. To exploit the rich spatial information, we propose a simple yet effective logit distillation approach. We add a local spatial pooling layer branch to the penultimate layer, thereby our method extends the standard logit distillation and enables learning of both finely-localized knowledge and holistic representation. Our proposed method shows favorable accuracy improvement against the state-of-the-art methods on several image classification datasets. We show that our distilled students trained on the image classification task can be successfully leveraged for object detection and semantic segmentation tasks; this result demonstrates our method’s high transferability.
Youmin Kim, Jinbae Park, Younho Jang, Muhammad Salman Ali, Tae-Hyun Oh, Sung-Ho Bae
ICCV6
2021 SaliencyMix: A Saliency Guided Data Augmentation Strategy for Better Regularization
A. F. M. Shahab Uddin, Mst. Sirazam Monira, Wheemyung Shin, TaeChoong Chung, Sung-Ho Bae
ICLR5
2021 MixACM: Mixup-Based Robustness Transfer via Distillation of Activated Channel Maps
abstract
Deep neural networks are susceptible to adversarially crafted, small, and imperceptible changes in the natural inputs. The most effective defense mechanism against these examples is adversarial training which constructs adversarial examples during training by iterative maximization of loss. The model is then trained to minimize the loss on these constructed examples. This min-max optimization requires more data, larger capacity models, and additional computing resources. It also degrades the standard generalization performance of a model. Can we achieve robustness more efficiently? In this work, we explore this question from the perspective of knowledge transfer. First, we theoretically show the transferability of robustness from an adversarially trained teacher model to a student model with the help of mixup augmentation. Second, we propose a novel robustness transfer method called Mixup-Based Activated Channel Maps (MixACM) Transfer. MixACM transfers robustness from a robust teacher to a student by matching activated channel maps generated without expensive adversarial perturbations. Finally, extensive experiments on multiple datasets and different learning scenarios show our method can transfer robustness while also improving generalization on natural images.
Fengwei Zhou, Chuanlong Xie, Sung-Ho Bae, Zhenguo Li
NeurIPS5
2021 Convolutional Network With Twofold Feature Augmentation for Diabetic Retinopathy Recognition From Multi-Modal Images
abstract
OBJECTIVE: With the scenario of limited labeled dataset, this paper introduces a deep learning-based approach that leverages Diabetic Retinopathy (DR) severity recognition performance using fundus images combined with wide-field swept-source optical coherence tomography angiography (SS-OCTA). METHODS: The proposed architecture comprises a backbone convolutional network associated with a Twofold Feature Augmentation mechanism, namely TFA-Net. The former includes multiple convolution blocks extracting representational features at various scales. The latter is constructed in a two-stage manner, i.e., the utilization of weight-sharing convolution kernels and the deployment of a Reverse Cross-Attention (RCA) stream. RESULTS: The proposed model achieves a Quadratic Weighted Kappa rate of 90.2% on the small-sized internal KHUMC dataset. The robustness of the RCA stream is also evaluated by the single-modal Messidor dataset, of which the obtained mean Accuracy (94.8%) and Area Under Receiver Operating Characteristic (99.4%) outperform those of the state-of-the-arts significantly. CONCLUSION: Utilizing a network strongly regularized at feature space to learn the amalgamation of different modalities is of proven effectiveness. Thanks to the widespread availability of multi-modal retinal imaging for each diabetes patient nowadays, such approach can reduce the heavy reliance on large quantity of labeled visual data. SIGNIFICANCE: Our TFA-Net is able to coordinate hybrid information of fundus photos and wide-field SS-OCTA for exhaustively exploiting DR-oriented biomarkers. Moreover, the embedded feature-wise augmentation scheme can enrich generalization ability efficiently despite learning from small-scale labeled data.
Cam-Hao Hua, Kiyoung Kim, Thien Huynh-The, Jong In You, Seung-Young Yu, Thuong Le-Tien, Sung-Ho Bae, Sungyoung Lee 0001
IEEE J. Biomed. Health Informatics7
2021 Revisiting Internal Covariate Shift for Batch Normalization
abstract
Despite the success of batch normalization (BatchNorm) and a plethora of its variants, the exact reasons for its success are still shady. The original BatchNorm article explained it as a mechanism that reduces the internal covariate shift (ICS), i.e., the distribution shifts in the input of the layers during training. Recently, some articles manifested skepticism on this hypothesis and provided alternative explanations for the success of BatchNorm, such as the applicability of very high learning rates and the ability to smooth the landscape in optimization. In this work, we counter these alternative arguments by demonstrating the importance of reduction in ICS following an empirical approach. We demonstrated various ways to achieve the abovementioned alternative properties without any performance boost. In this light, we explored the importance of different BatchNorm parameters (i.e., batch statistics and affine transformation parameters) by visualizing their effectiveness in the performance and analyzed their connections with ICS. Afterward, we showed a different normalization scheme that fulfills all the alternative explanations except reduction in ICS. Despite having all the alternative properties, we observed its poor performance, which nullifies the alternative claims, rather signifies the importance of the ICS reduction. We performed comprehensive experiments on many variants of BatchNorm, finding that all of them similarly reduce ICS.
Md. Tauhid Bin Iqbal, Sung-Ho Bae
IEEE Trans. Neural Networks Learn. Syst.3
2020 Effective Utilization of Hybrid Residual Modules in Deep Neural Networks for Super Resolution
Abdul Muqeet, Sung-Ho Bae
MMM (2)2
2020 Cross-Attentional Bracket-shaped Convolutional Network for semantic image segmentation
Cam-Hao Hua, Thien Huynh-The, Sung-Ho Bae, Sungyoung Lee 0002
Inf. Sci.3
2018 Learning-Based Just-Noticeable-Quantization- Distortion Modeling for Perceptual Video Coding
abstract
Conventional predictive video coding-based approaches are reaching the limit of their potential coding efficiency improvements, because of severely increasing computation complexity. As an alternative approach, perceptual video coding (PVC) has attempted to achieve high coding efficiency by eliminating perceptual redundancy, using just-noticeable-distortion (JND) directed PVC. The previous JNDs were modeled by adding white Gaussian noise or specific signal patterns into the original images, which were not appropriate in finding JND thresholds due to distortion with energy reduction. In this paper, we present a novel discrete cosine transform-based energy-reduced JND model, called ERJND, that is more suitable for JND-based PVC schemes. Then, the proposed ERJND model is extended to two learning-based just-noticeable-quantization-distortion (JNQD) models as preprocessing that can be applied for perceptual video coding. The two JNQD models can automatically adjust JND levels based on given quantization step sizes. One of the two JNQD models, called LR-JNQD, is based on linear regression and determines the model parameter for JNQD based on extracted handcraft features. The other JNQD model is based on a convolution neural network (CNN), called CNN-JNQD. To our best knowledge, our paper is the first approach to automatically adjust JND levels according to quantization step sizes for preprocessing the input to video encoders. In experiments, both the LR-JNQD and CNN-JNQD models were applied to high efficiency video coding (HEVC) and yielded maximum (average) bitrate reductions of 38.51% (10.38%) and 67.88% (24.91%), respectively, with little subjective video quality degradation, compared with the input without preprocessing applied.
Sehwan Ki, Sung-Ho Bae, Munchurl Kim, Hyunsuk Ko
IEEE Trans. Image Process.2
2017 A DCT-Based Total JND Profile for Spatiotemporal and Foveated Masking Effects
abstract
In image and video processing fields, Discrete Cosine Transform (DCT)-based just-noticeable difference (JND) profiles have effectively been utilized to remove perceptual redundancies in pictures for compression. In this paper, we solve two problems that are often intrinsic to the conventional DCT-based JND profiles: 1) no foveated masking (FM) JND model has been incorporated in modeling the DCT-based JND profiles and 2) the conventional temporal masking (TM) JND models assume that all moving objects in frames can be well tracked by the eyes and that they are projected on the fovea regions of the eyes, which is not a realistic assumption and may result in poor estimation of JND values for untracked moving objects (or image regions). To solve these two problems, we first propose a generalized JND model for joint effects between TM and FM effects. With this model, called the temporal-foveated masking (TFM) JND model, JND thresholds for any tracked/untracked and moving/still image regions can be elaborately estimated. Finally, the TFM-JND model is incorporated into a total DCT-based JND profile with a spatial contrast sensitivity function, luminance masking, and contrast masking JND models. In addition, we propose a JND adjustment method for our total JND profile to avoid overestimation of JND values for image blocks of fixed sizes with various image characteristics. To validate the effectiveness of the total JND profile, an experiment involving a subjective distortion-visibility assessment has been conducted. The experiment results show that the proposed total DCT-based JND profile yields significant performance improvement with much higher capability of distortion concealment (average 5.6-dB lower PSNR) compared with state-of-the-art JND profiles. The MATLAB source code of the proposed total DCT-based JND profile is publicly available online at https://sites.google.com/site/sunghobaecv/jnd.
Sung-Ho Bae, Munchurl Kim
IEEE Trans. Circuits Syst. Video Technol.1
2016 A Novel Image Quality Assessment With Globally and Locally Consilient Visual Quality Perception
abstract
Computational models for image quality assessment (IQA) have been developed by exploring effective features that are consistent with the characteristics of a human visual system (HVS) for visual quality perception. In this paper, we first reveal that many existing features used in computational IQA methods can hardly characterize visual quality perception for local image characteristics and various distortion types. To solve this problem, we propose a new IQA method, called the structural contrast-quality index (SC-QI), by adopting a structural contrast index (SCI), which can well characterize local and global visual quality perceptions for various image characteristics with structural-distortion types. In addition to SCI, we devise some other perceptually important features for our SC-QI that can effectively reflect the characteristics of HVS for contrast sensitivity and chrominance component variation. Furthermore, we develop a modified SC-QI, called structural contrast distortion metric (SC-DM), which inherits desirable mathematical properties of valid distance metricability and quasi-convexity. So, it can effectively be used as a distance metric for image quality optimization problems. Extensive experimental results show that both SC-QI and SC-DM can very well characterize the HVS's properties of visual quality perception for local image characteristics and various distortion types, which is a distinctive merit of our methods compared with other IQA methods. As a result, both SC-QI and SC-DM have better performances with a strong consilience of global and local visual quality perception as well as with much lower computation complexity, compared with the state-of-the-art IQA methods. The MATLAB source codes of the proposed SC-QI and SC-DM are publicly available online at https://sites.google.com/site/sunghobaecv/iqa.
Sung-Ho Bae, Munchurl Kim
IEEE Trans. Image Process.1
2016 DCT-QM: A DCT-Based Quality Degradation Metric for Image Quality Optimization Problems
abstract
Recent development of computational image quality assessment methods has shown to give very promising results in measuring perceptual visual quality for distorted images. However, most of them are difficult to be applied for optimization problems due to the lack of desirable mathematical properties, such as differentiability, convexity, and valid distance metricability. This paper proposes a novel Discrete Cosine Transform (DCT)-based quality degradation metric, called DCT-QM, which is based on the probability summation theory with a psychometric function for neural responses in the receptive fields of visual cortex in psychophysics. Consequently, the DCT-QM is formulated as a weighted mean L2norm in the DCT domain, which is very easy to implement and inherits the three desirable mathematical properties, that is, differentiability, convexity, and valid distance metricability, for image quality optimization problems. The extensive experimental results show that the proposed DCT-QM has promising results for many practical distortion types in image processing problems by showing high consistency with perceived visual quality.
Sung-Ho Bae, Munchurl Kim
IEEE Trans. Image Process.1
2016 HEVC-Based Perceptually Adaptive Video Coding Using a DCT-Based Local Distortion Detection Probability Model
abstract
Discrete Cosine Transform (DCT)-based just noticeable difference (JND) profiles have widely been applied into human perception-based video coding in order to reduce perceptual redundancy, which is one of the main goals of perceptual video coding (PVC). However, there are two problems for this approach: 1) the JND value of each transform coefficient is estimated for a fixed-sized DCT kernel (e.g., 8 × 8), but flexible coding structures with variable-sized transform units have been utilized in standard video coding frameworks, such high efficiency video coding (HEVC) and 2) the DCT transform coefficients are suppressed by the amounts of JND values for the removal of perceptual redundancy, but the DCT transform coefficients of residues are not sufficiently suppressed due to many small transform coefficient values in mid- and high-frequency regions below the JND values. In order to solve these problems, we propose a more generalized visibility model in the DCT domain, called the DCT-based local distortion detection probability (LDDP) model that can estimate a degree of distortion visibility for any distribution of the transform coefficients of any sized DCT kernel for residues. Furthermore, we propose an HEVC-compliant LDDP-based PVC scheme where transform coefficients are sufficiently suppressed based on the LDDP model. The proposed PVC scheme is implemented in the HEVC Test Model (HM 11.0) reference software to show the effectiveness of the LDDP-based PVC scheme. Objective and subjective tests for encoded test sequences are performed. The experimental results show that the LDDP-based PVC scheme achieves a significant performance improvement of bitrate reduction at the similar visual quality levels compared with the original HM 11.0.
Sung-Ho Bae, Jaeil Kim, Munchurl Kim
IEEE Trans. Image Process.1
2015 Single image super-resolution based on self-examples using context-dependent subpatches
abstract
Self-example-based super-resolution (SR) methods utilize internal dictionaries to reconstruct a high-resolution (HR) image from a single low-resolution (LR) input image. In general, a square-sized patch is used to find the LR-HR correspondences in the dictionaries. However, this may be a difficult issue because the LR input image and the dictionaries are of different scales. Inspired by this observation, we propose a novel self-example-based SR method, using context-dependent multi-shaped subpatches. Each LR input patch is segmented into multiple subpatches according to the context of the patch, enabling us to extract the better LR-HR correspondences. Our experimental results show that the proposed subpatch-based SR generates competitive high-quality HR images compared to state-of-the-art methods, with visually sharper edges that result in better visual quality.
Jae-Seok Choi, Sung-Ho Bae, Munchurl Kim
ICIP2
2015 A novel image quality assessment based on an adaptive feature for image characteristics and distortion types
abstract
In this paper, we reveal that many conventional features used in computational image quality assessment (IQA) methods can hardly characterize perceived distortions on various image characteristics and distortion types, thus resulting in relatively low prediction performance of visual quality scores. To solve this problem, we propose a new IQA method, called Structural Contrast-Quality Index (SC-QI) which is based on structural contrast index (SCI) as a very effective feature. SCI can adaptively quantify perceived distortions depending on various image characteristics and distortions types. In addition to SCI, some other perceptually important features that reflect effects of contrast sensitivity function and chrominance component variation are also combined into the proposed SC-QI. Our comprehensive experiments on three large IQA datasets verify that the proposed SC-QI outperforms the state-of-the-art ones while accompanying lower computational complexity.
Sung-Ho Bae, Munchurl Kim
VCIP1
2015 A novel SSIM index for image quality assessment using a new luminance adaptation effect model in pixel intensity domain
abstract
The Structural SIMimarity (SSIM) is one of the most prominent image quality assessment (IQA) methods due to its high prediction performance and wide applicability for image quality optimization problems. To reflect the luminance adaptation (LA) characteristic of human visual system (HVS), SSIM is modelled to have high consistency with Weber's law. However, it inevitably has some intrinsic faults that wrongly incorporate the LA effect into SSIM. In this paper, we firstly analyze that Weber's law in the conventional SSIM index cannot precisely reflect the LA effect due to two reasons: (i) it is reported that Weber's law model is not able to precisely be fitted in the measured experimental results for the LA effect; (ii) SSIM is calculated with pixel intensity values, but Weber's law is applied for luminance values which have non-linear relations with the pixel intensity values. To solve this problem, we first theoretically derive a new LA effect model in pixel intensity domain using a Gamma correction function and a power-law model. We then devise a weight function for the LA effect, called LA-based local weight function (LALF) which allows the proposed LA effect model to be precisely incorporated into SSIM index. To verify the effectiveness of the proposed LALF-based SSIM, we perform comprehensive experiments on four large IQA databases. Experimental results show that the proposed LALF helps performance improvement of the SSIM index.
Sung-Ho Bae, Munchurl Kim
VCIP1
2015 An HEVC-Compliant Perceptual Video Coding Scheme Based on JND Models for Variable Block-Sized Transform Kernels
abstract
In this paper, a High Efficiency Video Coding (HEVC)-compliant perceptual video coding (PVC) scheme is introduced based on just-noticeable difference (JND) models in both transform and pixel domains. We adopt an existing pixel-domain JND model for the transform skip mode of HEVC and propose a transform-domain JND model for the transform nonskip modes of HEVC. The proposed transform-domain JND model is designed by considering the spatial JND characteristics such as contrast sensitivity, luminance adaptation, and contrast masking effects as well as by considering the summation effects of variable block-sized transforms in HEVC. A temporal JND model is additionally incorporated into the proposed transform-domain JND model to further reduce perceptual redundancy. To incorporate the transform- and pixel-domain JND models into the encoding process in an HEVC-compliant manner, the transform coefficients and residues are suppressed in harmonization with the transform/quantization process and the quantization-only process of HEVC, respectively. To make the JND-based suppression effective, a distortion compensation factor is also proposed to reflect the perceptual distortion in the rate-distortion optimization-based encoding process. Based on subjective quality assessments of the encoded bit streams of test sequences, the proposed HEVC-compliant PVC scheme yields remarkable bitrate reductions of a maximum 49.10% and an average 16.10% with negligible subjective quality loss, compared with an HEVC reference software HEVC test model (HM 11.0). In addition, the proposed HEVC-compliant PVC scheme increases the encoding complexity of HM 11.0 only by an average of 11.25%.
Jaeil Kim, Sung-Ho Bae, Munchurl Kim
IEEE Trans. Circuits Syst. Video Technol.2
2014 Object tracking based on online partial instance learning with multiple local strong classifiers
abstract
In this paper, we propose a new appearance model based on Partial Instance Learning (PIL) with multiple local strong classifiers. The key idea of PIL is that image examples are divided into several partial image examples (or local-images), each of which is then independently trained with a local strong classifier. Finally, a tracker is updated for the optimal solution in the sense that the joint probability of partial image examples for each input image example becomes the largest. The proposed PIL method can be considered a risk diversification strategy for unpredictable partial occlusions or appearance changes of an object. Also, it can be regarded as a divide-and-conquer method of Online Boosting (OB), so that PIL only requires approximately 20% of computations compared with other OB methods in terms of iterations taken for learning process. Experiment results show that the proposed PIL-based object tracking method achieves better performance in tracking accuracy and much faster processing speed than other compared real-time based ones.
Sung-Ho Bae, Munchurl Kim
ICIP1
2014 A Novel Generalized DCT-Based JND Profile Based on an Elaborate CM-JND Model for Variable Block-Sized Transforms in Monochrome Images
abstract
In this paper, we propose a new DCT-based just noticeable difference (JND) profile incorporating the spatial contrast sensitivity function, the luminance adaptation effect, and the contrast masking (CM) effect. The proposed JND profile overcomes two limitations of conventional JND profiles: 1) the CM JND models in the conventional JND profiles employed simple texture complexity metrics, which are not often highly correlated with perceived complexity, especially for unstructured patterns. So, we proposed a new texture complexity metric that considers not only contrast intensity, but also structureness of image patterns, called the structural contrast index. We also newly found out that, as the structural contrast index of a background texture pattern increases, the modulation factors for CM-JND show a bandpass property in frequency. Based on this observation, a new CM-JND is modeled as a function of DCT frequency and the proposed structural contrast index, showing significantly high correlations with measured CM-JND values and 2) while the conventional DCT-based JND profiles are only applicable for specific transform block sizes, our proposed DCT-based JND profile is first designed to be applicable to any size of transform by deriving a new summation effect function, which can also be appropriately applied for quad-tree transform of high efficiency video coding. For the overall performance, the proposed DCT-based JND profile shows more tolerance for distortions with better perceptual quality than other JND profiles under comparison.
Sung-Ho Bae, Munchurl Kim
IEEE Trans. Image Process.1
2013 A Novel DCT-Based JND Model for Luminance Adaptation Effect in DCT Frequency
abstract
Many conventional DCT based Just Noticeable Distortion (JND) models incorporate luminance adaptation (LA) effect of the human visual system (HVS). The conventional LA-JND models exploit only background luminance to estimate JND values. In this letter, we reveal that the LA effect of HVS depends not only on background luminance but also on frequency in DCT domain. In addition, we first propose a novel DCT-based LA-JND model that takes into account its frequency characteristics. From our psychophysical experiment results, we found that the LA-JND threshold exhibits quasi-parabolic shapes with lower and higher curvatures in lower and higher frequency ranges, respectively. Our subjective evaluation shows that the proposed LA-JND model yields almost invisible distortions for test images with average PSNR of 29.11 dB, which is 2.63 dB lower than the other models for comparison, showing higher consistency with real visual perception for the LA in DCT domain.
Sung-Ho Bae, Munchurl Kim
IEEE Signal Process. Lett.1