Jong-Seok Lee

dblp:70/1152 · DBLP profile ↗
← Back
119ranked-venue papers
24as first author
39since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 70 · 12 first-author · 18 since 2021Artificial intelligence and machine learning · 50 · 5 first-author · 26 since 2021Human-computer interaction and ubiquitous computing · 13 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Noise-aware weight updating in AdaBoost for handling mislabeled data
Bonhyeok Ku, Jong-Seok Lee
Pattern Recognit.2
2025 Sketch to Stylized-Image: A Two-Stage Approach for Artistic Image Generation
abstract
Transforming sketches into artistic images is an essential task, especially for individuals without artistic expertise. Recent style transfer models struggle with applying styles to sketch images due to the absence of color information in the input image. To address this limitation, we propose a novel approach that leverages existing pre-trained models. Our method first colors the sketches using a well-performing pre-trained condition-based image generation model, and then applies style transfer to achieve the desired artistic effect. Experimental results demonstrate that our method effectively generates stylized sketch images with high quality, without the overhead of additional training.
Junha Park, Jong-Seok Lee
ICIP3
2025 Exploring the Camera Bias of Person Re-identification
abstract
We empirically investigate the camera bias of person re-identification (ReID) models. Previously, camera-aware methods have been proposed to address this issue, but they are largely confined to training domains of the models. We measure the camera bias of ReID models on unseen domains and reveal that camera bias becomes more pronounced under data distribution shifts. As a debiasing method for unseen domain data, we revisit feature normalization on embedding vectors. While the normalization has been used as a straightforward solution, its underlying causes and broader applicability remain unexplored. We analyze why this simple method is effective at reducing bias and show that it can be applied to detailed bias factors such as low-level image properties and body angle. Furthermore, we validate its generalizability across various models and benchmarks, highlighting its potential as a simple yet effective test-time postprocessing method for ReID. In addition, we explore the inherent risk of camera bias in unsupervised learning of ReID models. The unsupervised models remain highly biased towards camera labels even for seen domain data, indicating substantial room for improvement. Based on observations of the negative impact of camera-biased pseudo labels on training, we suggest simple training strategies to mitigate the bias. By applying these strategies to existing unsupervised learning algorithms, we show that significant performance improvements can be achieved with minor modifications.
Myungseo Song, Jong-Seok Lee
ICLR3
2025 Exploring Cross-Stage Adversarial Transferability in Class-Incremental Continual Learning
Jungwoo Kim 0003, Jong-Seok Lee
MMSP2
2025 Emotional EEG Classification using Upscaled Connectivity Matrices
abstract
Recent studies have demonstrated the effectiveness of using connectivity matrices as input to convolutional neural networks (CNNs) for emotional EEG classification, as they can effectively capture interregional interaction patterns. However, these matrices often suffer from loss of critical spatial information during convolution operations. To address this issue, we propose a simple yet effective approach: upscaling connectivity matrices to enhance local patterns. Experimental results show that this technique significantly improves classification performance, highlighting the importance of preserving spatial structures in early processing stages.
Chae-Won Lee, Jong-Seok Lee
SMC2
2025 Clean-to-clean: pretraining vision transformers without additional data
Hyeonjin Lee, Songkuk Kim, Jong-Seok Lee
Mach. Vis. Appl.3
2025 Network Fission Ensembles for low-cost self-ensembles
Hojung Lee, Jong-Seok Lee
Pattern Recognit. Lett.2
2025 Gently Sloped and Extended Classification Margin for Overconfidence Relaxation of Out-of-Distribution Samples
abstract
Recently, machine learning models are expected to be capable of detecting out-of-distribution (OOD) samples for safe use. However, the existing OOD detection methods have limitations. Post hoc calibration techniques used for OOD detection during the inference phase suffer from slow inference and low OOD detection accuracy because pretrained classifiers were not originally designed for this task. Training-phase methods require auxiliary data, entail slow training, and result in a decrease in classification accuracy. To address these issues, this article proposes jointly employing discriminative representation learning through angular margin loss and weight regularization during neural network training. Angular margin loss extends the classification margin, whereas weight regularization ensures a gently sloped margin in the learned embedding space. By constructing a classification margin that is both gently sloped and enlarged, the proposed approach mitigates the overconfidence of OOD samples and overcomes the shortcomings of previous methods. The experimental results demonstrate that the proposed method outperforms state-of-the-art detectors in identifying OOD samples without any side effects.
Jong-Seok Lee
IEEE Trans. Neural Networks Learn. Syst.2
2024 Curved Representation Space of Vision Transformers
abstract
Neural networks with self-attention (a.k.a. Transformers) like ViT and Swin have emerged as a better alternative to traditional convolutional neural networks (CNNs). However, our understanding of how the new architecture works is still limited. In this paper, we focus on the phenomenon that Transformers show higher robustness against corruptions than CNNs, while not being overconfident. This is contrary to the intuition that robustness increases with confidence. We resolve this contradiction by empirically investigating how the output of the penultimate layer moves in the representation space as the input data moves linearly within a small area. In particular, we show the following. (1) While CNNs exhibit fairly linear relationship between the input and output movements, Transformers show nonlinear relationship for some data. For those data, the output of Transformers moves in a curved trajectory as the input moves linearly. (2) When a data is located in a curved region, it is hard to move it out of the decision region since the output moves along a curved trajectory instead of a straight line to the decision boundary, resulting in high robustness of Transformers. (3) If a data is slightly modified to jump out of the curved region, the movements afterwards become linear and the output goes to the decision boundary directly. In other words, there does exist a decision boundary near the data, which is hard to find only because of the curved representation space. This explains the underconfident prediction of Transformers. Also, we examine mathematical properties of the attention operation that induce nonlinear response to linear perturbation. Finally, we share our additional findings, regarding what contributes to the curved representation space of Transformers, and how the curvedness evolves during training.
Juyeop Kim, Junha Park, Songkuk Kim, Jong-Seok Lee
AAAI4
2024 Anomaly Score: Evaluating Generative Models and Individual Generated Images Based on Complexity and Vulnerability
abstract
With the advancement of generative models, the assessment of generated images becomes increasingly more important. Previous methods measure distances between features of reference and generated images from trained vision models. In this paper, we conduct an extensive investigation into the relationship between the representation space and input space around generated images. We first propose two measures related to the presence of unnatural elements within images: complexity, which indicates how nonlinear the representation space is, and vulnerability, which is related to how easily the extracted feature changes by adversarial input changes. Based on these, we introduce a new metric to evaluating image-generative models called anomaly score (AS). Moreover, we propose AS-i (anomaly score for individual images) that can effectively evaluate generated images individually. Experimental results demonstrate the validity of the proposed approach.
Jaehui Hwang, Junghyuk Lee, Jong-Seok Lee
CVPR3
2024 Similarity of Neural Architectures Using Adversarial Attack Transferability
Jaehui Hwang, Dongyoon Han, Byeongho Heo, Song Park, Sanghyuk Chun, Jong-Seok Lee
ECCV (68)6
2024 Scalable Monotonic Neural Networks
abstract
In this research, we focus on the problem of learning monotonic neural networks, as preserving the monotonicity of a model with respect to a subset of inputs is crucial for practical applications across various domains. Although several methods have recently been proposed to address this problem, they have limitations such as not guaranteeing monotonicity in certain cases, requiring additional inference time, lacking scalability with increasing network size and number of monotonic inputs, and manipulating network weights during training. To overcome these limitations, we introduce a simple but novel architecture of the partially connected network which incorporates a 'scalable monotonic hidden layer' comprising three units: the exponentiated unit, ReLU unit, and confluence unit. This allows for the repetitive integration of the scalable monotonic hidden layers without other structural constraints. Consequently, our method offers ease of implementation and rapid training through the conventional error-backpropagation algorithm. We accordingly term this method as Scalable Monotonic Neural Networks (SMNN). Numerical experiments demonstrated that our method achieved comparable prediction accuracy to the state-of-the-art approaches while effectively addressing the aforementioned weaknesses.
Jong-Seok Lee
ICLR2
2024 Exploring Adversarial Robustness of Vision Transformers in the Spectral Perspective
abstract
The Vision Transformer has emerged as a powerful tool for image classification tasks, surpassing the performance of convolutional neural networks (CNNs). Recently, many researchers have attempted to understand the robustness of Transformers against adversarial attacks. However, previous researches have focused solely on perturbations in the spatial domain. This paper proposes an additional perspective that explores the adversarial robustness of Transformers against frequency-selective perturbations in the spectral domain. To facilitate comparison between these two domains, an attack framework is formulated as a flexible tool for implementing attacks on images in the spatial and spectral domains. The experiments reveal that Transformers rely more on phase and low frequency information, which can render them more vulnerable to frequency-selective attacks than CNNs. This work offers new insights into the properties and adversarial robustness of Transformers.
Gihyun Kim, Juyeop Kim, Jong-Seok Lee
WACV3
2024 PFSC: Parameter-free sphere classifier for imbalanced data classification
Yeontark Park, Jong-Seok Lee
Expert Syst. Appl.2
2024 Learning from class-imbalanced data using misclassification-focusing generative adversarial networks
Jaesub Yun, Jong-Seok Lee
Expert Syst. Appl.2
2024 Temporal shuffling for defending deep action recognition models against adversarial attacks
Jaehui Hwang, Huan Zhang 0001, Jun-Ho Choi, Cho-Jui Hsieh, Jong-Seok Lee
Neural Networks5
2023 Demystifying Randomly Initialized Networks for Evaluating Generative Models
abstract
Evaluation of generative models is mostly based on the comparison between the estimated distribution and the ground truth distribution in a certain feature space. To embed samples into informative features, previous works often use convolutional neural networks optimized for classification, which is criticized by recent studies. Therefore, various feature spaces have been explored to discover alternatives. Among them, a surprising approach is to use a randomly initialized neural network for feature embedding. However, the fundamental basis to employ the random features has not been sufficiently justified. In this paper, we rigorously investigate the feature space of models with random weights in comparison to that of trained models. Furthermore, we provide an empirical evidence to choose networks for random features to obtain consistent and reliable results. Our results indicate that the features from random networks can evaluate generative models well similarly to those from trained networks, and furthermore, the two types of features can be used together in a complementary way.
Junghyuk Lee, Jun-Hyuk Kim, Jong-Seok Lee
AAAI3
2023 ViPLO: Vision Transformer Based Pose-Conditioned Self-Loop Graph for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection, which localizes and infers relationships between human and objects, plays an important role in scene understanding. Although two-stage HOI detectors have advantages of high efficiency in training and inference, they suffer from lower performance than one-stage methods due to the old back-bone networks and the lack of considerations for the HOI perception process of humans in the interaction classifiers. In this paper, we propose Vision Transformer based Pose-Conditioned Self-Loop Graph (ViPLO) to resolve these problems. First, we propose a novel feature extraction method suitable for the Vision Transformer backbone, called masking with overlapped area (MOA) module. The MOA module utilizes the overlapped area between each patch and the given region in the attention function, which addresses the quantization problem when using the Vision Transformer backbone. In addition, we design a graph with a pose-conditioned self-loop structure, which updates the human node encoding with local features of human joints. This allows the classifier to focus on specific human joints to effectively identify the type of interaction, which is motivated by the human perception process for HOI. As a result, ViPLO achieves the state-of-the-art results on two public benchmarks, especially obtaining a +2.07 mAP performance gain on the HICO-DET dataset.
Jeeseung Park, Jong-Seok Lee
CVPR3
2023 Amicable Aid: Perturbing Images to Improve Classification Performance
abstract
While adversarial perturbation of images to attack deep image classification models pose serious security concerns in practice, this paper suggests a novel paradigm where the concept of image perturbation can benefit classification performance, which we call amicable aid. We show that by taking the opposite search direction of perturbation, an image can be modified to yield higher classification confidence and even a misclassified image can be made correctly classified. This can be also achieved with a large amount of perturbation by which the image is made unrecognizable by human eyes. The mechanism of the amicable aid is explained in the viewpoint of the underlying natural image manifold. Furthermore, we investigate the universal amicable aid, i.e., a fixed perturbation can be applied to multiple images to improve their classification results. While it is challenging to find such perturbations, we show that making the decision boundary as perpendicular to the image manifold as possible via training with modified data is effective to obtain a model for which universal amicable perturbations are more easily found.
Juyeop Kim, Jun-Ho Choi, Soobeom Jang, Jong-Seok Lee
ICASSP4
2023 Integer Quantized Learned Image Compression
abstract
Learned image compression (LIC) has shown remarkable improvement compared to traditional methods, but requires increased computational and memory complexity. Network quantization is an effective way to resolve the issue of complexity, but quantization of LIC has not been explored much. In this paper, we propose an integer quantized LIC (IQ-LIC) via static quantization of both weights and activations as integers. We design a quantized convolution layer involving a new Leaky-Clip module. We also propose a squared quantization error loss to help efficient quantization-aware training. Experimental results show that IQ-LIC achieves better rate-distortion performance compared to existing methods.
Geun-Woo Jeon, SeungEun Yu, Jong-Seok Lee
ICIP3
2023 LUT-LIC: Look-Up Table-Assisted Learned Image Compression
SeungEun Yu, Jong-Seok Lee
ICONIP (12)2
2023 Leveraging Feature Interaction for Modeling User-Item Interaction in Recommender Systems
abstract
Challenges in recent recommender systems include how to model high-order feature interaction and how to exploit user-item interaction, particularly for neural network-based recommender systems. While previous approaches have focused only on one aspect, this paper attempts to address both simultaneously by extracting augmented embeddings for users and items with feature interaction and modeling user-item interaction using graph neural networks. Real-world experimental results show that the proposed method outperforms state-of-the-art methods considering one type of interaction.
Soobeom Jang, Junha Park, Jong-Seok Lee
SMC4
2023 Enhancing Graph Structures for Node Classification: An Alternative View on Adversarial Attacks
abstract
Recently, graph neural networks (GNNs) have become a popular approach to deal with machine learning tasks for graph-structured data. To achieve reliable performance with a GNN-based approach, obtaining high-quality graph structures is crucial. However, the graph data in the real-world often contain noise from data themselves or during the collecting procedure, which leads to the performance degradation of GNNs. In this paper, we propose a novel approach to enhance graph structures for performance improvement of GNNs by reversely applying the concept of adversarial attacks on graph data. Experimental results demonstrate the effectiveness of our method in improving performance of GNNs. Furthermore, we investigate the changes in the graph structure induced by our method, taking into account the connectivity of both interclass and intra-class edges and measuring the extent of over-smoothing.
Soobeom Jang, Junha Park, Jong-Seok Lee
SMC3
2023 Maximizing AUC to learn weighted naive Bayes for imbalanced data classification
Taeheung Kim, Jong-Seok Lee
Expert Syst. Appl.2
2023 EEG-Based Emotional Video Classification via Learning Connectivity Structure
abstract
Electroencephalography (EEG) is a useful way to implicitly monitor the user's perceptual state during multimedia consumption. One of the primary challenges for the practical use of EEG-based monitoring is to achieve a satisfactory level of accuracy in EEG classification. Connectivity between different brain regions is an important property for the classification of EEG. However, how to define the connectivity structure for a given task is still an open problem, because there is no ground truth about how the connectivity structure should be in order to maximize the classification performance. In this paper, we propose an end-to-end neural network model for EEG-based emotional video classification, which can extract an appropriate multi-layer graph structure and signal features directly from a set of raw EEG signals and perform classification using them. Experimental results demonstrate that our method yields improved performance in comparison to the existing approaches where manually defined connectivity structures and signal features are used. Furthermore, we show that the graph structure extraction process is reliable in terms of consistency, and the learned graph structures make much sense in the viewpoint of emotional perception occurring in the brain.
Soobeom Jang, Seong-Eun Moon, Jong-Seok Lee
IEEE Trans. Affect. Comput.3
2022 Rethinking Online Knowledge Distillation with Multi-exits
Hojung Lee, Jong-Seok Lee
ACCV (6)2
2022 Joint Global and Local Hierarchical Priors for Learned Image Compression
abstract
Recently, learned image compression methods have out-performed traditional hand-crafted ones including BPG. One of the keys to this success is learned entropy models that estimate the probability distribution of the quantized latent representation. Like other vision tasks, most recent learned entropy models are based on convolutional neural networks (CNNs). However, CNNs have a limitation in modeling long-range dependencies due to their nature of local connectivity, which can be a significant bottleneck in image compression where reducing spatial redundancy is a key point. To overcome this issue, we propose a novel entropy model called Information Transformer (Informer) that exploits both global and local information in a content-dependent manner using an attention mechanism. Our experiments show that Informer improves rate-distortion performance over the state-of-the-art methods on the Kodak and Tecnick datasets without the quadratic computational complexity problem. Our source code is available at https://github.com/naver-ai/informer.
Jun-Hyuk Kim, Byeongho Heo, Jong-Seok Lee
CVPR3
2022 TREND: Truncated Generalized Normal Density Estimation of Inception Embeddings for GAN Evaluation
Junghyuk Lee, Jong-Seok Lee
ECCV (23)2
2022 Deep Image Destruction: Vulnerability of Deep Image-to-Image Models against Adversarial Attacks
abstract
Recently, the vulnerability of deep image classification models to adversarial attacks has been investigated. However, such an issue has not been thoroughly studied for image-to-image tasks that take an input image and generate an output image (e.g., colorization, denoising, deblurring, etc.) This paper presents comprehensive investigations into the vulnerability of deep image-to-image models to adversarial attacks. For five popular image-to-image tasks, 16 deep models are analyzed from various standpoints such as output quality degradation due to attacks, transferability of adversarial examples across different tasks, and characteristics of perturbations. We show that unlike image classification tasks, the performance degradation on image-to-image tasks largely differs depending on various factors, e.g., attack methods and task objectives. In addition, we analyze the effectiveness of conventional defense methods used for classification models in improving the robustness of the image-to-image models.
Jun-Ho Choi, Huan Zhang 0001, Jun-Hyuk Kim, Cho-Jui Hsieh, Jong-Seok Lee
ICPR5
2022 Adversarial Robustness of Flow-based Image Super-Resolution
abstract
This paper investigates the robustness of deep image super-resolution models using normalizing flow against adversarial attacks. Attack methods specific to flow-based super-resolution models are formulated, and the performance and influences of the attacks are analyzed. We show that flow-based super-resolution models are highly vulnerable to attacks, which are even more serious than other super-resolution models. Potential remedies to the vulnerability are also evaluated.
Junha Park, Jun-Ho Choi, Jong-Seok Lee
MMSP3
2022 Resting-State fNIRS Classification Using Connectivity and Convolutional Neural Networks
abstract
Functional near-infrared spectroscopy (fNIRS) is a brain imaging method introduced relatively recently, which is promising to implement brain-computer interfaces. However, there is still a lack of research on fNIRS signal classification, particularly that focusing on improved machine learning techniques for non-motor tasks. In this paper, we propose a novel deep learning method using brain connectivity for resting-state fNIRS signal classification. Our method is based on the powerful modeling capability of the convolutional neural network that learns the brain connectivity patterns residing in the fNIRS signal. In particular, we present a new data augmentation method that can overcome the scarcity of fNIRS data. Experimental results of subject-independent classification of flourishing levels demonstrate the superiority of our approach to conventional approaches. It is also shown that the data augmentation strategy is effective for improving classification performance.
Seohyun Moon, Seong-Eun Moon, Jong-Seok Lee
SMC3
2022 Successive learned image compression: Comprehensive analysis of instability
Jun-Hyuk Kim, Soobeom Jang, Jun-Ho Choi, Jong-Seok Lee
Neurocomputing4
2022 Convergence analysis of connection center evolution and faster clustering
Jaemin Lee 0003, Minseok Han, Jong-Seok Lee
Pattern Recognit.3
2022 Local Critic Training for Model-Parallel Learning of Deep Neural Networks
abstract
In this article, we propose a novel model-parallel learning method, called local critic training, which trains neural networks using additional modules called local critic networks. The main network is divided into several layer groups, and each layer group is updated through error gradients estimated by the corresponding local critic network. We show that the proposed approach successfully decouples the update process of the layer groups for both convolutional neural networks (CNNs) and recurrent neural networks (RNNs). In addition, we demonstrate that the proposed method is guaranteed to converge to a critical point. We also show that trained networks by the proposed method can be used for structural optimization. Experimental results show that our method achieves satisfactory performance, reduces training time greatly, and decreases memory consumption per machine. Code is available at https://github.com/hjdw2/Local-critic-training.
Hojung Lee, Cho-Jui Hsieh, Jong-Seok Lee
IEEE Trans. Neural Networks Learn. Syst.3
2021 Just One Moment: Structural Vulnerability of Deep Action Recognition against One Frame Attack
abstract
The video-based action recognition task has been extensively studied in recent years. In this paper, we study the structural vulnerability of deep learning-based action recognition models against the adversarial attack using the one frame attack that adds an inconspicuous perturbation to only a single frame of a given video clip. Our analysis shows that the models are highly vulnerable against the one frame attack due to their structural properties. Experiments demonstrate high fooling rates and inconspicuous characteristics of the attack. Furthermore, we show that strong universal one frame perturbations can be obtained under various scenarios. Our work raises the serious issue of adversarial vulnerability of the state-of-the-art action recognition models in various perspectives.
Jaehui Hwang, Jun-Hyuk Kim, Jun-Ho Choi, Jong-Seok Lee
ICCV4
2021 Edge Attention Network for Image Deblurring and Super-Resolution
abstract
While deep learning-based single image super-resolution has progressed significantly, super-resolution of images containing blurring artifacts is still challenging. This paper proposes a unified model that can simultaneously perform deblurring and super-resolution for a given image, which is called Edge Attention Network (EAN). Our model employs an attention mechanism using the edge information in order to enhance the sharpness of the super-resolved output image. Experimental results demonstrate that our method outperforms existing super-resolution methods and separate application of deblurring and super-resolution.
Jong-Wook Han, Jun-Ho Choi, Jun-Hyuk Kim, Jong-Seok Lee
SMC4
2021 Ambiguity of objective image quality metrics: A new methodology for performance evaluation
Manri Cheon, Toinon Vigier, Lukas Krasula, Junghyuk Lee, Patrick Le Callet, Jong-Seok Lee
Signal Process. Image Commun.6
2021 Prediction of Car Design Perception Using EEG and Gaze Patterns
abstract
In this paper, we deal with the issue of implicit monitoring of perceptual responses to product design through electroencephalography (EEG) and eye tracking. Four evaluation factors, namely, preference, luxury, complexity, and harmony are considered to investigate how people perceive the car design. In particular, the quantified perceptual responses are predicted based on EEG and gaze data. Average root-mean-square errors of 0.210 and 1.215 are obtained from subject-dependent and subject-independent regressions on a 7-point score scale, respectively, which demonstrates that perception of car design can be predicted via implicit monitoring.
Seong-Eun Moon, Jun-Hyuk Kim, Sun-Wook Kim, Jong-Seok Lee
IEEE Trans. Affect. Comput.4
2021 Wide Color Gamut Image Content Characterization: Method, Evaluation, and Applications
abstract
In this paper, we propose a novel framework to characterize a wide color gamut image content based on perceived quality due to the processes that change color gamut, and demonstrate two practical use cases where the framework can be applied. We first introduce the main framework and implementation details. Then, we provide analysis for understanding of existing wide color gamut datasets with quantitative characterization criteria on their characteristics, where four criteria, i.e., coverage, total coverage, uniformity, and total uniformity, are proposed. Finally, the framework is applied to content selection in a gamut mapping evaluation scenario in order to enhance reliability and robustness of the evaluation results. As a result, the framework fulfils content characterization for studies where quality of experience of wide color gamut stimuli is involved.
Junghyuk Lee, Toinon Vigier, Patrick Le Callet, Jong-Seok Lee
IEEE Trans. Multim.4
2020 Adversarially Robust Deep Image Super-Resolution Using Entropy Regularization
Jun-Ho Choi, Huan Zhang 0001, Jun-Hyuk Kim, Cho-Jui Hsieh, Jong-Seok Lee
ACCV (4)5
2020 Learning Conjunctive Information of Signals in Multi-Sensor Systems
abstract
This paper proposes a novel deep learning method for extraction of the conjunctive information that describes the relationship between signals in multi-sensor systems to enhance the performance of the given classification task. The signals obtained from different sensors included in the multi-sensor systems are closely related. Handcrafted metrics have been used to extract the relationship between the signals in some work, which is hardly optimal for the given task. Our proposed method learns the pair-wise relationship from data to maximize the performance of the given task, which is fully data-driven, multi-aspect, and target-oriented. We demonstrate the effectiveness of the proposed method on a toy example and two real-world problems, i.e., activity recognition using accelerometer signals and emotional video classification using brain signals.
Seong-Eun Moon, Jong-Seok Lee
ECAI2
2020 Dynamic Thresholding for Learning Sparse Neural Networks
abstract
This paper proposes a method called Dynamic Thresholding, which can dynamically adjust the size of deep neural networks by removing redundant weights during training. The key idea is to learn the pruning threshold values applied for weight removal, instead of fixing them manually. We approximate a discontinuous pruning function with a differentiable form involving the thresholds, which can be optimized via the gradient descent learning procedure. While previous sparsity-promoting methods perform pruning with manually determined thresholds, our method can directly obtain a sparse network at each training iteration and thus does not need a trial-and-error process to choose proper threshold values. We examine the performance of the proposed method on the image classification tasks including MNIST, CIFAR10, and ImageNet. It is demonstrated that our method achieves competitive results with existing methods and, at the same time, requires smaller numbers of training iterations in comparison to other approaches based on train-prune-retrain cycles.
Jong-Seok Lee
ECAI2
2020 Srzoo: An Integrated Repository For Super-Resolution Using Deep Learning
abstract
Deep learning-based image processing algorithms, including image super-resolution methods, have been proposed with significant improvement in performance in recent years. However, their implementations and evaluations are dispersed in terms of various deep learning frameworks and various evaluation criteria. In this paper, we propose an integrated repository for the super-resolution tasks, named SRZoo, to provide state-of-the-art super-resolution models in a single place. Our repository offers not only converted versions of existing pre-trained models, but also documentation and toolkits for converting other models. In addition, SRZoo provides platform-agnostic image reconstruction tools to obtain super-resolved images and evaluate the performance in place. It also brings the opportunity of extension to advanced image-based researches and other image processing models. The software, documentation, and pre-trained models are publicly available on GitHub1.
Jun-Ho Choi, Jun-Hyuk Kim, Jong-Seok Lee
ICASSP3
2020 Efficient Deep Learning-Based Lossy Image Compression Via Asymmetric Autoencoder and Pruning
abstract
Recently, deep learning-based lossy image compression methods have been proposed. However, their efficiency in terms of storage and computational costs has not been addressed adequately. In this paper, we propose efficient lossy image compression methods based on asymmetric autoencoder and decoder pruning. Experimental results demonstrate the effectiveness of our methods.
Jun-Hyuk Kim, Jun-Ho Choi, Jaehyuk Chang, Jong-Seok Lee
ICASSP4
2020 Instability of Successive Deep Image Compression
abstract
Successive image compression refers to the process of repeated encoding and decoding of an image. It frequently occurs during sharing, manipulation, and re-distribution of images. While deep learning-based methods have made significant progress for single-step compression, thorough analysis of their performance under successive compression has not been conducted. In this paper, we conduct comprehensive analysis of successive deep image compression. First, we introduce a new observation, instability of successive deep image compression, which is not observed in JPEG, and discuss causes of the instability. Then, we conduct a successive image compression benchmark for the state-of-the-art deep learning-based methods, and analyze the factors that affect the instability in a comparative manner. Finally, we propose a new loss function for training deep compression models, called feature identity loss, to mitigate the instability of successive deep image compression.
Jun-Hyuk Kim, Soobeom Jang, Jun-Ho Choi, Jong-Seok Lee
ACM Multimedia4
2020 Deep learning-based image super-resolution considering quantitative and perceptual quality
Jun-Ho Choi, Jun-Hyuk Kim, Manri Cheon, Jong-Seok Lee
Neurocomputing4
2020 MAMNet: Multi-path adaptive modulation network for image super-resolution
Jun-Hyuk Kim, Jun-Ho Choi, Manri Cheon, Jong-Seok Lee
Neurocomputing4
2020 Emotional EEG classification using connectivity features and convolutional neural networks
Seong-Eun Moon, Chun-Jui Chen, Cho-Jui Hsieh, Jane-Ling Wang, Jong-Seok Lee
Neural Networks5
2020 Objectivity and Subjectivity in Aesthetic Quality Assessment of Digital Photographs
abstract
Automatic prediction of the aesthetic quality of a photograph has been an important research problem in image processing and computer vision. While assessing the aesthetic quality of a photograph by human is highly subjective, most existing studies have considered only objective (or general) opinion of multiple viewers. In this paper, we provide a comprehensive investigation of the issue of subjectivity in aesthetic quality assessment using a large-scale database containing photos, user ratings, and user comments. First, we analyze how the mean aesthetic quality level and the level of subjectivity have evolved over time. Second, we examine the feasibility of automatic prediction of the level of subjectivity based on visual features, and identify which features are effective for the prediction. Third, we analyze the users' comments given to photos to understand the sources of subjectivity of aesthetic quality rating. Our results show that several factors are simultaneously involved in determining the level of subjectivity of a photo, but it can be predicted with reasonable accuracy. We believe that our results provide insight toward personalized aesthetic photo applications.
Jun-Ho Choi, Jong-Seok Lee
IEEE Trans. Affect. Comput.3
2019 Evaluating Robustness of Deep Image Super-Resolution Against Adversarial Attacks
abstract
Single-image super-resolution aims to generate a high-resolution version of a low-resolution image, which serves as an essential component in many image processing applications. This paper investigates the robustness of deep learning-based super-resolution methods against adversarial attacks, which can significantly deteriorate the super-resolved images without noticeable distortion in the attacked low-resolution images. It is demonstrated that state-of-the-art deep super-resolution methods are highly vulnerable to adversarial attacks. Different levels of robustness of different methods are analyzed theoretically and experimentally. We also present analysis on transferability of attacks, and feasibility of targeted attacks and universal attacks.
Jun-Ho Choi, Huan Zhang 0001, Jun-Hyuk Kim, Cho-Jui Hsieh, Jong-Seok Lee
ICCV5
2019 Local Critic Training of Deep Neural Networks
abstract
This paper proposes a novel approach to train deep neural networks by unlocking the layer-wise dependency of backpropagation training. The approach employs additional modules called local critic networks besides the main network model to be trained, which are used to obtain error gradients without complete feedforward and backward propagation processes. We propose a cascaded learning strategy for these local networks. In addition, the approach is also useful from multi-model perspectives, including structural optimization of neural networks, computationally efficient progressive inference, and ensemble classification for performance improvement. Experimental results show the effectiveness of the proposed approach and suggest guidelines for determining appropriate algorithm parameters.
Hojung Lee, Jong-Seok Lee
IJCNN2
2018 Eeg-Based Video Identification Using Graph Signal Modeling and Graph Convolutional Neural Network
abstract
This paper proposes a novel graph signal-based deep learning method for electroencephalography (EEG) and its application to EEG-based video identification. We present new methods to effectively represent EEG data as signals on graphs, and learn them using graph convolutional neural networks. Experimental results for video identification using EEG responses obtained while watching videos show the effectiveness of the proposed approach in comparison to existing methods. Effective schemes for graph signal representation of EEG are also discussed.
Soobeom Jang, Seong-Eun Moon, Jong-Seok Lee
ICASSP3
2018 Convolutional Neural Network Approach for Eeg-Based Emotion Recognition Using Brain Connectivity and its Spatial Information
abstract
Emotion recognition based on electroencephalography (EEG) has received attention as a way to implement human-centric services. However, there is still much room for improvement, particularly in terms of the recognition accuracy. In this paper, we propose a novel deep learning approach using convolutional neural networks (CNNs) for EEG-based emotion recognition. In particular, we employ brain connectivity features that have not been used with deep learning models in previous studies, which can account for synchronous activations of different brain regions. In addition, we develop a method to effectively capture asymmetric brain activity patterns that are important for emotion recognition. Experimental results confirm the effectiveness of our approach.
Seong-Eun Moon, Soobeom Jang, Jong-Seok Lee
ICASSP3
2018 A Perception-Based Framework for Wide Color Gamut Content Selection
abstract
Considering the content dependence of the perceived quality, selection of source content can significantly influence the results of studies related to Quality of Experience. In this paper, we propose an automated content selection method towards wide color gamut stimuli. The framework enables to objectively characterize the content according to its perceptual properties and thus allows to select a representative, diverse, and challenging subsets for various studies. Experimental results validate the reliability and robustness of the proposed framework.
Junghyuk Lee, Toinon Vigier, Patrick Le Callet, Jong-Seok Lee
ICIP4
2018 Evaluation of preference of multimedia content using deep neural networks for electroencephalography
abstract
Evaluation of quality of experience (Qo E) based on electroencephalography (EEG) has received great attention due to its capability of real-time Qo E monitoring of users. However, it still suffers from rather low recognition accuracy. In this paper, we propose a novel method using deep neural networks toward improved modeling of EEG and thereby improved recognition accuracy. In particular, we aim to model spatio-temporal characteristics relevant for QoE analysis within learning models. The results demonstrate the effectiveness of the proposed method.
Seong-Eun Moon, Soobeom Jang, Jong-Seok Lee
QoMEX3
2018 On the Repeatability of EEG-Based Image Quality Assessment
abstract
Electroencephalography (EEG) has attracted much attention because it allows to monitor user states in real time and is applicable to applications for improvement of multimedia user experience. The repeatability of EEG-based perceptual response analysis is critical for the reliability of such applications, which has not been sufficiently addressed in previous studies. In this paper, we evaluate the repeatability of EEG-based image quality assessment. We repeatedly perform the same experiment for three successive days. Then, we design a classification system that discriminates the quality of images based on EEG, whose performance is examined across the data collected in different days. We reveal that the repeatability of image quality evaluation using EEG decreases as time goes and the memory effect intensifies the loss of repeatability.
Jaehui Hwang, Seong-Eun Moon, Jong-Seok Lee
SMC3
2018 Subjective and Objective Quality Assessment of Compressed 4K UHD Videos for Immersive Experience
abstract
We present a study of subjective and objective quality assessment of compressed 4K ultra-high-definition (UHD) videos in an immersive viewing environment. First, we conduct a subjective quality evaluation experiment for 4K UHD videos compressed by three state-of-the-art video coding techniques, i.e., Advanced Video Coding, High Efficiency Video Coding, and VP9. In particular, we aim at investigating added values of UHD over conventional high definition (HD) in terms of perceptual quality. The results are systematically analyzed in various viewpoints, such as coding scheme, bitrate, and video content. Second, existing state-of-the-art objective quality assessment techniques are benchmarked using the subjective data in order to investigate their validity and limitation for 4K UHD videos. Finally, the video and subjective data are made publicly available for further research by the research community.
Manri Cheon, Jong-Seok Lee
IEEE Trans. Circuits Syst. Video Technol.2
2018 Music Popularity: Metrics, Characteristics, and Audio-Based Prediction
abstract
Understanding music popularity is important not only for the artists who create and perform music but also for the music-related industry. It has not been studied well how music popularity can be defined, what its characteristics are, and whether it can be predicted, which are addressed in this paper. We first define eight popularity metrics to cover multiple aspects of popularity. Then, the analysis of each popularity metric is conducted with long-term real-world chart data to deeply understand the characteristics of music popularity in the real world. We also build classification models for predicting popularity metrics using acoustic data. In particular, we focus on evaluating features describing music complexity together with other conventional acoustic features including MPEG-7 and Mel-frequency cepstral coefficient (MFCC) features. The results show that, although room still exists for improvement, it is feasible to predict the popularity metrics of a song significantly better than random chance based on its audio signal, particularly using both the complexity and MFCC features.
Junghyuk Lee, Jong-Seok Lee
IEEE Trans. Multim.2
2017 Influence of Video Quality on Multi-view Activity Recognition
abstract
This paper presents a study that evaluates the performance of multi-view human activity recognition with videos having degraded quality. For the activity recognition models, a support vector machine-based approach using spatiotemporal features and a deep learning-based approach using convolutional and recurrent layers are built. We investigate the recognition performance of the two models with respect to the bitrate of the compressed videos and the peak signal-to-noise ratio of the videos corrupted by additive Gaussian random noise. We analyze the robustness of the models for the degraded videos.
Jun-Ho Choi, Manri Cheon, Jong-Seok Lee
ISM3
2017 Instance categorization by support vector machines to adjust weights in AdaBoost for imbalanced data classification
Wonji Lee, Chi-Hyuck Jun, Jong-Seok Lee
Inf. Sci.3
2017 Blind single image super resolution with low computational complexity
Jong-Seok Lee
Multim. Tools Appl.2
2017 Implicit Analysis of Perceptual Multimedia Experience Based on Physiological Response: A Review
abstract
The exponential growth of popularity of multimedia has led to needs for user-centric adaptive applications that manage multimedia content more effectively. Implicit analysis, which examines users' perceptual experience of multimedia by monitoring physiological or behavioral cues, has potential to satisfy such demands. Particularly, physiological signals categorized into cerebral physiological signals (electroencephalography, functional magnetic resonance imaging, and functional near-infrared spectro-scopy) and peripheral physiological signals (heart rate, respiration, skin temperature, etc.) have recently received attention along with notable development of wearable physiological sensors. In this paper, we review existing studies on physiological signal analysis exploring perceptual experience of multimedia. Furthermore, we discuss current trends and challenges.
Seong-Eun Moon, Jong-Seok Lee
IEEE Trans. Multim.2
2016 Ambiguity-based evaluation of objective quality metrics for image compression
abstract
While results of subjective quality assessment are represented by mean opinion scores and corresponding confidence intervals, the output of an objective quality metric for a given stimulus is only a single estimated quality level. Accordingly, the performance of a metric is evaluated by measuring the accuracy of its outputs with respect to the corresponding subjective scores. However, the concept of the ambiguity interval for objective quality has been raised recently. In this paper, we propose to consider not only the accuracy but also the ambiguity of objective quality metrics for performance evaluation. In particular, we conduct benchmarking of the seven state-of-the-art image quality metrics for images compressed with JPEG and JPEG2000. It is demonstrated that the best metric in terms of accuracy may not be the best in terms of ambiguity.
Manri Cheon, Jong-Seok Lee
QoMEX2
2016 Evaluation of objective quality metrics for multidimensional video scalability
Manri Cheon, Jong-Seok Lee
J. Vis. Commun. Image Represent.2
2016 Subjective and objective quality assessment of videos in error-prone network environments
Soo-Jin Kim, Chan-Byoung Chae, Jong-Seok Lee
Multim. Tools Appl.3
2016 Modeling immersive media experiences by sensing impact on subjects
Eleni Kroupi, Philippe Hanhart, Jong-Seok Lee, Martin Rerábek, Touradj Ebrahimi
Multim. Tools Appl.3
2016 On Evaluating Perceptual Quality of Online User-Generated Videos
abstract
This paper deals with the issue of the perceptual quality evaluation of user-generated videos shared online, which is an important step toward designing video-sharing services that maximize users' satisfaction in terms of quality. We first analyze viewers' quality perception patterns by applying graph analysis techniques to subjective rating data. We then examine the performance of existing state-of-the-art objective metrics for the quality estimation of user-generated videos. In addition, we investigate the feasibility of metadata accompanied with videos in online video-sharing services for quality estimation. Finally, various issues in the quality assessment of online user-generated videos are discussed, including difficulties and opportunities.
Soobeom Jang, Jong-Seok Lee
IEEE Trans. Multim.2
2015 Gaze Analysis of Avatar-based Navigation with Different Perspectives in 3D Virtual Space
abstract
This paper explores the relations between perspectives of navigation and visual perception in 3D virtual space, by analyzing avatar-based navigation with the eye-gaze data. We examine how different perspectives and types of avatars affect the users' scopes of visual perception within 3D virtual environments. Throughout this research, we attempt to draw possible connections between the perspectives and cognitive patterns of visual perception. We propose that manipulating perspectives of avatars or those of users has immediate effects on the users' scopes of visual perception and patterns of visual attention.
Jooyeon Lee, Manri Cheon, Seong-Eun Moon, Jong-Seok Lee
HAI4
2015 EEG Analysis on 3D Navigation in Virtual Realty with Different Perspectives
abstract
In this paper we explore the relations between perspectives of navigation and electroencephalogram (EEG) in 3D virtual space. We analyze three types of navigation with EEG recordings and examine how the perspectives affect the users' electrical activities in their brains. Via a small-scale experiment, we find that the influence of peripersonal space is altered by the perspective, and it can be observed via EEG monitoring. These results have interesting implications on virtual reality applications where a sense of agency, or a peripersonal task takes important roles.
Jooyeon Lee, Seong-Eun Moon, Manri Cheon, Jong-Seok Lee
HAI4
2015 Automated Video Editing for Aesthetic Quality Improvement
abstract
In these days, a large number of videos is taken by various kinds of handheld devices, but many of them have poor aesthetic quality. In this paper, we present an automated video editing system that uses the shot length, camera motion, and color distribution as key aesthetic features. Given an amateur video, our system computes the original unrefined camera motion as homography and tries to remove some unreliable frames, which consequently splits the video into several shots. It then applies enhancement processes, including reconstruction of the overall camera motions and harmonization of color distributions. We apply our method to some amateur videos and evaluate the results through a subjective test. It is demonstrated that reducing the shot length in our method is a key point of editing that can lead enhanced satisfaction by viewers for the edited videos.
Jun-Ho Choi, Jong-Seok Lee
ACM Multimedia2
2015 Subjectivity in Aesthetic Quality Assessment of Digital Photographs: Analysis of User Comments
abstract
While most of the existing work in aesthetic image quality assessment focuses on the overall (or average) opinion of users, this paper raises the issue of subjectivity (or taste) of aesthetic quality. We argue that subjectivity differs among different images, and investigate what causes such difference. We first analyze statistics of the user ratings of photos in a photo contest website, DPChallenge, in the viewpoint of average and standard deviation values of the ratings. Then, more importantly, we analyze the users' comments in order to identify sources contributing to subjectivity. When considering the importance of personalization in photo applications, we believe that our findings will be a valuable first step in the relevant future research.
Jun-Ho Choi, Jong-Seok Lee
ACM Multimedia3
2015 EEG Connectivity Analysis in Perception of Tone-mapped High Dynamic Range Videos
abstract
High dynamic range (HDR) imaging has attracted attention as a new technology for immersive multimedia experience. In comparison to conventional low dynamic range (LDR) contents, HDR contents are expected to provide better quality of experience (QoE). In this paper, we investigate implicit QoE measurement of tone-mapped HDR videos by using connectivity-based EEG features that convey information about simultaneous activations of different brain regions and thus can explain better the cognitive process than the conventional features using single channel powers. Through the experiment classifying EEG signals into tone-mapped HDR and LDR, it is shown that the connectivity features, particularly those representing directed information flows between brain regions, are effective in both subject-dependent and subject-independent scenarios.
Seong-Eun Moon, Jong-Seok Lee
ACM Multimedia2
2015 A precise ranking method for outlier detection
Jihyun Ha, Seulgi Seok, Jong-Seok Lee
Inf. Sci.3
2015 Temporal resolution vs. visual saliency in videos: Analysis of gaze patterns and evaluation of saliency models
Manri Cheon, Jong-Seok Lee
Signal Process. Image Commun.2
2014 Quality assessment of on-line videos using metadata
abstract
As video consumption becomes popular, demand for high quality of experience of consumed videos is also increasing. While online video sharing is a popular video application, quality assessment of videos in such an application is challenging due to lack of reference videos and simultaneous involvement of diverse quality factors. In this paper, we take advantage of additional information of online videos, i.e., metadata, and explore the extent to which video quality can be estimated from metadata. Subjective quality assessment using crowdsourcing is conducted, based on which metadata-based quality models are constructed. It is shown that the estimated quality scores show fairly high correlation with the subjective quality.
Chul-Hee Han, Jong-Seok Lee
ICASSP2
2014 Predicting subjective sensation of reality during multimedia consumption based on EEG and peripheral physiological signals
abstract
Sensation of reality refers to the ability of users to feel present in a multimedia experience. As 3D technologies target to provide more immersive and higher quality multimedia experiences, it is important to understand Quality of Experience (QoE) and sensation of reality. Recently, there have been efforts to measure brain activity in order to understand implicitly QoE for various multimedia contents. However, brain activity accounting for sensation of reality has not been adequately investigated. The goal of this paper is twofold. First, we investigate how various aspects, such as perceived quality, perceived depth, and content preference affect subjective sensation of reality through explicit subjective ratings. Second, we construct subjective classification systems to predict sensation of reality from multimedia experiences based on electroencephalography (EEG) and peripheral physiological signals such as heart rate and respiration.
Eleni Kroupi, Philippe Hanhart, Jong-Seok Lee, Martin Rerábek, Touradj Ebrahimi
ICME3
2014 Objective Quality Comparison of 4K UHD and Up-Scaled 4K UHD Videos
abstract
In this paper, we perform objective quality comparison between 4K ultra-high definition (UHD) and up-scaled 4K UHD videos. We aim at investigating added values of 4K UHD over high definition (HD). We examine the quality of the two types of videos at the same bitrate conditions for two compression standards, AVC and HEVC. Using our own 4K UHD video data, we generate test sequences having various quality levels and bitrates using AVC and HEVC. Objective quality comparison is performed using multi-scale structural similarity (MS-SSIM), visual information fidelity (VIF), and visual signal-to-noise ratio (VSNR) metrics. The results show that superiority between the two UHD versions changes according to the bitrate, where content characteristics and the codec also have notable influences.
Manri Cheon, Jong-Seok Lee
ISM2
2014 On Unequal Power Allocation for Video Communications Using Scalable Video Coding in Massive MIMO Systems
abstract
In this paper, we investigate the effectiveness of radio resource allocation with H.264/AVC scalable video coding (SVC) in massive multiple-input multiple-output (MIMO) systems. During transmission of SVC-encoded videos in error prone network environments, packet losses of a certain layer may cause a severe reduction of quality or even prevent 1n2 correct decoding of other layers, since SVC layers are highly interdependent. It is generally said that the most important information should be preserved with the highest priority from packet losses. To investigate the validity of such a scheme, we apply unequal radio power allocation for SVC layers in massive MIMO systems. We first show that the error rate changes drastically with respect to transmit power in massive MIMO systems. Then, we conduct simulations to analyze the overall quality in terms of both the traditional peak signal-to-noise ratio (PSNR) and the structural similarity (SSIM) considering perceived quality. Our results show that layer prioritization in massive MIMO systems is not always beneficial in terms of quality and the content characteristics need to be considered for effective power allocation.
Soo-Jin Kim, Chan-Byoung Chae, Jong-Seok Lee
ISM3
2014 EvoTunes: Crowdsourcing-Based Music Recommendation
Jun-Ho Choi, Jong-Seok Lee
MMM (2)2
2014 Data-driven integration of multiple sentiment dictionaries for lexicon-based sentiment classification of product reviews
Heeryon Cho, Songkuk Kim, Jongseo Lee, Jong-Seok Lee
Knowl. Based Syst.4
2014 Robust outlier detection using the instability factor
Jihyun Ha, Seulgi Seok, Jong-Seok Lee
Knowl. Based Syst.3
2014 Visual-speech-pass filtering for robust automatic lip-reading
Jong-Seok Lee
Pattern Anal. Appl.1
2014 On Designing Paired Comparison Experiments for Subjective Multimedia Quality Assessment
abstract
This paper investigates the issue of designing paired comparison-based subjective quality assessment experiments for reliable results. In particular, the convergence behavior of the quality scores estimated from paired comparison results is considered. Via an extensive computer simulation experiment, the estimation performance in terms of the root mean squared error, the rank order correlation coefficient, and the change of the estimated scores with respect to the number of subjects are mathematically modeled. Furthermore, it is confirmed that the models coincide with the theoretical convergence behavior. Issues such as the effect of human errors and the underlying distribution of the true quality scores are also examined.
Jong-Seok Lee
IEEE Trans. Multim.1
2013 Enhancing Lexicon-Based Review Classification by Merging and Revising Sentiment Dictionaries
Heeryon Cho, Jong-Seok Lee, Songkuk Kim
IJCNLP2
2013 Paired comparison for subjective multimedia quality assessment: Theory and practice
abstract
For subjective multimedia quality assessment, paired comparison has been used less frequently than other methodologies such as single stimulus and double stimulus methodologies, mainly due to its increased time complexity. However, it has been shown that paired comparison provides improved discriminability between stimuli, which consequently improves reliability of the test results. Analyzing comparison results involves several challenges in quality score computation, outlier detection, etc. This paper presents theoretical backgrounds on these challenges and use cases where paired comparison are successfully applied. In particular, the issue of the increased time complexity is discussed along with recently proposed solutions.
Jong-Seok Lee
ISCAS1
2013 Gaze pattern analysis for video contents with different frame rates
abstract
This paper presents a study investigating the viewing behavior of human subjects for video contents having different frame rates. Frame rate variability arises when temporal video scalability is considered for adaptive video transmission, and the gaze pattern variation due to the frame rate variability would eventually affect the visual perception, which needs to be considered during perceptual optimization of such a system. We design an eye-tracking experiment using several high definition contents having a wide range of content characteristics. By comparing the gaze points for a normal frame rate condition and a low frame rate condition, it is shown that, although the overall viewing pattern remains quite similar, statistically significant difference is also observed for some time intervals. The difference is analyzed in terms of two factors, namely, overall gaze paths and subject-wise variability.
Manri Cheon, Jong-Seok Lee
VCIP2
2013 A meta-learning approach for determining the number of clusters with consideration of nearest neighbors
Jong-Seok Lee, Sigurdur Ólafsson
Inf. Sci.1
2013 Paired comparison-based subjective quality assessment of stereoscopic images
Jong-Seok Lee, Lutz Goldmann, Touradj Ebrahimi
Multim. Tools Appl.1
2012 TrueSkill-Based Pairwise Coupling for Multi-class Classification
Jong-Seok Lee
ICANN (2)1
2012 Comparison of objective quality metrics on the scalable extension of H.264/AVC
abstract
Due to the recent interest in video content delivery applications in networked environments, automatic quality assessment of scalable videos is becoming an important issue. This paper presents a study comparing the performance of objective metrics developed for quality assessment of scalable videos. A database containing video sequences produced by the scalable extension of H.264/AVC and corresponding subjective ratings is used for benchmarking. We aim at investigating the maximal capabilities of the metrics via optimization of their parameters on the database. Experimental results show that they outperform the peak signal-to-noise ratio (PSNR) and a correlation coefficient higher than 0.9 can be achieved, but there is still room for improvement in order to facilitate reliable objective quality assessment.
Jong-Seok Lee
ICIP1
2012 Shilling Attack Detection - A New Approach for a Trustworthy Recommender System
abstract
Recommender systems rely on the opinions of many users to predict the preferences of potential customers. These systems have been broadly used to make quality recommendations to increase sales. However, recommender systems are vulnerable to even small data inputs of malicious information. Inappropriate products can be offered to users by injecting a few unscrupulous “shilling” profiles into the recommender system. This research proposes to identify a cluster of profiles by focusing on “filler” ratings. We examine a number of properties of such profiles, followed by empirical evidence and detailed analysis of various characteristics of the shilling attacks. We then propose a hybrid two-phase procedure for shilling attack detection. First, a multidimensional scaling approach is adopted to identify distinct behaviors that help to detect and secure the recommendation activities. Clustering-based methods are subsequently proposed to discriminate attack users. Experimental studies are conducted to show the effectiveness of the proposed method.
Jong-Seok Lee
INFORMS J. Comput.1
2012 Geotag propagation in social networks based on user trust model
Peter Vajda, Jong-Seok Lee, Lutz Goldmann, Touradj Ebrahimi
Multim. Tools Appl.3
2012 DEAP: A Database for Emotion Analysis ;Using Physiological Signals
abstract
We present a multimodal data set for the analysis of human affective states. The electroencephalogram (EEG) and peripheral physiological signals of 32 participants were recorded as each watched 40 one-minute long excerpts of music videos. Participants rated each video in terms of the levels of arousal, valence, like/dislike, dominance, and familiarity. For 22 of the 32 participants, frontal face video was also recorded. A novel method for stimuli selection is proposed using retrieval by affective tags from the last.fm website, video highlight detection, and an online assessment tool. An extensive analysis of the participants' ratings during the experiment is presented. Correlates between the EEG signal frequencies and the participants' ratings are investigated. Methods and results are presented for single-trial classification of arousal, valence, and like/dislike ratings using the modalities of EEG, peripheral physiological signals, and multimedia content analysis. Finally, decision fusion of the classification results from different modalities is performed. The data set is made publicly available and we encourage other researchers to use it for testing their own affective state estimation methods.
Sander Koelstra, Christian Mühl, Mohammad Soleymani 0001, Jong-Seok Lee, Ashkan Yazdani, Touradj Ebrahimi, Thierry Pun, Anton Nijholt, Ioannis Patras
IEEE Trans. Affect. Comput.4
2012 Affect recognition based on physiological changes during the watching of music videos
abstract
Assessing emotional states of users evoked during their multimedia consumption has received a great deal of attention with recent advances in multimedia content distribution technologies and increasing interest in personalized content delivery. Physiological signals such as the electroencephalogram (EEG) and peripheral physiological signals have been less considered for emotion recognition in comparison to other modalities such as facial expression and speech, although they have a potential interest as alternative or supplementary channels. This article presents our work on: (1) constructing a dataset containing EEG and peripheral physiological signals acquired during presentation of music video clips, which is made publicly available, and (2) conducting binary classification of induced positive/negative valence, high/low arousal, and like/dislike by using the aforementioned signals. The procedure for the dataset acquisition, including stimuli selection, signal acquisition, self-assessment, and signal processing is described in detail. Especially, we propose a novel asymmetry index based on relative wavelet entropy for measuring the asymmetry in the energy distribution of EEG signals, which is used for EEG feature extraction. Then, the classification systems based on EEG and peripheral physiological signals are presented. Single-trial and single-run classification results indicate that, on average, the performance of the EEG-based classification outperforms that of the peripheral physiological signals. However, the peripheral physiological signals can be considered as a good alternative to EEG signals in the case of assessing a user's preference for a given music video clip (like/dislike) since they have a comparable performance to EEG signals while being more easily measured.
Ashkan Yazdani, Jong-Seok Lee, Jean-Marc Vesin, Touradj Ebrahimi
ACM Trans. Interact. Intell. Syst.2
2011 Audio-visual synchronization recovery in multimedia content
abstract
This paper proposes a method recovering audio-visual synchronization of multimedia content. It exploits the correlation between the acoustic and the visual signals in order to estimate the audio-visual drift existing in the content. By shifting the audio signal relative to the visual signal, the estimation of the drift is obtained by searching for the shift producing the maximal audio-visual correlation. We consider two correlation measures, namely, mutual information and canonical correlation, and compare their performance. Experimental results demonstrate that the method using the canonical correlation is effective in recovering the audio-visual synchronization for both speech and non-speech sequences.
Jong-Seok Lee, Touradj Ebrahimi
ICASSP1
2011 A new analysis method for paired comparison and its application to 3D quality assessment
abstract
Among various subjective quality evaluation methodologies, paired comparison has the advantage of improved simplicity of the subjects' evaluation task due to simplified rating scales and direct comparison of two stimuli. Thus, it may lead to more reliable results when individual quality levels are difficult to define, quality differences between stimuli are small or multiple quality factors are involved. This paper proposes a new method to analyze results of paired comparison-based subjective tests. By assuming that ties convey information about significant differences between two stimuli being compared, the confidence intervals for the quality scores are estimated using a maximum likelihood criterion, which enables us to intuitively examine the significance of quality score differences. We describe the complete test methodology including the test procedure, outlier detection and score analysis applied to quality assessment of 3D images acquired using varying camera distances. Experimental results demonstrate the usefulness of the proposed analysis method, as well as the enhanced quality discriminability of the paired comparison methodology in comparison to the conventional single stimulus methodology.
Jong-Seok Lee, Lutz Goldmann, Touradj Ebrahimi
ACM Multimedia1
2011 Motion parallax based restitution of 3D images on legacy consumer mobile devices
abstract
While 3D display technologies are already widely available for cinema and home or corporate use, only a few portable devices currently feature 3D display capabilities. Moreover, the large majority of 3D display solutions rely on binocular perception. In this paper, we study the alternative methods for restitution of 3D images on conventional 2D displays and analyze their respective performance. This particularly includes the extension of wiggle stereoscopy for portable devices which relies on motion parallax as an additional depth cue. The goal of this paper is to compare two different 3D display techniques, the anaglyph method which provides binocular depth cues and a method based on motion parallax, and to show that the motion parallax based approach to present 3D images on consumer 2D portable screen is an equivalent way in comparison to the above mentioned and well-known anaglyph method. The subsequently conducted subjective quality tests show that viewers even prefer wiggle over anaglyph stereoscopy mainly due to a better color reproduction and a comparable depth perception.
Martin Rerábek, Lutz Goldmann, Jong-Seok Lee, Touradj Ebrahimi
MMSP3
2011 Data clustering by minimizing disconnectivity
Jong-Seok Lee, Sigurdur Ólafsson
Inf. Sci.1
2011 Efficient video coding based on audio-visual focus of attention
Jong-Seok Lee, Francesca De Simone, Touradj Ebrahimi
J. Vis. Commun. Image Represent.1
2011 Towards high efficiency video coding: Subjective evaluation of potential coding technologies
Francesca De Simone, Lutz Goldmann, Jong-Seok Lee, Touradj Ebrahimi
J. Vis. Commun. Image Represent.3
2011 Subjective Quality Evaluation via Paired Comparison: Application to Scalable Video Coding
abstract
Scalable video coding is a powerful solution for content delivery in many interactive multimedia services due to its adaptability to varying terminal and network constraints. In order to successfully exploit such adaptability, it is necessary to understand users' preference among various scalability options and consequently develop an optimal bit rate adaptation strategy. In this paper, we present a study of subjective quality assessment of scalable video coding, which investigates the influence of the combination of scalability options on perceived quality with the goal of providing guidelines for an adaptive strategy that selects the optimal combination for a given bandwidth constraint. In particular, the study is based on paired comparison of stimuli that is suitable for our goal due to its simplicity and easiness. We propose a new method, called Paired Evaluation via Analysis of Reliability (PEAR), which analyzes paired comparison results and produces not only quality scores but also intuitive measures of confidence of the scores for significance analysis. Results and analysis of extensive subjective tests for two different scalable video codecs and high definition contents are described, from which general consistent conclusions are drawn. The video and subjective data used in the paper are publicly available to the research community.
Jong-Seok Lee, Francesca De Simone, Touradj Ebrahimi
IEEE Trans. Multim.1
2010 Temporal synchronization in stereoscopic video: Influence on quality of experience and automatic asynchrony detection
abstract
In this paper, we analyze the influence of temporal asynchrony on the subjective quality of stereoscopic video. Based on our recently created 3D video database, different levels of asynchrony were simulated and a comprehensive subjective test was conducted to determine the associated degradations in quality of experience. Furthermore, we developed a method to detect asynchrony between left and right video streams based on canonical correlation analysis. Experiments demonstrate the robustness of this method with respect to different amounts of asynchrony and scene depth, which makes it suitable to predict quality of experience or automatic resynchronization.
Lutz Goldmann, Jong-Seok Lee, Touradj Ebrahimi
ICIP2
2010 Subjective evaluation of scalable video coding for content distribution
abstract
This paper investigates the influence of the combination of the scalability parameters in scalable video coding (SVC) schemes on the subjective visual quality. We aim at providing guidelines for an adaptation strategy of SVC that can select the optimal scalability options for resource-constrained networks. Extensive subjective tests are conducted by using two different scalable video codecs and high definition contents. The results are analyzed with respect to five dimensions, namely, codec, content, spatial resolution, temporal resolution, and frame quality.
Jong-Seok Lee, Francesca De Simone, Naeem Ramzan, Zhijie Zhao, Engin Kurutepe, Thomas Sikora, Jörn Ostermann, Ebroul Izquierdo, Touradj Ebrahimi
ACM Multimedia1
2010 Hybrid Simulated Annealing and Its Application to Optimization of Hidden Markov Models for Visual Speech Recognition
abstract
We propose a novel stochastic optimization algorithm, hybrid simulated annealing (SA), to train hidden Markov models (HMMs) for visual speech recognition. In our algorithm, SA is combined with a local optimization operator that substitutes a better solution for the current one to improve the convergence speed and the quality of solutions. We mathematically prove that the sequence of the objective values converges in probability to the global optimum in the algorithm. The algorithm is applied to train HMMs that are used as visual speech recognizers. While the popular training method of HMMs, the expectation-maximization algorithm, achieves only local optima in the parameter space, the proposed method can perform global optimization of the parameters of HMMs and thereby obtain solutions yielding improved recognition performance. The superiority of the proposed algorithm to the conventional ones is demonstrated via isolated word recognition experiments.
Jong-Seok Lee, Cheol Hoon Park
IEEE Trans. Syst. Man Cybern. Part B1
2009 Two-Level Bimodal Association for Audio-Visual Speech Recognition
Jong-Seok Lee, Touradj Ebrahimi
ACIVS1
2009 Video coding based on audio-visual attention
abstract
This paper proposes an efficient video coding method based on audio-visual attention, which is motivated by the fact that cross-modal interaction significantly affects humans' perception of multimedia content. First, we propose an audio-visual source localization method to locate the sound source in a video sequence. Then, its result is used for applying spatial blurring to video frames in order to reduce redundant high-frequency information and achieve coding efficiency. We demonstrate the effectiveness of the proposed method for H.264/AVC coding along with the results of a subjective evaluation.
Jong-Seok Lee, Francesca De Simone, Touradj Ebrahimi
ICME1
2009 Efficient video coding in H.264/AVC by using audio-visual information
abstract
This paper proposes an efficient video coding method which utilizes audio-visual information, based on the observation that sound-emitting regions in a video sequence attract observer's attention. The regions responsible for the sound are identified by an audio-visual source localization algorithm. Then, the result is used for encoding different regions in the scene with different quality in such a way that a region far from the sound source is coded with a lesser quality than the sound-emitting regions. This is implemented by assigning different quantization parameter values for different regions in H.264/AVC. Experimental results demonstrate the effectiveness of the proposed approach.
Jong-Seok Lee, Touradj Ebrahimi
MMSP1
2009 Two-way cooperative prediction for collaborative filtering recommendations
Jong-Seok Lee, Sigurdur Ólafsson
Expert Syst. Appl.1
2008 Robust Audio-Visual Speech Recognition Based on Late Integration
abstract
Audio-visual speech recognition (AVSR) using acoustic and visual signals of speech has received attention because of its robustness in noisy environments. In this paper, we present a late integration scheme-based AVSR system whose robustness under various noise conditions is improved by enhancing the performance of the three parts composing the system. First, we improve the performance of the visual subsystem by using the stochastic optimization method for the hidden Markov models as the speech recognizer. Second, we propose a new method of considering dynamic characteristics of speech for improved robustness of the acoustic subsystem. Third, the acoustic and the visual subsystems are effectively integrated to produce final robust recognition results by using neural networks. We demonstrate the performance of the proposed methods via speaker-independent isolated word recognition experiments. The results show that the proposed system improves robustness over the conventional system under various noise conditions without a priori knowledge about the noise contained in the speech.
Jong-Seok Lee, Cheol Hoon Park
IEEE Trans. Multim.1
2007 Improving generalization capability of neural networks based on simulated annealing
abstract
This paper presents a single-objective and a multiobjective stochastic optimization algorithms for global training of neural networks based on simulated annealing. The algorithms overcome the limitation of local optimization by the conventional gradient-based training methods and perform global optimization of the weights of the neural networks. Especially, the multiobjective training algorithm is designed to enhance generalization capability of the trained networks by minimizing the training error and the dynamic range of the network weights simulataneously. For fast convergence and good solution quality of the algorithms, we suggest the hybrid simulated annealing algorithm with the gradient-based local optimization method. Experimental results show that the performance of the trained networks by the proposed methods is better than that by the gradient-based local training algorithm and, moreover, the generalization capability of the networks is significantly improved by preventing overfitting phenomena.
Yeejin Lee, Jong-Seok Lee, Sun-Young Lee, Cheol Hoon Park
IEEE Congress on Evolutionary Computation2
2007 Temporal filtering of visual speech for audio-visual speech recognition in acoustically and visually challenging environments
abstract
The use of visual information of speech has been shown to be effective for compensating for performance degradation of acoustic speech recognition in noisy environments. However, visual noise is usually ignored in most of audio-visual speech recognition systems, while it can be included in visual speech signals during acquisition or transmission of the signals. In this paper, we present a new temporal filtering technique for extraction of noise-robust visual features. In the proposed method, a carefully designed band-pass filter is applied to the temporal pixel value sequences of lip region images in order to remove unwanted temporal variations due to visual noise, illumination conditions or speakers' appearances. We demonstrate that the method can improve not only visual speech recognition performance for clean and noisy images but also audio-visual speech recognition performance in both acoustically and visually noisy conditions.
Jong-Seok Lee, Cheol Hoon Park
ICMI1
2006 Training Hidden Markov Models by Hybrid Simulated Annealing for Visual Speech Recognition
abstract
This paper presents a novel training algorithm of hidden Markov models (HMMs) for visual speech recognition based on a modified simulated annealing (SA) algorithm, hybrid simulated annealing, where SA is combined with a local optimization technique to improve the convergence speed and the solution quality. While the popular training method of HMMs, the expectation-maximization (EM) algorithm, only achieves local optima in the parameter space, the proposed algorithm performs global search and thus obtains solutions giving improved recognition performance. The effectiveness of the proposed method is demonstrated via isolated word recognition experiments.
Jong-Seok Lee, Cheol Hoon Park
SMC1
2005 Discriminative training of hidden Markov models by multiobjective optimization for visual speech recognition
abstract
This paper proposes a novel discriminative training algorithm of hidden Markov models (HMMs) based on the multiobjective optimization for visual speech recognition. We develop a new criterion composed of two minimization objectives for training HMMs discriminatively and a global multiobjective optimization algorithm based on the simulated annealing algorithm to find the Pareto solutions of the optimization problem. We demonstrate the effectiveness of the proposed method via an isolated digit recognition experiment. The results show that the proposed method is superior to the conventional maximum likelihood estimation and the popular discriminative training algorithms.
Jong-Seok Lee, Cheol Hoon Park
IJCNN1
2005 Classification-based collaborative filtering using market basket data
Jong-Seok Lee, Chi-Hyuck Jun, Jaewook Lee 0001
Expert Syst. Appl.1
2000 Speech/non-speech classification using multiple features for robust endpoint detection
abstract
In this paper, we describe a new speech/non-speech classification method that improves the endpoint detection performance for speech recognition in noisy environments. The proposed method uses multiple features to increase the robustness in noisy environments, and the classification and regression tree (CART) technique is applied to effectively combine these multiple features for classification of each frame. We evaluate the performance of the proposed method by conducting speech/non-speech classification experiments on noisy speech. We also investigate the importance of various features on speech/non-speech classification in noisy environments In particular, the proposed method is applied to the endpoint detection algorithm for isolated speech recognition of a voice-dialing cellular phone. We simulate the speech recognition experiments in various noise environments, and the effects of the proposed method on speech recognition performance are evaluated.
Won-Ho Shin, Byoung-Soo Lee, Jong-Seok Lee
ICASSP4
2000 Large vocabulary Korean continuous speech recognition using a one-pass algorithm
Ha-Jin Yu, Joon-Mo Hong, Jong-Seok Lee
INTERSPEECH5
1998 Smoothing and tying for Korean flexible vocabulary isolated word recognition
Jae-Seung Choi, Jong-Seok Lee, Hee-Youn Lee
ICSLP2
1998 Automatic recognition of Korean broadcast news speech
abstract
This paper describes preliminary results of automatic recognition of Korean broadcast-news speech. We have been working on flexible vocabulary isolated-word speech recognition, and the same HMM models are used for broadcast-news continuous speech recognition. The recognizer is trained by using phonetically balanced isolated words speech, rather than the broadcast news speech itself. In this research, we use several different lexica to investigate the recognition performance according to the length of the words. We also propose a long-distance bigram language model, which can be used at the first stage of the search, so that it can reduce the recognition errors caused by earlier pruning of correct hypothesis. 1.
Ha-Jin Yu, Jae-Seung Choi, Joon-Mo Hong, Kew-Suh Park, Jong-Seok Lee, Hee-Youn Lee
ICSLP6
1995 Minimum duration constrained non-keyword modeling and rejection for word spotting
Seung-Bae Lee, Lag-Yong Kim, Jong-Seok Lee, Shin-Wook Kang
EUROSPEECH4