EDBT 2026 Demo / reviewers in the wild / expert
Chul Lee
dblp:97/2693
· DBLP profile ↗
74ranked-venue papers
14as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 12 first-author · 31 since 2021Artificial intelligence and machine learning · 20 · 16 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unleashing Zoned UFS: Cross-Layer Optimizations for Next-Generation Mobile Storage
Jungae Kim, Jaegeuk Kim, Kyu-Jin Cho, Iksung Oh, Chul Lee, Bart Van Assche, Daeho Jeong, Konstantin Vyshetsky |
FAST | 8 |
| 2026 | Gradient-Guided Diffusion-Based Restoration of Extremely Compressed Backgrounds for Video Coding for MachinesabstractVideo coding for machines (VCM) is an emerging approach in video compression designed to optimize content for machine analysis tasks. Although VCM was initially developed for machine vision, scalable coding frameworks have been developed to support both machine-driven analysis and human viewing as required. In this work, we focus on scenarios where high-quality encoding of regions of interest (ROIs) for machine vision and low-bitrate encoding of the background (BG) for human vision. At the decoder, severely degraded BG quality in reconstructed frames makes them unsuitable for viewing; therefore, restoring the degraded BGs by leveraging high-quality ROIs is essential. To this end, we propose the Gradient-Guided Diffusion Restoration (GGDR) algorithm, which integrates a pretrained generative diffusion model with content-aware supervision and adaptive refinement mechanisms to restore severely degraded regions robustly while maintaining visual consistency across the entire frame. The GGDR algorithm consists of two key components: (i) a content-aware supervision mechanism that preserves salient features and structural information in the input image, ensuring superior performance even with challenging high-variance inputs and (ii) a refinement block that guides the generation process of the pretrained diffusion model based on a degradation model and structural guidance. Experimental results demonstrate that the proposed algorithm outperforms state-of-the-art algorithms both qualitatively and quantitatively. Le Thi Hue Dao, Vien Gia An, Jooyoung Lee 0004, Seyoon Jeong, Naeun Yang, Chul Lee |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | High-Resolution Screenshot Demoiréing With Auxiliary Negative Sample Generation-Based Contrastive Learning
Hai Duong Nguyen, Se-Ho Lee, Chul Lee |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Evaluation of engagement in online learning: insights based on human factor analysis
Sanga Park, Byeonghui Jeong, Wookho Son, Chul Lee, Young-Sik Jeong |
J. Supercomput. | 4 |
| 2025 | Physics-driven prior learning-based deep unrolling for underwater image enhancement
Thuy Thi Pham, Hansung Yu, Truong Thanh Nhat Mai, Chul Lee |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Dual-channel prior-based deep unfolding with contrastive learning for underwater image enhancement
Thuy Thi Pham, Truong Thanh Nhat Mai, Hansung Yu, Chul Lee |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | Transformer-guided exposure-aware fusion for single-shot HDR imaging
Vien Gia An, Chul Lee |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | 3D Face Tracking from 2D Video through Iterative Dense UV to Image FlowabstractWhen working with 3D facial data, improving fidelity and avoiding the uncanny valley effect is critically dependent on accurate 3D facial performance capture. Because such methods are expensive and due to the widespread availability of 2D videos, recent methods have focused on how to perform monocular 3D face tracking. However, these methods often fall short in capturing precise facial movements due to limitations in their network architecture, training, and evaluation processes. Addressing these challenges, we propose a novel face tracker, FlowFace, that in-troduces an innovative 2D alignment network for dense pervertex alignment. Unlike prior work, FlowFace is trained on high-quality 3D scan annotations rather than weak supervision or synthetic data. Our 3D model fitting module Jointly fits a 3D face model from one or many observations, integrating existing neutral shape priors for enhanced identity and expression disentanglement and per-vertex de-formations for detailed facial feature reconstruction. Additionally, we propose a novel metric and benchmark for assessing tracking accuracy. Our method exhibits superior performance on both custom and publicly available bench-marks. We further validate the effectiveness of our tracker by generating high-quality 3D data from 2D videos, which leads to performance gains on downstream tasks. Felix Taubner, Prashant Raina, Mathieu Tuli, Eu Wern Teh, Chul Lee, Jinmiao Huang |
CVPR | 5 |
| 2024 | Content-Aware Supervision For Diffusion-Based Restoration of Extremely Compressed Background For VCMabstractWe propose content-aware supervision (CAS) techniques for diffusion-based restoration of an extremely compressed background for video coding for machines (VCM). First, we develop a CAS block to exploit prior information in an input image to reconstruct the noisy image, which is used as the input for the pretrained diffusion model. Then, we construct a refinement block to guide the pretrained diffusion model at each diffusion step by incorporating a degradation model and correction gradient estimation. Experimental results demonstrate the proposed algorithm outperforms state-of-the-art algorithms. Le Thi Hue Dao, Vien Gia An, Jooyoung Lee 0004, Seyoon Jeong, Naeun Yang, Chul Lee |
ICIP | 6 |
| 2024 | Feature Decomposition Transformers for Infrared and Visible Image FusionabstractWe propose an infrared and visible image fusion algorithm using modality-shared and modality-specific feature decomposition transformers. First, the proposed algorithm extracts multiscale shallow features of infrared and visible images. Then, we develop modality-shared and modality-specific feature decomposition transformers that decompose the features into common and complementary components for each modality. For better decomposition, we develop a decomposition loss by constraining the common features to be correlated while the complementary features are uncorrelated. Finally, the reconstruction block generates the fused image by combining the common and complementary features. Experimental results show that the proposed algorithm significantly outperforms conventional algorithms on several datasets. Gahyeon Kim, Vien Gia An, Duong Hai Nguyen, Chul Lee |
ICIP | 4 |
| 2024 | H2O-SDF: Two-phase Learning for 3D Indoor Reconstruction using Object Surface FieldsabstractAdvanced techniques using Neural Radiance Fields (NeRF), Signed Distance Fields (SDF), and Occupancy Fields have recently emerged as solutions for 3D indoor scene reconstruction. We introduce a novel two-phase learning approach, H2O-SDF, that discriminates between object and non-object regions within indoor environments. This method achieves a nuanced balance, carefully preserving the geometric integrity of room layouts while also capturing intricate surface details of specific objects. A cornerstone of our two-phase learning framework is the introduction of the Object Surface Field (OSF), a novel concept designed to mitigate the persistent vanishing gradient problem that has previously hindered the capture of high-frequency details in other methods. Our proposed approach is validated through several experiments that include ablation studies. Mirae Do, YeonJae Shin, Jaeseok Yoo, Jongkwang Hong, Joongrock Kim, Chul Lee |
ICLR | 7 |
| 2024 | The interaction of inter-organizational diversity and team size, and the scientific impact of papersabstractLarge teams are known to be more likely to publish highly cited papers, while small teams are known to be better at publishing highly disruptive papers. However, there is a lack of adequate theoretical understanding of the mechanisms by which scientific collaboration among researchers is related to the scientific impact of their papers. We investigated the mechanisms more closely by focusing on the interaction of inter-organizational diversity and team size in the process of team formation and knowledge dissemination. We analyzed 12,010,102 Web of Science papers and examined how inter-organizational diversity is associated with the relationship of team size with disruption and citations. As a result, we found that not only small teams, but also large teams with great inter-organizational diversity were able to disrupt science and technology effectively. We also found that large teams with greater inter-organizational diversity were more likely to produce highly cited papers. Our findings are robust and consistently observed regardless of publication year, team size, the number of references, and the degree of multidisciplinarity. These results have significant implications for researchers in selecting collaborators to achieve greater impact and for improving the qualitative efficiency of public research investments. Hyoung Sun Yoo, Ye Lim Jung, June Young Lee, Chul Lee |
Inf. Process. Manag. | 4 |
| 2024 | Attention-Guided Low-Rank Tensor CompletionabstractLow-rank tensor completion (LRTC) aims to recover missing data of high-dimensional structures from a limited set of observed entries. Despite recent significant successes, the original structures of data tensors are still not effectively preserved in LRTC algorithms, yielding less accurate restoration results. Moreover, LRTC algorithms often incur high computational costs, which hinder their applicability. In this work, we propose an attention-guided low-rank tensor completion (AGTC) algorithm, which can faithfully restore the original structures of data tensors using deep unfolding attention-guided tensor factorization. First, we formulate the LRTC task as a robust factorization problem based on low-rank and sparse error assumptions. Low-rank tensor recovery is guided by an attention mechanism to better preserve the structures of the original data. We also develop implicit regularizers to compensate for modeling inaccuracies. Then, we solve the optimization problem by employing an iterative technique. Finally, we design a multistage deep network by unfolding the iterative algorithm, where each stage corresponds to an iteration of the algorithm; at each stage, the optimization variables and regularizers are updated by closed-form solutions and learned deep networks, respectively. Experimental results for high dynamic range imaging and hyperspectral image restoration show that the proposed algorithm outperforms state-of-the-art algorithms. Truong Thanh Nhat Mai, Edmund Y. Lam, Chul Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Cross-Modal Transformers for Infrared and Visible Image FusionabstractImage fusion techniques aim to generate more informative images by merging multiple images of different modalities with complementary information. Despite significant fusion performance improvements of recent learning-based approaches, most fusion algorithms have been developed based on convolutional neural networks (CNNs), which stack deep layers to obtain a large receptive field for feature extraction. However, important details and contexts of the source images may be lost through a series of convolution layers. In this work, we propose a cross-modal transformer-based fusion (CMTFusion) algorithm for infrared and visible image fusion that captures global interactions by faithfully extracting complementary information from source images. Specifically, we first extract the multiscale feature maps of infrared and visible images. Then, we develop cross-modal transformers (CMTs) to retain complementary information in the source images by removing redundancies in both the spatial and channel domains. To this end, we design a gated bottleneck that integrates cross-domain interaction to consider the characteristics of the source images. Finally, a fusion result is obtained by exploiting spatial-channel information in refined feature maps using a fusion block. Experimental results on multiple datasets demonstrate that the proposed algorithm provides better fusion performance than state-of-the-art infrared and visible image fusion algorithms, both quantitatively and qualitatively. Furthermore, we show that the proposed algorithm can be used to improve the performance of computer vision tasks, e.g., object detection and monocular depth estimation. Seonghyun Park 0003, Vien Gia An, Chul Lee |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Deep Unfolding Tensor Rank Minimization With Generalized Detail Injection for PansharpeningabstractPansharpening aims to generate a high-resolution multispectral (HRMS) image by merging a low-resolution multispectral (LRMS) image with a high-resolution panchromatic (PAN) image. While traditional model-based pansharpening algorithms have strong theoretical foundations, their performance and generalizability are limited by handcrafted formulations. In contrast, recent deep learning approaches outperform model-based algorithms but do not effectively consider the physical properties of multispectral (MS) images, such as their spatial and spectral dependencies. These physical properties facilitate the exploitation of the actual imaging process, leading to enhanced spatial and spectral fidelities. In this work, we propose a deep unfolded tensor rank minimization framework with generalized detail injection for pansharpening to overcome the weaknesses of both model- and learning-based approaches while leveraging their advantages. Specifically, we first formulate the pansharpening task as a tensor rank minimization problem to exploit the low-rankness of MS images, providing a robust theoretical foundation on the physical properties of MS data. We also develop a generalized detail injection component, which effectively exploits the information in the PAN images, and incorporate it into the optimization to improve generalizability and representation capability. Then, we define a data-driven regularizer to compensate for modeling inaccuracies in the low-rank model and solve the optimization problem using an iterative technique. Finally, the iterative algorithm is unfolded into a multistage deep network, in which the optimization variables are solved by closed-form solutions and a data-driven regularizer in each stage. Experimental results on various MS image datasets demonstrate that the proposed algorithm achieves better pansharpening performance and interpretability than state-of-the-art algorithms. Truong Thanh Nhat Mai, Edmund Y. Lam, Chul Lee |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Quantitative Manipulation of Custom Attributes on 3D-Aware Image SynthesisabstractWhile 3D-based GAN techniques have been successfully applied to render photo-realistic 3D images with a variety of attributes while preserving view consistency, there has been little research on how to fine-control 3D images without limiting to a specific category of objects of their properties. To fill such research gap, we propose a novel image manipulation model of 3D-based GAN representations for a fine-grained control of specific custom attributes. By extending the latest 3D-based GAN models (e.g., EG3D), our user-friendly quantitative manipulation model enables a fine yet normalized control of 3D manipulation of multi-attribute quantities while achieving view consistency. We validate the effectiveness of our proposed technique both qualitatively and quantitatively through various experiments. Hoseok Do, Eunkyung Yoo, Chul Lee, Jin Young Choi 0002 |
CVPR | 4 |
| 2023 | Blending-NeRF: Text-Driven Localized Editing in Neural Radiance FieldsabstractText-driven localized editing of 3D objects is particularly difficult as locally mixing the original 3D object with the intended new object and style effects without distorting the object’s form is not a straightforward process. To address this issue, we propose a novel NeRF-based model, Blending-NeRF, which consists of two NeRF networks: pre-trained NeRF and editable NeRF. Additionally, we introduce new blending operations that allow Blending-NeRF to properly edit target regions which are localized by text. By using a pretrained vision-language aligned model, CLIP, we guide Blending-NeRF to add new objects with varying colors and densities, modify textures, and remove parts of the original object. Our extensive experiments demonstrate that Blending-NeRF produces naturally and locally edited 3D objects from various text prompts. Hyeonseop Song, Seokhun Choi, Hoseok Do, Chul Lee |
ICCV | 4 |
| 2023 | Restoration of Extremely Compressed Background for VCM Using Guided Generative PriorsabstractWe propose a learning-based image restoration algorithm for a single decoded image with a high-quality foreground and an extremely degraded background for video coding for machines (VCM). First, we develop an encoder that extracts multiscale features and learns latent vectors. Then, a background generator with style and feature fusion blocks generates guided features that contain the prior background information in the input image. Finally, the decoder restores the degraded background region by merging the image features from the encoder and prior background information from the generator. Experimental results show that the proposed algorithm achieves better performance than state-of-the-art algorithms. Le Thi Hue Dao, Vien Gia An, Jooyoung Lee 0004, Seyoon Jeong, Chul Lee |
ICIP | 5 |
| 2023 | A Contrastive Learning Approach for Screenshot DemoiréingabstractWe propose a contrast learning-based approach for screenshot demoiréing based on the assumption that a moiré image can be separated into two layers in deep latent space: moiré artifacts and latent clean image. First, we develop a multiscale network, called SDN, that extracts multiscale feature maps of an input image and then separates them into moiré and clean image components. To improve the separation of the features, we develop a contrast learning approach that separates and clusters moiré and clean image features in the latent space in supervised and unsupervised manners, respectively. Experimental results on a misaligned real-world screenshot dataset show that the proposed algorithm provides better demoiréing performance than state-of-the-art algorithms. Duong Hai Nguyen, Chul Lee |
ICIP | 2 |
| 2023 | Deep Unfolding Network with Physics-Based Priors for Underwater Image EnhancementabstractWe propose an underwater image enhancement algorithm that leverages both model- and learning-based approaches by unfolding an iterative algorithm. We first formulate the underwater image enhancement task as a joint optimization problem, based on the image formation model with physical model and underwater-related priors. Then, we solve the optimization problem iteratively. Finally, we unfold the iterative algorithm so that, at each iteration, the optimization variables and regularizers for image priors are updated by closed-form solutions and learned deep networks, respectively. Experimental results demonstrate that the proposed algorithm outperforms state-of-the-art underwater image enhancement algorithms. Thuy Thi Pham, Truong Thanh Nhat Mai, Chul Lee |
ICIP | 3 |
| 2023 | Harmonic Neural NetworksabstractHarmonic functions are abundant in nature, appearing in limiting cases of Maxwell’s, Navier-Stokes equations, the heat and the wave equation. Consequently, there are many applications of harmonic functions from industrial process optimisation to robotic path planning and the calculation of first exit times of random walks. Despite their ubiquity and relevance, there have been few attempts to incorporate inductive biases towards harmonic functions in machine learning contexts. In this work, we demonstrate effective means of representing harmonic functions in neural networks and extend such results also to quantum neural networks to demonstrate the generality of our approach. We benchmark our approaches against (quantum) physics-informed neural networks, where we show favourable performance. Atiyo Ghosh, Antonio Andrea Gentile, Mario Dagrada, Chul Lee, Seong-Hyok Sean Kim, Hyukgeun Cha, Yunjun Choi, Jeong-Il Kye, Vincent E. Elfving |
ICML | 4 |
| 2023 | Transformer as a hippocampal memory consolidation model based on NMDAR-inspired nonlinearityabstractThe hippocampus plays a critical role in learning, memory, and spatial representation, processes that depend on the NMDA receptor (NMDAR). Inspired by recent findings that compare deep learning models to the hippocampus, we propose a new nonlinear activation function that mimics NMDAR dynamics. NMDAR-like nonlinearity shifts short-term working memory into long-term reference memory in transformers, thus enhancing a process that is similar to memory consolidation in the mammalian brain. We design a navigation task assessing these two memory functions and show that manipulating the activation function (i.e., mimicking the Mg$^{2+}$-gating of NMDAR) disrupts long-term memory processes. Our experiments suggest that place cell-like functions and reference memory reside in the feed-forward network layer of transformers and that nonlinearity drives these processes. We discuss the role of NMDAR-like nonlinearity in establishing this striking resemblance between transformer architecture and hippocampal spatial representation. Dong Kyum Kim, Jea Kwon, Meeyoung Cha, Chul Lee |
NeurIPS | 4 |
| 2023 | Multiple transformation function estimation for image enhancementabstractMost deep learning-based image enhancement algorithms have been developed based on the image-to-image translation approach, in which enhancement processes are difficult to interpret. In this paper, we propose a novel interpretable image enhancement algorithm that estimates multiple transformation functions to describe complex color mapping. First, we develop a histogram-based multiple transformation function estimation network (HMTF-Net) to estimate multiple transformation functions by exploiting both the spatial and statistical information of the input images. Second, we estimate pixel-wise weight maps, which indicate the contribution of each transformation function at each pixel, based on the local structures of the input image and the transformed images obtained by each transformation function. Finally, we obtain the enhanced image as the weighted sum of the transformed images using the estimated weight maps. Extensive experiments confirm the effectiveness of the proposed approach and demonstrate that the proposed algorithm outperforms state-of-the-art image enhancement algorithms for different image enhancement tasks. Vien Gia An, Minhee Cha, Thuy Thi Pham, Hanul Kim 0001, Chul Lee |
J. Vis. Commun. Image Represent. | 6 |
| 2023 | Multiscale Coarse-to-Fine Guided Screenshot DemoiréingabstractIn this letter, we propose a multiscale coarse-to-fine guided screenshot demoireing algorithm. We first extract the multiscale features of the input image. Then, we develop the multiscale guided restoration block (MGRB), which removes moire patterns with the guidance of multiscale information by exploiting the correlation between moire frequencies. To this end, we design two blocks for feature modulation and moire pattern removal. In addition, to further improve the performance, we develop an adaptive reconstruction loss to direct the network to focus on regions that are difficult to restore. Experimental results on multiple datasets demonstrate that the proposed algorithm provides comparable or even better demoireing performance than state-of-the-art algorithms. Hai Duong Nguyen, Se-Ho Lee, Chul Lee |
IEEE Signal Process. Lett. | 3 |
| 2022 | Exposure-Aware Dynamic Weighted Learning for Single-Shot HDR Imaging
Vien Gia An, Chul Lee |
ECCV (7) | 2 |
| 2022 | Depth Map Decomposition for Monocular Depth Estimation
Jinyoung Jun, Jaehan Lee, Chul Lee, Chang-Su Kim 0001 |
ECCV (2) | 3 |
| 2022 | OpenFEAT: Improving Speaker Identification by Open-Set Few-Shot Embedding Adaptation with TransformerabstractHousehold speaker identification with few enrollment utterances is an important yet challenging problem, especially when household members share similar voice characteristics and room acoustics. A common embedding space learned from a large number of speakers is not universally applicable for the optimal identification of every speaker in a household. In this work, we first formulate household speaker identification as a few-shot open-set recognition task and then propose a novel embedding adaptation framework to adapt speaker representations from the given universal embedding space to a household-specific embedding space using a set-to-set function, yielding better household speaker identification performance. With our algorithm, Open-set Few-shot Embedding Adaptation with Transformer (openFEAT), we observe that the speaker identification equal error rate (IEER) on simulated households with 2 to 7 hard-to-discriminate speakers is reduced by 23% to 31% relative. K. C. Kishan, Zhenning Tan, Long Chen 0027, Minho Jin, Eunjung Han, Andreas Stolcke, Chul Lee |
ICASSP | 7 |
| 2022 | Infrared and Visible Image Fusion Using Bimodal TransformersabstractWe propose an infrared and visible image fusion algorithm using bimodal transformers. First, the proposed algorithm extracts multiscale features of the input infrared and visible images. Then, we develop the bimodal transformers that refine the extracted features by estimating their irrelevance maps to exploit the complementary information of the source images. Finally, we develop a reconstruction block that generates the fusion result by merging the refined features in the frequency domain to exploit the global information of the source images. Experimental results show that the proposed algorithm outperforms state-of-the-art infrared and visible image fusion algorithms on several datasets. Seonghyun Park 0003, Vien Gia An, Chul Lee |
ICIP | 3 |
| 2022 | Histogram-Based Transformation Function Estimation for Low-Light Image EnhancementabstractWe propose a learning-based low-light image enhancement algorithm, called the histogram-based transformation function estimation network (HTFNet), that estimates transformation functions using the histogram of an input image. First, we obtain an attention image that indicates the pixel-wise information on the level of enhancement. Then, the proposed HTFNet generates the transformation functions by exploiting both the spatial and statistical information of the input image by combining two feature maps extracted from the input image and its histogram. Finally, the enhanced images are obtained via channel-wise intensity transformation. Experimental results show that the proposed algorithm provides higher image quality compared with the state-of-the-art algorithms. Vien Gia An, Jin-Hwan Kim, Chul Lee |
ICIP | 4 |
| 2022 | QbyE-MLPMixer: Query-by-Example Open-Vocabulary Keyword Spotting using MLPMixerabstractCurrent keyword spotting systems are typically trained with a large amount of pre-defined keywords.Recognizing keywords in an open-vocabulary setting is essential for personalizing smart device interaction.Towards this goal, we propose a pure MLP-based neural network that is based on MLPMixer -an MLP model architecture that effectively replaces the attention mechanism in Vision Transformers.We investigate different ways of adapting the MLPMixer architecture to the QbyE open-vocabulary keyword spotting task.Comparisons with the state-of-the-art RNN and CNN models show that our method achieves better performance in challenging situations (10dB and 6dB environments) on both the publicly available Hey-Snips dataset and a larger scale internal dataset with 400 speakers.Our proposed model also has a smaller number of parameters and MACs compared to the baseline models. Jinmiao Huang, Waseem Gharbieh, Qianhui Wan, Han Suk Shim, Chul Lee |
INTERSPEECH | 5 |
| 2022 | On joint training with interfaces for spoken language understandingabstractSpoken language understanding (SLU) systems extract both text transcripts and semantics associated with intents and slots from input speech utterances.SLU systems usually consist of (1) an automatic speech recognition (ASR) module, (2) an interface module that exposes relevant outputs from ASR, and (3) a natural language understanding (NLU) module.Interfaces in SLU systems carry information on text transcriptions or richer information like neural embeddings from ASR to NLU.In this paper, we study how interfaces affect joint-training for spoken language understanding.Most notably, we obtain the state-of-theart results on the publicly available 50-hr SLURP [1] dataset.We first leverage large-size pretrained ASR and NLU models that are connected by a text interface, and then jointly train both models via a sequence loss function.For scenarios where pretrained models are not utilized, the best results are obtained through a joint sequence loss training using richer neural interfaces.Finally, we show the overall diminishing impact of leveraging pretrained models with increased training data size. Anirudh Raju, Milind Rao, Gautam Tiwari, Pranav Dheram, Bryan Anderson, Chul Lee, Bach Bui, Ariya Rastrow |
INTERSPEECH | 7 |
| 2022 | Online Learning of Open-set Speaker Identification by Active User-registration
Eunkyung Yoo, Hyeonseop Song, Chul Lee |
INTERSPEECH | 4 |
| 2022 | Real-time image and video dehazing based on multiscale guided filtering
Thuong Van Nguyen, Vien Gia An, Chul Lee |
Multim. Tools Appl. | 3 |
| 2022 | Deep Unrolled Low-Rank Tensor Completion for High Dynamic Range ImagingabstractThe major challenge in high dynamic range (HDR) imaging for dynamic scenes is suppressing ghosting artifacts caused by large object motions or poor exposures. Whereas recent deep learning-based approaches have shown significant synthesis performance, interpretation and analysis of their behaviors are difficult and their performance is affected by the diversity of training data. In contrast, traditional model-based approaches yield inferior synthesis performance to learning-based algorithms despite their theoretical thoroughness. In this paper, we propose an algorithm unrolling approach to ghost-free HDR image synthesis algorithm that unrolls an iterative low-rank tensor completion algorithm into deep neural networks to take advantage of the merits of both learning- and model-based approaches while overcoming their weaknesses. First, we formulate ghost-free HDR image synthesis as a low-rank tensor completion problem by assuming the low-rank structure of the tensor constructed from low dynamic range (LDR) images and linear dependency among LDR images. We also define two regularization functions to compensate for modeling inaccuracy by extracting hidden model information. Then, we solve the problem efficiently using an iterative optimization algorithm by reformulating it into a series of subproblems. Finally, we unroll the iterative algorithm into a series of blocks corresponding to each iteration, in which the optimization variables are updated by rigorous closed-form solutions and the regularizers are updated by learned deep neural networks. Experimental results on different datasets show that the proposed algorithm provides better HDR image synthesis performance with superior robustness compared with state-of-the-art algorithms, while using significantly fewer training samples. Truong Thanh Nhat Mai, Edmund Y. Lam, Chul Lee |
IEEE Trans. Image Process. | 3 |
| 2021 | BW-EDA-EEND: streaming END-TO-END Neural Speaker Diarization for a Variable Number of SpeakersabstractWe present a novel online end-to-end neural diarization system, BW-EDA-EEND, that processes data incrementally for a variable number of speakers. The system is based on the Encoder-Decoder-Attractor (EDA) architecture of Horiguchi et al., but utilizes the incremental Transformer encoder, attending only to its left contexts and using block-level recurrence in the hidden states to carry information from block to block, making the algorithm complexity linear in time. We propose two variants: For unlimited-latency BW-EDA-EEND, which processes inputs in linear time, we show only moderate degradation for up to two speakers using a context size of 10 seconds compared to offline EDA-EEND. With more than two speakers, the accuracy gap between online and offline grows, but the algorithm still outperforms a baseline offline clustering diarization system for one to four speakers with unlimited context size, and shows comparable accuracy with context size of 10 seconds. For limited-latency BW-EDA-EEND, which produces diarization outputs block-by-block as audio arrives, we show accuracy comparable to the offline clustering-based system. Eunjung Han, Chul Lee, Andreas Stolcke |
ICASSP | 2 |
| 2021 | Learning Multiple Pixelwise Tasks Based on Loss Scale BalancingabstractWe propose a novel loss weighting algorithm, called loss scale balancing (LSB), for multi-task learning (MTL) of pixelwise vision tasks. An MTL model is trained to estimate multiple pixelwise predictions using an overall loss, which is a linear combination of individual task losses. The proposed algorithm dynamically adjusts the linear weights to learn all tasks effectively. Instead of controlling the trend of each loss value directly, we balance the loss scale — the product of the loss value and its weight — periodically. In addition, by evaluating the difficulty of each task based on the previous loss record, the proposed algorithm focuses more on difficult tasks during training. Experimental results show that the proposed algorithm outperforms conventional weighting algorithms for MTL of various pixelwise tasks. Codes are available at https://github.com/jaehanlee-mcl/LSB-MTL. Jae-Han Lee, Chul Lee, Chang-Su Kim 0001 |
ICCV | 2 |
| 2021 | Asymmetric Bilateral Motion Estimation for Video Frame InterpolationabstractWe propose a novel video frame interpolation algorithm based on asymmetric bilateral motion estimation (ABME), which synthesizes an intermediate frame between two input frames. First, we predict symmetric bilateral motion fields to interpolate an anchor frame. Second, we estimate asymmetric bilateral motions fields from the anchor frame to the input frames. Third, we use the asymmetric fields to warp the input frames backward and reconstruct the intermediate frame. Last, to refine the intermediate frame, we develop a new synthesis network that generates a set of dynamic filters and a residual frame using local and global information. Experimental results show that the proposed algorithm achieves excellent performance on various datasets. The source codes and pretrained models are available at https://github.com/JunHeum/ABME. Junheum Park, Chul Lee, Chang-Su Kim 0001 |
ICCV | 2 |
| 2021 | Ghost-Free HDR Imaging Via Unrolling Low-Rank Matrix CompletionabstractWe propose a ghost-free high dynamic range (HDR) image synthesis algorithm by unrolling low-rank matrix completion. By exploiting the low-rank structure of the irradiance maps from low dynamic range (LDR) images, we formulate ghost-free HDR imaging as a general low-rank matrix completion problem. Then, we solve the problem iteratively using the augmented Lagrange multiplier (ALM) method. At each iteration, the optimization variables are updated by closed-form solutions and the regularizers are updated by learned deep neural networks. Experimental results show that the proposed algorithm provides better image qualities with fewer visual artifacts compared to state-of-the-art algorithms. Truong Thanh Nhat Mai, Edmund Y. Lam, Chul Lee |
ICIP | 3 |
| 2021 | End-to-End Neural Diarization: From Transformer to ConformerabstractWe propose a new end-to-end neural diarization (EEND) system that is based on Conformer, a recently proposed neural architecture that combines convolutional mappings and Transformer to model both local and global dependencies in speech. We first show that data augmentation and convolutional subsampling layers enhance the original self-attentive EEND in the Transformer-based EEND, and then Conformer gives an additional gain over the Transformer-based EEND. However, we notice that the Conformer-based EEND does not generalize as well from simulated to real conversation data as the Transformer-based model. This leads us to quantify the mismatch between simulated data and real speaker behavior in terms of temporal statistics reflecting turn-taking between speakers, and investigate its correlation with diarization error. By mixing simulated and real data in EEND training, we mitigate the mismatch further, with Conformer-based EEND achieving 24% error reduction over the baseline SA-EEND system, and 10% improvement over the best augmented Transformer-based system, on two-speaker CALLHOME data. Yi-Chieh Liu, Eunjung Han, Chul Lee, Andreas Stolcke |
Interspeech | 3 |
| 2021 | An optimization-based approach to gamma correction parameter estimation for low-light image enhancement
Inho Jeong, Chul Lee |
Multim. Tools Appl. | 2 |
| 2020 | BMBC: Bilateral Motion Estimation with Bilateral Cost Volume for Video Interpolation
Junheum Park, Keunsoo Ko, Chul Lee, Chang-Su Kim 0001 |
ECCV (14) | 3 |
| 2020 | Optimized Color Contrast Enhancement For Dichromats Using Local And Global ContrastabstractWe propose an optimized color contrast enhancement algorithm for dichromats with color vision deficiency using local and global information. Based on the fact that dichromats perceive colors projected onto a 2D plane, we first measure the color differences on the plane. Then, we formulate an optimization problem to minimize the perceived color differences subject to constraints on the projected plane. Finally, by solving the optimization problem, we obtain the optimal plane and perform color conversion. Simulation results show that the proposed algorithm outperforms the state-of-the-art algorithms in preserving both local details and naturalness ofimages. Soo-Kyeong Kang, Chul Lee, Chang-Su Kim 0001 |
ICIP | 2 |
| 2020 | Cross-modal Non-linear Guided Attention and Temporal Coherence in Multi-modal Deep Video ModelsabstractVideos have data in multiple modalities, e.g., audio, video, text (captions). Understanding and modeling the interaction between different modalities is key for video analysis tasks like categorization, object detection, activity recognition, etc. However, data modalities are not always correlated --- so, learning when modalities are correlated and using that to guide the influence of one modality on the other is crucial. Another salient feature of videos is the coherence between successive frames due to continuity of video and audio, a property that we refer to as temporal coherence. We show how using non-linear guided cross-modal signals and temporal coherence can improve the performance of multi-modal machine learning (ML) models for video analysis tasks like categorization. Our experiments on the large-scale YouTube-8M dataset show how our approach significantly outperforms state-of-the-art multi-modal ML models for video categorization. The model trained on the YouTube-8M dataset also showed good performance on an internal dataset of video segments from actual Samsung TV Plus channels without retraining or fine-tuning, showing the generalization capabilities of our model. Saurabh Sahu, Palash Goyal, Shalini Ghosh, Chul Lee |
ACM Multimedia | 4 |
| 2018 | A Deep Multi-Modal Pairwise Ranking Model for User Generated Food DataabstractDue to the emergence of several nutrition-related mobile applications and websites in recent years, as well as the massive amount of crowd-sourced nutrition data, searching and finding relevant results has become increasingly difficult for users. This problem becomes even more challenging when dealing with crowd-sourced food names that are noisy and not well-structured. Because food names are short in length, it is difficult to incorporate existing methods to achieve an optimal matching quality. Despite several recent studies on nutrition data, these challenges remain. In this paper, we propose a novel learning-to-rank framework for crowd-sourced food names that has significant real-world applications, including food search and food recommendations. In particular, we propose a deep learning based, multi-modal learning-to-rank model that leverages the text describing a food name and the numerical values that represent its nutritional information. To this end, we also introduce a novel type of loss-function, which extends standard triplets hinge loss function into a multi-modal scenario. The proposed model is flexible and supports various data types as well as an arbitrary number of modalities. The effectiveness of our proposed model is demonstrated through several experiments on real-data, consisting of more than six million instances. Hesamoddin Salehian, Surender Reddy Yerva, Iman Barjasteh, Patrick D. Howell, Chul Lee |
ASONAM | 5 |
| 2018 | Photographic composition classification and dominant geometric element detection for outdoor scenes
Juntae Lee, Hanul Kim 0001, Chul Lee, Chang-Su Kim 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Semantic Line Detection and Its ApplicationsabstractSemantic lines characterize the layout of an image. Despite their importance in image analysis and scene understanding, there is no reliable research for semantic line detection. In this paper, we propose a semantic line detector using a convolutional neural network with multi-task learning, by regarding the line detection as a combination of classification and regression tasks. We use convolution and max-pooling layers to obtain multi-scale feature maps for an input image. Then, we develop the line pooling layer to extract a feature vector for each candidate line from the feature maps. Next, we feed the feature vector into the parallel classification and regression layers. The classification layer decides whether the line candidate is semant ic or not. In case of a semantic line, the regression layer determines the offset for refining the line location. Experimental results show that the proposed detector extracts semantic lines accurately and reliably. Moreover, we demonstrate that the proposed detector can be used successfully in three applications: horizon estimation, composition enhancement, and image simplification. Juntae Lee, Hanul Kim 0001, Chul Lee, Chang-Su Kim 0001 |
ICCV | 3 |
| 2017 | Matching Restaurant Menus to Crowdsourced Food Data: A Scalable Machine Learning ApproachabstractWe study the problem of how to match a formally structured restaurant menu item to a large database of less structured food items that has been collected via crowd-sourcing. At first glance, this problem scenario looks like a typical text matching problem that might possibly be solved with existing text similarity learning approaches. However, due to the unique nature of our scenario and the need for scalability, our problem imposes certain restrictions on possible machine learning approaches that we can employ. We propose a novel, practical, and scalable machine learning solution architecture, consisting of two major steps. First we use a query generation approach, based on a Markov Decision Process algorithm, to reduce the time complexity of searching for matching candidates. That is then followed by a re-ranking step, using deep learning techniques, to meet our required matching quality goals. It is important to note that our proposed solution architecture has already been deployed in a real application system serving tens of millions of users, and shows great potential for practical cases of user-entered text to structured text matching, especially when scalability is crucial. Hesamoddin Salehian, Patrick D. Howell, Chul Lee |
KDD | 3 |
| 2017 | Contrast enhancement of noisy low-light images based on structure-texture-noise decomposition
Jaemoon Lim, Minhyeok Heo, Chul Lee, Chang-Su Kim 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | A Maximum a Posteriori Estimation Framework for Robust High Dynamic Range Video SynthesisabstractHigh dynamic range (HDR) image synthesis from multiple low dynamic range exposures continues to be actively researched. The extension to HDR video synthesis is a topic of significant current interest due to potential cost benefits. For HDR video, a stiff practical challenge presents itself in the form of accurate correspondence estimation of objects between video frames. In particular, loss of data resulting from poor exposures and varying intensity makes conventional optical flow methods highly inaccurate. We avoid exact correspondence estimation by proposing a statistical approach via maximum a posterior estimation, and under appropriate statistical assumptions and choice of priors and models, we reduce it to an optimization problem of solving for the foreground and background of the target frame. We obtain the background through rank minimization and estimate the foreground via a novel multiscale adaptive kernel regression technique, which implicitly captures local structure and temporal motion by solving an unconstrained optimization problem. Extensive experimental results on both real and synthetic data sets demonstrate that our algorithm is more capable of delivering high-quality HDR videos than current state-of-the-art methods, under both subjective and objective assessments. Furthermore, a thorough complexity analysis reveals that our algorithm achieves better complexity-performance tradeoff than conventional methods. Chul Lee, Vishal Monga |
IEEE Trans. Image Process. | 2 |
| 2016 | Persistent Sharing of Fitness App Status on TwitterabstractAs the world becomes more digitized and interconnected, information that was once considered to be private such as one's health status is now being shared publicly. To understand this new phenomenon better, it is crucial to study what types of health information are being shared on social media and why, as well as by whom. In this paper, we study the traits of users who share their personal health and fitness related information on social media by analyzing fitness status updates that MyFitnessPal users have shared via Twitter. We investigate how certain features like user profile, fitness activity, and fitness network in social media can potentially impact the long-term engagement of fitness app users. We also discuss implications of our findings to achieve a better retention of these users and to promote more sharing of their status updates. Kunwoo Park, Ingmar Weber, Meeyoung Cha, Chul Lee |
CSCW | 4 |
| 2016 | High dynamic range imaging via truncated nuclear norm minimization of low-rank matrixabstractWe propose a ghost-free high dynamic range (HDR) image synthesis algorithm using a rank minimization framework. Based on the linear dependency among irradiance maps from low dynamic range (LDR) images, we formulate ghost-free HDR imaging as a low-rank matrix completion problem. The main contribution is to solve it efficiently via the augmented Lagrange multiplier (ALM) method, where the optimization variables are updated by closed-form solutions. Experiments on real image sets show that the proposed algorithm provides comparable or even better image qualities than state-of-the-art approaches, while demanding lower computational resources. Chul Lee, Edmund Y. Lam |
ICASSP | 1 |
| 2016 | Power-Constrained RGB-to-RGBW Conversion for Emissive Displays: Optimization-Based ApproachesabstractWe propose an optimization-based power-constrained red-green-blue (RGB)-to-red-green-blue-white (RGBW) conversion algorithm for emissive RGBW displays. We measure the perceived color distortion using a color difference model in a perceptually uniform color space, and compute the power consumption for displaying an RGBW pixel on an emissive display. The central contribution of this paper is to formulate the optimization problem to minimize the color distortion subject to a constraint on the power consumption. Subsequently, we solve the optimization problem efficiently to convert an image in real time. Furthermore, based on the properties of the human visual system, we extend the proposed algorithm to image-dependent conversion that can preserve spatial detail in an input image. The simulation results show that the proposed algorithm provides a significantly less color distortion than the conventional methods, while providing a graceful tradeoff with the amount of power consumed. Specifically, it is shown that the power consumption can be reduced by up to 20%, while providing about 50% less color distortion than the conventional algorithms. In addition, a subjective evaluation on a real RGBW display is performed, which reveals the merits of the proposed image-dependent conversion for improving the perceptual quality over state-of-the-art techniques. Chul Lee, Vishal Monga |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Computationally Efficient Truncated Nuclear Norm Minimization for High Dynamic Range ImagingabstractMatrix completion is a rank minimization problem to recover a low-rank data matrix from a small subset of its entries. Since the matrix rank is nonconvex and discrete, many existing approaches approximate the matrix rank as the nuclear norm. However, the truncated nuclear norm is known to be a better approximation to the matrix rank than the nuclear norm, exploiting a priori target rank information about the problem in rank minimization. In this paper, we propose a computationally efficient truncated nuclear norm minimization algorithm for matrix completion, which we call TNNM-ALM. We reformulate the original optimization problem by introducing slack variables and considering noise in the observation. The central contribution of this paper is to solve it efficiently via the augmented Lagrange multiplier (ALM) method, where the optimization variables are updated by closed-form solutions. We apply the proposed TNNM-ALM algorithm to ghost-free high dynamic range imaging by exploiting the low-rank structure of irradiance maps from low dynamic range images. Experimental results on both synthetic and real visual data show that the proposed algorithm achieves significantly lower reconstruction errors and superior robustness against noise than the conventional approaches, while providing substantial improvement in speed, thereby applicable to a wide range of imaging applications. Chul Lee, Edmund Y. Lam |
IEEE Trans. Image Process. | 1 |
| 2015 | A map estimation framework for HDR video synthesisabstractHigh dynamic range (HDR) image synthesis from multiple low dynamic range (LDR) exposures continues to be a topic of great interest. The extension to HDR video comprises a stiff challenge due to significant motion. In particular, loss of data due to poor exposures introduces great difficulty in exact motion estimation, and under such circumstances conventional optical flow calculation techniques usually fail. We propose a maximum a posterior (MAP) estimation framework for HDR video synthesis algorithm free of explicit optical flow calculation. We formulate HDR video synthesis as a MAP estimation problem, which subsequently can be reduced to an optimization problem based on meaningful statistical assumptions on foreground and background regions of the input video. In the background regions the underlying scenes are static, while in the foreground regions motion information is captured implicitly by a modified 3D steering kernel regression (3D SKR) approach. Solution to the optimization problem provides us with temporally coherent HDR video sequences without noticeable artifacts. Experimental results on challenging LDR video sets demonstrate that our proposed algorithm can achieve HDR video quality that is competitive with or better than state of the art alternatives. Chul Lee, Vishal Monga |
ICIP | 2 |
| 2014 | Power-constrained RGB-to-RGBW conversion for emissive displaysabstractWe propose a novel power-constrained RGB-to-RGBW conversion algorithm for emissive RGBW displays. We measure the perceived color distortion using a color difference model in a perceptually uniform color space, and compute the power consumption for displaying an RGBW pixel on an emissive display. The main contribution is to formulate the optimization problem to minimize the color distortion subject to the constraint on the power consumption. Then, we solve it efficiently to convert an image in real time. Simulation results show that the proposed algorithm provides significantly less color distortion than the conventional methods while providing a graceful trade-off with the amount of power consumed. Chul Lee, Vishal Monga |
ICASSP | 1 |
| 2014 | Ghost-Free High Dynamic Range Imaging via Rank MinimizationabstractWe propose a ghost-free high dynamic range (HDR) image synthesis algorithm using a low-rank matrix completion framework, which we call RM-HDR. Based on the assumption that irradiance maps are linearly related to low dynamic range (LDR) image exposures, we formulate ghost region detection as a rank minimization problem. We incorporate constraints on moving objects, i.e., sparsity, connectivity, and priors on under- and over-exposed regions into the framework. Experiments on real image collections show that the RM-HDR can often provide significant gains in synthesized HDR image quality over state-of-the-art approaches. Additionally, a complexity analysis is performed which reveals computational merits of RM-HDR over recent advances in deghosting for HDR. Chul Lee, Vishal Monga |
IEEE Signal Process. Lett. | 1 |
| 2014 | Optimized Brightness Compensation and Contrast Enhancement for Transmissive Liquid Crystal DisplaysabstractAn optimized brightness-compensated contrast enhancement (BCCE) algorithm for transmissive liquid crystal displays (LCDs) is proposed in this paper. We first develop a global contrast enhancement scheme to compensate for the reduced brightness when the backlight of an LCD device is dimmed for power reduction. We also derive a distortion model to describe the information loss due to the brightness compensation. Then, we formulate an objective function that consists of the contrast enhancement term and the distortion term. By minimizing the objective function, we maximize the backlight-scaled image contrast, subject to the constraint on the distortion. Simulation results show that the proposed BCCE algorithm provides high-quality images, even when the backlight intensity is reduced by up to 50-70% to save power. Chul Lee, Jin-Hwan Kim, Chulwoo Lee, Chang-Su Kim 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Single-image deraining using an adaptive nonlocal means filterabstractAn adaptive rain streak removal algorithm for a single image is proposed in this work. We observe that a typical rain streak has an elongated elliptical shape with a vertical orientation. Thus, we first detect rain streak regions by analyzing the rotation angle and the aspect ratio of the elliptical kernel at each pixel location. We then perform the nonlocal means filtering on the detected rain streak regions by selecting nonlocal neighbor pixels and their weights adaptively. Experimental results demonstrate that the proposed algorithm removes rain streaks more efficiently and provides higher restored image qualities than conventional algorithms. Jin-Hwan Kim, Chul Lee, Jae-Young Sim, Chang-Su Kim 0001 |
ICIP | 2 |
| 2013 | Probabilistic depth-guided multi-view image denoisingabstractA novel probabilistic depth-guided multi-view denoising (PDMD) algorithm is proposed in this work. We formulate the multi-view image denoising problem by considering the uncertainties in depth estimates in noisy environments. Specifically, we employ the geometric distributions of nonlocal neighbors, as well as the block similarities, to approximate the probabilities of depth estimates. We then use those probabilities to average all nonlocal neighbors and perform the minimum mean square error (MMSE) denoising. Simulation results show that the proposed PDMD algorithm provides better denoising performance than conventional algorithms. Chul Lee, Chang-Su Kim 0001, Sang Uk Lee |
ICIP | 1 |
| 2013 | Reliable optical flow estimation in motion-blurred regionsabstractA robust optical flow estimation algorithm for motion-blurred regions is proposed in this work. We first obtain initial optical flow vectors. Then, we detect motion-blurred regions that yield low contrast, low saturation, and inconsistent optical flow vectors. We replace the optical flow vectors at motion-blurred pixels with reliable vectors at nearby unblurred pixels. To this end, we develop an energy minimization framework. Simulation results demonstrate that the proposed algorithm refines optical flow vectors in motion-blurred regions accurately and provides better performance than conventional algorithms. Yeong Jun Koh, Chul Lee, Jae-Young Sim, Chang-Su Kim 0001 |
MMSP | 2 |
| 2013 | Motion-Compensated Frame Interpolation Based on Multihypothesis Motion Estimation and Texture OptimizationabstractA novel motion-compensated frame interpolation (MCFI) algorithm to increase video temporal resolutions based on multihypothesis motion estimation and texture optimization is proposed in this paper. Initially, we form multiple motion hypotheses for each pixel by employing different motion estimation parameters, i.e., different block sizes and directions. Then, we determine the best motion hypothesis for each pixel by solving a labeling problem and optimizing the parameters. In the labeling problem, the cost function is composed of color, shape, and smoothness terms. Finally, we refine the motion hypothesis field based on the texture optimization technique and blend multiple source pixels to interpolate each pixel in the intermediate frame. Simulation results demonstrate that the proposed algorithm provides significantly better MCFI performance than conventional algorithms. Seong-Gyun Jeong, Chul Lee, Chang-Su Kim 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Contrast Enhancement Based on Layered Difference Representation of 2D HistogramsabstractA novel contrast enhancement algorithm based on the layered difference representation of 2D histograms is proposed in this paper. We attempt to enhance image contrast by amplifying the gray-level differences between adjacent pixels. To this end, we obtain the 2D histogram h(k, k + l ) from an input image, which counts the pairs of adjacent pixels with gray-levels k and k + l , and represent the gray-level differences in a tree-like layered structure. Then, we formulate a constrained optimization problem based on the observation that the gray-level differences, occurring more frequently in the input image, should be more emphasized in the output image. We first solve the optimization problem to derive the transformation function at each layer. We then combine the transformation functions at all layers into the unified transformation function, which is used to map input gray-levels to output gray-levels. Experimental results demonstrate that the proposed algorithm enhances images efficiently in terms of both objective quality and subjective quality. Chulwoo Lee, Chul Lee, Chang-Su Kim 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Exemplar-based frame rate up-conversion with congruent segmentationabstractA novel motion-compensated frame interpolation algorithm to increase video frame rates is proposed in this work, which employs texture optimization techniques to refine inaccurate motion vector fields. We enforce the congruent segment constraint in the motion refinement so that matching objects have similar local image structures. More specifically, we use the congruence energy as well as the appearance energy in the motion optimization to estimate high quality motion vectors. Simulation results show that, by preserving complex object shapes and texture, the proposed algorithm provides more faithful intermediate frames than conventional algorithms. Seong-Gyun Jeong, Chul Lee, Chang-Su Kim 0001 |
ICIP | 2 |
| 2012 | Power-constrained backlight scaling and contrast enhancement for TFT-LCD displaysabstractWe propose a novel backlight-scaled contrast enhancement algorithm to maintain the image quality when the backlight of a TFT-LCD display is dimmed for power reduction. First, we derive the transformation function to maintain the perceived luminance and enhance the contrast. Second, we model the perceptual distortion, which describes the information loss due to the backlight scaling. Then, we propose a Lagrangian method to maximize the backlight-scaled image contrast subject to the constraint on the perceptual distortion. Simulation results show that the proposed algorithm provides better image qualities than the conventional algorithms. Chul Lee, Jin-Hwan Kim, Chulwoo Lee, Chang-Su Kim 0001 |
ICIP | 1 |
| 2012 | Contrast enhancement based on layered difference representationabstractA novel contrast enhancement algorithm based on the layered difference representation is proposed in this work. We first represent gray-level differences at multiple layers in a tree-like structure. Then, based on the observation that gray-level differences, occurring more frequently in the input image, should be more emphasized in the output image, we solve a constrained optimization problem to derive the transformation function at each layer. Finally, we aggregate the transformation functions at all layers into the overall transformation function. Simulation results demonstrate that the proposed algorithm enhances images efficiently in terms of both objective quality and subjective quality. Chulwoo Lee, Chul Lee, Chang-Su Kim 0001 |
ICIP | 2 |
| 2012 | Rate-distortion optimized layered coding of high dynamic range videos
Chul Lee, Chang-Su Kim 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2012 | An MMSE approach to nonlocal image denoising: Theory and practical implementation
Chul Lee, Chulwoo Lee, Chang-Su Kim 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2012 | Power-Constrained Contrast Enhancement for Emissive Displays Based on Histogram EqualizationabstractA power-constrained contrast-enhancement algorithm for emissive displays based on histogram equalization (HE) is proposed in this paper. We first propose a log-based histogram modification scheme to reduce overstretching artifacts of the conventional HE technique. Then, we develop a power-consumption model for emissive displays and formulate an objective function that consists of the histogram-equalizing term and the power term. By minimizing the objective function based on the convex optimization theory, the proposed algorithm achieves contrast enhancement and power saving simultaneously. Moreover, we extend the proposed algorithm to enhance video sequences, as well as still images. Simulation results demonstrate that the proposed algorithm can reduce power consumption significantly while improving image contrast and perceptual quality. Chulwoo Lee, Chul Lee, Young-Yoon Lee, Chang-Su Kim 0001 |
IEEE Trans. Image Process. | 2 |
| 2011 | Caching and Deferred Write of Metadata for Yaffs2 Flash File SystemabstractFlash file systems are used for flash memory-based storage systems for more efficient use of flash memory. With the dramatic capacity increase of NAND flash memory, some NAND flash memory-based file systems are developed. The conventional NAND flash file systems synchronously write their metadata in the NAND flash for reliability. However, the synchronous writing of the metadata generates excessive garbage in the aspect of flash memory operations. In this paper, we present caching and deferred writing for Yaffs2 file system's metadata. The designed caching and deferred write scheme makes the aggregation and merging chance of subsequent metadata operations. By merging subsequent metadata updates in the cache, the number of written pages in the NAND flash is reduced. We implemented the proposed scheme of the metadata on Linux OS. The evaluation results showed that our proposed schemes reduced the aspects of the overall application time and the number of written pages by 5% to 40%, compared to the conventional flash file system. Chul Lee, Seung-Ho Lim |
EUC | 1 |
| 2011 | MMSE nonlocal means denoising algorithm for Poisson noise removalabstractA nonlocal minimum mean square error (MMSE) image denoising algorithm to remove Poisson noise is proposed in this work. Based on the Bayesian estimation theory, we first derive the nonlocal MMSE denoising filter, which can minimize the mean square error (MSE) of a denoised block. Then, we develop an approximation of the filter for practical implementation. Simulation results show that the proposed algorithm provides significantly better denoising performance than the conventional nonlocal means filter and its recent extension for Poisson noise. Chul Lee, Chulwoo Lee, Chang-Su Kim 0001 |
ICIP | 1 |
| 2011 | Gradient domain contrast enhancement with histogram-guided boundary conditionsabstractA novel contrast enhancement algorithm with histogram-guided boundary conditions is proposed in this work. The proposed algorithm enhances details in local regions by boosting gradient components, while improving the overall contrast by imposing boundary conditions based on a global transformation function. Moreover, we develop an efficient masking scheme, called the soft masking, to strike the balance between the global enhancement and the local enhancement. Simulation results demonstrate that the proposed algorithm can yield high quality output images by improving both global and local contrast simultaneously. Chulwoo Lee, Chul Lee, Chang-Su Kim 0001 |
ICIP | 2 |
| 2010 | Power-constrained contrast enhancement for OLED displays based on histogram equalizationabstractA novel power-constrained contrast enhancement algorithm for organic light-emitting diode (OLED) displays is proposed in this work. We first develop the log-modified histogram equalization (LMHE) scheme, which reduces overstretching artifacts of the conventional histogram equalization technique. Then, we model the power consumption in OLED displays, and incorporate it into LMHE to achieve the optimal tradeoff between contrast enhancement and power saving. Simulation results demonstrate that the proposed algorithm can reduce the power consumption significantly, while preserving image qualities. Chulwoo Lee, Chul Lee, Chang-Su Kim 0001 |
ICIP | 2 |
| 2008 | A Hybrid Flash File System Based on NOR and NAND Flash Memories for Embedded DevicesabstractThis paper presents a hybrid flash file system (HFFS) based on both NOR flash and NAND flash memory. In a conventional NAND flash-based flash file system, there is a trade-off between life span and durability in the frequent writing of small amounts of data. Because NAND flash supports only a page-level I/O, at least one page is wasted in the synchronous writing of small amounts of data. The wasting of pages reduces the utilization and life span of the NAND flash. To alleviate the utilization problem, some NAND flash-based flash file systems write small amounts of data asynchronously with RAM buffers, though buffering in RAM decreases the durability of the system. Our HFFS eliminates the trade-off between life span and durability. It synchronously stores data as a log in the NOR flash, whenever we append small amounts of data to a file. The merged logs are then flushed to the NAND flash in a page-aligned fashion. The implementation of our HFFS is based on our previous NAND flash-based file system, called CFFS, The experimental results reveal that our HFFS provides a longer life span than a conventional NAND flash-based synchronous flash file system with a similar level of durability. Chul Lee, Sung Hoon Baek |
IEEE Trans. Computers | 1 |
| 2007 | Gradient Domain Tone Mapping of High Dynamic Range VideosabstractA gradient domain tone mapping algorithm is proposed to display high dynamic range (HDR) video sequences in low dynamic range (LDR) devices in this work. The proposed algorithm obtains a pixelwise motion vector field and incorporates the motion information into the Poisson equation. Then, by attenuating large spatial gradients, the proposed algorithm can yield a high-quality tone-mapped result without flickering artifacts. Simulation results show that the proposed algorithm provides a better performance than the frame-based method, which processes each frame independently. Chul Lee, Chang-Su Kim 0001 |
ICIP (3) | 1 |