EDBT 2026 Demo / reviewers in the wild / expert
Yan-Tsung Peng
dblp:01/4192
· DBLP profile ↗
31ranked-venue papers
11as first author
21since 2021 · last 2025
0000-0002-3802-1670ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 11 first-author · 16 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HVDualformer: Histogram-Vision Dual Transformer for White BalanceabstractCapturing images under different color temperatures can result in color casts, causing the color presented in photos to differ from what is perceived by the human eye. Correcting these color temperature shifts to achieve White Balance (WB) is a challenging task, requiring the identification of variations in color tones from diverse light sources and the removal of color casts. The advent of deep neural networks has significantly advanced the progress of WB methods, evolving from simply identifying the scene illumination color to directly producing a color-corrected image from the color-shifted input. To better map color distributions and scene information from the input to the WB image, we propose HVDualformer, an end-to-end histogram-vision dual transformer architecture that can rectify color temperature features from WB color histograms and exploit them to adjust image features to yield accurate WB results. Extensive experimental results on public benchmark datasets demonstrate that the proposed model performs favorably against state-of-the-art methods. Yan-Tsung Peng |
AAAI | 1 |
| 2025 | ABC-Former: Auxiliary Bimodal Cross-domain Transformer with Interactive Channel Attention for White BalanceabstractThe primary goal of white balance (WB) for sRGB images is to correct inaccurate color temperatures, ensuring that images display natural, neutral colors. While existing WB methods yield reasonable results, their effectiveness is limited. They either focus solely on global color adjustments applied before the camera-specific image signal processing pipeline or rely on end-to-end models that generate WB outputs without accounting for global color trends, leading to suboptimal correction. To address these limitations, we propose an Auxiliary Bimodal Cross-domain Transformer (ABC-Former) that enhances WB correction by leveraging complementary knowledge from global color information from CIELab and RGB histograms alongside sRGB inputs. By introducing an Interactive Channel Attention (ICA) module to facilitate cross-modality global knowledge transfer, ABC-Former achieves more precise WB correction. Experimental results on benchmark WB datasets show that ABC-Former performs favorably against state- of-the-art WB methods. The source code is available at https://github.com/ytpeng-aimlab/ABC-Former. Yu-Cheng Chiu, Zihao Chen 0003, Yan-Tsung Peng |
CVPR | 4 |
| 2025 | PHATNet: A Physics-Guided Haze Transfer Network for Domain-Adaptive Real-World Image DehazingabstractImage dehazing aims to remove unwanted hazy artifacts in images. Although previous research has collected paired real-world hazy and haze-free images to improve dehazing models' performance in real-world scenarios, these models often experience significant performance drops when handling unseen real-world hazy images due to limited training data. This issue motivates us to develop a flexible domain adaptation method to enhance dehazing performance during testing. Observing that predicting haze patterns is generally easier than recovering clean content, we propose the Physics-guided Haze Transfer Network (PHATNet) which transfers haze patterns from unseen target domains to source-domain haze-free images, creating domain-specific fine-tuning sets to update dehazing models for effective domain adaptation. Additionally, we introduce a Haze-Transfer-Consistency loss and a Content-Leakage Loss to enhance PHATNet's disentanglement ability. Experimental results demonstrate that PHATNet significantly boosts state-of-the-art dehazing models on benchmark real-world image dehazing datasets. Fu-Jen Tsai, Yan-Tsung Peng, Yen-Yu Lin, Chia-Wen Lin |
ICCV | 2 |
| 2025 | Enhancing Image Deraining Through VLM-Based Data Refinement and ClassificationabstractImage deraining has gained significant attention in recent years due to its essential role in applications like autonomous driving and surveillance systems. Despite their good performance in rain removal, most image-deraining models are trained on synthetic datasets, leading to a performance gap when applied to real-world scenarios due to differences between synthetic and actual rain patterns. Although some real-world datasets exist, they may contain unsuitable images for training image-deraining models, such as those without clear rain or containing haze, which could lead to suboptimal model performance. Therefore, we propose Automated Understanding and Refinement Agents (AURA) to refine deraining datasets by leveraging large vision-language models to filter out inappropriate training data and categorize images based on the rain density (e.g., light, moderate, and heavy rain). This automated refinement and annotation process enhances dataset quality, resulting in improved deraining model performance in real-world scenarios. Experimental results demonstrate that training deraining models on datasets refined and annotated by AURA can significantly enhance their deraining performance. Shih-Jui Liang, Zihao Chen 0003, Jun-Cheng Chen, Yan-Tsung Peng |
ICIP | 4 |
| 2025 | ATARS: An Aerial Traffic Atomic Activity Recognition and Temporal Segmentation DatasetabstractTraffic Atomic Activity, which describes traffic patterns for topological intersection dynamics, is a crucial topic for the advancement of intelligent driving systems. However, existing atomic activity datasets are collected from an egocentric view, which cannot support the scenarios where traffic activities in an entire intersection are required. Moreover, existing datasets only provide video-level atomic activity annotations, which require exhausting efforts to manually trim the videos for recognition and limit their applications to untrimmed videos. To bridge this gap, we introduce the Aerial Traffic Atomic Activity Recognition and Segmentation (ATARS) dataset, the first aerial dataset designed for multilabel atomic activity analysis. We offer atomic activity labels for each frame, which accurately record the intervals for traffic activities. Moreover, we propose a novel task, Multi-label Temporal Atomic Activity Recognition, enabling the study of accurate temporal localization for atomic activity and easing the burden of manual video trimming for recognition. We conduct extensive experiments to evaluate existing state-of-theart models on both atomic activity recognition and temporal atomic activity segmentation. The results highlight the unique challenges of our ATARS dataset, such as recognizing extremely small objects’ activities. We further provide a comprehensive discussion analyzing these challenges and offer valuable insights for future direction to improve recognition of atomic activity in an aerial view. Our source code and dataset are available at https://github.com/magecliff96/ATARS/. Zihao Chen 0003, Hsuanyu Wu, Chi-Hsi Kung, Yi-Ting Chen 0001, Yan-Tsung Peng |
IROS | 5 |
| 2025 | Restore Anything Anywhere: Targeted Image Restoration with Object Segmentation and Text GuidanceabstractImage restoration techniques are widely used in various fields, such as autonomous driving, medical imaging, and satellite imagery. These techniques typically aim to restore an entire degraded image to its original state. However, in many cases, users may wish to focus on restoring specific areas to achieve desired effects in the image, tailored to their preferences. In this work, we propose Restore Anything anyWhere (RAW), a framework that enables users to restore specific types of degradation on any selected object in an image by specifying point or text prompts. Our framework first employs the Multimodal Segmentation module to generate mask priors for the target objects. Then, the text-guided restoration model performs targeted restoration on the areas defined by the object mask priors. To improve segmentation performance, we propose a Text-To-Segmentation Refinement method that combines existing techniques with the capabilities of Segment Anything, enhancing text-to-segmentation accuracy. Additionally, we introduce a streamlined text-guided restoration method to provide better control over degradation restoration, delivering highly effective results. RAW provides excellent object-level control in restoration and significantly improves the Laion Aesthetics score. Yen-Ku Yeh, Chun-Hao Yang, Kun-Tai Wu, Yan-Tsung Peng, Chun-Rong Huang, Jun-Cheng Chen |
MMSP | 4 |
| 2025 | BlurDM: A Blur Diffusion Model for Image DeblurringabstractDiffusion models show promise for dynamic scene deblurring; however, existing studies often fail to leverage the intrinsic nature of the blurring process within diffusion models, limiting their full potential. To address it, we present a Blur Diffusion Model (BlurDM), which seamlessly integrates the blur formation process into diffusion for image deblurring. Observing that motion blur stems from continuous exposure, BlurDM implicitly models the blur formation process through a dual-diffusion forward scheme, diffusing both noise and blur onto a sharp image. During the reverse generation process, we derive a dual denoising and deblurring formulation, enabling BlurDM to recover the sharp image by simultaneously denoising and deblurring, given pure Gaussian noise conditioned on the blurred image as input. Additionally, to efficiently integrate BlurDM into deblurring networks, we perform BlurDM in the latent space, forming a flexible prior generation network for deblurring. Extensive experiments demonstrate that BlurDM significantly and consistently enhances existing deblurring methods on four benchmark datasets. The project page is available at https://jin-ting-he.github.io/BlurDM/. Jin-Ting He, Fu-Jen Tsai, Yan-Tsung Peng, Min-Hung Chen, Chia-Wen Lin, Yen-Yu Lin |
NeurIPS | 3 |
| 2025 | FAME: a lightweight spatio-temporal network for model attribution of face-swap deepfakes
Yan-Tsung Peng, Yuan-Hao Chang 0001 |
Expert Syst. Appl. | 2 |
| 2025 | AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face VideosabstractMultimodal manipulations (also known as audio-visual deepfakes) make it difficult for unimodal deepfake detectors to detect forgeries in multimedia content. To avoid the spread of false propaganda and fake news, timely detection is crucial. The damage to either modality (i.e., visual or audio) can only be discovered through multimodal models that can exploit both pieces of information simultaneously. However, previous methods mainly adopt unimodal video forensics and use supervised pretraining for forgery detection. This study proposes a new method based on a multimodal self-supervised-learning (SSL) feature extractor to exploit inconsistency between audio and visual modalities for multimodal video forgery detection. We use the transformer-based SSL pretrained Audio-Visual HuBERT (AV-HuBERT) model as a visual and acoustic feature extractor and a multiscale temporal convolutional neural network to capture the temporal correlation between the audio and visual modalities. Since AV-HuBERT only extracts visual features from the lip region, we also adopt another transformer-based video model to exploit facial features and capture spatial and temporal artifacts caused during the deepfake generation process. Experimental results show that our model outperforms all existing models and achieves new state-of-the-art performance on the FakeAVCeleb and DeepfakeTIMIT datasets. Sahibzada Adil Shahzad, Ammarah Hashmi, Yan-Tsung Peng, Yu Tsao 0001, Hsin-Min Wang |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2025 | Rain2Avoid: Learning Deraining by Self-SupervisionabstractImages captured on rainy days often contain rain streaks that can obscure important scenery and degrade the performance of high-level vision tasks, such as image segmentation in autonomous vehicles. As a result, image deraining, a low-level vision task focused on removing rain streaks from images, has gained popularity over the past decade. Recent advancements have primarily concentrated on supervised image deraining methods, which rely on paired rain-clean image datasets to train deep neural network models. However, collecting such paired real data is challenging and time-consuming. To address this, our method introduces a novel self-supervised approach that leverages the proposed locally dominant gradient prior and non-local self-similarity stochastic sampling. This approach extracts potential rain streaks and generates stochastic derained references for image deraining. Experimental results on public benchmark image-deraining datasets show that our proposed method performs favorably against state-of-the-art few-shot and self-supervised image deraining methods. Yan-Tsung Peng, Zihao Chen 0003 |
IEEE Trans. Multim. | 1 |
| 2025 | CapST: Leveraging Capsule Networks and Temporal Attention for Accurate Model Attribution in Deep-fake VideosabstractDeep-fake videos, generated through AI face-swapping techniques, have garnered considerable attention due to their potential for impactful impersonation attacks. While existing research primarily distinguishes real from fake videos, attributing a deep-fake to its specific generation model or encoder is crucial for forensic investigation, enabling precise source tracing and tailored countermeasures. This approach not only enhances detection accuracy by leveraging unique model-specific artifacts but also provides insights essential for developing proactive defenses against evolving deep-fake techniques. Addressing this gap, this article investigates the model attribution problem for deep-fake videos using two datasets—Deepfakes from Different Models (DFDM) and GANGen-Detection, which comprise deep-fake videos and images generated by GAN models. We select only fake images from the GANGen-Detection dataset to align with the DFDM dataset, which specifies the goal of this study, focusing on model attribution rather than real/fake classification. This study formulates deep-fake model attribution as a multiclass classification task, introducing a novel Capsule-Spatial-Temporal (CapST) model that effectively integrates a modified VGG19 (utilizing only the first 26 out of 52 layers) for feature extraction, combined with Capsule Networks and a Spatio-Temporal attention mechanism. The Capsule module captures intricate feature hierarchies, enabling robust identification of deep-fake attributes, while a video-level fusion technique leverages temporal attention mechanisms to process concatenated feature vectors and capture temporal dependencies in deep-fake videos. By aggregating insights across frames, our model achieves a comprehensive understanding of video content, resulting in more precise predictions. Experimental results on the DFDM and GANGen-Detection datasets demonstrate the efficacy of CapST, achieving substantial improvements in accurately categorizing deep-fake videos over baseline models, all while demanding fewer computational resources. Yan-Tsung Peng, Yuan-Hao Chang 0001, Gaddisa Olani Ganfure, Sarwar Khan |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | ID-Blau: Image Deblurring by Implicit Diffusion-Based reBLurring AUgmentationabstractImage deblurring aims to remove undesired blurs from an image captured in a dynamic scene. Much research has been dedicated to improving deblurring performance through model architectural designs. However, there is little work on data augmentation for image deblurring. Since continuous motion causes blurred artifacts during image exposure, we aspire to develop a groundbreaking blur augmentation method to generate diverse blurred images by simulating motion trajectories in a continuous space. This paper proposes Implicit Diffusion-based reBLurring AUgmentation (ID-Blau), utilizing a sharp image paired with a controllable blur condition map to produce a corresponding blurred image. We parameterize the blur patterns of a blurred image with their orientations and magnitudes as a pixel-wise blur condition map to simulate motion trajectories and implicitly represent them in a continuous space. By sampling diverse blur conditions, ID-Blau can generate various blurred images unseen in the training set. Experimental results demonstrate that ID-Blau can produce realistic blurred images for training and thus significantly improve performance for state-of-the-art deblurring models. The source code is available at https://github.com/plusgood-steven/ID-Blau. Fu-Jen Tsai, Yan-Tsung Peng, Chung-Chi Tsai, Chia-Wen Lin, Yen-Yu Lin |
CVPR | 3 |
| 2024 | Domain-Adaptive Video Deblurring via Test-Time Blurring
Jin-Ting He, Fu-Jen Tsai, Yan-Tsung Peng, Chung-Chi Tsai, Chia-Wen Lin, Yen-Yu Lin |
ECCV (30) | 4 |
| 2023 | Rain2Avoid: Self-Supervised Single Image DerainingabstractThe single image deraining task aims to remove rain from a single image, attracting much attention in the field. Recent research on this topic primarily focuses on discriminative deep learning methods, which train models on rainy images with their clean counterparts. However, collecting such paired images for training takes much work. Thus, we present Rain2Avoid (R2A), a training scheme that requires only rainy images for image deraining. We propose a locally dominant gradient prior to reveal possible rain streaks and overlook those rain pixels while training with the input rainy image directly. Understandably, R2A may not perform as well as deraining methods that supervise their models with rain-free ground truth. However, R2A favors when training image pairs are unavailable and can self-supervise only one rainy image for deraining. Experimental results show that the proposed method performs favorably against state-of-the-art few-shot deraining and self-supervised denoising methods. Yan-Tsung Peng |
ICASSP | 1 |
| 2022 | Meta Transferring for Deblurring
Po-Sheng Liu, Fu-Jen Tsai, Yan-Tsung Peng, Chung-Chi Tsai, Chia-Wen Lin, Yen-Yu Lin |
BMVC | 3 |
| 2022 | Stripformer: Strip Transformer for Fast Image Deblurring
Fu-Jen Tsai, Yan-Tsung Peng, Yen-Yu Lin, Chung-Chi Tsai, Chia-Wen Lin |
ECCV (19) | 2 |
| 2022 | Single Image Reflection Removal Based on Knowledge-Distilling Content DisentanglementabstractWhen we shoot pictures through transparent media, such as glass, reflection can undesirably occur, obscuring the scene we intended to capture. Therefore, removing reflection is practical in image restoration. However, a reflective scene mixed with that behind the glass is challenging to be separated, considered significantly ill-posed. This letter addresses the single image reflection removal (SIRR) problem by proposing a knowledge-distilling-based content disentangling model that can effectively decompose the transmission and reflection layers. The experiments on benchmark SIRR datasets demonstrate that our method performs favorably against state-of-the-art SIRR methods. Yan-Tsung Peng, Kai-Han Cheng, I-Sheng Fang, Wen-Yi Peng, Jr-Shian Wu |
IEEE Signal Process. Lett. | 1 |
| 2022 | BANet: A Blur-Aware Attention Network for Dynamic Scene DeblurringabstractImage motion blur results from a combination of object motions and camera shakes, and such blurring effect is generally directional and non-uniform. Previous research attempted to solve non-uniform blurs using self-recurrent multi-scale, multi-patch, or multi-temporal architectures with self-attention to obtain decent results. However, using self-recurrent frameworks typically leads to a longer inference time, while inter-pixel or inter-channel self-attention may cause excessive memory usage. This paper proposes a Blur-aware Attention Network (BANet), that accomplishes accurate and efficient deblurring via a single forward pass. Our BANet utilizes region-based self-attention with multi-kernel strip pooling to disentangle blur patterns of different magnitudes and orientations and cascaded parallel dilated convolution to aggregate multi-scale content features. Extensive experimental results on the GoPro and RealBlur benchmarks demonstrate that the proposed BANet performs favorably against the state-of-the-arts in blurred image restoration and can provide deblurred results in real-time. Fu-Jen Tsai, Yan-Tsung Peng, Chung-Chi Tsai, Yen-Yu Lin, Chia-Wen Lin |
IEEE Trans. Image Process. | 2 |
| 2022 | Two Exposure Fusion Using Prior-Aware Generative Adversarial NetworkabstractProducing a high dynamic range (HDR) image from two low dynamic range (LDR) images with extreme exposures is challenging due to the lack of well-exposed contents. Existing works either use pixel fusion based on weighted quantization or conduct feature fusion using deep learning techniques. In contrast to these methods, our core idea is to progressively incorporate the pixel domain knowledge of LDR images into the feature fusion process. Specifically, we propose a novel Prior-Aware Generative Adversarial Network (PA-GAN), along with a new dual-level loss for two exposure fusion. The proposed PA-GAN is composed of a content-prior-guided encoder and a detail-prior-guided decoder, respectively in charge of content fusion and detail calibration. We further train the network using a dual-level loss that combines the semantic-level loss and pixel-level loss. Extensive qualitative and quantitative evaluations on diverse image datasets demonstrate that our proposed PA-GAN has superior performance than state-of-the-art methods. Jia-Li Yin, Yan-Tsung Peng |
IEEE Trans. Multim. | 3 |
| 2022 | Automatic Intermediate Generation With Deep Reinforcement Learning for Robust Two-Exposure Image FusionabstractFusing low dynamic range (LDR) for high dynamic range (HDR) images has gained a lot of attention, especially to achieve real-world application significance when the hardware resources are limited to capture images with different exposure times. However, existing HDR image generation by picking the best parts from each LDR image often yields unsatisfactory results due to either the lack of input images or well-exposed contents. To overcome this limitation, we model the HDR image generation process in two-exposure fusion as a deep reinforcement learning problem and learn an online compensating representation to fuse with LDR inputs for HDR image generation. Moreover, we build a two-exposure dataset with reference HDR images from a public multiexposure dataset that has not yet been normalized to train and evaluate the proposed model. By assessing the built dataset, we show that our reinforcement HDR image generation significantly outperforms other competing methods under different challenging scenarios, even with limited well-exposed contents. More experimental results on a no-reference multiexposure image dataset demonstrate the generality and effectiveness of the proposed model. To the best of our knowledge, this is the first work to use a reinforcement-learning-based framework for an online compensating representation in two-exposure image fusion. Jia-Li Yin, Yan-Tsung Peng, Hau Hwang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Deep Battery Saver: End-to-End Learning for Power Constrained Contrast EnhancementabstractDue to the problems of power-hungry displays and limited battery life in electronic devices, the concept of “green computing,” which entails a reduction in power consumption, is proposed. One often seen green computing is the power-constrained contrast enhancement (PCCE), yet it is much more challenging because of the noticeable local intensity suppressions in images. This paper aims at developing an image-quality-lossless end-to-end learning network called deep battery saver to achieve power savings in emissive displays, i.e., produce power-saved images with high perceptual quality and less power consumption. Built upon the end-to-end network of the displayed image, we propose a variational loss function for enhancing the visual quality and suppressing the power consumption, simultaneously. The basic idea is to integrate both high-level perceptual losses and low-level pixel losses by a deep residual convolutional neural network (CNN) over a devised variational loss function with strong human perceptual consistency. Such deep residual CNN network leads to a visually pleasing image representation during the suppression of power consumption. Experimental results demonstrated the superiority of our deep battery saver to existing PCCE methods. Jia-Li Yin, Yan-Tsung Peng, Chung-Chi Tsai |
IEEE Trans. Multim. | 3 |
| 2020 | Deep Prior Guided Network For High-Quality Image FusionabstractHigh dynamic range imaging requires fusing a set of low dynamic range (LDR) images at different exposure levels. Existing works combine the LDRs by either assigning each LDR a weighting map based on texture metrics at the pixel level or transferring the images into semantic space at the feature level while neglecting the fact that both texture calibration and semantic consistency are required. In this paper, we propose a novel encoder-decoder network consisting of a content prior guided (CPG) encoder and a detail prior guided (DPG) decoder for fusing the images at both the pixel level and feature level. Explicitly, the encoder constructed by the CPG layers includes the pyramid content prior to blend at the pixel level to transform the feature maps in the encoding layers. Correspondingly, the decoder comprises the DPG layers incorporated with the Laplacian pyramid detail prior to further boost the fusion performance. As the content and the detail priors are added to the network in a pyramid-structure manner, which provides fine-grained control to the features, both semantic consistency and texture calibration can be assured. Extensive experiments demonstrated the superiority of our method over existing state-of-the-art methods. Jia-Li Yin, Yan-Tsung Peng, Chung-Chi Tsai |
ICME | 3 |
| 2020 | Image Haze Removal Using Airlight White Correction, Local Light Filter, and Aerial Perspective PriorabstractLight is scattered and absorbed when travelling through atmosphere particles, leading to visibility attenuation for images captured, especially in hazy scenes. In addition, hazy images may suffer from color distortion caused by haze or sandstorm, resulting in a poor visual quality. In order to effectively enhance visibility and correct possible color casts for such images, we propose a new image dehazing algorithm based on an improved haze optical model, which consists of three modules: airlight white correction (AWC), local light filter (LLF), and aerial perspective prior (APP). In the proposed algorithm, the AWC module detects and corrects possible color cast, the LLF module downplays non-hazy bright pixels (e.g., headlight and white objects) for more accurate airlight estimation, and the APP module uses the minimum/maximum channel and their difference for scene transmission estimation. The experimental results demonstrate that the proposed method outperforms other state-of-the-art dehazing methods in three ways: 1) our results have better visual quality; 2) our method performs the best in terms of color restoration; and 3) our method is very efficient at removing haze and color casts. Yan-Tsung Peng, Zhihui Lu 0002, Fan-Chieh Cheng, Yalun Zheng, Shih-Chia Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Enhancing object detection in the dark using U-Net based restoration moduleabstractIn recent years, we have witnessed the widespread application of deep-learning techniques to various surveillance tasks, including human tracking and counting, abnormal behavior detection, and video segmentation. In most cases, the input images/videos are assumed to possess adequate visual quality to guarantee satisfactory performance. However, accuracy may be adversely affected when the input data are degraded by factors such as excessive noise or poor lighting conditions. In the paper, we develop a deep neural network based on the U-Net architecture that acts as a pre-processing module to restore images/videos with nonuniform light sources to ensure the accuracy of the subsequent object detection process. Experimental results on the VisDrone20 19 dataset [1] demonstrate the effectiveness of the proposed method, achieving a remarkable 5% increase in average recall. We expect the framework to be universally applicable to situations that call for the enhancement of raw input data. Yen-Ting Huang, Yan-Tsung Peng, Wen-Hung Liao |
AVSS | 2 |
| 2019 | Image Denoising Based on Overlapped and Adaptive Gaussian Smoothing and Convolutional Refinement NetworksabstractWe propose to use overlapped and adaptive Gaussian smoothing (OAGS) and convolutional refinement networks (CRN) to recover images corrupted by salt-and-pepper noise. First, the OAGS method identifies noise pixels and recover them. Then, CRN further improve and restore the recovered results with sharper and clearer edges. Experimental results demonstrate the proposed OAGS+CRN method significantly outperforms state-of-the-art denoising methods. Yan-Tsung Peng, Ming-Hao Lin, Chun-Lin Tang, Chin-Hsien Wu |
ISM | 1 |
| 2019 | Color Shifting-Aware Image DehazingabstractBuilt upon an image formation model for a single hazy image, existing image dehazing methods typically restore hazed pixels by estimating the unknown transmission map and global ambient light via exploiting image priors. They often produce visually unpleasing results when hazy images are with unwanted color shifts due to inaccurate estimation about the actual ambient light of hazy images with color shifts. To address the problem, we propose a novel color shifting-aware image dehazing model that explicitly disentangles the inference of the image formation model. Specifically, our model attempts to calibrate color fading and shifting first, and then restores the hazed pixels via the scene depth based gamma correction using the color-corrected image as the guidance. Extensive experiments show that the proposed dehazing model significantly outperforms existing dehazing methods and achieves superior dehazing results on challenging cases with unwanted color casts. Jia-Li Yin, Yan-Tsung Peng |
ISM | 3 |
| 2018 | Generalization of the Dark Channel Prior for Single Image RestorationabstractImages degraded by light scattering and absorption, such as hazy, sandstorm, and underwater images, often suffer color distortion and low contrast because of light traveling through turbid media. In order to enhance and restore such images, we first estimate ambient light using the depth-dependent color change. Then, via calculating the difference between the observed intensity and the ambient light, which we call the scene ambient light differential, scene transmission can be estimated. Additionally, adaptive color correction is incorporated into the image formation model (IFM) for removing color casts while restoring contrast. Experimental results on various degraded images demonstrate the new method outperforms other IFM-based methods subjectively and objectively. Our approach can be interpreted as a generalization of the common dark channel prior (DCP) approach to image restoration, and our method reduces to several DCP variants for different special cases of ambient lighting and turbid medium conditions. Yan-Tsung Peng, Keming Cao, Pamela C. Cosman |
IEEE Trans. Image Process. | 1 |
| 2017 | Underwater Image Restoration Based on Image Blurriness and Light AbsorptionabstractUnderwater images often suffer from color distortion and low contrast, because light is scattered and absorbed when traveling through water. Such images with different color tones can be shot in various lighting conditions, making restoration and enhancement difficult. We propose a depth estimation method for underwater scenes based on image blurriness and light absorption, which can be used in the image formation model (IFM) to restore and enhance underwater images. Previous IFM-based image restoration methods estimate scene depth based on the dark channel prior or the maximum intensity prior. These are frequently invalidated by the lighting conditions in underwater images, leading to poor restoration results. The proposed method estimates underwater scene depth more accurately. Experimental results on restoring real and synthesized underwater images demonstrate that the proposed method outperforms other IFM-based underwater image restoration methods. Yan-Tsung Peng, Pamela C. Cosman |
IEEE Trans. Image Process. | 1 |
| 2016 | Single image restoration using scene ambient light differentialabstractIn this paper, we restore images degraded by scattering and absorption such as hazy, sandstorm, and underwater images. By calculating the difference between the observed intensity and the ambient light in a degraded image scene, which we call the scene ambient light differential, we estimate the transmission map. In the restoration process, we first enhance the degraded images based on the proposed transmission estimation using the image formation model, and then use an adaptive color correction method to restore color. Experimental results on various degraded images demonstrate the proposed method outperforms other enhancement and restoration methods. Yan-Tsung Peng, Pamela C. Cosman |
ICIP | 1 |
| 2015 | Single underwater image enhancement using depth estimation based on blurrinessabstractIn this paper, we propose to use image blurriness to estimate the depth map for underwater image enhancement. It is based on the observation that objects farther from the camera are more blurry for underwater images. Adopting image blurriness with the image formation model (IFM), we can estimate the distance between scene points and the camera and thereby recover and enhance underwater images. Experimental results on enhancing such images in different lighting conditions demonstrate the proposed method performs better than other IFM-based enhancement methods. Yan-Tsung Peng, Xiangyun Zhao, Pamela C. Cosman |
ICIP | 1 |
| 2013 | Histogram shrinking for power-saving contrast enhancementabstractIn this paper, a power-saving method for emissive display by shrinking histogram is proposed. Based on a modern pixel-level power model of an OLED module, the power consumption factor can be employed in the objective function. Nevertheless, contrast enhancement intrinsically contradicts saving power. In order to solve this problem, we formulate a new objective function which is subject to the constant entropy. By minimizing the distance between two near non-empty bins of image histogram, the power reduction and entropy preservation are simultaneously achieved. To further enhance the perceptional quality, the proposed method is also integrated with other related algorithms. Experimental results show that the proposed method is capable of reducing display power, while the performance of contrast enhancement is also improved. Yan-Tsung Peng, Fan-Chieh Cheng, Li-Ming Jan, Shanq-Jang Ruan |
ICIP | 1 |