EDBT 2026 Demo / reviewers in the wild / expert
Yakun Chang
dblp:192/8449
· DBLP profile ↗
22ranked-venue papers
9as first author
16since 2021 · last 2026
0009-0007-2384-386XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seam-Guided Unsupervised Image Stitching With Parallax-Aware Mask GenerationabstractImage stitching under large parallax remains a challenging task due to the conflict of content alignment and shape preservation. Most methods focus on precisely aligning overlapping regions via spatially varying transformations, often causing unexpected distortions in large-parallax areas. Differently, we aim to produce stitched images that are both visually natural and free of artifacts. To this end, we present a parallax-aware unsupervised warping model for seam-guided image stitching. To preserve natural content, we first design an edge-enhanced mask generation module to distinguish large-parallax regions and suppress excessive deformation around these areas. It is constrained by a comprehensive objective function that integrates masked photometric difference, nontrivial mask learning, and adaptive regularization, simultaneously ensuring mask reliability and alignment robustness. Besides, to eliminate parallax artifacts, we incorporate a seam-guided alignment strategy into our warping network, which iteratively registers local regions with the assistance of optimal seam estimation. Through adaptively finetuning the warping model, we progressively improve the stitching quality with improved seam quality. To facilitate the learning process of perceiving parallax, we construct a new image stitching dataset with larger parallax than that of UDISD, which could benefit the model’s generalization in challenging scenarios. Experiments show our solution not only removes misaligned regions but also maintains shape consistency especially in challenging parallax scenarios. Yuzhu Tao, Lang Nie, Yakun Chang, Shikui Wei |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | SpikeDiff: Zero-Shot High-Quality Video Reconstruction from Chromatic Spike Camera and Sub-Millisecond Spike Streams
Jinxiu Liang, Zhaojun Huang, Yeliduosi Xiaokaiti, Yakun Chang, Zhaofei Yu, Boxin Shi |
ICCV | 5 |
| 2025 | Self-supervised Image Flicker Removal for Rolling-shutter CamerasabstractAlternating current (AC)-powered artificial lighting systems often introduce high-frequency flickering, which manifests as banding-pattern flickers in images captured by rolling-shutter cameras. Existing methods rely heavily on prior knowledge of lighting systems, synthetic data, or specialized hardware, limiting their practicality in real-world scenarios. To address these limitations, we propose a self-supervised framework for flicker removal using real-world data. Our method is grounded in a theoretical analysis demonstrating that flicker-induced luminance fluctuations follow a zero-mean distribution in the temporal domain, enabling the adoption of a self-supervised strategy. We design a modified U-Net architecture and introduce the Banding Exclusion (BE) loss to suppress residual artifacts along edges while preserving structural details. To support robust training and evaluation, we curate a comprehensive dataset comprising synthetic images generated via a physics-based flicker simulation and 160 real-world RAW sequences captured under flickering illumination. Experimental results demonstrate that our framework outperforms existing methods. Shuoxin Shan, Yakun Chang, Yujia Liu 0005, Renshuai Tao, Shikui Wei, Yao Zhao 0001 |
VCIP | 2 |
| 2025 | Several Points Are All It Takes: Saluting User-Assisted Single Image Reflection RemovalabstractReflection removal is essential for applications in photography, object detection, and augmented reality. Single-image reflection removal (SIRR) offers greater flexibility and applicability than multi-image methods, making it ideal for real-time scenarios. However, strong reflections obscure large portions of the transmission layer, limiting the performance of existing methods. We propose a novel user-assisted approach for SIRR, where users annotate occluded objects by selecting their categories and locations. This provides critical semantic information to guide accurate transmission layer recovery. Additionally, we design a hybrid CNN-Transformer network that leverages local feature extraction and global context modeling to address strong reflection challenges. Experiments on the strong reflection datasets demonstrate the effectiveness of our method, achieving significant improvements in transmission layer recovery and outperforming existing advanced methods across multiple metrics. Lingzhi He, Yakun Chang, Yao Zhao 0001 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Unbiased Sample Selection and Label Improvement for Mitigating Noisy Labels in Class-Imbalanced DatasetsabstractReal-world datasets often suffer from both noisy labels and imbalanced class distribution, presenting significant challenges for the effective deployment of deep neural networks (DNNs). Existing studies typically address these challenges separately and struggle to perform effectively when they occur simultaneously. In this paper, we introduce an unbiased Sample Selection method based on the Graph Attention Network (GAT), namely GSS. GSS can effectively divide the training set into clean and noisy subsets while avoiding sample selection bias by analyzing the intrinsic relationships between the training set and a small clean validation set. For the clean subset, we propose an Adaptive Label Refinement (ALR) strategy to improve the reliability of the labels within the clean subset. ALR dynamically integrates the network’s predictions with the given labels, mitigating the adverse impacts of misidentification. For the noisy subset, we introduce a Class-Balanced Pseudo Labeling (CBPL) method. CBPL addresses the cognitive bias in model predictions caused by class imbalance by integrating class distribution information into the pseudo-label generation process, resulting in more accurate pseudo-labels. Comprehensive evaluations on both synthetic and real-world datasets highlight the effectiveness and superiority of our approach, especially in scenarios characterized by noisy labels and imbalanced class distributions. Yuan Wang 0078, Yakun Chang, Yao Zhao 0001, Shikui Wei |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | SAGNet: Decoupling Semantic-Agnostic Artifacts From Limited Training Data for Robust Generalization in Deepfake DetectionabstractDeepfake detection presents a significant challenge, particularly when the available training data is constrained to a limited set of semantic categories—a common and realistic scenario. In deepfake detection, the training labels typically indicate whether an image is real or fake, without specifying the semantic content, such as object classes. Moreover, we cannot know in advance the object categories present in an image to be detected. Ideally, a deepfake detection model should perform consistently across different semantic categories during inference, irrespective of the content. However, existing methods often exhibit significant performance bias between seen and unseen classes, struggling to generalize effectively. To address this issue, we propose Semantic-AGnostic artifact Network (SAGNet), an innovative and efficient approach designed to decouple semantic-agnostic artifacts from content-specific distributions in the training data. Our method eliminates semantic-specific biases, ensuring that the model focuses on universal artifacts related to image authenticity rather than content-dependent features. By employing this decoupling strategy, SAGNet greatly enhances the model’s generalization capacity, even when trained on limited data. Remarkably, through experiments, we demonstrate that SAGNet achieves performance comparable to models trained with 10 times more data, despite being trained on only 2 classes (comparing SAGNet trained on 2 classes in Table I with Ojha [1] trained on 20 categories in Table IV). Furthermore, through extensive experiments, we show that SAGNet’s improvements are not only evident across different semantic categories but also extend to various generative methods, including multiple GAN-based and diffusion-based models. This cross-method generalization emphasizes SAGNet’s versatility and effectiveness in diverse generative scenarios. Overall, our method represents a significant advancement in deepfake detection, particularly in realistic situations where the training data is limited. The code is released at https://github.com/rstao-bjtu/SAGNet/. Renshuai Tao, Chuangchuang Tan, Huan Liu 0030, Jiakai Wang, Haotong Qin, Yakun Chang, Wei Wang 0108, Yao Zhao 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Rethinking Depth Guided Reflection RemovalabstractWhen photographing through glass, reflections are often observed, which negatively impact the quality of the captured images or videos. In this article, we summarize and rethink depth guided reflection removal methods and, inspired by the human binocular vision system, investigate how to utilize depth for effective binocular video reflection removal. We propose an end-to-end learning-based reflection removal method that learns the transmission depth and designs a unified structure to achieve depth guided, cross-view, and cross-frame feature enhancement in a cascaded manner. Within the unified structure, different gating controllers are custom-designed to emphasize the direction of feature interaction. A dataset containing synthetic and real binocular mixture video dataset is built for network training and testing. Experimental results on both synthetic and real data from the proposed dataset demonstrate that the proposed method achieves superior performance in binocular video reflection removal. Lingzhi He, Yakun Chang, Runmin Cong, Hongyu Liu 0003, Renshuai Tao, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Colorizing Monochromatic Radiance FieldsabstractThough Neural Radiance Fields (NeRF) can produce colorful 3D representations of the world by using a set of 2D images, such ability becomes non-existent when only monochromatic images are provided. Since color is necessary in representing the world, reproducing color from monochromatic radiance fields becomes crucial. To achieve this goal, instead of manipulating the monochromatic radiance fields directly, we consider it as a representation-prediction task in the Lab color space. By first constructing the luminance and density representation using monochromatic images, our prediction stage can recreate color representation on the basis of an image colorization module. We then reproduce a colorful implicit model through the representation of luminance, density, and color. Extensive experiments have been conducted to validate the effectiveness of our approaches. Our project page: https://liquidammonia.github.io/color-nerf. Yean Cheng, Renjie Wan, Shuchen Weng, Chengxuan Zhu, Yakun Chang, Boxin Shi |
AAAI | 5 |
| 2024 | Towards HDR and HFR Video from Rolling-Mixed-Bit SpikingsabstractThe spiking cameras offer the benefits of high dynamic range (HDR), high temporal resolution, and low data redundancy. However, reconstructing HDR videos in high-speed conditions using single-bit spikings presents challenges due to the limited bit depth. Increasing the bit depth of the spikings is advantageous for boosting HDR performance, but the readout efficiency will be decreased, which is unfavorable for achieving a high frame rate (HFR) video. To address these challenges, we propose a readout mechanism to obtain rolling-mixed-bit (RMB) spikings, which involves inter-leaving multi-bit spikings within the single-bit spikings in a rolling manner, thereby combining the characteristics of high bit depth and efficient readout. Furthermore, we introduce RMB-Net for reconstructing HDR and HFR videos. RMB-Net comprises a cross-bit attention block for fusing mixed-bit spikings and a cross-time attention block for achieving temporal fusion. Extensive experiments conducted on synthetic and real-synthetic data demonstrate the superiority of our method. For instance, pure 3 -bit spikings result in 3 times of data volume, whereas our method achieves comparable performance with less than 2% increase in data volume. Yakun Chang, Yeliduosi Xiaokaiti, Yujia Liu 0005, Bin Fan 0002, Zhaojun Huang, Tiejun Huang 0001, Boxin Shi |
CVPR | 1 |
| 2024 | Real-Data-Driven 2000 FPS Color Video from Mosaicked Chromatic Spikes
Zhaojun Huang, Yakun Chang, Bin Fan 0002, Zhaofei Yu, Boxin Shi |
ECCV (12) | 3 |
| 2024 | Light Flickering Guided Reflection Removal
Yuchen Hong, Yakun Chang, Jinxiu Liang, Lei Ma 0008, Tiejun Huang 0001, Boxin Shi |
Int. J. Comput. Vis. | 2 |
| 2024 | Model-Free Rectification via Cascaded Distortion Model and Enhanced Backward Flow NetworkabstractModel-free rectification methods are limited by poor rectification quality and low generalization. This paper introduces a novel framework for enhancing model-free distortion rectification by addressing the limitations of existing methods. Our proposed method incorporates a Cascaded Distortion Model (CDM) inspired by fisheye lenses, which combines multiple reversible distortion models to create a versatile and comprehensive framework. By utilizing backward warping instead of forward warping, our approach overcomes the limitations of non-integer pixel positions and grid artifacts. Furthermore, our data synthesis method facilitates the fusion of different distortion models, bridging the distribution gap and improving generalization. To improve flow prediction accuracy, we introduce a two-stream network that incorporates both forward and backward flow branches. This approach enhances the prediction of backward flow and improves overall distortion rectification performance. We evaluate our method on large-scale synthetic datasets and real distorted images, and the results demonstrate its superior performance in both qualitative and quantitative experiments. Jie Zhao 0035, Shikui Wei, Yakun Chang, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | 1000 FPS HDR Video with a Spike-RGB Hybrid CameraabstractCapturing high frame rate and high dynamic range (HFR&HDR) color videos in high-speed scenes with conventional frame-based cameras is very challenging. The increasing frame rate is usually guaranteed by using shorter exposure time so that the captured video is severely interfered by noise. Alternating exposures can alleviate the noise issue but sacrifice frame rate due to involving long-exposure frames. The neuromorphic spiking camera records high-speed scenes of high dynamic range without colors using a completely different sensing mechanism and visual representation. We introduce a hybrid camera system composed of a spiking and an alternating-exposure RGB camera to capture HFR&HDR scenes with high fidelity. Our insight is to bring each camera's superiority into full play. The spike frames, with accurate fast motion information encoded, are firstly reconstructed for motion representation, from which the spike-based optical flows guide the recovery of missing temporal information for long-exposure RGB images while retaining their reliable color appearances. With the strong temporal constraint estimated from spike trains, both missing and distorted colors cross RGB frames are recovered to generate time-consistent and HFR color frames. We collect a new Spike-RGB dataset that contains 300 sequences of synthetic data and 20 groups of real-world data to demonstrate 1000 FPS HDR videos outperforming HDR video reconstruction methods and commercial high-speed cameras. Yakun Chang, Chu Zhou, Yuchen Hong, Liwen Hu 0002, Chao Xu 0002, Tiejun Huang 0001, Boxin Shi |
CVPR | 1 |
| 2022 | Head pose estimation using deep neural networks and 3D point clouds
Yuanquan Xu, Cheolkon Jung, Yakun Chang |
Pattern Recognit. | 3 |
| 2021 | FinerPCN: High fidelity point cloud completion network using pointwise convolution
Yakun Chang, Cheolkon Jung, Yuanquan Xu |
Neurocomputing | 1 |
| 2021 | Joint Reflection Removal and Depth Estimation From a Single ImageabstractReflection caused by glass often degrades the quality of an image and further makes it difficult to estimate depth. In this article, we propose joint reflection removal and depth estimation from a single image. We perform reflection removal (transmission recovery) and depth estimation jointly using a collaborative neural network that consists of four blocks: 1) encoder for feature extraction; 2) reflection removal subnetwork (RRN); 3) depth estimation subnetwork (DEN); and 4) depth refinement guided by the transmission layer. We achieve collaboration between reflection removal and depth estimation by concatenating intermediate features of DEN with RRN. Since the recovered transmission layer contains accurate edges of objects behind glass, we refine the estimated depth with its guidance by guided image filtering. The experimental results demonstrate that the proposed method achieves both reflection removal and depth estimation even for images with dominant reflections. Besides, this article offers a new way of treating reflections in images to introduce depth estimation into reflection removal and achieve reflection removal and depth estimation simultaneously. Yakun Chang, Cheolkon Jung |
IEEE Trans. Cybern. | 1 |
| 2020 | Siamese Dense Network for Reflection Removal with Flash and No-Flash Image Pairs
Yakun Chang, Cheolkon Jung, Fengqiao Wang |
Int. J. Comput. Vis. | 1 |
| 2020 | Automatic cardiac MRI segmentation and permutation-invariant pathology classification using deep neural networks and point clouds
Yakun Chang, Cheolkon Jung |
Neurocomputing | 1 |
| 2019 | Multi-Modal Reflection Removal Using Convolutional Neural NetworksabstractAlthough color images are easily interfered by glass, depth images captured by infrared sensors are robust to reflection. In this letter, we propose multi-modal reflection removal using convolutional neural networks (CNNs). We build a multi-modal CNN for reflection removal to separate transmission from reflection using depth information. The proposed network consists of two sub-networks: image restoration and depth adaptation. Image restoration sub-network (IRN) recovers transmission layer from the input image with reflection, whereas depth adaptation sub-network (DAN) guides reflection removal of the IRN. Moreover, to extract image details for reflection removal, we present a multi-scale loss function that penalizes non-similarity for multi-scale outputs. Experimental results demonstrate that the proposed method is robust to dominant reflections and outperforms state-of-the-art methods in terms of both peak signal-to-noise ratio (PSNR) and structural similarity. Yakun Chang, Cheolkon Jung |
IEEE Signal Process. Lett. | 2 |
| 2019 | Single Image Reflection Removal Using Convolutional Neural NetworksabstractWhen people take a picture through glass, the scene behind the glass is often interfered by specular reflection. Due to relatively easy implementation, most studies have tried to recover the transmitted scene from multiple images rather than single image. However, the use of multiple images is not practical for common users in real situations due to the critical shooting conditions. In this paper, we propose single image reflection removal using convolutional neural networks. We provide a ghosting model that causes reflection effects in captured images. First, we synthesize multiple reflection images from the input single one based on ghosting model and relative intensity. Then, we construct an end-to-end network that consists of encoder and decoder. To optimize the network parameters, we use a joint training strategy to learn the layer separation knowledge from the synthesized reflection images. For the loss function, we utilize both internal and external losses in optimization. Finally, we apply the proposed network to single image reflection removal. Compared with the previous work, the proposed method does not need handcrafted features and specular filters for reflection removal. Experimental results show that the proposed method successfully removes reflection from both synthetic and real images as well as achieves the highest scores in PSNR, SSIM and FSIM. Yakun Chang, Cheolkon Jung |
IEEE Trans. Image Process. | 1 |
| 2018 | Automatic Segmentation and Cardiopathy Classification in Cardiac Mri Images Based on Deep Neural NetworksabstractSegmentation of cardiac MRI images plays a key role in clinical diagnosis. In the traditional diagnostic process, clinical experts manually segment left ventricle (LV), right ventricle (RV) and myocardium to obtain guideline for cardiopathy diagnosis. However, manual segmentation is time-consuming and labor-intensive. In this paper, we propose automatic segmentation and cardiopathy classification in cardiac MRI images based on deep neural networks. First, we perform object detection based on a YOLO-based network to get region of interest (ROI) from the whole sequence of diastolic and systolic MRI. Then, we obtain a pixel-wise segmentation mask automatically by feeding ROI into fully convolutional neural networks (FCN). Finally, we construct a fully connected network for cardiopathy diagnosis to decide a heart disease from the given MRI. Experimental results show that the proposed method successfully segments LV, RV and myocardium as well as achieves 90% accuracy in heart disease classification. Yakun Chang, Baoyu Song, Cheolkon Jung |
ICASSP | 1 |
| 2016 | Perceptual contrast enhancement of dark images based on textural coefficientsabstractWe propose perceptual contrast enhancement of dark images based on textural coefficients. The textural coefficient indicates textural degree of intensity and adaptively stretches the dynamic range in an image. First, we calculate gray level difference between a central pixel and its adjacent ones. Because some differences are obviously noticeable by human eyes, we only use unnoticeable differences to obtain the textural coefficient. We apply the just noticeable difference (JND) of the human visual system (HVS) to obtain the proper threshold. Then, we apply a Gaussian kernel to texture coefficients for avoiding excessive differences between adjacent ones. Finally, we perform optimal contrast tone mapping to obtain a mapping function. Experimental results show that the proposed method successfully enhances dark regions while avoiding over-enhancement in bright regions without halo artifact and tone distortion. Yakun Chang, Cheolkon Jung |
VCIP | 1 |