Zikun Liu 0001

dblp:172/9824-1 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-8697-6351ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching
abstract
Image retouching aims to enhance visual quality while aligning with users' personalized aesthetic preferences. To address the challenge of balancing controllability and subjectivity, we propose a unified diffusion-based image retouching framework called PerTouch. Our method supports semantic-level image retouching while maintaining global aesthetics. Using parameter maps containing attribute values in specific semantic regions as input, PerTouch constructs an explicit parameter-to-image mapping for fine-grained image retouching. To improve semantic boundary perception, we introduce semantic replacement and parameter perturbation mechanisms during training. To connect natural language instructions with visual control, we develop a VLM-driven agent to handle both strong and weak user instructions. Equipped with mechanisms of feedback-driven rethinking and scene-aware memory, PerTouch better aligns with user intent and captures long-term preferences. Extensive experiments demonstrate each component’s effectiveness and the superior performance of PerTouch in personalized image retouching.
Zewei Chang, Zheng-Peng Duan, Jianxing Zhang, Chunle Guo, Hyungju Chun, Hyunhee Park, Zikun Liu 0001, Chongyi Li
AAAI8
2025 FaceMe: Robust Blind Face Restoration with Personal Identification
abstract
Blind face restoration is a highly ill-posed problem due to the lack of necessary context. Although existing methods produce high-quality outputs, they often fail to faithfully preserve the individual's identity. In this paper, we propose a personalized face restoration method, FaceMe, based on a diffusion model. Given a single or a few reference images, we use an identity encoder to extract identity-related features, which serve as prompts to guide the diffusion model in restoring high-quality and identity-consistent facial images. By simply combining identity-related features, we effectively minimize the impact of identity-irrelevant features during training and support any number of reference image inputs during inference. Additionally, thanks to the robustness of the identity encoder, synthesized images can be used as reference images during training, and identity changing during inference does not require fine-tuning the model. We also propose a pipeline for constructing a reference image training pool that simulates the poses and expressions that may appear in real-world scenarios. Experimental results demonstrate that our FaceMe can restore high-quality facial images while maintaining identity consistency, achieving excellent performance and robustness.
Zheng-Peng Duan, Jia Ouyang, Jiayi Fu, Hyunhee Park, Zikun Liu 0001, Chunle Guo, Chongyi Li
AAAI6
2025 Iterative Predictor-Critic Code Decoding for Real-World Image Dehazing
abstract
We propose a novel Iterative Predictor-Critic Code Decoding framework for real-world image dehazing, abbreviated as IPC-Dehaze, which leverages the high-quality codebook prior encapsulated in a pre-trained VQGAN. Apart from previous codebook-based methods that rely on oneshot decoding, our method utilizes high-quality codes obtained in the previous iteration to guide the prediction of the Code-Predictor in the subsequent iteration, improving code prediction accuracy and ensuring stable dehazing performance. Our idea stems from the observations that 1) the degradation of hazy images varies with haze density and scene depth, and 2) clear regions play crucial cues in restoring dense haze regions. However, it is nontrivial to progressively refine the obtained codes in subsequent iterations, owing to the difficulty in determining which codes should be retained or replaced at each iteration. Another key insight of our study is to propose CodeCritic to capture interrelations among codes. The CodeCritic is used to evaluate code correlations and then resample a set of codes with the highest mask scores, i.e., a higher score indicates that the code is more likely to be rejected, which helps retain more accurate codes and predict difficult ones. Extensive experiments demonstrate the superiority of our method over state-of-the-art methods in real-world dehazing. Our project page can be found at https://github.com/Jiayi-Fu/IPC-Dehaze.
Jiayi Fu, Zikun Liu 0001, Chunle Guo, Hyunhee Park, Guoqing Wang 0001, Chongyi Li
CVPR3
2024 Restore Anything with Masks: Leveraging Mask Image Modeling for Blind All-in-One Image Restoration
Chu-Jie Qin, Zikun Liu 0001, Chunle Guo, Hyun Hee Park, Chongyi Li
ECCV (45)3
2024 Blind Face Video Restoration with Temporal Consistent Generative Prior and Degradation-Aware Prompt
abstract
Within the domain of blind face restoration (BFR), approaches lacking facial priors frequently result in excessively smoothed visual outputs. Exiting BFR methods predominantly utilize generative facial priors to achieve realistic and authentic details. However, these methods, primarily designed for images, encounter challenges in maintaining temporal consistency when applied to face video restoration. To tackle this issue, we introduce StableBFVR, an innovative Blind Face Video Restoration method based on Stable Diffusion that incorporates temporal information into the generative prior. This is achieved through the introduction of temporal layers in the diffusion process. These temporal layers consider both long-term and short-term information aggregation. Moreover, to improve generalizability, BFR methods employ complex, large-scale degradation during training, but it often sacrifices accuracy. Addressing this, StableBFVR features a novel mixed-degradation-aware prompt module, capable of encoding specific degradation information to dynamically steer the restoration process. Comprehensive experiments demonstrate that our proposed StableBFVR outperforms state-of-the-art methods.
Jingfan Tan, Hyunhee Park, Tao Wang 0052, Kaihao Zhang, Pengwen Dai, Zikun Liu 0001, Wenhan Luo
ACM Multimedia8
2024 EnsIR: An Ensemble Algorithm for Image Restoration via Gaussian Mixture Models
abstract
Image restoration has experienced significant advancements due to the development of deep learning. Nevertheless, it encounters challenges related to ill-posed problems, resulting in deviations between single model predictions and ground-truths. Ensemble learning, as a powerful machine learning technique, aims to address these deviations by combining the predictions of multiple base models. Most existing works adopt ensemble learning during the design of restoration models, while only limited research focuses on the inference-stage ensemble of pre-trained restoration models. Regression-based methods fail to enable efficient inference, leading researchers in academia and industry to prefer averaging as their choice for post-training ensemble. To address this, we reformulate the ensemble problem of image restoration into Gaussian mixture models (GMMs) and employ an expectation maximization (EM)-based algorithm to estimate ensemble weights for aggregating prediction candidates. We estimate the range-wise ensemble weights on a reference set and store them in a lookup table (LUT) for efficient ensemble inference on the test set. Our algorithm is model-agnostic and training-free, allowing seamless integration and enhancement of various pre-trained image restoration models. It consistently outperforms regression-based methods and averaging ensemble approaches on 14 benchmarks across 3 image restoration tasks, including super-resolution, deblurring and deraining. The codes and all estimated weights have been released in Github.
Shangquan Sun, Wenqi Ren, Zikun Liu 0001, Hyunhee Park, Rui Wang 0032, Xiaochun Cao
NeurIPS3
2024 Self-Supervised Recovery and Guide for Low-Resolution Person Re-Identification
abstract
Low-resolution person re-identification is a challenging task to match low-resolution (LR) probes with high-resolution (HR) gallery images. To address the resolution gap, existing methods typically recover missing details for LR probes by super-resolution, and then match the recovered HR images (instead of the original LR probes) with gallery images. However, they usually pre-specify fixed scale factors for all LR images, and ignore that choosing a preferable scale factor for each image can recover more discriminative content and accordingly benefit the re-id performance. Moreover, these methods do not focus on learning LR representations themselves and always resort to extra recovery to handle LR probes, which is quite time-consuming during inference. To tackle these problems, we propose a Self-supervised Recovery and Guide (SRG) re-id model in this paper. Given LR images during training, our model firstrecoversmore discriminative HR images by finding out preferable scale factors, and further leverages them asguideto improve original LR representations. Through enforcing LR representations to approach the self-recovered HR guide in a self-supervised manner, our model can learn more discriminative representations for LR images. As a result, our model is able to directly handle LR probes without requiring recovery during inference, thereby reducing inference time significantly. Extensive experiments demonstrate the effectiveness of our method on four datasets.
Yan Huang 0008, Liang Wang 0001, Zikun Liu 0001
IEEE Trans. Inf. Forensics Secur.4
2023 Multi-Frequency Representation Enhancement with Privilege Information for Video Super-Resolution
abstract
CNN’s limited receptive field restricts its ability to capture long-range spatial-temporal dependencies, leading to unsatisfactory performance in video super-resolution (VSR). To tackle this challenge, this paper presents a novel multi-frequency representation enhancement module (MFE) that performs spatial-temporal information aggregation in the frequency domain. Specifically, MFE mainly includes a spatial-frequency representation enhancement branch which captures the long-range dependency in the spatial dimension, and an energy frequency representation enhancement branch to obtain the inter-channel feature relationship. Moreover, a novel model training method named privilege training is proposed to encode the privilege information from high-resolution videos to facilitate model training. With these two methods, we introduce a new VSR model named MFPI, which outperforms state-of-the-art methods by a large margin while maintaining good efficiency on various datasets, including REDS4, Vimeo, Vid4, and UDM10.
Fei Li 0022, Linfeng Zhang 0001, Zikun Liu 0001, Juan Lei
ICCV3
2019 Classification Assisted Segmentation Network for Human Parsing
abstract
In human parsing task, it is important to fully exploit global and local structure information and get accurate and coherent results. In this paper, we propose a classification assisted segmentation network, in which a multi-label classification task can obtain the probability of each class in an image that used to learn better weights for parsing task. Our method takes advantages of both the global information from classification and the detail information from segmentation. Experiments demonstrate that our method could efficiently avoid the confusion between similar categories and get more reasonable results. Particularly, it significantly boosts performances of rare categories such as scarf, belt and sunglasses with mean IoU increased by 6.29%.
Zikun Liu 0001, Yinglu Liu, Zifeng Lian, Yihong Wu 0002
ICIP1
2017 Rotated region based CNN for ship detection
abstract
The state-of-the-art object detection networks for natural images have recently demonstrated impressive performances. However the complexity of ship detection in high resolution satellite images exposes the limited capacity of these networks for strip-like rotated assembled object detection which are common in remote sensing images. In this paper, we embrace this observation and introduce the rotated region based CNN (RR-CNN), which can learn and accurately extract features of rotated regions and locate rotated objects precisely. RR-CNN has three important new components including a rotated region of interest (RRoI) pooling layer, a rotated bounding box regression model and a multi-task method for non-maximal suppression (NMS) between different classes. Experimental results on the public ship dataset HRSC2016 confirm that RR-CNN outperforms baselines by a large margin.
Zikun Liu 0001, Jingao Hu, Lubin Weng
ICIP1
2017 A High Resolution Optical Satellite Image Dataset for Ship Recognition and Some New Baselines
abstract
Institute of Automation Chinese Academy of Sciences, 95 Zhongguancun East Road, 100190, Beijing, China
Zikun Liu 0001, Liu Yuan 0002, Lubin Weng
ICPRAM1
2016 Ship Rotated Bounding Box Space for Ship Extraction From High-Resolution Optical Satellite Images With Complex Backgrounds
abstract
Extracting ships from complex backgrounds is the bottleneck of ship detection in high-resolution optical satellite images. In this letter, we propose a nearly closed-form ship rotated bounding box space used for ship detection and design a method to generate a small number of highly potential candidates based on this space. We first analyze the possibility of accurately covering all ships by labeling rotated bounding boxes. Moreover, to reduce search space, we construct a nearly closed-form ship rotated bounding box space. Then, by scoring for each latent candidate in the space using a two-cascaded linear model followed by binary linear programming, we select a small number of highly potential candidates. Moreover, we also propose a fast version of our method. Experiments on our data set validate the effectiveness of our method and the efficiency of its fast version, which achieves a close detection rate in near real time.
Zikun Liu 0001, Hongzhen Wang, Lubin Weng
IEEE Geosci. Remote. Sens. Lett.1
2015 Objectness estimation using edges
abstract
Generating object proposals before object detection has become a common way. In this paper, we present a novel method to measure the objectness of bounding boxes using edges. The contours play an important role in object localization and detection. The number of edges that are close to the boundary of a box has strong relationship with the likelihood of the box covering an object. In our method, we adopt a two-step scheme to generate object proposals. In the first step, we count the number of contours close to the box, where we use the proposed “Tile Algorithm” to wipe off the inner edges of a box. In the second step we re-rank the object proposals with a linear SVM classifier across all aspect-ratios for calibration. Experiments on the VOC2007 dataset show that we achieve 96.47% object detection rate with 1000 proposals.
Hongzhen Wang, Zikun Liu 0001, Lingfeng Wang 0002, Lubin Weng, Chunhong Pan
ICIP2