Jiayi Fu

dblp:189/4508 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Image and video processing · 88% Visual content generation and editing · 12%
Artificial intelligence
2 papers
Face, body and person analysis · 70% Language models and text generation · 30%
Network and information security
1 paper
Digital forensics and information hiding · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing
image restoration
1.722025
Iterative Predictor-Critic Code Decoding for Real-World Image Dehazing · CVPR 2025
FaceMe: Robust Blind Face Restoration with Personal Identification · AAAI 2025
Image and video processing › video frame interpolation
high-resolution video frame interpolation
1.012026
VTinker: Guided Flow Upsampling and Texture Mapping for High-Resolution Video Frame Interpolation · AAAI 2026
Image and video processing › motion estimation
optical flow
1.012026
VTinker: Guided Flow Upsampling and Texture Mapping for High-Resolution Video Frame Interpolation · AAAI 2026
Image and video processing
video frame interpolation
1.012026
VTinker: Guided Flow Upsampling and Texture Mapping for High-Resolution Video Frame Interpolation · AAAI 2026
Computer vision › Face, body and person analysis
face recognition
0.912025
FaceMe: Robust Blind Face Restoration with Personal Identification · AAAI 2025
Computer vision › Face, body and person analysis › face recognition
identity preservation
0.912025
FaceMe: Robust Blind Face Restoration with Personal Identification · AAAI 2025
Image and video processing › image restoration › face restoration
blind face restoration
0.912025
FaceMe: Robust Blind Face Restoration with Personal Identification · AAAI 2025
Image and video processing › image restoration
image dehazing
0.912025
Iterative Predictor-Critic Code Decoding for Real-World Image Dehazing · CVPR 2025
Natural language and speech › Language models and text generation
watermarking
0.812024
GumbelSoft: Diversified Language Model Watermarking via the GumbelMax-trick · ACL (1) 2024
Digital forensics and information hiding › synthetic media detection
machine-generated text detection
0.812024
GumbelSoft: Diversified Language Model Watermarking via the GumbelMax-trick · ACL (1) 2024

Methods — techniques the papers use, named apart from their topics

identity encoder · 1.7diffusion model · 1.7logits-addition · 1.5texture mapping · 1.0reconstruction · 1.0guided flow upsampling · 1.0codebook prior · 0.9code predictor-critic · 0.9VQGAN · 0.9gumbelmax-trick · 0.8gumbel-max trick · 0.8
YearPublicationVenuePosition
2026 VTinker: Guided Flow Upsampling and Texture Mapping for High-Resolution Video Frame Interpolation
abstract
Due to large pixel movement and high computational cost, estimating the motion of high-resolution frames is challenging. Thus, most flow-based Video Frame Interpolation (VFI) methods first predict bidirectional flows at low resolution and then use high-magnification upsampling (e.g., bilinear) to obtain the high-resolution ones. However, this kind of upsampling strategy may cause blur or mosaic at the flows' edges. Additionally, the motion of fine pixels at high resolution cannot be adequately captured in motion estimation at low resolution, which leads to the misalignment of task-oriented flows. With such inaccurate flows, input frames are warped and combined pixel-by-pixel, resulting in ghosting and discontinuities in the interpolated frame. In this study, we propose a novel VFI pipeline, VTinker, which consists of two core components: guided flow upsampling (GFU) and Texture Mapping. After motion estimation at low resolution, GFU introduces input frames as guidance to alleviate the blurring details in bilinear upsampling flows, which makes flows' edges clearer. Subsequently, to avoid pixel-level ghosting and discontinuities, Texture Mapping generates an initial interpolated frame, referred to as the intermediate proxy. The proxy serves as a cue for selecting clear texture blocks from the input frames, which are then mapped onto the proxy to facilitate producing the final interpolated frame via a reconstruction module. Extensive experiments demonstrate that VTinker achieves state-of-the-art performance in VFI.
Jiayi Fu, Chunle Guo, Shuhao Han, Chongyi Li
AAAI2
2026 Enhancing cross-modal retrieval through element-level semantic enrichment and momentum contrast
Jiayi Fu, Guangyun Lu
Vis. Comput.1
2025 FaceMe: Robust Blind Face Restoration with Personal Identification
abstract
Blind face restoration is a highly ill-posed problem due to the lack of necessary context. Although existing methods produce high-quality outputs, they often fail to faithfully preserve the individual's identity. In this paper, we propose a personalized face restoration method, FaceMe, based on a diffusion model. Given a single or a few reference images, we use an identity encoder to extract identity-related features, which serve as prompts to guide the diffusion model in restoring high-quality and identity-consistent facial images. By simply combining identity-related features, we effectively minimize the impact of identity-irrelevant features during training and support any number of reference image inputs during inference. Additionally, thanks to the robustness of the identity encoder, synthesized images can be used as reference images during training, and identity changing during inference does not require fine-tuning the model. We also propose a pipeline for constructing a reference image training pool that simulates the poses and expressions that may appear in real-world scenarios. Experimental results demonstrate that our FaceMe can restore high-quality facial images while maintaining identity consistency, achieving excellent performance and robustness.
Zheng-Peng Duan, Jia Ouyang, Jiayi Fu, Hyunhee Park, Zikun Liu 0001, Chunle Guo, Chongyi Li
AAAI4
2025 Iterative Predictor-Critic Code Decoding for Real-World Image Dehazing
abstract
We propose a novel Iterative Predictor-Critic Code Decoding framework for real-world image dehazing, abbreviated as IPC-Dehaze, which leverages the high-quality codebook prior encapsulated in a pre-trained VQGAN. Apart from previous codebook-based methods that rely on oneshot decoding, our method utilizes high-quality codes obtained in the previous iteration to guide the prediction of the Code-Predictor in the subsequent iteration, improving code prediction accuracy and ensuring stable dehazing performance. Our idea stems from the observations that 1) the degradation of hazy images varies with haze density and scene depth, and 2) clear regions play crucial cues in restoring dense haze regions. However, it is nontrivial to progressively refine the obtained codes in subsequent iterations, owing to the difficulty in determining which codes should be retained or replaced at each iteration. Another key insight of our study is to propose CodeCritic to capture interrelations among codes. The CodeCritic is used to evaluate code correlations and then resample a set of codes with the highest mask scores, i.e., a higher score indicates that the code is more likely to be rejected, which helps retain more accurate codes and predict difficult ones. Extensive experiments demonstrate the superiority of our method over state-of-the-art methods in real-world dehazing. Our project page can be found at https://github.com/Jiayi-Fu/IPC-Dehaze.
Jiayi Fu, Zikun Liu 0001, Chunle Guo, Hyunhee Park, Guoqing Wang 0001, Chongyi Li
CVPR1
2024 GumbelSoft: Diversified Language Model Watermarking via the GumbelMax-trick
abstract
Large language models (LLMs) excellently generate human-like text, but also raise concerns about misuse in fake news and academic dishonesty.Decoding-based watermark, particularly the GumbelMax-trick-based watermark (GM watermark), is a standout solution for safeguarding machine-generated texts due to its notable detectability.However, GM watermark encounters a major challenge with generation diversity, always yielding identical outputs for the same prompt, negatively impacting generation diversity and user experience.To overcome this limitation, we propose a new type of GM watermark, the Logits-Addition watermark, and its three variants, specifically designed to enhance diversity.Among these, the GumbelSoft watermark (a softmax variant of the Logits-Addition watermark) demonstrates superior performance in high diversity settings, with its AUROC score outperforming those of the two alternative variants by 0.1 to 0.3 and surpassing other decoding-based watermarking methods by a minimum of 0.1.1
Jiayi Fu, Xuandong Zhao, Ruihan Yang, Yuansen Zhang, Jiangjie Chen, Yanghua Xiao
ACL (1)1
2022 Improving Spoken Language Understanding with Cross-Modal Contrastive Learning
Jingjing Dong, Jiayi Fu, Hao Li 0078
INTERSPEECH2
2019 A Dropout-Based Single Model Committee Approach for Active Learning in ASR
abstract
In this paper, we proposed a new committee-based approach for active learning (AL) in automatic speech recognition (ASR). This approach can achieve lower recognition word error rate (WER) with fewer transcription by selecting the most informative samples. Different from previous committee-based AL approaches, the committee construction process of this approach needs to train only one acoustic model(AM) with dropout. Since only one model needs to be trained, this approach is simpler and faster. At the same time, the AM will be improved continuously, we also found this approach is more robust to its improvement. In experiments, we compared our approach with the random sampling and another state-of-the-art committee-based approach: heterogeneous neural networks (HNN) based approach. We examined our approach in WER, the time to construct committee and the robustness of model improvement in the Mandarin ASR task with 1600 hours speech data. The results showed that it achieves 2-3 times relative WER reduction compare with the random sampling, and it only uses 75% the time to achieve close WER with HNN-based approach.
Jiayi Fu, Kuang Ru
ASRU1
2016 A non-parametric approach for learning from crowds
abstract
Learning from crowds, which the labels of the instances are collected through crowdsourcing ways, has become an important research topic recently. Personal Classifier (PC) approach is a representative approach for learning from crowds due to its convex optimization formulation. PC approach makes assumptions about parameters' distribution, thus it is a parametric approach. However, these assumptions may not always hold, especially for real-world data sets. In this paper, we propose a new non-parametric approach, called NP approach, for learning from crowds. NP approach has a convex optimization formulation but without assumptions about parameters' distribution. In addition, NP approach can be generalized to the non-linear case directly, while the PC approach does not. Experimental studies show that NP approach outperforms other two compared approaches.
Jiayi Fu, Jinhong Zhong, Ke Tang 0001
IJCNN1