Meiguang Jin

dblp:142/0036 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0003-3796-2310ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Emotion-Conditioned Motion Sub-spaces with Flow Matching for Real-Time Audio-Driven Talking Heads
abstract
Recent advances in audio-driven talking-head synthesis have brought lip-sync precision close to human perception, yet emotional fidelity and real-time inference remain open challenges. Existing pipelines typically disentangle lip articulation, facial expression, and head pose in latent space; this rigid factorization ignores the intrinsic coupling between articulation and affect — e.g., downward lip corners when sad—thus limiting expressiveness. We cast speech-conditioned facial motion as a sample from an emotion-conditioned distribution in a motion latent space. Concretely, we (i) learn a motion dictionary of orthogonal bases with an autoencoder via self-supervision, (ii) construct emotion-conditioned sub-spaces within the latent space, and (iii) design a layer-progressive cross-attention fusion module that modulates a flow-matching sampler with both audio and emotion signals. Only ten reverse ODE steps are required to generate a motion-latent trajectory, enabling real-time end-to-end latency. Extensive experiments on MEAD and RAVDESS show that our method outperforms recent GAN- and diffusion-based baselines in emotion accuracy while running at around 75 FPS on a single desktop GPU. The proposed framework delivers the first emotionally expressive Audio2Face system that simultaneously achieves lip-sync accuracy, affective realism, and real-time performance.
Haoyu Wang 0009, Xiaozhe Xin, Xiaoyu Qin 0001, Meiguang Jin, Junfeng Ma, Jia Jia 0001
AAAI4
2026 Attention Grounded Enhancement for Visual Document Retrieval
abstract
Visual document retrieval requires understanding heterogeneous and multi-modal content to satisfy implicit information needs. Recent advances use screenshot-based document encoding with fine-grained late interaction to encode holistic information and capture nuanced alignments, significantly improving retrieval performance. However, retrievers are still trained with coarse global relevance labels, without revealing which regions support the match. As a result, retrievers tend to rely on surface-level cues and struggle to capture implicit semantic connections, hindering their ability to handle non-extractive queries. To improve fine-grained relevance modeling, we propose a Attention-Grounded REtriever Enhancement (AGREE) framework. AGREE leverages cross-modal attention from multimodal large language models (MLLMs) as proxy supervision to guide the retriever in identifying relevant document regions. Specifically, AGREE extracts attention maps from the MLLM that highlight which document regions are attended to based on the query. These attention scores serve as local, region-level relevance signals. During training, AGREE combines local signals with the global document-level relevance label to jointly optimize the retriever. This dual-level supervision enables the model to learn not only whether documents match, but also which content drives relevance. Experiments on the challenging visual document retrieval benchmark, ViDoRe V2, show that AGREE significantly outperforms the global-supervision-only baseline by 12.82% and 5.03% in terms of average nDCG@1 and nDCG@5. Quantitative and qualitative analyses further demonstrate that AGREE promotes deeper alignment between query terms and document regions, moving beyond surface-level matching toward more accurate and interpretable retrieval. Our code is available at: https://github.com/VickiCui/AGREE.
Wanqing Cui, Yazhi Guo, Yibo Hu 0001, Meiguang Jin, Junfeng Ma, Keping Bi
SIGIR5
2026 Self-distilled learning of adaptive interval 3D lookup tables on real-time image enhancement
Ruikai Zhou, Canqian Yang, Meiguang Jin, Xu Jia 0012, Ying Chen 0011, Yi Xu 0001
Pattern Recognit.4
2024 Toward Tiny and High-Quality Facial Makeup with Data Amplify Learning
Qiaoqiao Jin, Xuanhong Chen, Meiguang Jin, Yucheng Zheng, Yupeng Zhu, Bingbing Ni
ECCV (27)3
2023 Lightweight Network Towards Real-Time Image Denoising On Mobile Devices
abstract
Deep convolutional neural networks have achieved great progress in image denoising tasks. However, their complicated architectures and heavy computational cost hinder their deployments on mobile devices. Some recent efforts in designing lightweight denoising networks focus on reducing either FLOPs (floating-point operations) or the number of parameters. However, these metrics are not directly correlated with the on-device latency. In this paper, we identify the real bottlenecks that affect the CNN-based models’ runtime performance on mobile devices: memory access cost and NPU-incompatible operations, and build the model based on these. To further improve the denoising performance, the mobile-friendly attention module MFA and the model reparameterization module RepConv are proposed, which enjoy both low latency and excellent denoising performance. To this end, we propose a mobile-friendly denoising network, namely MFDNet. The experiments show that MFDNet achieves state-of-the-art performance on real-world denoising benchmarks SIDD and DND under real-time latency on mobile devices. The code and pre-trained models will be released.
Zhuoqun Liu, Meiguang Jin, Ying Chen 0011, Huaida Liu, Canqian Yang, Hongkai Xiong
ICIP2
2022 AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on Real-time Image Enhancement
abstract
The 3D Lookup Table (3D LUT) is a highly-efficient tool for real-time image enhancement tasks, which models a non-linear 3D color transform by sparsely sampling it into a discretized 3D lattice. Previous works have made efforts to learn image-adaptive output color values of LUTs for flexible enhancement but neglect the importance of sampling strategy. They adopt a sub-optimal uniform sampling point allocation, limiting the expressiveness of the learned LUTs since the (tri-)linear interpolation between uniform sampling points in the LUT transform might fail to model local non-linearities of the color transform. Focusing on this problem, we present AdaInt (Adaptive Intervals Learning), a novel mechanism to achieve a more flexible sampling point allocation by adaptively learning the non-uniform sampling intervals in the 3D color space. In this way, a 3D LUT can increase its capability by conducting dense sampling in color ranges requiring highly non-linear transforms and sparse sampling for near-linear transforms. The proposed AdaInt could be implemented as a compact and efficient plug-and-play module for a 3D LUT-based method. To enable the end-to-end learning of AdaInt, we design a novel differentiable operator called AiLUT-Transform (Adaptive Interval LUT Transform) to locate input colors in the non-uniform 3D LUT and provide gradients to the sampling intervals. Experiments demonstrate that methods equipped with AdaInt can achieve state-of-the-art performance on two public benchmark datasets with a negligible overhead increase. Our source code is available at https://github.com/ImCharlesY/AdaInt.
Canqian Yang, Meiguang Jin, Xu Jia 0012, Yi Xu 0001, Ying Chen 0011
CVPR2
2022 SepLUT: Separable Image-Adaptive Lookup Tables for Real-Time Image Enhancement
Canqian Yang, Meiguang Jin, Yi Xu 0001, Rui Zhang 0052, Ying Chen 0011, Huaida Liu
ECCV (18)2
2019 Learning to Extract Flawless Slow Motion From Blurry Videos
abstract
In this paper, we introduce the task of generating a sharp slow-motion video given a low frame rate blurry video. We propose a data-driven approach, where the training data is captured with a high frame rate camera and blurry images are simulated through an averaging process. While it is possible to train a neural network to recover the sharp frames from their average, there is no guarantee of the temporal smoothness for the formed video, as the frames are estimated independently. To address the temporal smoothness requirement we propose a system with two networks: One, DeblurNet, to predict sharp keyframes and the second, InterpNet, to predict intermediate frames between the generated keyframes. A smooth transition is ensured by interpolating between consecutive keyframes using InterpNet. Moreover, the proposed scheme enables further increase in frame rate without retraining the network, by applying InterpNet recursively between pairs of sharp frames. We evaluate the proposed method on several datasets, including a novel dataset captured with a Sony RX V camera. We also demonstrate its performance of increasing the frame rate up to 20 times on real blurry videos.
Meiguang Jin, Paolo Favaro
CVPR1
2018 Learning to Extract a Video Sequence From a Single Motion-Blurred Image
abstract
We present a method to extract a video sequence from a single motion-blurred image. Motion-blurred images are the result of an averaging process, where instant frames are accumulated over time during the exposure of the sensor. Unfortunately, reversing this process is nontrivial. Firstly, averaging destroys the temporal ordering of the frames. Secondly, the recovery of a single frame is a blind deconvolution task, which is highly ill-posed. We present a deep learning scheme that gradually reconstructs a temporal ordering by sequentially extracting pairs of frames. Our main contribution is to introduce loss functions invariant to the temporal order. This lets a neural network choose during training what frame to output among the possible combinations. We also address the ill-posedness of deblurring by designing a network with a large receptive field and implemented via resampling to achieve a higher computational efficiency. Our proposed method can successfully retrieve sharp image sequences from a single motion blurred image and can generalize well on synthetic and real datasets captured with different cameras.
Meiguang Jin, Givi Meishvili, Paolo Favaro
CVPR1
2018 Normalized Blind Deconvolution
Meiguang Jin, Stefan Roth 0001, Paolo Favaro
ECCV (7)1
2018 Learning to see through reflections
abstract
Pictures of objects behind a glass are difficult to interpret and understand due to the superposition of two real images: a reflection layer and a background layer. Separation of these two layers is challenging due to the ambiguities in assigning texture patterns and the average color in the input image to one of the two layers. In this paper, we propose a novel method to reconstruct these layers given a single input image by explicitly handling the ambiguities of the reconstruction. Our approach combines the ability of neural networks to build image priors on large image regions with an image model that accounts for the brightness ambiguity and saturation. We find that our solution generalizes to real images even in the presence of strong reflections. Extensive quantitative and qualitative experimental evaluations on both real and synthetic data show the benefits of our approach over prior work. Moreover, our proposed neural network is computationally and memory efficient.
Meiguang Jin, Sabine Süsstrunk, Paolo Favaro
ICCP1
2018 Plenoptic Image Motion Deblurring
abstract
We propose a method to remove motion blur in a single light field captured with a moving plenoptic camera. Since motion is unknown, we resort to a blind deconvolution formulation, where one aims to identify both the blur point spread function and the latent sharp image. Even in the absence of motion, light field images captured by a plenoptic camera are affected by a non-trivial combination of both aliasing and defocus, which depends on the 3D geometry of the scene. Therefore, motion deblurring algorithms designed for standard cameras are not directly applicable. Moreover, many state of the art blind deconvolution algorithms are based on iterative schemes, where blurry images are synthesized through the imaging model. However, current imaging models for plenoptic images are impractical due to their high dimensionality. We observe that plenoptic cameras introduce periodic patterns that can be exploited to obtain highly parallelizable numerical schemes to synthesize images. These schemes allow extremely efficient GPU implementations that enable the use of iterative methods. We can then cast blind deconvolution of a blurry light field image as a regularized energy minimization to recover a sharp high-resolution scene texture and the camera motion. Furthermore, the proposed formulation can handle non-uniform motion blur due to camera shake as demonstrated on both synthetic and real light field data.
Paramanand Chandramouli, Meiguang Jin, Daniele Perrone, Paolo Favaro
IEEE Trans. Image Process.2
2017 Noise-Blind Image Deblurring
abstract
We present a novel approach to noise-blind deblurring, the problem of deblurring an image with known blur, but unknown noise level. We introduce an efficient and robust solution based on a Bayesian framework using a smooth generalization of the 0-1 loss. A novel bound allows the calculation of very high-dimensional integrals in closed form. It avoids the degeneracy of Maximum a-Posteriori (MAP) estimates and leads to an effective noise-adaptive scheme. Moreover, we drastically accelerate our algorithm by using Majorization Minimization (MM) without introducing any approximation or boundary artifacts. We further speed up convergence by turning our algorithm into a neural network termed GradNet, which is highly parallelizable and can be efficiently trained. We demonstrate that our noise-blind formulation can be integrated with different priors and significantly improves existing deblurring algorithms in the noise-blind and in the known-noise case. Furthermore, GradNet leads to state-of-the-art performance across different noise levels, while retaining high computational efficiency.
Meiguang Jin, Stefan Roth 0001, Paolo Favaro
CVPR1
2017 Deep Mean-Shift Priors for Image Restoration
abstract
In this paper we introduce a natural image prior that directly represents a Gaussian-smoothed version of the natural image distribution. We include our prior in a formulation of image restoration as a Bayes estimator that also allows us to solve noise-blind image restoration problems. We show that the gradient of our prior corresponds to the mean-shift vector on the natural image distribution. In addition, we learn the mean-shift vector field using denoising autoencoders, and use it in a gradient descent approach to perform Bayes risk minimization. We demonstrate competitive results for noise-blind deblurring, super-resolution, and demosaicing.
Siavash Arjomand Bigdeli, Matthias Zwicker, Paolo Favaro, Meiguang Jin
NIPS4
2014 Adaptive Propagation-Based Color-Sampling for Alpha Matting
abstract
Image matting refers to the problem of foreground extraction from an image and transparency determination of the pixels. Although other matting algorithms have been proposed, most are not sufficiently robust to obtain satisfactory matting results in different regions of an image, such as smooth regions, nonuniform color distribution regions and isolated color regions. This paper proposes a novel matting algorithm that can extract high-quality mattes from different regions of an image. Our proposed algorithm combines propagation and color-sampling methods. Unlike previous propagation-based approaches that use either local or nonlocal propagation methods, our propagation framework adaptively uses both local and nonlocal processes according to the detection results of the different regions in the image. Our color-sampling strategy, which is based on the characteristics of the superpixel, uses a simple sample selection criterion and requires significantly less computational cost than previous color-sampling methods. Experimental results show that our adaptive propagation framework, alone, outperforms the state-of-the-art propagation-based approaches. Combined with our color-sampling method, it can effectively handle different regions in the image and produce both visually and quantitatively high-quality matting results.
Meiguang Jin, Byoung-Kwang Kim, Woo-Jin Song
IEEE Trans. Circuits Syst. Video Technol.1
2013 KNN-based color line model for image matting
abstract
Image matting is the extraction of a foreground object from an image and determination of the transparency of each pixel. Matting is inherently an ill-posed and underconstrained problem. Therefore, some assumptions need to be made to solve it. Recent methods that provide a closed-form solution to this problem are based on the assumption of either local smoothness or the nonlocal principle, but they cannot always produce satisfactory matting results. In this paper, we propose a K-nearest neighbors (KNN)-based color line model that combines and preserves the advantages of both the above assumptions. The experimental matting results indicate that they are of comparable or higher quality than those obtained by the existing methods based on the above two assumptions.
Meiguang Jin, Byoung-Kwang Kim, Woo-Jin Song
ICIP1