Longhai Wu

dblp:294/4946 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Unified Arbitrary-Time Video Frame Interpolation and Prediction
abstract
Video frame interpolation and prediction aim to synthesize frames in-between and subsequent to existing frames, respectively. Despite being closely-related, these two tasks are traditionally studied with different model architectures, or same architecture but individually trained weights. Furthermore, while arbitrary-time interpolation has been extensively studied, the value of arbitrary-time prediction has been largely overlooked. In this work, we present uniVIP - unified arbitrary-time Video Interpolation and Prediction. Technically, we firstly extend an interpolation-only network for arbitrary-time interpolation and prediction, with a special input channel for task (interpolation or prediction) encoding. Then, we show how to train a unified model on common triplet frames. Our uniVIP provides competitive results for video interpolation, and outperforms existing state-of-the-arts for video prediction. Codes will be available at: https://github.com/srcn-ivl/uniVIP
Xin Jin 0023, Longhai Wu, Ilhyun Cho, Cheul-Hee Hahm
ICASSP2
2025 Exploring Simple Siamese Network for High-Resolution Video Quality Assessment
abstract
In the research of video quality assessment (VQA), two-branch network [1] has emerged as a promising solution. It decouples VQA with separate technical and aesthetic branches to measure the perception of low-level distortions and high-level semantics respectively. However, we argue that while technical and aesthetic perspectives are complementary, the technical perspective itself should be measured in semantic-aware manner. We hypothesize that existing technical branch struggles to perceive the semantics of high-resolution videos, as it is trained on local mini-patches sampled from videos. This issue can be hidden by apparently good results on low-resolution videos, but indeed becomes critical for high-resolution VQA. This work introduces SiamVQA, a simple but effective Siamese network for high-resolution VQA. SiamVQA shares weights between technical and aesthetic branches, enhancing the semantic perception ability of technical branch to facilitate technical-quality representation learning. Furthermore, it integrates a dual cross-attention layer for fusing technical and aesthetic features. SiamVQA achieves state-of-the-art accuracy on high-resolution benchmarks, and competitive results on lower-resolution benchmarks. Codes will be available at: https://github.com/srcn-ivl/SiamVQA
Guotao Shen, Ziheng Yan, Xin Jin 0023, Longhai Wu, Ilhyun Cho, Cheul-Hee Hahm
ICASSP4
2025 UPR-Net: A Unified Pyramid Recurrent Network for Video Frame Interpolation
Xin Jin 0023, Longhai Wu, Youxin Chen, Jayoon Koo, Cheul-Hee Hahm
Int. J. Comput. Vis.2
2024 Alignment-aware Patch-level Routing for Dynamic Video Frame Interpolation
Ban Chen, Xin Jin 0023, Longhai Wu, Ilhyun Cho, Cheul-Hee Hahm
BMVC3
2024 Dynamic Video Frame Interpolation with Integrated Difficulty Pre-Assessment
abstract
Video frame interpolation (VFI) has witnessed great progress in recent years. However, existing VFI models still struggle to achieve a good trade-off between accuracy and efficiency. Accurate VFI models typically rely on heavy compute to process all samples, ignoring the fact that easy samples with small motion or clear texture can be well addressed by a fast VFI model and do not require such heavy compute. In this paper, we present a dynamic VFI pipeline with integrated pre-assessment of interpolation difficulty. Specifically, it leverages a difficulty pre-assessment model to measure the difficulty level of interpolating input frames, and then dynamically selects an accurate or a fast VFI model for frame interpolation. Furthermore, we contribute a large-scale annotated dataset to train our VFI difficulty pre-assessment model. Extensive experiments show that our dynamic VFI pipeline can achieve an excellent trade-off between accuracy and efficiency, by feeding hard samples to accurate model, and passing easy samples through fast model.
Ban Chen, Xin Jin 0023, Youxin Chen, Longhai Wu, Jayoon Koo, Cheul-Hee Hahm
ICASSP4
2023 A Unified Pyramid Recurrent Network for Video Frame Interpolation
abstract
Flow-guided synthesis provides a common framework for frame interpolation, where optical flow is estimated to guide the synthesis of intermediate frames between consecutive inputs. In this paper, we present UPR-Net, a novel Unified Pyramid Recurrent Network for frame interpolation. Cast in a flexible pyramid framework, UPR-Net exploits lightweight recurrent modules for both bi-directional flow estimation and intermediate frame synthesis. At each pyramid level, it leverages estimated bi-directional flow to generate forward-warped representations for frame synthesis; across pyramid levels, it enables iterative refinement for both optical flow and intermediate frame. In particular, we show that our iterative synthesis strategy can significantly improve the robustness of frame interpolation on large motion cases. Despite being extremely lightweight (1.7M parameters), our base version of UPR-Net achieves excellent performance on a large range of benchmarks. Code and trained models of our UPR-Net series are available at: https://github.com/srcn-iv1/UPR-Net.
Xin Jin 0023, Longhai Wu, Youxin Chen, Jayoon Koo, Cheul-Hee Hahm
CVPR2
2023 Enhanced Bi-directional Motion Estimation for Video Frame Interpolation
abstract
We propose a simple yet effective algorithm for motion-based video frame interpolation. Existing motion-based interpolation methods typically rely on an off-the-shelf optical flow model or a U-Net based pyramid network for motion estimation, which either suffer from large model size or limited capacity in handling various challenging motion cases. In this work, we present a novel compact model to simultaneously estimate the bi-directional motions between input frames. It is designed by carefully adapting the ingredients (e.g., warping, correlation) in optical flow research for simultaneous bi-directional motion estimation within a flexible pyramid recurrent framework. Our motion estimator is extremely lightweight (15x smaller than PWC-Net), yet enables reliable handling of large and complex motion cases. Based on estimated bi-directional motions, we employ a synthesis network to fuse forward-warped representations and predict the intermediate frame. Our method achieves excellent performance on a broad range of frame interpolation benchmarks. Code and trained models are available at https://github.com/srcn-ivl/EBME.
Xin Jin 0023, Longhai Wu, Guotao Shen, Youxin Chen, Jayoon Koo, Cheul-Hee Hahm
WACV2
2021 Dual attention autoencoder for all-weather outdoor lighting estimation
Piaopiao Yu, Jie Guo 0001, Longhai Wu, Yanwen Guo 0001
Sci. China Inf. Sci.3
2021 SADRNet: Self-Aligned Dual Face Regression Networks for Robust 3D Dense Face Alignment and Reconstruction
abstract
Three-dimensional face dense alignment and reconstruction in the wild is a challenging problem as partial facial information is commonly missing in occluded and large pose face images. Large head pose variations also increase the solution space and make the modeling more difficult. Our key idea is to model occlusion and pose to decompose this challenging task into several relatively more manageable subtasks. To this end, we propose an end-to-end framework, termed as Self-aligned Dual face Regression Network (SADRNet), which predicts a pose-dependent face, a pose-independent face. They are combined by an occlusion-aware self-alignment to generate the final 3D face. Extensive experiments on two popular benchmarks, AFLW2000-3D and Florence, demonstrate that the proposed method achieves significant superior performance over existing state-of-the-art methods.
Zeyu Ruan, Changqing Zou, Longhai Wu, Gangshan Wu, Limin Wang 0002
IEEE Trans. Image Process.3