Qianrui Wang

dblp:316/4169 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0001-9058-0253ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2024 STDC-Net: A spatial-temporal deformable convolution network for conference video frame interpolation
abstract
Abstract Video conference communication can be seriously affected by dropped frames or reduced frame rates due to network or hardware restrictions. Video frame interpolation techniques can interpolate the dropped frames and generate smoother videos. However, existing methods can not generate plausible results in video conferences due to the large motions of the eyes, mouth and head. To address this issue, we propose a Spatial-Temporal Deformable Convolution Network (STDC-Net) for conference video frame interpolation. The STDC-Net first extracts shallow spatial-temporal features by an embedding layer. Secondly, it extracts multi-scale deep spatial-temporal features through Spatial-Temporal Representation Learning (STRL) module, which contains several Spatial-Temporal Feature Extracting (STFE) blocks and downsample layers. To extract the temporal features, each STFE block splits feature maps along the temporal pathway and processes them with Multi-Layer Perceptron (MLP). Similarly, the STFE block splits the temporal features along horizontal and vertical pathways and processes them by another two MLPs to get spatial features. By splitting the feature maps into segments of varying lengths in different scales, the STDC-Net can extract both local details and global spatial features, allowing it to effectively handle large motions. Finally, Frame Synthesis (FS) module predicts weights, offsets and masks using the spatial-temporal features, which are used in deformable convolution to generate the intermediate frames. Experimental results demonstrate the STDC-Net outperforms state-of-the-art methods in both quantitative and qualitative evaluations. Compared to the baseline, the proposed method achieved a PSNR improvement of 0.13 dB and 0.17 dB on the Voxceleb2 and HDTF datasets, respectively.
Qianrui Wang, Dengshi Li, Yu Gao 0018
Multim. Tools Appl.2
2024 SVMFI: speaker video multi-frame interpolation with the guidance of audio
Qianrui Wang, Dengshi Li, Yu Gao 0018, Aolei Chen
Multim. Tools Appl.1
2023 ASVFI: Audio-Driven Speaker Video Frame Interpolation
abstract
Due to limited data transmission, the video frame rate is low during the online conference, severely affecting user experience. Video frame interpolation can solve the problem by interpolating intermediate frames to increase the video frame rate. Generally, most existing video frame interpolation methods are based on the linear motion assumption. However, the mouth motion is nonlinear, and these methods can not generate superior intermediate frames in speaker video. Considering the strong correlation between mouth shape and vocalization, a new method is proposed, named Audio-driven Speaker Video Frame Interpolation(ASVFI). First, we extract the audio feature from Audio Net(ANet). Second, we use Video Net(VNet) encoder to extract the video feature. Finally, we fuse the audio and video features by AVFusion and decode out the intermediate frame in the VNet decoder. The experimental results show that the PSNR is nearly 0.13dB higher than the baseline of interpolating one frame. When interpolating seven frames, the PSNR is 0.33dB higher than the baseline.
Qianrui Wang, Dengshi Li, Jing Xiao 0004
ICIP1
2023 Self-guided Transformer for Video Super-Resolution
Tong Xue, Qianrui Wang, Xinyi Huang 0005, Dengshi Li
PRCV (10)2
2022 Adaptive Speech Intelligibility Enhancement for Far-and-Near-end Noise Environments Based on Self-attention StarGAN
Dengshi Li, Lanxin Zhao, Jing Xiao 0004, Duanzheng Guan, Qianrui Wang
MMM (2)6
2022 Speech Intelligibility Enhancement By Non-Parallel Speech Style Conversion Using CWT and iMetricGAN Based CycleGAN
Jing Xiao 0004, Dengshi Li, Lanxin Zhao, Qianrui Wang
MMM (1)5