EDBT 2026 Demo / reviewers in the wild / expert
Wenyi Wang 0005
dblp:79/2722-5
· DBLP profile ↗
16ranked-venue papers
1as first author
7since 2021 · last 2024
0000-0003-1619-0294ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Perception-Oriented Video Frame Interpolation via Asymmetric BlendingabstractPrevious methods for Video Frame Interpolation (VFI) have encountered challenges, notably the manifestation of blur and ghosting effects. These issues can be traced back to two pivotal factors: unavoidable motion errors and misalignment in supervision. In practice, motion estimates often prove to be error-prone, resulting in misaligned features. Furthermore, the reconstruction loss tends to bring blurry results, particularly in misaligned regions. To mitigate these challenges, we propose a new paradigm called PerVFI (Perception-oriented Video Frame Interpolation). Our approach incorporates an Asymmetric Synergistic Blending module (ASB) that utilizes features from both sides to synergistically blend intermediate features. One reference frame emphasizes primary content, while the other contributes complementary information. To impose a stringent constraint on the blending process, we introduce a self-learned sparse quasi-binary mask which effectively mitigates ghosting and blur artifacts in the output. Additionally, we employ a normalizing flow-based generator and utilize the negative log-likelihood loss to learn the conditional distribution of the output, which further facilitates the generation of clear and fine details. Experimental results validate the superiority of PerVFI, demonstrating significant improvements in perceptual quality compared to existing methods. Codes are available at https://github.com/mulns/PerVFI Guangyang Wu, Xin Tao 0001, Wenyi Wang 0005, Xiaohong Liu 0001, Qingqing Zheng |
CVPR | 4 |
| 2024 | Face animation based on multiple sources and perspective alignmentabstractFace image animation generates a synthetic human face video that harmoniously integrates the identity derived from the source image and facial motion obtained from the driving video. This technology could be beneficial in multiple medical fields, such as diagnosis and privacy protection. Previous studies on face animation often relied on a single source image to generate an output video. With a significant pose difference between the source image and the driving frame, the quality of the generated video is likely to be suboptimal because the source image may not provide sufficient features for the warped feature map. In this study, we propose a novel face-animation scheme based on multiple sources and perspective alignment to address these issues. We first introduce a multiple-source sampling and selection module to screen the optimal source image set from the provided driving video. We then propose an inter-frame interpolation and alignment module to further eliminate the misalignment between the selected source image and the driving frame. The proposed method exhibits superior performance in terms of objective metrics and visual quality in large-angle animation scenes compared to other state-of-the-art face animation methods. It indicates the effectiveness of the proposed method in addressing the distortion issues in large-angle animation. Yuanzong Mei, Wenyi Wang 0005, Wei Yong, Weijie Wu, Yifan Zhu 0005 |
Virtual Real. Intell. Hardw. | 2 |
| 2023 | AccFlow: Backward Accumulation for Long-Range Optical FlowabstractRecent deep learning-based optical flow estimators have exhibited impressive performance in generating local flows between consecutive frames. However, the estimation of long-range flows between distant frames, particularly under complex object deformation and large motion occlusion, remains a challenging task. One promising solution is to accumulate local flows explicitly or implicitly to obtain the desired long-range flow. Nevertheless, the accumulation errors and flow misalignment can hinder the effectiveness of this approach. This paper proposes a novel recurrent framework called AccFlow, which recursively backward accumulates local flows using a deformable module called as AccPlus. In addition, an adaptive blending module is designed along with AccPlus to alleviate the occlusion effect by backward accumulation and rectify the accumulation error. Notably, we demonstrate the superiority of backward accumulation over conventional forward accumulation, which to the best of our knowledge has not been explicitly established before. To train and evaluate the proposed AccFlow, we have constructed a large-scale high-quality dataset named CVO, which provides ground-truth optical flow labels between adjacent and distant frames. Extensive experiments validate the effectiveness of AccFlow in handling long-range optical flow estimation. Codes are available at https://github.com/mulns/AccFlow. Guangyang Wu, Xiaohong Liu 0001, Kunming Luo, Qingqing Zheng, Shuaicheng Liu, Xinyang Jiang, Guangtao Zhai, Wenyi Wang 0005 |
ICCV | 9 |
| 2023 | FastLLVE: Real-Time Low-Light Video Enhancement with Intensity-Aware Look-Up TableabstractLow-Light Video Enhancement (LLVE) has received considerable attention in recent years. One of the critical requirements of LLVE is inter-frame brightness consistency, which is essential for maintaining the temporal coherence of the enhanced video. However, most existing single-image-based methods fail to address this issue, resulting in flickering effect that degrades the overall quality after enhancement. Moreover, 3D Convolution Neural Network (CNN)-based methods, which are designed for video to maintain inter-frame consistency, are computationally expensive, making them impractical for real-time applications. To address these issues, we propose an efficient pipeline named FastLLVE that leverages the Look-Up-Table (LUT) technique to maintain inter-frame brightness consistency effectively. Specifically, we design a learnable Intensity-Aware LUT (IA-LUT) module for adaptive enhancement, which addresses the low-dynamic problem in low-light scenarios. This enables FastLLVE to perform low-latency and low-complexity enhancement operations while maintaining high-quality results. Experimental results on benchmark datasets demonstrate that our method achieves the State-Of-The-Art (SOTA) performance in terms of both image quality and inter-frame brightness consistency. More importantly, our FastLLVE can process 1,080p videos at 50+ Frames Per Second (FPS), which is 2 X faster than SOTA CNN-based methods in inference time, making it a promising solution for real-time applications. The code is available at https://github.com/Wenhao-Li-777/FastLLVE. Wenhao Li 0018, Guangyang Wu, Wenyi Wang 0005, Peiran Ren, Xiaohong Liu 0001 |
ACM Multimedia | 3 |
| 2022 | Adversarial Examples Detection Based on Error Level Analysis and Space MappingabstractDeep neural network (DNN) shows impressive performance on many tasks but they usually suffer from adversarial examples with human eyes invisible slight perturbation. Such examples can not be distinguished by human but can mislead DNN classifiers leading to its important role in DNN attack and defense. Many adversarial examples detection methods perform well in identifying global perturbation adversarial examples but less efficiently for local perturbation ones. We observe both global perturbation and local perturbation adversarial examples have similar BOF histogram distribution after JPEG compression and Error Level Analysis (ELA) while these distributions are clearly different to clean example’s distribution. Meanwhile, researchers have found that the stability of adversarial example after space mapping is worse than that of the clean example. Therefore, we propose a two-branch architecture to detect adversarial examples based on the aforementioned strategies. Experiments show that our method has achieved better or similar performance compared to several state-of-the-art methods in terms of the detection accuracy and generation property for adversarial examples with global and local perturbation. Sizhao Huang, Guozhi Li, Wenyi Wang 0005 |
ICASSP | 5 |
| 2022 | Rangeinet: Fast Lidar Point Cloud Temporal InterpolationabstractDue to the low scan rate of LiDAR sensors, LiDAR point cloud streams usually have a low frame rate, which is far below that of other sensors such as cameras. This could incur frame rate mismatch while conducting multi-sensor data fusion. LiDAR point cloud temporal interpolation aims to synthesize the non-existing intermediate frame between input frames to improve the frame rate of point clouds. However, the existing methods heavily depend on 3D scene flow or 2D flow estimation, which yield huge computational complexity and obstacles in real-time applications. To resolve this issue, we propose a fast and non-flow involved method, which analyzes the LiDAR point cloud by exploiting its corresponding 2D range images (RIs). Specifically, we develop a Siamese context extractor containing asymmetrical convolution kernels to learn the shape context and spatial feature of RIs, and the 3D space-time convolutions are introduced to precisely capture the temporal characteristics. Experimental results have clearly shown that our method is much faster than the state-of-the-art LiDAR point cloud temporal interpolation methods on various datasets, while delivering either comparable or superior frame interpolation performance. Lili Zhao 0001, Xuhu Lin, Wenyi Wang 0005, Kai-Kuang Ma |
ICASSP | 3 |
| 2021 | Regularized Intermediate Layers Attack: Adversarial Examples With High TransferabilityabstractAlthough convolutional neural network (CNN) models have demonstrated state-of-the-art performance, especially in many image classification and recognition tasks, the classification accuracy would significantly decreased in the adversarial image samples set by adding slight perturbations in the input images. Currently, many adversarial examples were designed for specific CNNs but they were not universal valid across different CNNs. In this paper, we proposed a new intermediate layer optimization method to ensure that the adversarial examples are effective across different CNN models. Given one image, the proposed algorithm can derive multiple adversarial examples from just one white-box adversarial example by analyzing its regularized features in the intermediate layers of the attacked CNN. The adversarial examples derived from the intermediate layers showed better transferability compared with the original white-box adversarial example. According to the experiments on multiple CNN models, our algorithm promotes the averaging transfer attacking success rate (ASR) by 10.5% and 4.81%, compared to the baseline white-box attacking methods and the recent intermediate layer based attacking method ILA respectively. Weiyu Cui, Wenyi Wang 0005 |
ICIP | 4 |
| 2020 | Substitute Model Generation for Black-Box Adversarial Attack Based on Knowledge DistillationabstractAlthough deep convolutional neural network (CNN) performs well in many computer vision tasks, its classification mechanism is very vulnerable when it is exposed to the perturbation of adversarial attacks. In this paper, we proposed a new algorithm to generate the substitute model of black-box CNN models by using knowledge distillation. The proposed algorithm distills multiple CNN teacher models to a compact student model as the substitution of other black-box CNN models to be attacked. The black-box adversarial samples can be consequently generated on this substitute model by using various white-box attacking methods. According to our experiments on ResNet18 and DenseNet121, our algorithm boosts the attacking success rate (ASR) by 20% by training the substitute model based on knowledge distillation. Weiyu Cui, Wenyi Wang 0005 |
ICIP | 4 |
| 2019 | PRED: A Parallel Network for Handling Multiple Degradations via Single Model in Single Image Super-ResolutionabstractExisting SISR (single image super-resolution) methods mostly assume that a low-resolution (LR) image is bicubicly down-sampled from its high-resolution (HR) counterpart, which inevitably give rise to poor performance when the degradation is out of assumption. To address this issue, we propose a framework PRED (parallel residual and encoder-decoder network) with an innovative training strategy to enhance the robustness to multiple degradations. Consequently, the network can handle spatially variant degradations, which significantly improves the practicability of the proposed method. Extensive experimental results on real LR images show that the proposed method can not only produce favorable results on multiple degradations, but also reconstruct visually plausible HR images. Guangyang Wu, Lili Zhao 0001, Wenyi Wang 0005, Liaoyuan Zeng |
ICIP | 3 |
| 2019 | Efficient Screen Content Coding Based on Convolutional Neural Network Guided by a Large-Scale DatabaseabstractScreen content videos (SCVs) are becoming popular in many applications. Compared with natural content videos (NCVs), the SCVs have different characteristics. Therefore, the screen content coding (SCC) based on HEVC adopts some new coding tools (intra block copy and palette mode etc.) to improve coding efficiency, but these tools increase the computational complexity as well. In this paper, we propose to predict the CU partition of the SCVs by a convolutional neural network (CNN) which is trained by the large-scale database that we firstly established for screen content coding. The proposed approach is implemented in SCC reference software SCM-6.1. Experimental results show that our proposed approach can save 53.2% encoding time with 2.67% BD-rate increase on average in All Intra (AI) configurations. Lili Zhao 0001, Zhiwen Wei, Weitong Cai, Wenyi Wang 0005, Liaoyuan Zeng |
ICIP | 4 |
| 2019 | Adaptive illumination normalization via adaptive illumination preprocessing and modified weber-face
Rumin Zhang, Wenyi Wang 0005 |
Appl. Intell. | 4 |
| 2018 | High Efficient VR Video Coding Based on Auto Projection Selection Using Transferable FeaturesabstractGiven multiple texture projection methods from the sphere surface to the planar surface, this paper proposes an adaptive selection mode that automatically chooses the appropriate projection method to obtain high compression efficiency of the VR video. The video compression efficiency is inherently affected by the video content, which is closely related to the projection method in the case of VR video encoding. In order to represent the VR video content in a compact manner, a feature vector (transferable feature) for each frame is extracted by a Res-CNN which is pre-trained by a large scale data set for general classification. Afterwards, the relation between the feature and the optimal projection method is investigated by using PCA-KNN, which can project the initial feature vector to a subspace where the VR videos can be efficiently classified with low ambiguity. The experimental results show that the proposed method can select the appropriate projection method that generates the best BD rate. Lili Zhao 0001, Wenyi Wang 0005, Rumin Zhang, Liaoyuan Zeng |
VCIP | 3 |
| 2018 | Robust Multi-Frame Super-Resolution Based on Spatially Weighted Half-Quadratic Estimation and Adaptive BTV RegularizationabstractMulti-frame image super-resolution focuses on reconstructing a high-resolution image from a set of low-resolution images with high similarity. Combining image prior knowledge with fidelity model, the Bayesian-based methods have been considered as an effective technique in super-resolution. The minimization function derived from maximum a posteriori probability (MAP) is composed of a fidelity term and a regularization term. In this paper, based on the MAP estimation, we propose a novel initialization method for super-resolution imaging. For the fidelity term in our proposed method, the half-quadratic estimation is used to choose error norm adaptively instead of using fixed and norms. Besides, a spatial weight matrix is used as a confidence map to scale the estimation result. For the regularization term, we propose a novel regularization method based on adaptive bilateral total variation (ABTV). Both the fidelity term and the ABTV regularization guarantee the robustness of our framework. The fidelity term is mainly responsible for dealing with misregistration, blur, and other kinds of large errors, while the ABTV regularization aims at edge preservation and noise removal. The proposed scheme is tested on both synthetic data and real data. The experimental results illustrate the superiority of our proposed method in terms of edge preservation and noise removal over the state-of-the-art algorithms. Xiaohong Liu 0001, Lei Chen 0034, Wenyi Wang 0005, Jiying Zhao |
IEEE Trans. Image Process. | 3 |
| 2016 | Real-time video chroma keying: a parallel approach based on local texture and global colour distributionabstractThis study presents an automatic, human perception based chroma‐keying algorithm that extracts the objects of interest (i.e. foreground) from monochromatic background. Given an image to be chroma keyed, the global colour distribution and the local texture property are analysed in CIECAM02 colour appearance model. After the analysis, input image is automatically segmented into three parts: foreground, background, and uncertain regions. Afterwards, the background colour is propagated from known background to uncertain region by using interpolation functions; and the foreground colour is estimated based on global colour distribution and a linear cost criteria. The quantitative and perceptual comparisons on the matting results show that the proposed method can reliably remove the background region, correctly restore the intrinsic foreground colour, and accurately keep the fine details. In addition, the authors implement the proposed method on a heterogeneous parallel computing architecture which efficiently distributes the workload among different processors. The simulation results show that the foreground objects can be accurately extracted from high‐definition and/or ultra‐high‐definition videos in real time. Wenyi Wang 0005, Jiying Zhao |
IET Image Process. | 2 |
| 2016 | Video chroma keying via global sampling and trimap propagation
Chengcheng Hao, Wenyi Wang 0005, Jiying Zhao |
Multim. Syst. | 2 |
| 2016 | Hiding depth information in compressed 2D image/video using reversible watermarking
Wenyi Wang 0005, Jiying Zhao |
Multim. Tools Appl. | 1 |