EDBT 2026 Demo / reviewers in the wild / expert
Qiuxia Lai
dblp:210/4586
· DBLP profile ↗
23ranked-venue papers
9as first author
18since 2021 · last 2025
0000-0001-6872-5540ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cycle-Consistent Learning for Joint Layout-to-Image Generation and Object Detection
Xinhao Cai, Qiuxia Lai, Gensheng Pei, Xiangbo Shu, Yazhou Yao, Wenguan Wang |
ICCV | 2 |
| 2025 | A Conditional Probability Framework for Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to recognize unseen combinations of known objects and attributes by leveraging knowledge from previously seen compositions. Traditional approaches primarily focus on disentangling attributes and objects, treating them as independent entities during learning. However, this assumption overlooks the semantic constraints and contextual dependencies inside a composition. For example, certain attributes naturally pair with specific objects (e.g., "striped" applies to "zebra" or "shirts" but not "sky" or "water"), while the same attribute can manifest differently depending on context (e.g., "young" in "young tree" vs. "young dog"). Thus, capturing attribute-object interdependence remains a fundamental yet long-ignored challenge in CZSL. In this paper, we adopt a Conditional Probability Framework (CPF) to explicitly model attribute-object dependencies. We decompose the probability of a composition into two components: the likelihood of an object and the conditional likelihood of its attribute. To enhance object feature learning, we incorporate textual descriptors to highlight semantically relevant image regions. These enhanced object features then guide attribute learning through a cross-attention mechanism, ensuring better contextual alignment. By jointly optimizing object likelihood and conditional attribute likelihood, our method effectively captures compositional dependencies and generalizes well to unseen compositions. Extensive experiments on multiple CZSL benchmarks demonstrate the superiority of our approach. Code is available at here. Peng Wu 0014, Qiuxia Lai, Hao Fang 0010, Guosen Xie, Yilong Yin, Xiankai Lu, Wenguan Wang |
ICCV | 2 |
| 2025 | FDTest: Prioritizing Test Inputs for Object Detection Models via Foundation Model ExploitationabstractTesting Deep Neural Networks (DNNs) often incurs high labeling costs due to the need for extensive labeled data. Hence, it is critical to strategically prioritize test inputs that can uncover more model errors for efficient labeling. Existing test input prioritization methods focus on image classification. However, these methods may not be directly applicable to object detection (OD) models, as testing of OD models confronts unique challenges involving complex model errors related to object localization and counting. To address the above challenges, in this paper, we introduce FDTest, a black-box test input prioritization framework for OD models that incorporates foundation models trained on diverse datasets to provide supplementary information. Guided by the foundation model, we categorize the target model’s predictions as reliable and suspicious objects and conduct confidence calibration on them to improve the estimation of wrongly detected objects, i.e., false positives (FPs). Additionally, we use the foundation model to supply missed objects and cross-verify them to improve the estimation of missed objects, i.e., false negatives (FNs). Then, we prioritize test images according to the total estimated number of FPs and FNs. Experiments on two standard OD benchmarks demonstrate the superiority of FDTest in test input prioritization for OD models. Qiuxia Lai, Yu Li 0007 |
IJCNN | 2 |
| 2024 | Poly Kernel Inception Network for Remote Sensing DetectionabstractObject detection in remote sensing images (RSIs) often suffers from several increasing challenges, including the large variation in object scales and the diverse-ranging context. Prior methods tried to address these challenges by expanding the spatial receptive field of the backbone, either through large-kernel convolution or dilated convolution. However, the former typically introduces considerable background noise, while the latter risks generating overly sparse feature representations. In this paper, we introduce the Poly Kernel Inception Network (PKINet) to handle the above challenges. PKINet employs multi-scale convolution kernels without dilation to extract object features of varying scales and capture local context. In addition, a Context Anchor Attention (CAA) module is introduced in parallel to capture long-range contextual information. These two components work jointly to advance the performance of PKINet on four challenging remote sensing detection benchmarks, namely DOTA-v1.0, DOTA-v1.5, HRSC2016, and DIOR-R. Xinhao Cai, Qiuxia Lai, Wenguan Wang, Zeren Sun, Yazhou Yao |
CVPR | 2 |
| 2024 | Information Bottleneck-Inspired Spatial Attention for Robust Semantic Segmentation in Autonomous DrivingabstractSemantic segmentation is critical for autonomous driving systems that rely on accurate scene understanding to ensure safety and reliability. However, domain shifts between training and testing datasets, as well as challenges in nighttime conditions, significantly hinder the performance of semantic segmentation models when deployed in real-world driving scenarios. This paper presents an approach for addressing these challenges in semantic segmentation by leveraging an Information Bottleneck (IB)-inspired spatial attention mechanism. By integrating IB principles into the spatial attention framework, the model could capture essential features while filtering out irrelevant information, thereby improving its generalization capability across diverse domains and enhancing performance in nighttime segmentation. Extensive experiments demonstrate that IB-inspired attention consistently enhances domain generation and nighttime performance, showing robust semantic segmentation for autonomous driving. Qiuxia Lai, Qipeng Tang |
ICARCV | 1 |
| 2024 | Vector Quantization Prompting for Continual LearningabstractContinual learning requires to overcome catastrophic forgetting when training a single model on a sequence of tasks. Recent top-performing approaches are prompt-based methods that utilize a set of learnable parameters (i.e., prompts) to encode task knowledge, from which appropriate ones are selected to guide the fixed pre-trained model in generating features tailored to a certain task. However, existing methods rely on predicting prompt identities for prompt selection, where the identity prediction process cannot be optimized with task loss. This limitation leads to sub-optimal prompt selection and inadequate adaptation of pre-trained features for a specific task. Previous efforts have tried to address this by directly generating prompts from input queries instead of selecting from a set of candidates. However, these prompts are continuous, which lack sufficient abstraction for task knowledge representation, making them less effective for continual learning. To address these challenges, we propose VQ-Prompt, a prompt-based continual learning method that incorporates Vector Quantization (VQ) into end-to-end training of a set of discrete prompts. In this way, VQ-Prompt can optimize the prompt selection process with task loss and meanwhile achieve effective abstraction of task knowledge for continual learning. Extensive experiments show that VQ-Prompt outperforms state-of-the-art continual learning methods across a variety of benchmarks under the challenging class-incremental setting. Qiuxia Lai, Yu Li 0007, Qiang Xu 0001 |
NeurIPS | 2 |
| 2024 | Spatial attention for human-centric visual understanding: An Information Bottleneck method
Qiuxia Lai, Yongwei Nie, Yu Li 0007, Hanqiu Sun, Qiang Xu 0001 |
Comput. Vis. Image Underst. | 1 |
| 2024 | Self-Supervised Video Representation Learning via Capturing Semantic Changes Indicated by SaccadesabstractIn this paper, we propose a self-supervised video representation learning (video SSL) method by taking inspiration from cognitive science and neuroscience on human visual perception. Different from previous methods that focus on the inherent properties of videos, we argue that humans learn to perceive the world through the self-awareness of the semantic changes or consistency in the input stimuli in the absence of labels, accompanied by representation reorganization during the post-learning rest periods. To this end, we first exploit the presence of saccades as an indicator of semantic changes in a contrastive learning framework, mimicking self-awareness in human representation learning. The saccades are generated by alternating the fixations following the predicted scanpath. Second, we model the semantic consistency in eye fixation by minimizing the prediction error between the predicted and the true state of another time point. Finally, we incorporate prototypical contrastive learning to reorganize the learned representations to enhance the associations among perceptually similar ones. Compared to previous video SSL solutions, our method can capture finer-grained semantics from video instances and further associate similar ones together. Experiments show that the proposed bio-inspired video SSL method significantly improves the Top-1 video retrieval accuracy on UCF101 and achieves superior performance on downstream tasks such as action recognition under comparable settings. Qiuxia Lai, Ailing Zeng, Ye Wang 0011, Lihong Cao, Yu Li 0007, Qiang Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | BIFRNet: A Brain-Inspired Feature Restoration DNN for Partially Occluded Image RecognitionabstractThe partially occluded image recognition (POIR) problem has been a challenge for artificial intelligence for a long time. A common strategy to handle the POIR problem is using the non-occluded features for classification. Unfortunately, this strategy will lose effectiveness when the image is severely occluded, since the visible parts can only provide limited information. Several studies in neuroscience reveal that feature restoration which fills in the occluded information and is called amodal completion is essential for human brains to recognize partially occluded images. However, feature restoration is commonly ignored by CNNs, which may be the reason why CNNs are ineffective for the POIR problem. Inspired by this, we propose a novel brain-inspired feature restoration network (BIFRNet) to solve the POIR problem. It mimics a ventral visual pathway to extract image features and a dorsal visual pathway to distinguish occluded and visible image regions. In addition, it also uses a knowledge module to store classification prior knowledge and uses a completion module to restore occluded features based on visible features and prior knowledge. Thorough experiments on synthetic and real-world occluded image datasets show that BIFRNet outperforms the existing methods in solving the POIR problem. Especially for severely occluded images, BIRFRNet surpasses other methods by a large margin and is close to the human brain performance. Furthermore, the brain-inspired design makes BIFRNet more interpretable. Jiahong Zhang, Lihong Cao, Qiuxia Lai, Yunxiao Qin |
AAAI | 3 |
| 2023 | On EDA-Driven Learning for SAT SolvingabstractWe present DeepSAT, a novel end-to-end learning framework for the Boolean satisfiability (SAT) problem. Unlike existing solutions trained on random SAT instances with relatively weak supervision, we propose applying the knowledge of the well-developed electronic design automation (EDA) field for SAT solving. Specifically, we first resort to logic synthesis algorithms to pre-process SAT instances into optimized and-inverter graphs (AIGs). By doing so, the distribution diversity among various SAT instances can be dramatically reduced, which facilitates improving the generalization capability of the learned model. Next, we regard the distribution of SAT solutions being a product of conditional Bernoulli distributions. Based on this observation, we approximate the SAT solving procedure with a conditional generative model, leveraging a novel directed acyclic graph neural network (DAGNN) with two polarity prototypes for conditional SAT modeling. To effectively train the generative model, with the help of logic simulation tools, we obtain the probabilities of nodes in the AIG being logic ‘1’ as rich supervision. We conduct comprehensive experiments on various SAT problems. Our results show that, DeepSAT achieves significant accuracy improvements over state-of-the-art learning-based SAT solutions, especially when generalized to SAT instances that are relatively large or with diverse distributions. Min Li 0019, Zhengyuan Shi, Qiuxia Lai, Sadaf Khan, Shaowei Cai 0001, Qiang Xu 0001 |
DAC | 3 |
| 2022 | T-WaveNet: A Tree-Structured Wavelet Neural Network for Time Series Signal Analysis
Minhao Liu, Ailing Zeng, Qiuxia Lai, Ruiyuan Gao 0001, Min Li 0019, Harry Qin, Qiang Xu 0001 |
ICLR | 3 |
| 2022 | What You See is Not What the Network Infers: Detecting Adversarial Examples Based on Semantic Contradiction
Ruiyuan Gao 0001, Yu Li 0007, Qiuxia Lai, Qiang Xu 0001 |
NDSS | 4 |
| 2022 | SCINet: Time Series Modeling and Forecasting with Sample Convolution and InteractionabstractOne unique property of time series is that the temporal relations are largely preserved after downsampling into two sub-sequences. By taking advantage of this property, we propose a novel neural network architecture that conducts sample convolution and interaction for temporal modeling and forecasting, named SCINet. Specifically, SCINet is a recursive downsample-convolve-interact architecture. In each layer, we use multiple convolutional filters to extract distinct yet valuable temporal features from the downsampled sub-sequences or features. By combining these rich features aggregated from multiple resolutions, SCINet effectively models time series with complex temporal dynamics. Experimental results show that SCINet achieves significant forecasting accuracy improvements over both existing convolutional models and Transformer-based solutions across various real-world time series forecasting datasets. Our codes and data are available at https://github.com/cure-lab/SCINet. Minhao Liu, Ailing Zeng, Muxi Chen, Qiuxia Lai, Lingna Ma, Qiang Xu 0001 |
NeurIPS | 5 |
| 2022 | Salient Object Detection in the Deep Learning Era: An In-Depth SurveyabstractAs an essential problem in computer vision, salient object detection (SOD) has attracted an increasing amount of research attention over the years. Recent advances in SOD are predominantly led by deep learning-based solutions (named deep SOD). To enable in-depth understanding of deep SOD, in this paper, we provide a comprehensive survey covering various aspects, ranging from algorithm taxonomy to unsolved issues. In particular, we first review deep SOD algorithms from different perspectives, including network architecture, level of supervision, learning paradigm, and object-/instance-level detection. Following that, we summarize and analyze existing SOD datasets and evaluation metrics. Then, we benchmark a large group of representative SOD models, and provide detailed analyses of the comparison results. Moreover, we study the performance of SOD algorithms under different attribute settings, which has not been thoroughly explored previously, by constructing a novel SOD dataset with rich attribute annotations covering various salient object types, challenging factors, and scene categories. We further analyze, for the first time in the field, the robustness of SOD models to random input perturbations and adversarial attacks. We also look into the generalization and difficulty of existing SOD datasets. Finally, we discuss several open issues of SOD and outline future research directions. All the saliency prediction maps, our constructed dataset with annotations, and codes for evaluation are publicly available at https://github.com/wenguanwang/SODsurvey. Wenguan Wang, Qiuxia Lai, Huazhu Fu, Jianbing Shen, Haibin Ling, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Weakly Supervised Visual Saliency PredictionabstractThe success of current deep saliency models heavily depends on large amounts of annotated human fixation data to fit the highly non-linear mapping between the stimuli and visual saliency. Such fully supervised data-driven approaches are annotation-intensive and often fail to consider the underlying mechanisms of visual attention. In contrast, in this paper, we introduce a model based on various cognitive theories of visual saliency, which learns visual attention patterns in a weakly supervised manner. Our approach incorporates insights from cognitive science as differentiable submodules, resulting in a unified, end-to-end trainable framework. Specifically, our model encapsulates the following important components motivated from biological vision. (a) As scene semantics are closely related to visually attentive regions, our model encodes discriminative spatial information for scene understanding through spatial visual semantics embedding. (b) To model the objectness factors in visual attention deployment, we incorporate object-level semantics embedding and object relation information. (c) Considering the "winner-take-all" mechanism in visual stimuli processing, we model the competition mechanism among objects with softmax based neural attention. (d) Lastly, a conditional center prior is learned to mimic the spatial distribution bias of visual attention. Furthermore, we propose novel loss functions to utilize supervision cues from image-level semantics, saliency prior knowledge, and self-information compression. Experiments show that our method achieves promising results, and even outperforms many of its fully supervised counterparts. Overall, our weakly supervised saliency method makes an essential step towards reducing the annotation budget of current approaches, as well as providing a more comprehensive understanding of the visual attention mechanism. Our code is available at: https://github.com/ashleylqx/WeakFixation.git. Qiuxia Lai, Tianfei Zhou, Salman Khan 0001, Hanqiu Sun, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Information Bottleneck Approach to Spatial Attention LearningabstractThe selective visual attention mechanism in the human visual system (HVS) restricts the amount of information to reach visual awareness for perceiving natural scenes, allowing near real-time information processing with limited computational capacity. This kind of selectivity acts as an ‘Information Bottleneck (IB)’, which seeks a trade-off between information compression and predictive accuracy. However, such information constraints are rarely explored in the attention mechanism for deep neural networks (DNNs). In this paper, we propose an IB-inspired spatial attention module for DNN structures built for visual recognition. The module takes as input an intermediate representation of the input image, and outputs a variational 2D attention map that minimizes the mutual information (MI) between the attention-modulated representation and the input, while maximizing the MI between the attention-modulated representation and the task label. To further restrict the information bypassed by the attention map, we quantize the continuous attention scores to a set of learnable anchor values during training. Extensive experiments show that the proposed IB-inspired spatial attention mechanism can yield attention maps that neatly highlight the regions of interest while suppressing backgrounds, and bootstrap standard DNN structures for visual recognition tasks (e.g., image classification, fine-grained recognition, cross-domain classification). The attention maps are interpretable for the decision making of the DNNs as verified in the experiments. Our code is available at this https URL. Qiuxia Lai, Yu Li 0007, Ailing Zeng, Minhao Liu, Hanqiu Sun, Qiang Xu 0001 |
IJCAI | 1 |
| 2021 | TestRank: Bringing Order into Unlabeled Test Instances for Deep Learning TasksabstractDeep learning (DL) systems are notoriously difficult to test and debug due to the lack of correctness proof and the huge test input space to cover. Given the ubiquitous unlabeled test data and high labeling cost, in this paper, we propose a novel test prioritization technique, namely TestRank, which aims at revealing more model failures with less labeling effort. TestRank brings order into the unlabeled test data according to their likelihood of being a failure, i.e., their failure-revealing capabilities. Different from existing solutions, TestRank leverages both intrinsic and contextual attributes of the unlabeled test data when prioritizing them. To be specific, we first build a similarity graph on both unlabeled test samples and labeled samples (e.g., training or previously labeled test samples). Then, we conduct graph-based semi-supervised learning to extract contextual features from the correctness of similar labeled samples. For a particular test instance, the contextual features extracted with the graph neural network and the intrinsic features obtained with the DL model itself are combined to predict its failure-revealing capability. Finally, TestRank prioritizes unlabeled test inputs in descending order of the above probability value. We evaluate TestRank on three popular image classification datasets, and results show that TestRank significantly outperforms existing test prioritization techniques. Yu Li 0007, Min Li 0019, Qiuxia Lai, Yannan Liu, Qiang Xu 0001 |
NeurIPS | 3 |
| 2021 | Understanding More About Human and Machine Attention in Deep Neural NetworksabstractHuman visual system can selectively attend to parts of a scene for quick perception, a biological mechanism known asHuman attention. Inspired by this, recent deep learning models encode attention mechanisms to focus on the most task-relevant parts of the input signal for further processing, which is calledMachine/Neural/Artificial attention. Understanding the relation between human and machine attention is important for interpreting and designing neural networks. Many works claim that the attention mechanism offers an extra dimension of interpretability by explaining where the neural networks look. However, recent studies demonstrate that artificial attention maps do not always coincide with common intuition. In view of these conflicting evidence, here we make a systematic study on using artificial attention and human attention in neural network design. With three example computer vision tasks (i.e., salient object segmentation, video action recognition, and fine-grained image classification), diverse representative backbones (i.e., AlexNet, VGGNet, ResNet) and famous architectures (i.e., Two-stream, FCN), corresponding real human gaze data, and systematically conducted large-scale quantitative studies, we quantify the consistency between artificial attention and human visual attention and offer novel insights into existing artificial attention mechanisms by giving preliminary answers to several key questions related to human and artificial attention mechanisms. Overall results demonstrate that human attention can benchmark the meaningful ‘ground-truth’ in attention-driven tasks, where the more the artificial attention is close to human attention, the better the performance; for higher-level vision tasks, it is case-by-case. It would be advisable for attention-driven tasks to explicitly force a better alignment between artificial and human attention to boost the performance; such alignment would also improve the network explainability for higher-level computer vision tasks. Qiuxia Lai, Salman Khan 0001, Yongwei Nie, Hanqiu Sun, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | DeepFuse: An IMU-Aware Network for Real-Time 3D Human Pose Estimation from Multi-View ImageabstractIn this paper, we propose a two-stage fully 3D network, namely DeepFuse, to estimate human pose in 3D space by fusing body-worn Inertial Measurement Unit (IMU) data and multi-view images deeply. The first stage is designed for pure vision estimation. To preserve data primitiveness of multi-view inputs, the vision stage uses multi-channel volume as data representation and 3D soft-argmax as activation layer. The second one is the IMU refinement stage which introduces an IMU-bone layer to fuse the IMU and vision data earlier at data level. without requiring a given skeleton model a priori, we can achieve a mean joint error of 28.9mm on TotalCapture dataset and 13.4mm on Human3.6M dataset under protocol 1, improving the SOTA result by a large margin. Finally, we discuss the effectiveness of a fully 3D network for 3D pose estimation experimentally which may benefit future research. Fuyang Huang, Ailing Zeng, Minhao Liu, Qiuxia Lai, Qiang Xu 0001 |
WACV | 4 |
| 2020 | Video super-resolution via pre-frame constrained and deep-feature enhanced sparse reconstruction
Qiuxia Lai, Yongwei Nie, Hanqiu Sun, Qiang Xu 0001, Zhensong Zhang, Mingyu Xiao 0001 |
Pattern Recognit. | 1 |
| 2020 | Video Saliency Prediction Using Spatiotemporal Residual Attentive NetworksabstractThis paper proposes a novel residual attentive learning network architecture for predicting dynamic eye-fixation maps. The proposed model emphasizes two essential issues, i.e, effective spatiotemporal feature integration and multi-scale saliency learning. For the first problem, appearance and motion streams are tightly coupled via dense residual cross connections, which integrate appearance information with multi-layer, comprehensive motion features in a residual and dense way. Beyond traditional two-stream models learning appearance and motion features separately, such design allows early, multi-path information exchange between different domains, leading to a unified and powerful spatiotemporal learning architecture. For the second one, we propose a composite attention mechanism that learns multi-scale local attentions and global attention priors end-to-end. It is used for enhancing the fused spatiotemporal features via emphasizing important features in multi-scales. A lightweight convolutional Gated Recurrent Unit (convGRU), which is flexible for small training data situation, is used for long-term temporal characteristics modeling. Extensive experiments over four benchmark datasets clearly demonstrate the advantage of the proposed video saliency model over other competitors and the effectiveness of each component of our network. Our code and all the results will be available at https://github.com/ashleylqx/STRA-Net. Qiuxia Lai, Wenguan Wang, Hanqiu Sun, Jianbing Shen |
IEEE Trans. Image Process. | 1 |
| 2020 | Multi-View Video Synopsis via Simultaneous Object-Shifting and View-Switching OptimizationabstractWe present a method for synopsizing multiple videos captured by a set of surveillance cameras with some overlapped field-of-views. Currently, object-based approaches that directly shift objects along the time axis are already able to compute compact synopsis results for multiple surveillance videos. The challenge is how to present the multiple synopsis results in a more compact and understandable way. Previous approaches show them side by side on the screen, which however is difficult for user to comprehend. In this paper, we solve the problem by joint object-shifting and camera view-switching. Firstly, we synchronize the input videos, and group the same object in different videos together. Then we shift the groups of objects along the time axis to obtain multiple synopsis videos. Instead of showing them simultaneously, we just show one of them at each time, and allow to switch among the views of different synopsis videos. In this view switching way, we obtain just a single synopsis results consisting of content from all the input videos, which is much easier for user to follow and understand. To obtain the best synopsis result, we construct a simultaneous object-shifting and view-switching optimization framework instead of solving them separately. We also present an alternative optimization strategy composed of graph cuts and dynamic programming to solve the unified optimization. Experiments demonstrate that our single synopsis video generated from multiple input videos is compact, complete, and easy to understand. Zhensong Zhang, Yongwei Nie, Hanqiu Sun, Qing Zhang 0006, Qiuxia Lai, Guiqing Li, Mingyu Xiao 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | Temporal Coherent Video Super-resolution via Pre-frame-constrained Sparse ReconstructionabstractIn this paper, we extend the sparse representation based image super-resolution method to process videos, mainly aiming at obtaining temporally consistent consecutive high-resolution (HR) video frames. In our formulation, the previous estimated HR frame is used to guide the sparse reconstruction of current low-resolution (LR) frame, which is able to obtain more consistent representations. We show that such guidance is robust and effective by incorporating with a non-rigid dense correspondence based motion compensation schema. We also propose a dictionary updating strategy which regularly updates the dictionaries that are critical for the sparse representation procedure using the newly reconstructed HR frames. To further preserve sharp edges and remove reconstruction errors, once a HR image is recovered, we refine it with a L0-norm based optimization that constrains the final HR output with relatively sparse gradients. Experimental results on natural videos demonstrated the effectiveness of our proposed method. Qiuxia Lai, Yongwei Nie, Zhensong Zhang, Hanqiu Sun |
CGI | 1 |