VLDB 2026 Research / reviewers in the wild / expert
Xiaopeng Guo 0001
dblp:131/9316-1
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-1111-2035ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-View Interaction and Collaboration for Semi-supervised Time Series Classification
Maojing Shu, Xiaopeng Guo 0001, Jun Sun 0012 |
KSEM (1) | 2 |
| 2026 | Enhancing Multivariate Time-Series Class Incremental Learning with Informative Spatiotemporal Consolidation
Maojing Shu, Xiaopeng Guo 0001, Jun Sun 0012 |
KSEM (1) | 2 |
| 2025 | STR-Saliency: Decomposition-based Perturbations to Generate Saliency Maps for Temporal Black-box Model InterpretationabstractThe interpretability of deep black-box temporal models is crucial in modern machine learning. Identifying crucial time steps and temporal patterns is an important way in understanding how a black-box model makes a decision on a time series instance. Saliency methods are widely used for the interpretation of deep models. However, its application on temporal models faces two challenges: handling temporal relations and the selection of perturbation functions. In order to overcome these challenges, we propose Seasonal-Trend-Remainder Saliency (STR-Saliency), a new interpretation framework which decomposes a time series into three components then generates saliency maps on each component using a learning-based perturbation process. Our new method not only addresses the two issues above, but also produces more human-understandable saliency maps than previous methods. We evaluate our method on both synthetic and real-world datasets and it outperforms various baselines. Our code and detailed results are available at https://github.com/chenrunkai/STRSaliency. Runkai Chen, Yuedi Chen, Haozhe Xu, Xiaopeng Guo 0001, Jun Sun 0012 |
ICASSP | 4 |
| 2024 | Programming Knowledge Tracing with Context and Structure Integration
Xiaopeng Guo 0001, Maojing Shu, Jun Sun 0012 |
KSEM (1) | 1 |
| 2023 | Generalized Compressed Video Restoration by Multi-Scale Temporal Fusion and Hierarchical Quality Score EstimationabstractLearning-based methods have achieved excellent performance for compressed video restoration (CVR) in recent years. However, existing networks aggregate multi-frame information inefficiently and are usually developed for specific quantization parameters (QPs), which are not convenient for practical usage. Moreover, current works only consider compressed video restoration in Constant QP (CQP) setting, but do not discuss the performance of the model in more realistic scenarios, e.g., Constant Rate Factor (CRF) and Constant Bitrate (CBR). In this paper, we propose a generalized quality-aware compressed video restoration network, namely QCRN. Specifically, to achieve multi-frame aggregation efficiently, we propose a multi-scale deformable temporal fusion. Meanwhile, QCRN decouples the global quality and local quality representations from input via the hierarchical quality score estimator, and then employs them to adjust the feature enhancement. Extensive experiments on compressed videos in various settings demonstrate that our proposed QCRN achieves favorable performance against state-of-the-art methods in terms of both quantitative metrics and visual quality. Xiaopeng Guo 0001, Jun Sun 0012 |
ICME | 3 |
| 2023 | FastCNN: Towards Fast and Accurate Spatiotemporal Network for HEVC Compressed Video EnhancementabstractDeep neural networks have achieved remarkable success in HEVC compressed video quality enhancement. However, most existing multiframe-based methods either deliver unsatisfactory results or consume a significant amount of resources to leverage temporal information of neighboring frames. For the sake of practicality, a thorough investigation of the architecture design of the video quality enhancement network regarding enhancement performance, model parameters, and running speed is essential. In this article, we first propose an efficient alignment module that can quickly and accurately aggregate the spatiotemporal information of neighboring frames. The proposed module estimates deformable offsets progressively in lower-resolution space motivated by the observation of offset correlations between adjacent pixels. Then, the quantization parameter (QP) that represents compression level prior knowledge is utilized to guide aligned feature enhancement. By combining alignment feature distillation with residual feature correction, we obtain an efficient QP attention block. To save the storage space of the network, we design a hash buffer to store QP embedding features. These efficient components allow our network to effectively exploit temporal redundancies and obtain favorable enhancement capability while maintaining a lightweight structure and fast running speed. Extensive experiments demonstrate that the proposed approach outperforms state-of-the-art methods over different QPs by up to 0.09 to 0.11 dB, whereas the inference time can be reduced by up to 69%. Jun Sun 0012, Xiaopeng Guo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | One-for-All: An Efficient Variable Convolution Neural Network for In-Loop Filter of VVCabstractRecently, many researches on convolution neural network (CNN) based in-loop filters have been proposed to improve coding efficiency. However, most existing CNN based filters tend to train and deploy multiple networks for various quantization parameters (QP) and frame types (FT), which drastically increases resources in training these models and the memory burdens for video codec. In this paper, we propose a novel variable CNN (VCNN) based in-loop filter for VVC, which can effectively handle the compressed videos with different QPs and FTs via a single model. Specifically, an efficient and flexible attention module is developed to recalibrate features according to QPs or FTs. Then we embed the module into the residual block so that these informative features can be continuously utilized in the residual learning process. To minimize the information loss in the learning process of the entire network, we utilize a residual feature aggregation module (RFA) for more efficient feature extraction. Based on it, an efficient network architecture VCNN is designed that can not only effectively reduce compression artifacts, but also can be adaptive to various QPs and FTs. To address training data imbalance on various QPs and FTs and improve the robustness of the model, a focal mean square error loss function is employed to train the proposed network. Then we integrate the VCNN into VVC as an additional tool of in-loop filters after the deblocking filter. Extensive experimental results show that our VCNN approach obtains on average 3.63%, 4.36%, 4.23%, 3.56% under all intra, low-delay P, low-delay, and random access configurations, respectively, which is even better than QP-Separate models. Jun Sun 0012, Xiaopeng Guo 0001, Mingyu Shang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | An Efficient QP Variable Convolutional Neural Network Based In-loop Filter for Intra CodingabstractIn this paper, a novel QP variable convolutional neural network based in-loop filter is proposed for VVC intra coding. To avoid training and deploying multiple networks, we develop an efficient QP attention module (QPAM) which can capture compression noise levels for different QPs and emphasize meaningful features along channel dimension. Then we embed QPAM into the residual block, and based on it, we design a network architecture that is equipped with controllability for different QPs. To make the proposed model focus more on examples that have more compression artifacts or is hard to restore, a focal mean square error (MSE) loss function is employed to fine tune the network. Experimental results show that our approach achieves 4.03% BD-Rate saving on average for all intra configuration, which is even better than QP-separate CNN models while having less model parameters. Xiaopeng Guo 0001, Mingyu Shang, Jun Sun 0012 |
DCC | 2 |
| 2021 | Enhancing Knowledge Tracing via Adversarial TrainingabstractWe study the problem of knowledge tracing (KT) where the goal is to trace the students' knowledge mastery over time so as to make predictions on their future performance. Owing to the good representation capacity of deep neural networks (DNNs), recent advances on KT have increasingly concentrated on exploring DNNs to improve the performance of KT. However, we empirically reveal that the DNNs based KT models may run the risk of overfitting, especially on small datasets, leading to limited generalization. In this paper, by leveraging the current advances in adversarial training (AT), we propose an efficient AT based KT method (ATKT) to enhance KT model's generalization and thus push the limit of KT. Specifically, we first construct adversarial perturbations and add them on the original interaction embeddings as adversarial examples. The original and adversarial examples are further used to jointly train the KT model, forcing it is not only to be robust to the adversarial examples, but also to enhance the generalization over the original ones. To better implement AT, we then present an efficient attentive-LSTM model as KT backbone, where the key is a proposed knowledge hidden state attention module that adaptively aggregates information from previous knowledge hidden states while simultaneously highlighting the importance of current knowledge hidden state to make a more accurate prediction. Extensive experiments on four public benchmark datasets demonstrate that our ATKT achieves new state-of-the-art performance. Code is available at: https://github.com/xiaopengguo/ATKT. Xiaopeng Guo 0001, Mingyu Shang, Maojing Shu, Jun Sun 0012 |
ACM Multimedia | 1 |
| 2021 | Adaptive Deep Reinforcement Learning-Based In-Loop Filter for VVCabstractDeep learning-based in-loop filters have recently demonstrated great improvement for both coding efficiency and subjective quality in video coding. However, most existing deep learning-based in-loop filters tend to develop a sophisticated model in exchange for good performance, and they employ a single network structure to all reconstructed samples, which lack sufficient adaptiveness to the various video content, limiting their performances to some extent. In contrast, this paper proposes an adaptive deep reinforcement learning-based in-loop filter (ARLF) for versatile video coding (VVC). Specifically, we treat the filtering as a decision-making process and employ an agent to select an appropriate network by leveraging recent advances in deep reinforcement learning. To this end, we develop a lightweight backbone and utilize it to design a network set S containing networks with different complexities. Then a simple but efficient agent network is designed to predict the optimal network from S , which makes the model adaptive to various video contents. To improve the robustness of our model, a two-stage training scheme is further proposed to train the agent and tune the network set. The coding tree unit (CTU) is seen as the basic unit for the in-loop filtering processing. A CTU level control flag is applied in the sense of rate-distortion optimization (RDO). Extensive experimental results show that our ARLF approach obtains on average 2.17%, 2.65%, 2.58%, 2.51% under all-intra, low-delay P, low-delay, and random access configurations, respectively. Compared with other deep learning-based methods, the proposed approach can achieve better performance with low computation complexity. Jun Sun 0012, Xiaopeng Guo 0001, Mingyu Shang |
IEEE Trans. Image Process. | 3 |
| 2020 | Page-Level Handwritten Word Spotting via Discriminative Feature Learning
Xiaopeng Guo 0001, Mingyu Shang, Jun Sun 0012 |
KSEM (1) | 2 |
| 2020 | Multi-focus image fusion with Siamese self-attention networkabstractRecently, convolutional neural networks (CNNs) have achieved impressive progress in multi‐focus image fusion (MFF). However, it always fails to capture sufficient discrimination features due to the local receptive field limitations of the convolutional operator, restricting most current CNN‐based methods’ performance. To address this issue, by leveraging self‐attention (SA) mechanism, the authors propose Siamese SA network (SSAN) for MFF. Specifically, two kinds of SA modules, position SA (PSA) and channel SA (CSA) are utilised to model the long‐range dependencies across focused and defocused regions in the multi‐focus image, alleviating the local receptive field limitations of convolution operators in CNN. To search a better feature representation of the input image for MFF, the captured features obtained by PSA and CSA are further merged through a learnable 1 × 1 convolution operator. The whole pipeline is in a Siamese network fashion to reduce the complexity. After training, the authors SSAN can accomplish well the fusion task with no post‐processing. Experiments demonstrate that their approach outperforms other current state‐of‐the‐art methods, not only in subjective visual perception but also in the quantitative assessment. Xiaopeng Guo 0001, Lingyu Meng, Liye Mei, Yueyun Weng, Hengqing Tong |
IET Image Process. | 1 |
| 2019 | FuseGAN: Learning to Fuse Multi-Focus Image via Conditional Generative Adversarial NetworkabstractWe study the problem of multi-focus image fusion, where the key challenge is detecting the focused regions accurately among multiple partially focused source images. Inspired by the conditional generative adversarial network (cGAN) to image-to-image task, we propose a novel FuseGAN to fulfill the images-to-image for multi-focus image fusion. To satisfy the requirement of dual input-to-one output, the encoder of the generator in FuseGAN is designed as a Siamese network. The least square GAN objective is employed to enhance the training stability of FuseGAN, resulting in an accurate confidence map for focus region detection. Also, we exploit the convolutional conditional random fields technique on the confidence map to reach a refined final decision map for better focus region detection. Moreover, due to the lack of a large-scale standard dataset, we synthesize a large enough multi-focus image dataset based on a public natural image dataset PASCAL VOC 2012, where we utilize a normalized disk point spread function to simulate the defocus and separate the background and foreground in the synthesis for each image. We conduct extensive experiments on two public datasets to verify the effectiveness of the proposed method. Results demonstrate that the proposed method presents accurate decision maps for focus regions in multi-focus images, such that the fused images are superior to 11 recent state-of-the-art algorithms, not only in visual perception, but also in quantitative analysis in terms of five metrics. Xiaopeng Guo 0001, Rencan Nie, Jinde Cao, Dongming Zhou 0001, Liye Mei, Kangjian He |
IEEE Trans. Multim. | 1 |
| 2018 | Fully Convolutional Network-Based Multifocus Image FusionabstractAs the optical lenses for cameras always have limited depth of field, the captured images with the same scene are not all in focus. Multifocus image fusion is an efficient technology that can synthesize an all-in-focus image using several partially focused images. Previous methods have accomplished the fusion task in spatial or transform domains. However, fusion rules are always a problem in most methods. In this letter, from the aspect of focus region detection, we propose a novel multifocus image fusion method based on a fully convolutional network (FCN) learned from synthesized multifocus images. The primary novelty of this method is that the pixel-wise focus regions are detected through a learning FCN, and the entire image, not just the image patches, are exploited to train the FCN. First, we synthesize 4500 pairs of multifocus images by repeatedly using a gaussian filter for each image from PASCAL VOC 2012, to train the FCN. After that, a pair of source images is fed into the trained FCN, and two score maps indicating the focus property are generated. Next, an inversed score map is averaged with another score map to produce an aggregative score map, which take full advantage of focus probabilities in two score maps. We implement the fully connected conditional random field (CRF) on the aggregative score map to accomplish and refine a binary decision map for the fusion task. Finally, we exploit the weighted strategy based on the refined decision map to produce the fused image. To demonstrate the performance of the proposed method, we compare its fused results with several start-of-the-art methods not only on a gray data set but also on a color data set. Experimental results show that the proposed method can achieve superior fusion performance in both human visual quality and objective assessment. Xiaopeng Guo 0001, Rencan Nie, Jinde Cao, Dongming Zhou 0001, Wenhua Qian |
Neural Comput. | 1 |