EDBT 2026 Demo / reviewers in the wild / expert
Yanbo Gao
dblp:44/4817
· DBLP profile ↗
45ranked-venue papers
10as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 8 first-author · 20 since 2021Artificial intelligence and machine learning · 15 · 8 since 2021Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Frequency-Guided Alignment Transformer for Compressed Video Quality EnhancementabstractDuring the video encoding process, the original spatial domain signal is first transformed into the frequency domain, followed by quantization and compression. As a result, the quality degradation in compressed videos primarily stems from distortions in the frequency domain information. However, existing video enhancement methods typically directly fuse information from adjacent frames in the spatial domain, making it difficult for models to effectively compensate for frequency domain distortions, which leads to suboptimal detail restoration. To address this issue, we propose a Hierarchical Frequency-Guided Alignment Transformer. Additionally, by analyzing the characteristics of the frequency domain, we find that different frequency bands exhibit both correlations and a certain degree of independence. Based on this, we introduce a Frequency-Aware Transformer module that employs a combination of independent and mixed processing to optimize information exchange across different frequency domains, effectively mitigating cross-interference from irrelevant information. Experimental results demonstrate that, compared to existing methods, our approach achieves state-of-the-art performance in objective metrics (PSNR/SSIM), perceptual quality (LPIPS), and subjective visual effects, while reducing model complexity. Liuhan Peng, Shuai Li 0005, Yanbo Gao, Mao Ye 0001, Chong Lv |
AAAI | 3 |
| 2026 | A Noise Constrained Diffusion (NC-Diffusion) Framework for High-Fidelity Image CompressionabstractWith the great success of diffusion models in image generation, diffusion-based image compression is attracting increasing interests. However, due to the random noise introduced in the diffusion learning, they usually produce reconstructions with deviation from the original images, leading to suboptimal compression results. To address this problem, in this paper, we propose a Noise Constrained Diffusion (NC-Diffusion) framework for high fidelity image compression. Unlike existing diffusion-based compression methods that add random Gaussian noise and direct the noise into the image space, the proposed NC-Diffusion formulates the quantization noise originally added in the learned image compression as the noise in the forward process of diffusion. Then a noise constrained diffusion process is constructed from the ground-truth image to the initial compression result generated with quantization noise. The NC-Diffusion overcomes the problem of noise mismatch between compression and diffusion, significantly improving the inference efficiency. In addition, an adaptive frequency-domain filtering module is developed to enhance the skip connections in the U-Net based diffusion architecture, in order to enhance high-frequency details. Moreover, a zero-shot sample-guided enhancement method is designed to further improve the fidelity of the image. Experiments on multiple benchmark datasets demonstrate that our method can achieve the best performance compared with existing methods. Yanbo Gao, Shuai Li 0005, Hui Yuan 0001, Mao Ye 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | CWRNN-INVR: A Coupled WarpRNN Based Implicit Neural Video RepresentationabstractImplicit Neural Video Representation (INVR) has emerged as a novel approach for video representation and compression, using learnable grids and neural networks. Existing methods focus on developing new grid structures efficient for latent representation and neural network architectures with large representation capability, lacking the study on their roles in video representation. In this paper, the difference between INVR based on neural network and INVR based on grid is first investigated from the perspective of video information composition to specify their own advantages, i.e., neural network for general structure while grid for specific detail. Accordingly, an INVR based on mixed neural network and residual grid framework is proposed, where the neural network is used to represent the regular and structured information and the residual grid is used to represent the remaining irregular information in a video. A Coupled WarpRNN-based multi-scale motion representation and compensation module is specifically designed to explicitly represent the regular and structured information, thus terming our method as CWRNN-INVR. For the irregular information, a mixed residual grid is learned where the irregular appearance and motion information are represented together. The mixed residual grid can be combined with the coupled WarpRNN in a way that allows for network reuse. Experiments show that our method achieves the best reconstruction results compared with the existing methods, with an average PSNR of 33.73 dB on the UVG dataset under the 3M model and outperforms existing INVR methods in other downstream tasks. The code can be found athttps://github.com/yiyang-sdu/CWRNN-INVR.git. Yanbo Gao, Shuai Li 0005, Jinglin Zhang 0001, Hui Yuan 0001, Mao Ye 0001, Xingyu Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | MetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural RepresentationabstractThis paper presents MetricGrids, a novel grid-based neural representation that combines elementary metric grids in various metric spaces to approximate complex nonlinear signals. While grid-based representations are widely adopted for their efficiency and scalability, the existing feature grids with linear indexing for continuous-space points can only provide degenerate linear latent space representations, and such representations cannot be adequately compensated to represent complex nonlinear signals by the following compact decoder. To address this problem while keeping the simplicity of a regular grid structure, our approach builds upon the standard grid-based paradigm by constructing multiple elementary metric grids as high-order terms to approximate complex nonlinearities, following the Taylor expansion principle. Furthermore, we enhance model compactness with hash encoding based on different sparsities of the grids to prevent detrimental hash collisions, and a high-order extrapolation decoder to reduce explicit grid storage requirements. experimental results on both 2D and 3D reconstructions demonstrate the superior fitting and rendering accuracy of the proposed method across diverse signal types, validating its robustness and generalizability. Code is available at https://github.com/wangshu31/MetricGrids. Yanbo Gao, Shuai Li 0005, Chong Lv, Chuankun Li, Hui Yuan 0001, Jinglin Zhang 0001 |
CVPR | 2 |
| 2025 | A Fourier priors-Guided Diffusion Model for Image Harmonization with Structure-Preservation and Illumination-ConsistencyabstractImage harmonization is a crucial computer vision task that adjusts the appearance of foreground regions in composite images to match the background. Existing methods face challenges in effectively separating illumination from structure and preserving content consistency. In this paper, we propose a Fourier priors-Guided Diffusion Model for Image Harmonization with Structure Preservation and Illumination Consistency (FGDIH), which leverages frequency domain information and diffusion models. The key insight is that illumination is primarily concentrated in low-frequency amplitude, while structure is preserved in phase. Leveraging this insight, a Background-guided Amplitude Transfer Module (BATM) is proposed to transfer illumination from background to foreground. Meanwhile, a Phase Preservation Module (PPM) is developed to keep the original structure. Furthermore, FGDIH introduces a dual-constraint training strategy that incorporates both content and frequency domain supervision, ensuring stable harmonization during inference. Extensive experiments on iHarmony4 datasets demonstrate that our algorithm achieves SOTA performance and delivers competitive visual quality. Tianyou Wang, Yanbo Gao, Shuai Li 0005 |
ICME | 3 |
| 2025 | An Optimization based on Cyclic Check Network for the Process Simulation of PolyacrylonitrileabstractProcess simulation serves as a vital technical means in the manufacturing industry to enhance production efficiency and improve product quality. However, each step in complex processes will have a variety of physical and chemical reactions that involve a large number of parameters and variables. Meanwhile, the prolonged production cycle and high cost associated with complex processes lead to insufficient authentic samples for training the simulation model. To tackle these challenges, this paper proposes an optimization method based on Cyclic Check Network. The network offers a forward mapping from raw material to finished product and a backward mapping from finished product to raw material. During the training phase, the initial input undergoes an additional cyclic check. Besides, to ensure logical consistency between these two mappings, cycle-consistency loss is introduced to measure the change before and after the cycle. Experimental results on the polyacrylonitrile dataset and other public datasets demonstrate that the proposed method outperforms conventional simulation approaches in terms of accuracy and stability for complex processes on small datasets. Siwei Xue, Yanbo Gao, Tianyou Wang, Jianguo Lv, Kin-Tak Lau |
IJCNN | 3 |
| 2025 | Diverse Information Aggregation with Adaptive Graph Construction and prompts for deepfake detection
Zhenhua Bai, Qiangchang Wang, Lu Yang 0005, Xinxin Zhang 0004, Yanbo Gao, Yilong Yin |
Image Vis. Comput. | 5 |
| 2025 | Adaptive Depth-Converted-Scale Convolution for Self-Supervised Monocular Depth EstimationabstractSelf-supervised monocular depth estimation (MDE) has received increasing interests in the last few years. The objects in the scene, including the object size and relationship among different objects, are the main clues to extract the scene structure. However, previous works lack the explicit handling of the changing sizes of the object due to the change of its depth. Especially in a monocular video, the size of the same object is continuously changed, resulting in size and depth ambiguity. To address this problem, we propose a Depth-converted-Scale Convolution (DcSConv) enhanced monocular depth estimation framework, by incorporating the prior relationship between the object depth and object scale to extract features from appropriate scales of the convolution receptive field. The proposed DcSConv focuses on the adaptive scale of the convolution filter instead of the local deformation of its shape. It establishes that the scale of the convolution filter matters no less (or even more in the evaluated task) than its local deformation. Moreover, a Depth-converted-Scale aware Fusion (DcS-F) is developed to adaptively fuse the DcSConv features and the conventional convolution features. Our DcSConv enhanced monocular depth estimation framework can be applied on top of existing CNN based methods as a plug-and-play module to enhance the conventional convolution block. Extensive experiments with different baselines have been conducted on the KITTI benchmark and our method achieves the best results with an improvement up to 11.6% in terms of SqRel reduction. Ablation study also validates the effectiveness of each proposed module. Yanbo Gao, Huibin Bai, Huasong Zhou, Xingyu Gao 0001, Shuai Li 0005, Hui Yuan 0001, Wei Hua 0002, Tian Xie 0011 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Unsupervised Feature Enrichment and Fidelity Preservation Learning Framework for Skeleton-Based Action RecognitionabstractUnsupervised skeleton-based action recognition has achieved remarkable progress recently. Existing unsupervised learning methods suffer from severe overfitting problem, and thus small networks are used, significantly reducing the representation capability. To address this problem, the overfitting mechanism behind the unsupervised learning for skeleton-based action recognition is first investigated. It is observed that skeleton is already a relatively high-level and low-dimension feature, but not in the same manifold as the features for action recognition. Simply applying the existing unsupervised learning method tends to produce features that discriminate the different samples rather than action classes, resulting in the overfitting problem. To address this problem, this paper proposes an Unsupervised spatial-temporal Feature Enrichment and Fidelity Preservation (U-FEFP) learning framework to generate rich distributed features that contain all the information of a skeleton sample. A spatial-temporal feature transformation subnetwork is developed using channel-wise topology refinement graph convolutional block and graph convolutional gated recurrent unit block as the basic feature extraction network. The unsupervised Bootstrap Your Own Latent-based learning is utilized to generate rich distributed features, and the unsupervised pretext task-based learning is employed to preserve the information contained in the skeleton. The two unsupervised learning ways are collaborated as U-FEFP to produce robust and discriminative representations. Experimental results on four widely used benchmarks, namely NTU-RGB+D-60, PKU-MMD, NTU-RGB+D-120 and AAV-Human dataset, demonstrate that the proposed U-FEFP obtains the best result compared with the state-of-the-art unsupervised learning methods. Chuankun Li, Shuai Li 0005, Yanbo Gao, Xingyu Gao 0001, Ping Chen 0004, Wanqing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Approximately Invertible Neural Network for Learned Image CompressionabstractLearned image compression has attracted considerable interests in recent years. An analysis transform and a synthesis transform, which can be regarded as coupled transforms, are used to encode an image to latent feature and decode the feature after quantization to reconstruct the image. Inspired by the success of invertible neural networks in generative modeling, invertible modules can be used to construct the coupled analysis and synthesis transforms. Considering the noise introduced in the feature quantization invalidates the invertible process, this paper proposes an Approximately Invertible Neural Network (A-INN) framework for learned image compression. It formulates the rate-distortion optimization in lossy image compression when using INN with quantization, which differentiates from using INN for generative modelling. Generally speaking, A-INN can be used as the theoretical foundation for any INN based lossy compression method. Based on this formulation, A-INN with a progressive denoising module (PDM) is developed to effectively reduce the quantization noise in the decoding. Moreover, a Cascaded Feature Recovery Module (CFRM) is designed to learn high-dimensional feature recovery from low-dimensional ones to further reduce the noise in feature channel compression. In addition, a Frequency-enhanced Decomposition and Synthesis Module (FDSM) is developed by explicitly enhancing the high-frequency components in an image to address the loss of high-frequency information inherent in neural network based image compression, thereby enhancing the reconstructed image quality. Extensive experiments demonstrate that the proposed A-INN framework achieves better or comparable compression efficiency than the conventional image compression approach and state-of-the-art learned image compression methods. Yanbo Gao, Shuai Li 0005, Chong Lv, Hui Yuan 0001, Mao Ye 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | LiftFormer: Lifting and Frame Theory Based Monocular Depth Estimation Using Depth and Edge Oriented Subspace RepresentationabstractMonocular depth estimation (MDE) has attracted increasing interest in the past few years, owing to its important role in 3D vision. MDE is the estimation of a depth map from a monocular image/video to represent the 3D structure of a scene, which is a highly ill-posed problem. To solve this problem, in this paper, we propose a LiftFormer based on lifting theory topology, for constructing an intermediate subspace that bridges the image color features and depth values, and a subspace that enhances the depth prediction around edges. MDE is formulated by transforming the depth value prediction problem into depth-oriented geometric representation (DGR) subspace feature representation, thus bridging the learning from color values to geometric depth values. A DGR subspace is constructed based on frame theory by using linearly dependent vectors in accordance with depth bins to provide a redundant and robust representation. The image spatial features are transformed into the DGR subspace, where these features correspond directly to the depth values. Moreover, considering that edges usually present sharp changes in a depth map and tend to be erroneously predicted, an edge-aware representation (ER) subspace is constructed, where depth features are transformed and further used to enhance the local features around edges. The experimental results demonstrate that our LiftFormer achieves state-of-the-art performance on widely used datasets, and an ablation study validates the effectiveness of both proposed lifting modules in our LiftFormer. Shuai Li 0005, Huibin Bai, Yanbo Gao, Chong Lv, Hui Yuan 0001, Chuankun Li, Wei Hua 0002, Tian Xie 0011 |
IEEE Trans. Multim. | 3 |
| 2024 | A Bi-Directional Prediction based on Multi-Stage Deep Multilayer Perceptron for the Productive Process of Polyacrylonitrile PrecursorabstractThe precipitating bath is a vital process of producing Polyacrylonitrile precursor in the production of carbon fiber, involving multi-stages with many independent controlling parameters. In this paper, a bi-directional simulation model of precipitating bath process is proposed, which can not only simulate the whole process of precipitating bath but also predict the controllable parameters of every sub process of precipitating bath. The proposed model consists of two parts: forward cascade multi-level parameter prediction based on multi-layer perceptron and hierarchical backward multi-level parameter prediction based on genetic algorithm. The experimental results show that this model can not only simulate the multistage precipitating bath process but also can predict the process parameters efficiently. Siwei Xue, Tianyou Wang, Yeqian Yang, Jianguo Lv, Kin-Tak Lau, Yanbo Gao |
CSCWD | 7 |
| 2024 | ReCIDE: robust estimation of cell type proportions by integrating single-reference-based deconvolutionsabstractIn this study, we introduce Robust estimation of Cell type proportions by Integrating single-reference-based DEconvolutions (ReCIDE), an innovative framework for robust estimation of cell type proportions by integrating single-reference-based deconvolutions. ReCIDE outperforms existing approaches in benchmark and real datasets, particularly excelling in estimating rare cell type proportions. Through exploratory analysis on public bulk data of triple-negative breast cancer (TNBC) patients using ReCIDE, we demonstrate a significant correlation between the prognosis of TNBC patients and the proportions of both T cell and perivascular-like cell subtypes. Built upon this discovery, we develop a prognostic assessment model for TNBC patients. Our contribution presents a novel framework for enhancing deconvolution accuracy, showcasing its effectiveness in medical research. Yuqing Su, Yanbo Gao |
Briefings Bioinform. | 3 |
| 2024 | Static graph convolution with learned temporal and channel-wise graph topology generation for skeleton-based action recognition
Chuankun Li, Shuai Li 0005, Yanbo Gao, Lijuan Zhou 0002, Wanqing Li 0001 |
Comput. Vis. Image Underst. | 3 |
| 2024 | Color and Geometric Contrastive Learning Based Intra-Frame Supervision for Self-Supervised Monocular Depth EstimationabstractIn recent years, self-supervised monocular depth estimation has become popular due to its advantage in estimating the depth without the need of groundtruth depth labels. Instead, it takes an inter-frame supervision using depth based view synthesis to reconstruct temporal adjacent frames to indirectly supervise the generated depth. However, such supervision weakens the depth estimation at temporal incoherent regions containing small changes among consecutive frames. To overcome the above problem, we propose a color and geometric contrastive learning based intra-frame supervision framework to enhance self-supervised monocular depth estimation. Color-contrastive learning is proposed to guide the network to learn color invariant features considering color information is irrelevant to depth data. To improve the local details of the learned feature, a pixel-level contrastive learning is further used to optimize the learning. In view that the depth estimation, as a pixel-level task, is sensitive to the geometric transformation, geometric-contrastive learning is developed using an inverse geometric transformation to learn features that are equivariant to the geometric data augmentation. A local plane guidance layer (LPG) with contrastive learning is further used to decompose the geometric information and enhance the geometric contrastive learning. Experiments demonstrate that the proposed method achieves the best result compared to the state-of-the-art methods in all tested quality metrics, with the largest improvement of 22.8% over baseline Monodepth2 and 3.2% over Monovit, in terms of SqRel reduction. Yanbo Gao, Xianye Wu, Shuai Li 0005, Chuankun Li |
IEEE Signal Process. Lett. | 1 |
| 2024 | Aligned Intra Prediction and Hyper Scale Decoder Under Multistage Context Model for JPEG AIabstractLearning-based image compression has raised increasing interests in the last few years. Currently, Joint Photographic Experts Group (JPEG) is working on the standardization of learning-based image compression as JPEG AI. It adopts a deep neural network based encoder-decoder architecture with hyperprior based probability formulation for entropy coding. JPEG AI currently contains two coding profiles, including the Base Operating Point (BaseOP) and High Operating Point (HighOP). Among the various techniques developed in JPEG AI, Multistage Context Model (MCM) was adopted as the context model to perform intra prediction in HighOP. It transforms the spatially progressive context prediction into sub-image feature prediction among channels via feature down-shuffling. However, in this prediction process, sub-image features are not spatially aligned to each other, and directly using the neighboring sub-image features cannot provide accurate prediction. Moreover, the distributions of residual features generated by MCM are also not consistent with that of the hyper scale decoder, which is used to construct the probability model in the entropy coding of residual features, leading to suboptimal residual coding. To address the above problems, we propose an Aligned Intra Prediction (AIP) and Aligned Hyper Scale Decoder (AHSD) under Multistage Context Model for JPEG AI coding. AIP aligns the reference sub-image features to the to-be-predicted feature in MCM with an offset prediction network and deformable convolution. AHSD further generates hyper scale features with matched distributions to the residual features, in order to enhance the probability formulation in its entropy coding. Experimental results demonstrate that the proposed method improves the coding performance by 1.3% in terms of BD-rate saving over the JPEG AI reference software and the effectiveness of each module is verified in ablation study. Shuai Li 0005, Yanbo Gao, Chuankun Li, Hui Yuan 0001 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Geometric Warping Error Aware Spatial-Temporal Enhancement for DIBR Oriented View Synthesis
Rui Peng 0009, Shuai Li 0005, Yanbo Gao, Chuankun Li |
IEEE Signal Process. Lett. | 4 |
| 2024 | A Structure-Preserving and Illumination-Consistent Cycle Framework for Image HarmonizationabstractIna composite image, the foreground and background are filmed under different scenarios, such as different lighting conditions, causing inconsistency and reducing the overall realism of the image. Image harmonization aims to generate visually realistic composite images by adjusting the foreground to the background conditions while maintaining the structure. Existing methods focus on adjusting the foreground object by directly training the foreground generation network with the ground truth, neglecting the different roles of the illumination and structure of the foreground in image harmonization. Moreover, the use of background, except for providing illumination, is not thoroughly investigated in this task. In this paper, we propose a structure-preserving and illumination-consistent cycle (SP-IC cycle) framework for image harmonization by exploring the illumination and structure of both the foreground and background. It achieves image harmonization by specifically changing the illumination and keeping the structure instead of ambiguously changing the foreground. Then, an illumination-consistent foreground harmonization cycle is developed to change the foreground illumination, while a structure-preserving cycle is designed to keep the foreground structure. Background information is explored in both cycles to assist in decomposing the illumination and structure of the foreground. In addition, the proposed SP-IC cycle framework can be applied to any image harmonization method to further boost its performance. Experimental results demonstrate that our method achieves better harmonious image quality than state-of-the-art methods, especially on an illumination-varying dataset. Qingjie Shi, Yanbo Gao, Shuai Li 0005, Wei Hua 0002, Tian Xie 0011 |
IEEE Trans. Multim. | 3 |
| 2023 | Improved Shift Graph Convolutional Network for Action Recognition With SkeletonabstractShift graph convolutional network (Shift-GCN) achieves remarkable performance for skeleton based action recognition with lower computational complexity than other GCN based methods. However, the current Shift-GCN, with one spatial shift, a static mask and a local temporal convolution, cannot fully explore the spatial-temporal features among skeleton joints of different frames. In order to address these problems, an improved shift graph convolutional network (Ishift-GCN) is proposed in this letter. The Ishift-GCN consists of two parts including a bidirectional spatial shift graph convolution with a dynamic mask, and a multi-scale temporal shift graph convolution. The bidirectional spatial shift graph convolution exploits more spatial information among joints, and the dynamic mask with stronger generalization ability can learn different correlations among features of different joints for different actions. The multi-scale temporal shift graph convolution captures more temporal information by complementing the shifted features with multi-scale convolution. Furthermore, knowledge distillation is used to reduce computational complexity. Compared with Shift-GCN, the proposed Ishift-GCN achieves better results with less computation complexity on two widely used benchmarks, namely the NTU-RGB+D and UAV-Human dataset. Chuankun Li, Shuai Li 0005, Yanbo Gao, Wanqing Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | A Multiscale Gradient-Backpropagation Optimization Framework for Deformable Convolution Based Compressed Video EnhancementabstractDeep learning based compressed video quality enhancement has raised lots of interest recently. To explore the information over multiple frames, deformable convolution has been used for temporal alignment. However, in the existing methods, the deformable convolution is used in a relatively naïve way, without differing the characteristics of offset and features, and their behavior in gradient backpropagation. In this paper, a multiscale gradient-backpropagation optimization framework is proposed for the deformable convolution based compressed video quality enhancement. By analyzing the gradient backpropagation mechanism of deformable convolution, a multi-scale deformable convolution alignment structure is developed to facilitate the gradient backpropagation at all scales. Moreover, a progressive offset prediction module is developed, which decouples the offset prediction from the feature up-sampling, thus reducing the noise flow over scales. Experimental results show that the proposed method achieves the state-of-the-art performance, with 25.6% BD-rate saving compared to the HEVC reference software (HM). Yanbo Gao, Menghu Jia, Shuai Li 0005, Mao Ye 0001, Frédéric Dufaux |
ICASSP | 1 |
| 2022 | Light Field Integral Image Coding Optimization under 2D Hierarchical Coding StructureabstractIn integral photography, one important acquisition method of light field, a 3D scene is captured from different perspectives and the data with repetitive views is contained in one integral image. To encode such integral images, 2D hierarchical coding structure (2D-HCS) is used to exploit the dependency among images of different views. However, the fixed 2D-HCS cannot fully explore the inter-view dependency, making the coding not optimal. Therefore, in this paper, an inter-view dependent rate-distortion optimization (RDO) method is proposed for light field integral image coding under 2D-HCS. Experimental results demonstrate that, compared to 2D-HCS, the proposed method can achieve averaged {Y, U, V} BD-rate savings of {13.2%, 9.2%, 12.0%} only by adapting the Lagrange Multiplier and up to {21.6%, 18.4%, 21.8%} together with QP adaptation. Yanbo Gao, Ce Zhu |
ICIP | 1 |
| 2022 | Content Adaptive Compressed Screen Content Video Quality EnhancementabstractIn recent years, with the rise of various online learning plat-forms and game live broadcasting industry, screen content video is explosively increasing. There is an urgent demand to reduce the inevitable compression artifacts produced by traditional lossy compression method. However, there do not exist any research on compressed Screen Content Video (SCV) quality enhancement. Since SCV frames always con-sist of two main types of contents with different characteris-tics, i.e., text and graphic, a Content Adaptive model based on Two branches (CAT) is proposed in this paper. For en-hancing graphic, we utilize temporal information by motion compensation, while for enhancing text, we explore spatial relevance in horizontal and vertical directions to recover sharp edges. Furthermore, a content adaptive block is used to select content related features for the collaborative enhancement of both contents. We build a large SCV dataset compressed by H.266/VVC Test Model (VTM 12.1). On this dataset, experi-mental results demonstrate that the proposed method achieves the state-of-the-art performance on 10 different kinds of SCV test sequences. Mao Ye 0001, Yanbo Gao, Shuai Li 0005, Xue Li 0001 |
ICME | 3 |
| 2022 | Geometric Warping Error Aware CNN for DIBR Oriented View SynthesisabstractDepth Image based Rendering (DIBR) oriented view synthesis is an important virtual view generation technique. It warps the reference view images to the target viewpoint based on their depth maps, without requiring many available viewpoints. However, in the 3D warping process, pixels are warped to fractional pixel locations and then rounded (or interpolated) to integer pixels, resulting in geometric warping error and reducing the image quality. This resembles, to some extent, the image super-resolution problem, but with unfixed fractional pixel locations. To address this problem, we propose a geometric warping error aware CNN (GWEA) framework to enhance the DIBR oriented view synthesis. First, a deformable convolution based geometric warping error aware alignment (GWEA-DCA) module is developed, by taking advantage of the geometric warping error preserved in the DIBR module. The offset learned in the deformable convolution can account for the geometric warping error to facilitate the mapping from the fractional pixels to integer pixels. Moreover, in view that the pixels in the warped images are of different qualities due to the different strengths of warping errors, an attention enhanced view blending (GWEA-AttVB) module is further developed to adaptively fuse the pixels from different warped images. Finally, a partial convolution based hole filling and refinement module fills the remaining holes and improves the quality of the overall image. Experiments show that our model can synthesize higher-quality images than the existing methods, and ablation study is also conducted, validating the effectiveness of each proposed module. Shuai Li 0005, Yanbo Gao, Mao Ye 0001 |
ACM Multimedia | 3 |
| 2022 | An explicit self-attention-based multimodality CNN in-loop filter for versatile video coding
Menghu Jia, Yanbo Gao, Shuai Li 0005, Jian Yue, Mao Ye 0001 |
Multim. Tools Appl. | 2 |
| 2022 | Passivity-based Bipartite Synchronization of Coupled Delayed Inertial Neural Networks via Non-reduced Order Method
Xinyu Zhong, Jie Ren 0002, Yanbo Gao |
Neural Process. Lett. | 3 |
| 2021 | Deep Marginal Fisher Analysis based CNN for Image Representation and ClassificationabstractDeep Convolutional Neural Networks (CNNs) have achieved great success in image classification. While conventional CNNs optimized with iterative gradient descent algorithms with large data have been widely used and investigated, there is also research focusing on learning CNNs with non-iterative optimization methods such as the principle component analysis network (PCANet). It is very simple and efficient but achieves competitive performance for some image classification tasks especially on tasks with only a small amount of data available. This paper further extends this line of research and proposes a deep Marginal Fisher Analysis (MFA) based CNN, termed as DMNet. It addresses the limitation of PCANet like CNNs when the samples do not follow Gaussian distribution, by using a local MFA for CNN filter optimization. It uses a graph embedding framework for convolution filter optimization by maximizing the inter-class discriminability among marginal points while minimizing intra-class distance. Cascaded MFA convolution layers can be used to construct a deep network. Moreover, a binary stochastic hashing is developed by randomly selecting features with a probability based on the importance of feature maps for binary hashing. Experimental results demonstrate that the proposed method achieves state-of-the-art result in non-iterative optimized CNN methods, and ablation studies have been conducted to verify the effectiveness of the proposed modules in our DMNet. Jiajing Chai, Yanbo Gao, Shuai Li 0005 |
ACM Multimedia | 3 |
| 2021 | Leader-following consensus of delayed neural networks under multi-layer signed graphs
Jie Ren 0002, Qiang Song 0001, Yanbo Gao, Guoping Lu |
Neurocomputing | 3 |
| 2021 | Synchronization of Inertial Neural Networks With Time-Varying Delays via Quantized Sampled-Data ControlabstractThis article addresses the quantized sampled-data (QSD) synchronization for inertial neural networks (INNs) with heterogeneous time-varying delays, in which the sampled-data control and state quantization effect have been considered. By utilizing a proper variable substitution to transform the original system into a first-order differential system, choosing a new Lyapunov-Krasovskii functional (LKF) containing both the continuous terms and the discontinuous terms, and applying Jensen inequality and an improved reciprocally convex inequality to estimate the derivative of the LKF, the sufficient conditions for QSD synchronization for INNs are newly obtained in terms of linear matrix inequalities (LMIs), and the desired QSD controllers are designed by solving a set of LMIs. Finally, three numerical examples are provided to validate the effectiveness and benefit of the proposed results. Xinyu Zhong, Yanbo Gao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Bidirectional Independently Recurrent Neural Network for Skeleton-Based Hand Gesture RecognitionabstractGestures are a common form of human communication and important for Human-Computer Interaction (HCI). In this paper, we propose a new approach for skeleton-based hand gesture recognition based on the Independently Recurrent Neural Network (IndRNN). First, a bidirectional IndRNN (Bi-IndRNN) is developed to extend the IndRNN with the capability of bidirectional processing. Then, a deep Bi-IndRNN network is constructed for gesture recognition, where, in addition to the joint coordinates, the temporal displacement of each joint is also used to enhance the input features. Experimental results demonstrate that the proposed method achieves the state-of-the-art performance on the widely used DHG dataset with an accuracy of 93.15% for the 14 gesture classes case and 91.13% for the 28 gesture classes case. Shuai Li 0005, Longfei Zheng, Ce Zhu, Yanbo Gao |
ISCAS | 4 |
| 2020 | A Mixed Appearance-based and Coding Distortion-based CNN Fusion Approach for In-loop Filtering in Video CodingabstractWith the success of the convolutional neural networks (CNNs) in image denoising and other computer vision tasks, CNNs have been investigated for in-loop filtering in video coding. Many existing methods directly use CNNs as powerful tools for filtering without much analysis on its effect. Considering the in-loop filters process the reconstructed video frames produced from a fixed line of video coding operations, the coding distortion in the reconstructed frames may share similar properties that can be learned by CNNs in addition to being a noisy image. Therefore, in this paper, we first categorize the CNN based filtering into two types of processes: appearance-based CNN filtering and coding distortion-based CNN filtering, and develop a two-stream CNN fusion framework accordingly. In the appearance-based CNN filtering, a CNN processes the reconstructed frame as a distorted image and extracts the global appearance information to restore the original image. In order to extract the global information, a CNN with pooling is used first to increase the receptive field and up-sampling is added in the late stage to produce pixel-level frame information. On the contrary, in the coding distortion-based filtering, a CNN processes the reconstructed frame as blocks with certain types of distortions by focusing on the local information to learn the coding distortion resulted by the fixed video coding pipeline. Finally, the appearance-based filtering stream and the coding distortion-based filtering stream are fused together to combine the two aspects of CNN filtering, and also the global and local information. To further reduce the complexity, the similar initial and last convolutional layers are shared over two streams to generate a mixed CNN. Experiments demonstrate that the proposed method achieves better performance than the existing CNN-based filtering methods, with 11.26% BD-rate saving under the All Intra configuration. Jian Yue, Yanbo Gao, Shuai Li 0005, Menghu Jia |
VCIP | 2 |
| 2020 | Leader-following bipartite consensus of second-order time-delay nonlinear multi-agent systems with event-triggered pinning control under signed digraph
Jie Ren 0002, Qiang Song 0001, Yanbo Gao, Guoping Lu |
Neurocomputing | 3 |
| 2019 | A fully trainable network with RNN-based pooling
Shuai Li 0005, Wanqing Li 0001, Chris Cook, Ce Zhu, Yanbo Gao |
Neurocomputing | 5 |
| 2019 | Exponential stability in Lagrange sense for inertial neural networks with time-varying delays
Shuang Lu, Yanbo Gao |
Neurocomputing | 2 |
| 2019 | Source Distortion Temporal Propagation Analysis for Random-Access Hierarchical Video Coding OptimizationabstractDue to the widely used inter prediction in the current video coding standards, encoding units in different frames is of temporal dependency in that the rate-distortion optimization (RDO) of one unit may affect the coding performance of the following units in the temporal domain. To achieve optimal coding solution for a given video sequence, temporal dependency among units needs to be considered in the RDO process, which is known as the temporally dependent RDO (TD-RDO). The hierarchical coding structure (HCS) employed in the High Efficiency Video Coding (HEVC) standard further complicates this problem by grouping frames into different layers of varying coding strategies, leading to a more complex temporal relationship. In our earlier work, we addressed TD-RDO for the low delay HCS (LD-HCS), where only uni-prediction is considered. This paper aims to address more complicated TD-RDO under random access HCS (RA-HCS), where both uni-prediction and bi-prediction are considered, making the temporal relationship even more intricate. The temporal dependency introduced in the RA-HCS is thoroughly examined and an RA-based TD-RDO scheme is formulated for each layer by modeling temporal propagation of distortion under different prediction types. Based on the formulation, the global Lagrange multiplier can be obtained analytically. Moreover, the effect of random access point pictures is considered in the RA-based TD-RDO scheme. The proposed method can be simply realized by updating the Lagrange multiplier as in the independent RDO formulation or combined with adjusting quantization parameter (QP) for better results in terms of BD-rate saving. Experimental results show that under RA-HCS, the proposed method, by adapting the Lagrange multiplier only, can achieve about 2.2% bitrate savings in average. With multi-QP optimization, an average BD-rate gain of 5.2% can be obtained. Yanbo Gao, Ce Zhu, Shuai Li 0005, Tianwu Yang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Independently Recurrent Neural Network (IndRNN): Building a Longer and Deeper RNNabstractRecurrent neural networks (RNNs) have been widely used for processing sequential data. However, RNNs are commonly difficult to train due to the well-known gradient vanishing and exploding problems and hard to learn long-term patterns. Long short-term memory (LSTM) and gated recurrent unit (GRU) were developed to address these problems, but the use of hyperbolic tangent and the sigmoid action functions results in gradient decay over layers. Consequently, construction of an efficiently trainable deep network is challenging. In addition, all the neurons in an RNN layer are entangled together and their behaviour is hard to interpret. To address these problems, a new type of RNN, referred to as independently recurrent neural network (IndRNN), is proposed in this paper, where neurons in the same layer are independent of each other and they are connected across layers. We have shown that an IndRNN can be easily regulated to prevent the gradient exploding and vanishing problems while allowing the network to learn long-term dependencies. Moreover, an IndRNN can work with non-saturated activation functions such as relu (rectified linear unit) and be still trained robustly. Multiple IndRNNs can be stacked to construct a network that is deeper than the existing RNNs. Experimental results have shown that the proposed IndRNN is able to process very long sequences (over 5000 time steps), can be used to construct very deep networks (21 layers used in the experiment) and still be trained robustly. Better performances have been achieved on various tasks by using IndRNNs compared with the traditional RNN and LSTM. Shuai Li 0005, Wanqing Li 0001, Chris Cook, Ce Zhu, Yanbo Gao |
CVPR | 5 |
| 2017 | A frame-level rate control scheme for low delay video coding in HEVCabstractR-λ rate control scheme is recommended in the High Efficiency Video Coding (HEVC) standard, which shows high accuracy of bit rate control but lower rate distortion performance. In order to minimize the distortion subject to a target bit rate, a λ domain frame-level rate control scheme for low delay coding of HEVC was proposed. Firstly, an improved parameter updating method is presented for the frame level rate distortion model, which full uses the information of encoded frames in the previous group of pictures (GOP). Then, an adaptive dynamic frame level bit allocation scheme is proposed by employing global rate distortion optimization theory. Finally, to further improve the coding efficiency, the bits allocation adjustment is made for the first several GOPs in video sequence according to the rate distortion dependency. The experimental results show that the proposed method can greatly improve the rate distortion performance under the condition of high accuracy of bit rate control. Hongwei Guo 0001, Ce Zhu, Yanbo Gao, Shichang Song |
MMSP | 3 |
| 2017 | Synchronization of coupled neural networks with time-varying delay
Yanbo Gao |
Neurocomputing | 2 |
| 2017 | Temporally Dependent Rate-Distortion Optimization for Low-Delay Hierarchical Video CodingabstractLow-delay hierarchical coding structure (LD-HCS), as one of the most important components in the latest High Efficiency Video Coding (HEVC) standard, greatly improves coding performance. It groups consecutive P/B frames into different layers and encodes them with different quantization parameters (QPs) and reference mechanisms in such a way that temporal dependency among frames can be exploited. However, due to varying characteristics of video contents, temporal dependency among coding units differs significantly from each other in the same or different layers, while a fixed LD-HCS scheme cannot take full advantage of the dependency, leading to a substantial loss in coding performance. This paper addresses the temporally dependent rate distortion optimization (RDO) problem by attempting to exploit varying temporal dependency of different units. First, the temporal relationship of different frames under the LD-HCS is examined, and hierarchical temporal propagation chains are constructed to represent the temporal dependency among coding units in different frames. Then, a hierarchical temporally dependent RDO scheme is developed specifically for the LD-HCS based on a source distortion propagation model. Experimental results show that our proposed scheme can achieve 2.5% and 2.3% BD-rate gain in average compared with the HEVC codec under the same configuration of P and B frames, respectively, with a negligible increase in encoding time. Furthermore, coupled with QP adaption, our proposed method can achieve higher coding gains, e.g., with multi-QP optimization, about 5.4% and 5.0% BD-rate saving in average over the HEVC codec under the same setting of P and B frames, respectively. Yanbo Gao, Ce Zhu, Shuai Li 0005, Tianwu Yang |
IEEE Trans. Image Process. | 1 |
| 2016 | Hierarchical temporal dependent rate-distortion optimization for low-delay codingabstractHierarchical coding structure (HCS) is one of the most important components in High Efficiency Video Coding (HEVC) that improves the coding performance greatly, especially for Low-Delay (LD) coding. It groups frames into different layers and enc odes them with different quantization parameters (QP) and different reference mechanisms. Due to the extensively used inter-prediction, the coding of frames in different layers is highly dependent and an appropriate QP and reference selection scheme may significantly improve the performance by taking advantage of such temporal dependency. However in the current HEVC codec, a predefined HCS, such as the Low-Delay HCS (LD-HCS), is performed without considering the different characteristic of different video contents, thus leading to a suboptimal coding solution. In this paper, the hierarchical temporal relationship under LD-HCS is first investigated and a hierarchical temporal propagation chain is constructed to describe the temporal dependency among frames. Then a hierarchical temporal dependent rate-distortion optimization scheme is developed specifically for the LD-HCS in HEVC. Experiments results show that the proposed scheme achieves BD-rate saving of 2.9% and 2.8% in average against HEVC codec under LD-HCS of P and B frames, respectively, with a negligible increase in encoding time. Yanbo Gao, Ce Zhu, Shuai Li 0005 |
ISCAS | 1 |
| 2016 | Layer-based temporal dependent rate-distortion optimization in Random-Access hierarchical video codingabstractRate-distortion optimization (RDO) plays an important part in improving the coding efficiency of High Efficiency Video Coding (HEVC), especially for the hierarchical coding structure defined in the Random-Access (RA) configuration, noted as Random-Access Hierarchical Video Coding (RA-HVC), where different frames are assigned to different temporal layers and further coded with different coding parameters. Due to the inter-frame prediction, coding result of one unit may affect the coding performance of the following temporally related units. Therefore, the temporal dependency among units needs to be considered in the coding process. However, the RDO process in the current video codec is performed without considering the varying temporal dependency, thus compromising the rate-distortion performance significantly. To address this problem, a layer-based temporal dependent RDO method is proposed in this paper where the temporal dependency among different frames in the same or different layers is examined. By reformulating the temporal dependent RDO for the RA-HVC, we show that it can be implemented in a way of simply refining the Lagrange multiplier. Experimental results show that the proposed method achieves, in average, about 1.4% BD-rate savings with a negligible increase in encoding time for the random-access configuration. Yanbo Gao, Ce Zhu, Shuai Li 0005, Tianwu Yang |
MMSP | 1 |
| 2016 | Lagrangian Multiplier Adaptation for Rate-Distortion Optimization With Inter-Frame DependencyabstractRate-distortion optimization (RDO) is widely used in video coding, which plays a critical role in enhancing the coding efficiency substantially. Currently, the RDO process is performed in a way that coding efficiency of each coding unit (CU) is maximized independently without considering the dependency among CUs. As we know, in the current hybrid video coding structure, spatial/temporal prediction techniques are extensively used, which introduce strong dependency among CUs. In this paper, we investigate RDO with inter-frame dependency, where the impact of coding performance of the current CU on that of the following frames is considered. Accordingly, an RDO scheme taking the inter-frame dependency into account is proposed by adapting the Lagrangian multiplier. The experimental results show that the proposed scheme can achieve about 3.22% and 3.19% BD-rate saving in average over the state-of-the-art High Efficiency Video Coding (HEVC) reference software HM15.0 in the low-delay $P$ (LDP) and low-delay $B$ (LDB) coding structures, respectively, with no extra encoding time. The proposed scheme can obtain a significantly higher coding gain than the multiple quantization parameter (MQP) (±3) optimization technique that would greatly increase the encoding time by a factor of about six. Coupled with MQP optimization, the proposed scheme can further achieve about 5.96% and 5.57% BD-rate savings in average over the HEVC and about 4.03% and 4.07% over the HEVC with MQP optimization, under the specified common test conditions for LDP and LDB coding structures, respectively. Shuai Li 0005, Ce Zhu, Yanbo Gao, Yimin Zhou 0002, Frédéric Dufaux, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Inter-frame dependent rate-distortion optimization using lagrangian multiplier adaptionabstractIt is known that, in the current hybrid video coding structure, spatial and temporal prediction techniques are extensively used which introduce strong dependency among coding units. Such dependency poses a great challenge to perform a global rate-distortion optimization (RDO) when encoding a video sequence. RDO is usually performed in a way that coding efficiency of each coding unit is optimized independently without considering dependeny among coding units, leading to a suboptimal coding result for the whole sequence. In this paper, we investigate the inter-frame dependent RDO, where the impact of coding performance of the current coding unit on that of the following frames is considered. Accordingly, an inter-frame dependent rate-distortion optimization scheme is proposed and implemented on the newest video coding standard High Efficiency Video Coding (HEVC) platform. Experimental results show that the proposed scheme can achieve about 3.19% BD-rate saving in average over the state-of-the-art HEVC codec (HM15.0) in the low-delay B coding structure, with no extra encoding time. It obtains a significantly higher coding gain than the multiple QP (±3) optimization technique which would greatly increase the encoding time by a factor of about 6. Coupled with the multiple QP optimization, the proposed scheme can further achieve a higher BD-rate saving of 5.57% and 4.07% in average than the HEVC codec and the multiple QP optimization enabled HEVC codec, respectively. Shuai Li 0005, Ce Zhu, Yanbo Gao, Yimin Zhou 0001, Frédéric Dufaux, Ming-Ting Sun |
ICME | 3 |
| 2015 | A multi-variable grey model with a self-memory component and its application on engineering prediction
Sifeng Liu, Lifeng Wu 0001, Yanbo Gao, Yingjie Yang |
Eng. Appl. Artif. Intell. | 4 |
| 2013 | Dissipative synchronization of nonlinear chaotic systems under information constraints
Yanbo Gao, Guoping Lu |
Inf. Sci. | 1 |
| 2007 | Automatic Construction of a Lexical Attribute Knowledge Base
Jinglei Zhao, Yanbo Gao, Hui Liu 0002, Ruzhan Lu |
KSEM | 2 |