VLDB 2026 Research / reviewers in the wild / expert
Gaoxing Chen
dblp:121/6897
· DBLP profile ↗
10ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0005-5801-0812ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Perceptually Driven Spatial-Temporal Adaptive Quantization AlgorithmabstractIn this paper, we derive a theoretical model for the temporal propagation of perceptual distortion and propose a novel adaptive Quantization Parameter (QP) adjustment strategy to enhance the video coding performance. Temporal adaptive quantization (AQ) models and exploits the complex temporal dependencies introduced by inter-frame prediction, with the objective of Peak Signal-to-Noise Ratio (PSNR) optimization. In contrast to temporal AQ, spatial AQ primarily focuses on intra-frame characteristics and targets Structural Similarity (SSIM) optimization. Motivated by these complementary properties, we propose an algorithm that integrates both aspects. Yadong Shao, Gaoxing Chen, Fuzheng Yang 0001 |
DCC | 3 |
| 2026 | Text and Non-Text Latent Feature Disentanglement for Screen Content Image CompressionabstractWith the growing prevalence of screen content images in multimedia communication, efficient compression has become increasingly crucial. Unlike natural scene images, screen content typically contains rich text regions that exhibit unique characteristics and low correlation with surrounding non-text elements. The intricate mixture of text and non-text within images poses significant challenges for existing learned compression networks, as the text and non-text features are severely entangled in the latent domain along the channel dimension, leading to compromised reconstruction quality and suboptimal entropy estimation. In this paper, we propose a novel Disentangled Image Compression Architecture (DICA) that enhances the analysis module and the entropy model of existing compression architectures to address these limitations. First, we introduce a Disentangled Analysis Module (DAM) by augmenting original analysis modules with an additional text approximation branch and a disentangling network. They work in concert to disentangle latent features into text and non-text classes along the channel dimension, resulting in a more structured feature distribution that better aligns with compression requirements. Second, we propose a Disentangled Channel-Conditional Entropy Model (DCEM) that efficiently leverages the feature distribution bias introduced by DAM, thereby further improving compression performance. Experimental results demonstrate that the proposed DICA, along with DAM and DCEM can be integrated into various channel-conditional compression backbones, significantly improving their performance in screen content compression—particularly in hard-to-compress text regions. When integrated with an advanced WACNN backbone, our method achieves a 13% overall BD-Rate gain and a 16% BD-Rate gain in text regions on the SIQAD dataset. Hao Wang 0184, Junyan Huo, Fei Yang 0004, Shuai Wan, Gaoxing Chen, Luis Herranz, Fuzheng Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Customizing Image Codecs for Text-Rich Screen Content with Plugin Processing NetworksabstractWith the rapid growth of remote education, telemedicine, and cloud gaming, screen content images have become prevalent in these applications. They differ significantly from natural scene images, making learning-based image codecs optimized with natural scenes inefficient when compressing them. Through empirical analysis, we observe the textual region in screen content is not only hard to compress in itself but also impacts the compression efficiency of the non-textual region. To customize the image codecs to screen content without altering their parameters, we introduced plugin pre- and post-processing modules. Specifically, we designed a filtering network in the pre-processing module to remove compression-unfriendly information from textual regions and a restoration network in the post-processing module to recover it. Additionally, we implemented a multi-scale fuse approach to enhance the high-frequency details in images. Experiments on public datasets demonstrated that our plugin solution can be seamlessly integrated into learning-based image codecs, significantly improving compression performance. Hao Wang 0184, Junyan Huo, Shuai Wan, Gaoxing Chen, Fuzheng Yang 0001 |
ICME | 5 |
| 2024 | Proposal With Alignment: A Bi-Directional Transformer for 360° Video Viewport ProposalabstractPeople normally watch 360 ° videos through a head-mounted display, inside which only the content of viewports can be seen. Therefore, viewport proposal, referring to detecting potential viewport candidates, plays an important role in many 360 ° video processing tasks. In this paper, we advance the viewport proposal by further aligning the predicted viewports across frames for individual subject. This provides a better methodology and a deeper perspective to learn the human perceptual behaviours on 360 ° videos. Specifically, we first analyze three 360 ° video datasets and obtain several findings on human consistency, objectness and motion of viewports. Inspired by these findings, we propose a bi-directional transformer approach, named BiT, for 360 ° video viewport proposal and alignment. Specifically, BiT is composed of a multi-level residual module, a bi-directional encoder-decoder module and a spherical matching module. This way, the viewports can be well proposed and aligned via considering multi-level, bi-directional and non-local information. Moreover, the aligned viewports by BiT are used to refine the viewports and improve viewport proposal accuracy in return. Finally, we validate that our BiT approach is superior on viewport proposal, compared with the state-of-the-art approaches. Besides, the aligned viewports from BiT is verified to be effective in multiple applications, such as saliency prediction, trajectory prediction and perceptual video compression. Mai Xu, Lai Jiang 0004, Xin Deng 0002, Gaoxing Chen, Leonid Sigal |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Enhancing Video Encoding for Cloud Virtual Reality Gaming Based on User TypesabstractCloud Virtual Reality (VR) gaming is a novel technology that allows users to enjoy complex games on their thin clients by offloading the graphics rendering to cloud servers. The thin clients only need to perform basic decoding functions, which reduces the hardware requirements and costs. However, cloud VR gaming also faces the challenge of high bandwidth consumption when transmitting high-resolution game video streams. This paper presents a cloud VR gaming system that can transmit users’ gaze point data to the server in real time to identify users’ regions of interest. With this system, we verify the difference in spatial visual sensitivity caused by the different types of users. Then, a user-type-based video encoding method is proposed. Through conducting the subjective test experiment, the proposed video encoding method can reduce the bitrate for players and viewers by at least 71% and 69%, respectively, without compromising the perceptual quality. Junyan Huo, Fuzheng Yang 0001, Gaoxing Chen |
VCIP | 6 |
| 2015 | Deblocking strength prediction based CTU-level SAO category determination in HEVC encoderabstractHigh efficiency video coding (HEVC) is a video compression standard that outperforms the predecessor H.264/AVC by doubling the compression efficiency. To enhance the coding accuracy, HEVC adopts sample adaptive offset (SAO), which reduces the distortion of reconstructed pixels using classification based non-linear filtering. In the traditional coding tree unit (CTU) based VLSI encoder implementation, during the pixel classification stage, SAO cannot use the raw samples in the boundary of the current CTU because these pixels have not been processed by deblocking filter (DF). This paper proposes a category determination algorithm based on estimating the deblocking strengths on CTU boundaries and selectively adopting the promising samples in these areas during SAO classification. Compared with HEVC test mode (HM11.0), experimental results indicate that the proposed method achieves an average 0.15% BD-bitrate reduction (equivalent to 0.0084 dB increases in P-SNR). Gaoxing Chen, Zhenyu Pei, Zhenyu Liu 0001, Takeshi Ikenaga |
VCIP | 1 |
| 2014 | A multiple scattering reflectance model for vegetation canopy based on recollision probabilityabstractThe physically-based vegetation canopy reflectance model is the basis for accurate inversion of important vegetation parameters. According the interactive process between photons and canopy, canopy reflectance could be divided into two parts: single scattering and multiple scattering reflectance. Recollision probability is a useful tool linking leaf optical properties to canopy reflectance or absorption. In this paper, based on the recollision-probability, a new and practical multiple scattering model was approached. To estimate the accuracy of this model, Monte-Carlo simulation and 3D radiosity model were used and the results showed high accuracy of this model. Then the difference between first recollision probability (p1) and multiple recollision probability(pm) was discussed. Both of p1and pmwere needed in modeling multiple scattering model. At last, the contributions of soil reflectance, leaf albedo and proportion of sky radiation to multiple scattering reflectance were discussed. Gaoxing Chen, Beitong Zhang, Wenjie Fan 0001, Xiru Xu, Yuan Liu 0016 |
IGARSS | 1 |
| 2013 | Monitoring of degrading grassland based on HJ-1A-HSI imageabstractGrassland is one of the most important parts of ecosystem on the earth. In China, one of the best grass lands is degrading because of draught or effects of human activity. It is important to monitor the growing condition of degrading grasslands. LAI is an important variable which can accurately represent the growing situation of grass. With DSD method, hyper-spectral images from HJ-1A satellite are used to inverse LAI accurately by restraining the effect of background. In this paper, multi spectral image was used to retrieve LAI with BRDF method for comparing. The maps of LAI and degradation level in study area were made. According to the results, the DSD method can retrieve LAI of grassland accurately and the HSI image of HJ-1A is a potential ideal data source for monitoring grassland. Gaoxing Chen, Wenjie Fan 0001, Xiru Xu, Mengzhi Deng |
IGARSS | 1 |
| 2013 | A new FAPAR retrieval model for continuous vegetationabstractThe Fraction of Absorbed Photosynthetically Active Radiation (FAPAR) is the fraction of incoming solar radiation that is absorbed by green vegetation in the spectral range from 400 nm to 700 nm. FAPAR reflects the energy absorption ability of vegetation canopy. It is a critical input in many land surface models, such as crop growth models, net primary productivity models, climate models and ecological models. Existing models for FAPAR retrieval are complex and difficult to retrieve, most of them cannot be used under cloudy weather. In this paper, a new quantitative FAPAR retrieval model considering the diffuse skylight and multiple scattering between canopy and background is introduced to retrieve FAPAR of vegetation canopy. The model was used to continuous vegetation and was validated by Monte Carlo (MC) simulation and field tests. The conclusion shows that the error is less than 0.32%. Yuan Liu 0016, Wenjie Fan 0001, Xiru Xu, Gaoxing Chen |
IGARSS | 4 |
| 2012 | Estimating clumping index of sparse forest using hemispherical photographs combined with Geoeye-1 dataabstractClumping index is a critical physical parameter used to describe the clumping effect of vegetation canopy. In many studies, foliage elements are assumed to distribute randomly in the canopy and the clumping index is 1. However, for the irregularly or artificially spaced discrete vegetation canopies, such as savanna and sparse forests, the assumption is not in accordance with the actual case, the clumping index varies in the range of 0 to 1. As a result, the Leaf Area Index (LAI) retrieved directly from remote sensing data is always underestimated. Optical instruments, such as LAI-2000 canopy analyzer, TRAC, fish-eye camera, are difficult to measure the clumping index for sparse forests directly. In this paper, taking populus euphratica sparse forest in Heihe Basin as the research object, a new method combining hemispherical photography and high resolution images is established to estimate the clumping index. The results show that the method can accurately calculate clumping index and LAI for sparse forests and improve the validation of LAI products in water stressed regions. Yuan Liu 0016, Yingying Gai, Gaoxing Chen, Wenjie Fan 0001, Xiru Xu, Binyan Yan, Yanran Liao |
IGARSS | 3 |