VLDB 2026 Research / reviewers in the wild / expert
Jiawen Gu
dblp:192/8587
· DBLP profile ↗
19ranked-venue papers
7as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MG-VLQA: Multi-Granularity Quality Assessment for Image Compression via Visual Language ModelsabstractDespite significant advances in image compression, existing evaluation metrics remain poorly aligned with human visual perception-particularly under extremely low bitrates, where reconstructed images often suffer from abstract distortions or semantic degradation that are difficult for conventional metrics to capture. To address this limitation, we propose MG-VLQA, a novel multi-granularity quality assessment framework that leverages VisionLanguage Models (VLMs) to evaluate image reconstruction fidelity through the lens of semantic consistency with the original caption. Our method formulates a suite of captionderived questions spanning three complementary dimensions: (1) entity presence (semantic completeness), (2) detail fidelity (local appearance accuracy), and (3) inter-entity interactions (relational coherence). By simulating human-like perceptual judgment via VLMbased question answering and semantic similarity scoring, MG-VLQA provides a more interpretable, fine-grained, and perceptually relevant assessment of compression quality. Extensive experiments across multiple datasets and codecs demonstrate that our metric achieves higher correlation with human judgment and offers superior discriminative power. Hanfei Li, Anle Ke, Jiawen Gu, Tong Chen 0004, Zhan Ma 0001 |
DCC | 3 |
| 2026 | Deep Network-Based Adaptive Quantization for Practical Video CodingabstractThe optimization of block-level quantization parameters (QP) is critical to improving the performance of practical block-based video compression encoders, but the extremely large optimization space makes it challenging to solve. Existing solutions, e.g. HEVC encoder x265, usually add some optimization constraints of the block-independent assumption and linear distortion propagation model, which limits compression efficiency improvement to a certain extent. To address this problem, a deep learning-based encoder-only adaptive quantization method (DAQ) is proposed in this paper, where a deep network is designed to adaptively model the joint temporal propagation relationship of quantization among blocks. Specifically, DAQ consists of two phases: in the training phase, considering the heavy searching cost of the traditional codec, we introduce a well-designed end-to-end learned block-based video compression network as an effective training proxy tool for the deep encoder-side network. While in the deployment phase, the trained deep network is applied to jointly predict all block QPs in a frame for the traditional encoder. Besides, our network deploys only on the encoder side without changing the standard decoder and has very low inference complexity, making it able to apply in practice. At last, we deploy DAQ in HEVC and VVC encoder for performance comparison, and the experimental results demonstrate that DAQ significantly outperforms practically used x265 with on average 15.0%, 10.9% BD-rate reduction under the SSIM and PSNR, and also achieves 12.5%, 5.0% coding gain than VTM. Moreover, for deploying deep video codec in practice, this work provides a new insight for optimizing the encoder parameters with a large space. Hewei Liu, Jiawen Gu, Dengchao Jin, Meng Lei, Chao Zhou 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Deep Adaptive Quantization for Practical Video CompressionabstractIn this work, we propose a deep learning-based adaptive quantization method to promote video coding performance. Due to inter-prediction and reference mechanism, the block-level quantization parameter (QP) not only influences current block distortion but also has complex temporal propagation effects on subsequent coding frames. Our idea is to utilize a deep network to model the complex temporal propagation relationship of quantization. As shown in Fig. 1, the deep network directly predicts all block-level QPs of the frame for the traditional encoder without changing the standard decoder. Since our network deploys only on the encoder side and has low inference complexity, it can be easily applied in practice. In addition, we use a learned coding network as a proxy of the traditional codec to train our network. Hewei Liu, Jiawen Gu, Dengchao Jin, Meng Lei, Chao Zhou 0003 |
DCC | 3 |
| 2025 | GPPT: Gaussian Process-infused Prompt Tuning for Vision-language ModelsabstractPre-trained vision-language models (VLMs) have achieved remarkable success in image classification tasks, leveraging efficient prompt tuning methods. However, the reliability of fine-tuned VLMs in safety-critical scenarios remains a concern due to the under-explored issue of confidence calibration. To address this limitation, we introduce Gaussian Process-infused Prompt Tuning (GPPT), a novel framework that integrates a Gaussian process into the hidden representations of images. By utilizing random Fourier features (RFF) and Laplace approximation, GPPT enables end-to-end training and seamless integration with existing prompt tuning methods. Our extensive experiments on 11 diverse downstream datasets demonstrate that GPPT achieves competitive performance with deep ensembles in both prediction accuracy and calibration, while requiring only a fraction of the inference time. Shijing Si, Haixia Sun 0001, Jiawen Gu |
ICASSP | 3 |
| 2025 | Ultra Lowrate Image Compression with Semantic Residual Coding and Compression-aware DiffusionabstractExisting multimodal large model-based image compression frameworks often rely on a fragmented integration of semantic retrieval, latent compression, and generative models, resulting in suboptimal performance in both reconstruction fidelity and coding efficiency. To address these challenges, we propose a residual-guided ultra lowrate image compression named ResULIC, which incorporates residual signals into both semantic retrieval and the diffusion-based generation process. Specifically, we introduce Semantic Residual Coding (SRC) to capture the semantic disparity between the original image and its compressed latent representation. A perceptual fidelity optimizer is further applied for superior reconstruction quality. Additionally, we present the Compression-aware Diffusion Model (CDM), which establishes an optimal alignment between bitrates and diffusion time steps, improving compression-reconstruction synergy. Extensive experiments demonstrate the effectiveness of ResULIC, achieving superior objective and subjective performance compared to state-of-the-art diffusion-based methods with -80.7%, -66.3% BD-rate saving in terms of LPIPS and FID. Anle Ke, Xu Zhang 0027, Tong Chen 0004, Ming Lu 0003, Jiawen Gu, Zhan Ma 0001 |
ICML | 6 |
| 2025 | Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time ControlabstractReinforcement learning (RL) has achieved significant success across a wide range of domains, however, most existing methods are formulated in discrete time. In this work, we introduce a novel RL method for continuous-time control, where stochastic differential equations govern state-action dynamics. Departing from traditional value function-based approaches, our key contribution is the characterization of continuous-time Q-functions via a martingale condition and the linking of diffusion policy scores to the action gradient of a learned continuous Q-function by the dynamic programming principle. This insight motivates Continuous Q-Score Matching (CQSM), a score-based policy improvement algorithm. Notably, our method addresses a long-standing challenge in continuous-time RL: preserving the action-evaluation capability of Q-functions without relying on time discretization. We further provide theoretical closed-form solutions for linear-quadratic (LQ) control problems within our framework. Numerical results in simulated environments demonstrate the effectiveness of our proposed method and compare it to popular baselines. Chengxiu Hua, Jiawen Gu, Yushun Tang |
NeurIPS | 2 |
| 2025 | Neural B-frame Video Compression with Bi-directional Reference HarmonizationabstractNeural video compression (NVC) has made significant progress in recent years, while neural B-frame video compression (NBVC) remains underexplored compared to P-frame compression. NBVC can adopt bi-directional reference frames for better compression performance. However, NBVC's hierarchical coding may complicate continuous temporal prediction, especially at some hierarchical levels with a large frame span, which could cause the contribution of the two reference frames to be unbalanced. To optimize reference information utilization, we propose a novel NBVC method, termed Bi-directional Reference Harmonization Video Compression (BRHVC), with the proposed Bi-directional Motion Converge (BMC) and Bi-directional Contextual Fusion (BCF). BMC converges multiple optical flows in motion compression, leading to more accurate motion compensation on a larger scale. Then BCF explicitly models the weights of reference contexts under the guidance of motion compensation accuracy. With more efficient motions and contexts, BRHVC can effectively harmonize bi-directional references. Experimental results indicate that our BRHVC outperforms previous state-of-the-art NVC methods, even surpassing the traditional coding, VTM-RA (under random access configuration), on the HEVC datasets. The source code will be released. The source code is released at https://github.com/kwai/NVC. Dengchao Jin, Jiawen Gu, Ming Lu 0003, Zhan Ma 0001 |
NeurIPS | 4 |
| 2023 | Weighted Multivariate Mean Reversion for Online Portfolio Selection
Boqian Wu, Benmeng Lyu, Jiawen Gu |
ECML/PKDD (5) | 3 |
| 2019 | Mid-depth Based Block Structure Determination for AV1abstractAV1 is an emerging open-source and royalty-free video compression format as a successor to VP9.The increase in coding efficiency and complexity over VP9 is due to the time required to find the optimal partition structure among the more flexible encoding modes for the coding units (CUs) and prediction units (PUs). Due to differences in the frame structure, existing fast block structure determination algorithm cannot be directly applied to AV1. To tackle this problem, we proposed a novel mid-depth based fast block structure determination algorithm for AV1. It checks the partition from mid-depth to provide information for estimating the posterior probabilistic distribution of the partition decisions as well as fast pruning in the PU prediction. Experimental results show that the proposed method can save up to 29.06% time saving with only 0.95% BD-Rate increase. Jiawen Gu, Jiangtao Wen |
ICASSP | 1 |
| 2019 | High Efficiency Light Field Compression via Virtual Reference and Hierarchical MV-HEVCabstractEfficient storage and delivery of the light field (LF) information rely on high performance compression. In this paper, we propose a high efficiency light field compression algorithm that utilizes a hierarchical coding structure with synthetic virtual references. Specifically, a LF image are interpreted as a multi-view sequence that is efficiently compressed using the multi-view extension of high efficiency video coding (MV-HEVC). Using deep neural networks, we synthesize virtual references from reconstructed neighbor frames, they serve as extra reference candidates in our novel hierarchical coding structure. Compared with previous work, the proposed algorithm further exploits the intrinsic similarities in LF images. Experimental results show that the proposed algorithms demonstrate a superior performance that achieves up to 55.2% BD-rate reduction and 2.55dB BD-PSNR improvement compared with the HEVC benchmark and outperforms the state-of-the-art. Jiawen Gu, Bichuan Guo, Jiangtao Wen |
ICME | 1 |
| 2019 | A Universal Optical Flow Based Real-Time Low-Latency Omnidirectional Stereo Video SystemabstractOmnidirectional stereoscopic video (ODSV) is a key element of creating an immersive experience for virtual reality that has attracted extensive interest while presenting many technical challenges. Two such key challenges are real-time, low-latency high-quality seamless video stitching from multiple cameras, and faithful reconstruction of 3-D information. Even though various attempts have been made to achieve different combinations of real-time, low-latency, automation, and high output resolution in stereoscopic panoramic video communication, achieving these characteristics simultaneously remains a challenge to be tackled. In this paper, we present a universally applicable and practical end-to-end system based on a novel real-time optical flow algorithm to produce high-quality real-time ODSV with reconstructed 3-D depth information at low latency. Through a configurable process, various camera systems can be calibrated and seamlessly stitched together using the proposed system. The stitched 3-D panoramic video is encoded with a standard compliant video encoder that is optimized for panoramic video. Thanks to various optimizations introduced in this paper, the proposed system is capable of producing real-time ODSV of ultra High definition resolution with a glass to glass latency of 2.2 s using a desktop computer with a single Nvidia graphic card. Experiments show that the proposed system achieves an encoding performance superior to existing open-source HEVC implementations and an optical flow estimation performance better than the Facebook algorithm while running two orders of magnitudes faster. Minhao Tang, Jiangtao Wen, Jiawen Gu, Philip Junker, Bichuan Guo, Guansyun Jhao, Yuxing Han 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | A Bayesian Approach to Block Structure Inference in AV1-Based Multi-Rate Video EncodingabstractDue to differences in frame structure, existing multi-rate video encoding algorithms cannot be directly adapted to encoders utilizing special reference frames such as AV1 without introducing substantial rate-distortion loss. To tackle this problem, we propose a novel bayesian block structure inference model inspired by a modification to an HEVC-based algorithm. It estimates the posterior probabilistic distributions of block partitioning, and adapts early terminations in the RDO procedure accordingly. Experimental results show that the proposed method provides flexibility for controlling the tradeoff between speed and coding efficiency, and can achieve an average time saving of 36.1% (up to 50.6%) with negligible bitrate cost. Bichuan Guo, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen |
DCC | 3 |
| 2018 | Multi-Representations Encoding Framework for Adaptive Http StreamingabstractAdaptive HTTP streaming requires a video to be encoded at multiple representations of different target bitrates. To achieve both smooth streaming and good quality, the representations need to be encoded with accurate rate control and the best quality possible for the target bitrates. However, in practical applications, accurate rate control and good video quality are very difficult to achieve at the same time with one-pass and real-time encoding required by low latency applications. In this paper, we proposed a multi-representation encoding framework that reused the encoding information from low-quality representations to accelerate and optimize higher bitrate encodings. The proposed framework is implemented on two different rate control models, namely R-λ model in HM-16.3 and complexity model in x264, to demonstrate the universality. To be best of our knowledge, the optimization in compression performance of multi - representation is first proposed in this paper. The proposed algorithm can achieve better video quality with only small latency in video coding. Results show that up to 49.6% BDRate savings and 4.43dB BDPSNR improvement are achieved in HM as compared with independent one-pass encodings. Meanwhile, 12.5% time and 14.65% BDRate savings were observed in x264. Jiawen Gu, Jiangtao Wen, Bichuan Guo, Yuxing Han 0001 |
ICIP | 1 |
| 2018 | Adaptive Intra Candidate Selection With Early Depth Decision for Fast Intra Prediction in HEVCabstractTo better exploit spatial correlations in a video frame, the High Efficiency Video Coding (HEVC) standard has adopted a great many more intra prediction modes than H.264/AVC. As a result, the complexity of rate-distortion-optimized (RDO) HEVC intra mode selection is very high. Many techniques have been proposed to expedite the intra mode selection process to achieve a better overall tradeoff between complexity and RD performance. In this paper, two novel techniques for adaptive intra mode candidate selection and bidirectional depth search algorithms are utilized to accelerate intra prediction. The proposed techniques show an average of 63% (up to 67%) time saving with only 1% BD-Rate increase, outperforming most of existing intra prediction algorithms. Jiawen Gu, Minhao Tang, Jiangtao Wen, Yuxing Han 0001 |
IEEE Signal Process. Lett. | 1 |
| 2018 | Accelerating HEVC Encoding Using Early-SplitabstractThe increase in coding efficiency and complexity of high efficiency video coding (HEVC) over H.264 is due to, among other factors, the time needed to find the optimal partition structure among the more flexible encoding modes for the coding units (CUs) and prediction units (PUs). Although many classification-based algorithms have been proposed to expedite the partition decision, the features that can be acquired from current HEVC encoding order are not sufficient to minimize the loss in coding efficiency. In this letter, we proposed an early-split (ES) order for HEVC CU-level encoding, where the encoder checks the split mode before the nonsquare PU partition modes and utilizes the encoding output of the subCUs to expedite subsequent encoding. Experiments show that the proposed algorithm can save 48% of encoding time on average with only about 0.8% loss in coding performance. Minhao Tang, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen, Shiqiang Yang |
IEEE Signal Process. Lett. | 3 |
| 2017 | SATD Based Fast Intra Prediction for HEVCabstractSummary form only given. To better exploit spatial correlations in a video frame, the HEVC video coding standard has introduced many intra prediction modes and a recursive quadtree-based coding unit (CU) structure. As a result, the complexity of Rate-distortion optimized (RDO) HEVC intra mode selection is significantly higher. Many techniques have been proposed to expedite the intra mode selection process to achieve a good overall trade-off between complexity and RD performance. In this paper, we proposed a fast intra decision algorithm based on Hadamard Transform. The algorithm consist of three parts: SATD calculation reduction, adaptive intra candidate selection, and SATD based early termination. Experiments conducted using the HEVC common test conditions show an average of 56.4% (up to 64.1%) time saving with only 1.2% increase in Bjontegaard delta rate (BD-rate) using the proposed algorithm. Jiawen Gu, Minhao Tang, Jiangtao Wen |
DCC | 1 |
| 2017 | Early-Split Based Fast HEVC EncodingabstractThe High Efficiency Video Coding (HEVC) standard achieves 50% improvement incompression efficiency over the widely used H.264/AVC standard at a cost of much higher complexity. The increase in complexity is due to, among other factors, the time needed to findthe optimal partition structure among the more flexible possibilities for the coding units (CUs) and prediction units (PUs). Many classification based algorithms have been proposed to reduce this partition decision time, but the features that can be acquired from current HEVC encoding order may not be sufficient to control the loss in coding efficiency. In this paper, we proposed an Early-Split (ES) order for HEVC encoding, where the encoder checks the split mode before the non-square PU partition modes and utilizes the encoding output of the subCUs to expedite subsequent encoding. Experiments show that the proposed algorithm achieved an average of 48% saving in encoding time with only 0.92% loss in the coding performance. Minhao Tang, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen |
DCC | 3 |
| 2017 | A novel satd based fast intra prediction for HEVCabstractTo better exploit spatial correlations in a video frame, the HEVC video coding standard has introduced many intra prediction modes and a recursive quadtree-based coding unit (CU) structure. As a result, the complexity of Rate-distortion optimized (RDO) HEVC intra mode selection is significantly higher. Many techniques have been proposed to expedite the intra mode selection process to achieve a good overall trade-off between complexity and RD performance. In this paper, we proposed a fast intra decision algorithm consisting of three parts: calculation reduction, adaptive intra candidate selection, and fast depth decision used early termination. Experiments conducted using the HEVC common test conditions show an average of 61.1% (up to 67.7%) time saving with only 1.03% increase in Bjontegaard delta rate (BD-rate) using the proposed algorithm, out-performing existing state-of-the-art algorithms. Jiawen Gu, Minhao Tang, Jiangtao Wen |
ICIP | 1 |
| 2016 | A novel low delay in-loop filtering WPP process for parallel HEVC encodingabstractWavefront parallel processing (WPP) is a parallelization technique that enables processing of several rows of Largest Coding Units (LCUs) in parallel. It achieves a relatively high level of parallelism with reasonable loss in compression performance. At the same time, the HEVC standard specifies two in-loop filters, namely the deblocking filter and the sample adaptive offset (SAO), to improve subjective quality as well as coding efficiency. Because SAO parameters cannot be precisely determined until the lower right deblocked samples are available, a delay of several rows is introduced to implement the filtering process. Since the reconstruction of the LCUs will not start until finish filtering, the total delay of WPP and filtering is at least two rows, which obviously influences the performance of parallel processing. In this paper, we propose a novel In-Loop Filtering WPP method that reduces the row delay into four LCUs and significantly improves the parallelism with little rate-distortion (RD) performance loss. Experimental results show that the proposed algorithms can achieve up to the 2.89× speedup compared with the existing WPP method with 24-core server, where the speedup improves with the increase of core number. Jiawen Gu, Yuxing Han 0001, Jiangtao Wen |
VCIP | 1 |