VLDB 2026 Research / reviewers in the wild / expert
Chia-Hao Kao
dblp:299/4433
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CSGaussian: Progressive Rate-Distortion Compression and Segmentation for 3D Gaussian SplattingabstractWe present the first unified framework for rate-distortion-optimized compression and segmentation of 3D Gaussian Splatting (3DGS). While 3DGS has proven effective for both real-time rendering and semantic scene understanding, prior works have largely treated these tasks independently, leaving their joint consideration unexplored. Inspired by recent advances in rate-distortion-optimized 3DGS compression, this work integrates semantic learning into the compression pipeline to support decoder-side applications–such as scene editing and manipulation–that extend beyond traditional scene reconstruction and view synthesis. Our scheme features a lightweight implicit neural representation-based hyperprior, enabling efficient entropy coding of both color and semantic attributes while avoiding costly grid-based hyperprior as seen in many prior works. To facilitate compression and segmentation, we further develop compression-guided segmentation learning, consisting of quantization-aware training to enhance feature separability and a quality-aware weighting mechanism to suppress unreliable Gaussian primitives. Extensive experiments on the LERF and 3D-OVS datasets demonstrate that our approach significantly reduces transmission cost while preserving high rendering quality and strong segmentation performance. Yu-Jen Tseng, Chia-Hao Kao, Jing-Zhong Chen, Alessandro Gnutti, Shao-Yuan Lo, Yen-Yu Lin, Wen-Hsiao Peng |
WACV | 2 |
| 2025 | Learning Optimal Linear Block Transform by Rate Distortion MinimizationabstractThe rise of deep learning has spurred advancements in image compression, with end-to-end learned systems gaining traction. However, their adoption in standard frameworks is limited, as they require a major overhaul of existing hardware designed for traditional methods. Moreover, their computational complexity, especially on the decoder side, remains significantly higher than conventional codecs. Consequently, optimizing traditional codecs remains a key research focus. Alessandro Gnutti, Chia-Hao Kao, Wen-Hsiao Peng, Riccardo Leonardi |
DCC | 2 |
| 2025 | Bridging Compressed Image Latents and Multimodal Large Language ModelsabstractThis paper presents the first-ever study of adapting compressed image latents to suit the needs of downstream vision tasks that adopt Multimodal Large Language Models (MLLMs). MLLMs have extended the success of large language models to modalities (e.g. images) beyond text, but their billion scale hinders deployment on resource-constrained end devices. While cloud-hosted MLLMs could be available, transmitting raw, uncompressed images captured by end devices to the cloud requires an efficient image compression system. To address this, we focus on emerging neural image compression and propose a novel framework with a lightweight transform-neck and a surrogate loss to adapt compressed image latents for MLLM-based vision tasks.
Given the huge scale of MLLMs, our framework excludes the entire downstream MLLM except part of its visual encoder from training our system. This stands out from most existing coding for machine approaches that involve downstream networks in training and thus could be impractical when the networks are MLLMs. The proposed framework is general in that it is applicable to various MLLMs, neural image codecs, and multiple application scenarios, where the neural image codec can be (1) pre-trained for human perception without updating, (2) fully updated for joint human and machine perception, or (3) fully updated for only machine perception.
Extensive experiments on different neural image codecs and various MLLMs show that our method achieves great rate-accuracy performance with much less complexity. Chia-Hao Kao, Cheng Chien, Yu-Jen Tseng, Yi-Hsin Chen, Alessandro Gnutti, Shao-Yuan Lo, Wen-Hsiao Peng, Riccardo Leonardi |
ICLR | 1 |
| 2025 | Rate Distortion Learned Transform For Image Compression
Alessandro Gnutti, Chia-Hao Kao, Wen-Hsiao Peng, Riccardo Leonardi |
PCS | 2 |
| 2025 | Learned Transcoding for Neural Image Compression
Chia-Hao Kao, Andrea Migliorati, Alessandro Gnutti, Enrico Magli, Riccardo Leonardi |
PCS | 1 |
| 2023 | TransTIC: Transferring Transformer-based Image Compression from Human Perception to Machine PerceptionabstractThis work aims for transferring a Transformer-based image compression codec from human perception to machine perception without fine-tuning the codec. We propose a transferable Transformer-based image compression framework, termed TransTIC. Inspired by visual prompt tuning, TransTIC adopts an instance-specific prompt generator to inject instance-specific prompts to the encoder and task-specific prompts to the decoder. Extensive experiments show that our proposed method is capable of transferring the base codec to various machine tasks and outperforms the competing methods significantly. To our best knowledge, this work is the first attempt to utilize prompting on the low-level image compression task. Yi-Hsin Chen, Ying-Chieh Weng, Chia-Hao Kao, Cheng Chien, Walon Wei-Chen Chiu, Wen-Hsiao Peng |
ICCV | 3 |
| 2023 | Transformer-Based Variable-Rate Image Compression with Region-of-Interest ControlabstractThis paper proposes a transformer-based learned image compression system. It is capable of achieving variable-rate compression with a single model while supporting the region-of-interest (ROI) functionality. Inspired by prompt tuning, we introduce prompt generation networks to condition the transformer-based autoencoder of compression. Our prompt generation networks generate content-adaptive tokens according to the input image, an ROI mask, and a rate parameter. The separation of the ROI mask and the rate parameter allows an intuitive way to achieve variable-rate and ROI coding simultaneously. Extensive experiments validate the effectiveness of our proposed method and confirm its superiority over the other competing methods. Chia-Hao Kao, Ying-Chieh Weng, Yi-Hsin Chen, Walon Wei-Chen Chiu, Wen-Hsiao Peng |
ICIP | 1 |
| 2022 | Neural Frank-Wolfe Policy Optimization for Region-of-Interest Intra-Frame Coding with HEVC/H.265abstractThis paper presents a reinforcement learning (RL) framework that utilizes Frank-Wolfe policy optimization to solve Coding- Tree-Unit (CTU) bit allocation for Region-of-Interest (ROI) intra-frame coding. Most previous RL-based methods employ the single-critic design, where the rewards for distortion minimization and rate regularization are weighted by an empirically chosen hyper-parameter. Recently, the dual-critic design is proposed to update the actor by alternating the rate and distortion critics. However, its convergence is not guaranteed. To address these issues, we introduce Neural Frank-Wolfe Policy Op-timization (NFWPO) in formulating the CTU-level bit allocation as an action-constrained RL problem. In this new framework, we exploit a rate critic to predict a feasible set of actions. With this feasible set, a distortion critic is invoked to update the actor to maximize the ROI-weighted image quality subject to a rate constraint. Experimental results produced with x265 confirm the superiority of the proposed method to the other baselines. Yung-Han Ho, Chia-Hao Kao, Wen-Hsiao Peng, Ping-Chun Hsieh |
VCIP | 2 |