EDBT 2026 Demo / reviewers in the wild / expert
Junlong Gao
dblp:172/2413
· DBLP profile ↗
3ranked-venue papers in the field
1as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 3 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prompt-Optimization with Contextual Mining for Cross-Modal Image CompressionabstractRecent advances in cross-modal compression(CMC) have opened new horizons for perceptual image coding at ultra-low bitrates (below 0.1 bpp) within a generative compression paradigm, but reconstruction fidelity is often compromised, yielding visually plausible yet semantically inconsistent reconstructions. While prompt engineering with contextual optimization has been extensively explored in generative models, its potential for controlling perception-fidelity trade-offs in image compression remains largely under-explored. To address these challenges, we propose PO-CMC, a novel diffusion-based cross-modal image compression approach that introduces contextual prompt optimization to achieve efficient and perceptually faithful reconstruction. The proposed method comprises three synergistic components: an optimized image codec that produces a compact structural prior, a contextual prompt module that adaptively encodes semantic cues into compact textual embeddings, and a diffusion-based decoder that fuses the structural and semantic priors to reconstruct high-fidelity images. Extensive experiments show that PO-CMC achieves superior perceptual quality while maintaining comparable reconstruction fidelity, yielding an average BD-rate saving of 72.5 % and 79.8 % over VVC at equivalent LPIPS and DISTS levels, respectively. Shenpeng Song, Zhimeng Huang, Junlong Gao, Chuanmin Jia, Siwei Ma 0001 |
DCC | 3 |
| 2025 | Image Coding for Machine with Visual-Language Mimic Feature LearningabstractThis paper propose a Image Coding for Machine (ICM) framework with Visual-Language Mimic Feature Learning (VLM-ICM). VLM-ICM decouples the position and semantic information into language modality and extracts universal features from the input image. Language, inherently more semantically compact, helps reduce the bitrate. Meanwhile, the universal features in VLM-ICM, guided by the language at the decoder side, allow for flexible domain adaptation, thereby enhancing versatility and practicality. Zhimeng Huang, Junlong Gao, Jiaqi Zhang 0007, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001, Chuanmin Jia |
DCC | 2 |
| 2023 | Rate-Distortion Optimization for Cross Modal CompressionabstractRecently, cross modal compression (CMC) is proposed to compress highly redundant visual data into a compact, common, human-comprehensible domain (such as text) to preserve semantic fidelity for semantic-related applications. However, CMC only achieves a certain level of semantic fidelity at a constant rate, and the model aims to optimize the probability of the ground truth text but not directly semantic fidelity. To tackle the problems, we propose a novel scheme named rate-distortion optimized CMC (RDO-CMC). Specifically, we model the text generation process as a Markov decision process and propose rate-distortion reward which is used in reinforcement learning to optimize text generation. In rate-distortion reward, the distortion measures both the semantic fidelity and naturalness of the encoded text. The rate for the text is estimated by the sum of the amount of information of all the tokens in the text since the amount of information of each token is a lower bound of coding bits. Experimentally, RDO-CMC effectively controls the rate in the CMC framework and achieves competitive performance on MSCOCO dataset. Junlong Gao, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
DCC | 1 |