Tong Chen 0004

dblp:22/1512-4 · DBLP profile ↗
← Back
2ranked-venue papers in the field
0as first author
2since 2021 · last 2026
0000-0001-5020-6099ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2026 MG-VLQA: Multi-Granularity Quality Assessment for Image Compression via Visual Language Models
abstract
Despite significant advances in image compression, existing evaluation metrics remain poorly aligned with human visual perception-particularly under extremely low bitrates, where reconstructed images often suffer from abstract distortions or semantic degradation that are difficult for conventional metrics to capture. To address this limitation, we propose MG-VLQA, a novel multi-granularity quality assessment framework that leverages VisionLanguage Models (VLMs) to evaluate image reconstruction fidelity through the lens of semantic consistency with the original caption. Our method formulates a suite of captionderived questions spanning three complementary dimensions: (1) entity presence (semantic completeness), (2) detail fidelity (local appearance accuracy), and (3) inter-entity interactions (relational coherence). By simulating human-like perceptual judgment via VLMbased question answering and semantic similarity scoring, MG-VLQA provides a more interpretable, fine-grained, and perceptually relevant assessment of compression quality. Extensive experiments across multiple datasets and codecs demonstrate that our metric achieves higher correlation with human judgment and offers superior discriminative power.
Hanfei Li, Anle Ke, Jiawen Gu, Tong Chen 0004, Zhan Ma 0001
DCC5
2024 Variable-rate Neural Speech Compression with Multi-scale Feature Extraction and Improved Entropy Modeling
abstract
Speech coding serves as a means of data compression, aiming to decrease the expenses related to data storage and transmission. The efficacy of compressing speech efficiently through neural networks has been demonstrated in methods using vector quantization (VQ). However, the complex procedure of VQ makes it challenging to fit into frameworks and limits compression at discrete bitrate points. This paper proposes a neural speech compression framework, which achieves flexible bitrate speech reconstruction through compact latent representation and better entropy estimation.
Shaohan Sun, Yuzhuo Kong, Tong Chen 0004, Zhan Ma 0001
DCC3