EDBT 2026 Demo / reviewers in the wild / expert
Jinming Liu 0001
dblp:135/4552-1
· DBLP profile ↗
20ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0002-8714-073XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Token Compression for the Understanding and Generation Unified MLLMs
Junyan Lin, Jinming Liu 0001, Shengyang Zhao, Xin Jin 0014 |
ISCAS | 2 |
| 2026 | Revisiting MLLM Token Technology through the Lens of Classical Visual CodingabstractClassical visual coding and Multimodal Large Language Model (MLLM) token technology share the core objective - maximizing information fidelity while minimizing computational cost. Therefore, this paper reexamines MLLM token technology, including tokenization, token compression, and token reasoning, through the established principles of long-developed visual coding area. From this perspective, we (1) establish a unified formulation bridging token technology and visual coding, enabling a systematic, module-by-module comparative analysis; (2) synthesize bidirectional insights, exploring how visual coding principles can enhance MLLM token techniques' efficiency and robustness, and conversely, how token technology paradigms can inform the design of next-generation semantic visual codecs; (3) prospect for promising future research directions and critical unsolved challenges. In summary, this study presents the first comprehensive and structured technology comparison of MLLM token and visual coding, paving the way for more efficient multimodal models and more powerful visual codecs simultaneously. Jinming Liu 0001, Junyan Lin, Yuntao Wei, Kele Shao, Keda Tao, Jianguo Huang, Zhibo Chen 0001, Huan Wang 0014, Xin Jin 0014 |
ISCAS | 1 |
| 2025 | Representation Disentanglement for Semantic CodingabstractThe learned image compression methods have achieved advances in both human perception and machine vision. However, previous methods focus on transmitting visual symbols losslessly instead of precisely conveying semantic meaning, resulting in bandwidth waste, especially for AI applications. In this work, we study "Semantic Coding" and propose a novel compression method based on representation disentanglement, which understands images at the attribute level and separates semantic factors into different parts, achieving a semantically structured bitstream for transmission. Specifically, we first leverage a conditional generative diffusion procedure for a disentangled representation learning, which learns meaningful semantic attribute factors in the latent space of the image assisted by the extra language inductive bias. Furthermore, we employ a learned codec to compress the inner disentangled representations as a bitstream, where each part represents a specific semantic and can be used for purposely decoding. Experiments show that our semantic coding method could reconstruct high-quality images and enable encryption by shifting the inner semantics. Jinming Liu 0001, Junhao Geng, Lexiang Lv, Wenjun Zeng 0001, Xin Jin 0014 |
ICME | 1 |
| 2025 | Multi-Attribute Continual Learning for Blind Image Quality AssessmentabstractBlind image quality assessment (BIQA) has evolved into a critical task in visual computing, requiring effective evaluation across multiple quality attributes such as brightness, sharpness, contrast, and colorfulness. Traditional BIQA methods based on single-task learning often suffer from catastrophic forgetting and struggle to generalize across diverse IQ attributes. To address these challenges, we propose a novel Multi-attribute Continual Learning framework, MaC-BIQA, which integrates Gated Attention Mechanism and Knowledge Graph Embedding (KGE) with the Learning without Forgetting (LwF) approach. Specifically, the Gated Attention Mechanism dynamically adjusts attention distribution by focusing on task-specific key regions, while the integration of Knowledge Graph Embedding (KGE) complements LwF by improving the understanding of inter-task relationships, ensuring that relevant information from previous tasks is preserved more effectively, even as the model learns new tasks. Together, they enhance the model’s adaptability and robustness in handling complex multi-attribute tasks, effectively mitigating catastrophic forgetting and significantly improving overall performance in multi-task environments. Extensive experiments on the KonIQ-10K and SPAQ datasets show that our method significantly reduces forgetting rates and improves robustness and generalization across multiple quality attributes. This study presents a key technical contribution by addressing the limitations of catastrophic forgetting and offering a scalable, adaptive solution for real-world multi-attribute BIQA. Yunhao Luo 0005, Jinming Liu 0001, Wei Zhou 0021, Xin Jin 0014 |
ISCAS | 2 |
| 2025 | Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative PriorabstractImage compression methods are usually optimized isolatedly for human perception or machine analysis tasks. We reveal fundamental commonalities between these objectives: preserving accurate semantic information is paramount, as it directly dictates the integrity of critical information for intelligent tasks and aids human understanding. Concurrently, enhanced perceptual quality not only improves visual appeal but also, by ensuring realistic image distributions, benefits semantic feature extraction for machine tasks.
Based on this insight, we propose Diff-ICMH, a generative image compression framework aiming for harmonizing machine and human vision in image compression. It ensures perceptual realism by leveraging generative priors and simultaneously guarantees semantic fidelity through the incorporation of Semantic Consistency loss (SC loss) during training.
Additionally, we introduce the Tag Guidance Module (TGM) that leverages highly semantic image-level tags to stimulate the pre-trained diffusion model's generative capabilities, requiring minimal additional bit rates. Consequently, Diff-ICMH supports multiple intelligent tasks through a single codec and bitstream without any task-specific adaptation, while preserving high-quality visual experience for human perception. Extensive experimental results demonstrate Diff-ICMH's superiority and generalizability across diverse tasks, while maintaining visual appeal for human perception. Ruoyu Feng 0001, Yunpeng Qi, Jinming Liu 0001, Xin Li 0082, Xin Jin 0014, Zhibo Chen 0001 |
NeurIPS | 3 |
| 2025 | Standard Codec is Enough: A Training-Free 4D Gaussian Compression with Dynamic UV Mappingabstract4D Gaussian Splatting (4DGS) has demonstrated advances in the dynamic scene representation. However, the time-varying attributes across frames introduce considerable storage and transmission costs, making 4DGS challenging to widely deploy. Existing compression methods struggle to obtain inter-frame residuals due to the unstructured nature of Gaussian representations, making explicit motion estimation and residual modeling inherently challenging. To address these, we propose a Training-Free 4D Gaussian Compression framework, TF4DGC, which transforms 4D Gaussian into a well-structured 2D representation, easy to estimate motion for coding, via a UV mapping. Specifically, we project 3D Gaussians onto a canonical sphere to obtain temporally consistent UV coordinates, and organize per-frame Gaussian attributes into multi-channel video sequences. This design enables the direct use of standard video codecs (e.g., AVC, HEVC) for compression, which is compatible with widespread hardware decoder support on laptops and mobile devices. Experimental results show that our method efficiently compresses both reconstructed and generated Gaussian scenarios, highlighting its general applicability. Our method offers a scalable and practical solution for 4DGS compression and facilitates real-time deployment in bandwidth constrained environments. Jinming Liu 0001, Shengyang Zhao, Qiang Hu 0003, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014 |
VCIP | 2 |
| 2025 | Quadtree Partitioning-based Visual Token Pruning for MLLMs Considering Information DensityabstractMultimodal Large Language Models (MLLMs) excel at comprehensive understanding by integrating visual and textual information. However, their inference speed is often bottlenecked by redundant visual token inputs. Existing methods tend to alleviate this issue with a heuristic pruning strategy based on token importance, tailored to certain commonly adopted vision encoders like CLIP. In this paper, we propose a novel training-free token pruning method based on a well-designed metric of information density, where we decide which tokens are retained according to their entropy, following the classic information theory. Based on that, we further propose a quadtree partitioning strategy, in which we retain these tokens with higher entropy so as to preserve the visual spatial structure while allocating more tokens to more informative regions. Experiments on LLaVA-v1.5-7B and 13B across six benchmarks show our method achieves state-of-the-art performance—retaining over 90% of full-token accuracy even at a 6.25% token budget—while cutting TFLOPs by up to 20% compared to FastV and by 81% compared to the original LLaVA-v1.5. Yuntao Wei, Jinming Liu 0001, Shengyang Zhao, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014 |
VCIP | 2 |
| 2025 | MDLPCC: Misalignment-aware dynamic LiDAR point cloud compressionabstractLiDAR point cloud plays an important role in various real-world areas. It is usually generated as sequences by LiDAR on moving vehicles. Regarding the large data size of LiDAR point clouds, Dynamic Point Cloud Compression (DPCC) methods are developed to reduce transmission and storage data costs. However, most existing DPCC methods neglect the intrinsic misalignment in LiDAR point cloud sequences, limiting the rate–distortion (RD) performance. This paper proposes a Misalignment-aware Dynamic LiDAR Point Cloud Compression method (MDLPCC), which alleviates the misalignment problem in both macroscope and microscope. MDLPCC exploits a global transformation (GlobTrans) method to eliminate the macroscopic misalignment problem, which is the obvious gap between two continuous point cloud frames. MDLPCC also uses a spatial–temporal mixed structure to alleviate the microscopic misalignment, which still exists in the detailed parts of two point clouds after GlobTrans. The experiments on our MDLPCC show superior performance over existing point cloud compression methods. Ao Luo, Linxin Song, Keisuke Nonaka, Jinming Liu 0001, Kyohei Unno, Kohei Matsuzaki, Heming Sun, Jiro Katto |
J. Vis. Commun. Image Represent. | 4 |
| 2024 | One at a Time: Progressive Multi-Step Volumetric Probability Learning for Reliable 3D Scene PerceptionabstractNumerous studies have investigated the pivotal role of reliable 3D volume representation in scene perception tasks, such as multi-view stereo (MVS) and semantic scene completion (SSC). They typically construct 3D probability volumes directly with geometric correspondence, attempting to fully address the scene perception tasks in a single forward pass. However, such a single-step solution makes it hard to learn accurate and convincing volumetric probability, especially in challenging regions like unexpected occlusions and complicated light reflections. Therefore, this paper proposes to decompose the complicated 3D volume representation learning into a sequence of generative steps to facilitate fine and reliable scene perception. Considering the recent advances achieved by strong generative diffusion models, we introduce a multi-step learning framework, dubbed as VPD, dedicated to progressively refining the Volumetric Probability in a Diffusion process. Specifically, we first build a coarse probability volume from input images with the off-the-shelf scene perception baselines, which is then conditioned as the basic geometry prior before being fed into a 3D diffusion UNet, to progressively achieve accurate probability distribution modeling. To handle the corner cases in challenging areas, a Confidence-Aware Contextual Collaboration (CACC) module is developed to correct the uncertain regions for reliable volumetric learning based on multi-scale contextual contents. Moreover, an Online Filtering (OF) strategy is designed to maintain representation consistency for stable diffusion sampling. Extensive experiments are conducted on scene perception tasks including multi-view stereo (MVS) and semantic scene completion (SSC), to validate the efficacy of our method in learning reliable volumetric representations. Notably, for the SSC task, our work stands out as the first to surpass LiDAR-based methods on the SemanticKITTI dataset. Bohan Li 0015, Yasheng Sun, Jingxin Dong 0002, Jinming Liu 0001, Xin Jin 0014, Wenjun Zeng 0001 |
AAAI | 5 |
| 2024 | Closed-Loop Unsupervised Representation Disentanglement with β-VAE Distillation and Diffusion Probabilistic Feedback
Xin Jin 0014, Bohan Li 0015, Baao Xie, Jinming Liu 0001, Tao Yang 0032, Wenjun Zeng 0001 |
ECCV (45) | 5 |
| 2024 | Rate-Distortion-Cognition Controllable Versatile Neural Image Compression
Jinming Liu 0001, Ruoyu Feng 0001, Yunpeng Qi, Qiuyu Chen, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014 |
ECCV (56) | 1 |
| 2024 | Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMsabstractWe present a new image compression paradigm to achieve "intelligently coding for machine" by cleverly leveraging the common sense of Large Multimodal Models (LMMs). We are motivated by the evidence that large language/multimodal models are powerful general-purpose semantics predictors for understanding the real world. Different from traditional image compression typically optimized for human eyes, the image coding for machines (ICM) framework we focus on requires the compressed bitstream to more comply with different downstream intelligent analysis tasks. To this end, we employ LMM to${\text{tell codec what to compress}}$: 1) first utilize the powerful semantic understanding capability of LMMs w.r.t object grounding, identification, and importance ranking via prompts, to disentangle image content before compression, 2) and then based on these semantic priors we accordingly encode and transmit objects of the image in order with a structured bitstream. In this way, diverse vision benchmarks including image classification, object detection, instance segmentation, etc., can be well supported with such a semantically structured bitstream. We dub our method "SDComp" for "Semantically Disentangled Compression", and compare it with state-of-the-art codecs on a wide variety of different vision tasks. SDComp codec leads to more flexible reconstruction results, promised decoded visual quality, and a more generic/satisfactory intelligent task-supporting ability. Jinming Liu 0001, Yuntao Wei, Junyan Lin, Shengyang Zhao, Heming Sun, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014 |
VCIP | 1 |
| 2023 | Learned Image Compression with Mixed Transformer-CNN ArchitecturesabstractLearned image compression (LIC) methods have exhibited promising progress and superior rate-distortion performance compared with classical image compression standards. Most existing LIC methods are Convolutional Neural Networks-based (CNN-based) or Transformer-based, which have different advantages. Exploiting both advantages is a point worth exploring, which has two challenges: 1) how to effectively fuse the two methods? 2) how to achieve higher performance with a suitable complexity? In this paper, we propose an efficient parallel Transformer-CNN Mixture (TCM) block with a controllable complexity to incorporate the local modeling ability of CNN and the non-local modeling ability of transformers to improve the overall architecture of image compression models. Besides, inspired by the recent progress of entropy estimation models and attention modules, we propose a channel-wise entropy model with parameter-efficient swin-transformer-based attention (SWAtten) modules by using channel squeezing. Experimental results demonstrate our proposed method achieves state-of-the-art rate-distortion performances on three different resolution datasets (i.e., Kodak, Tecnick, CLIC Professional Validation) compared to existing LIC methods. The code is at https://github.com/jmliu206/LIC_TCM. Jinming Liu 0001, Heming Sun, Jiro Katto |
CVPR | 1 |
| 2023 | Multistage Spatial Context Models for Learned Image CompressionabstractRecent state-of-the-art Learned Image Compression methods feature spatial context models, achieving great rate-distortion improvements over hyperprior methods. However, the autoregressive context model requires serial decoding, limiting run-time performance. The Checkerboard context model allows parallel decoding at a cost of reduced RD performance. We present a series of multistage spatial context models allowing both fast decoding and better RD performance. We split the latent space into square patches and decode serially within each patch while different patches are decoded in parallel. The proposed method features a comparable decoding speed to Checkerboard while reaching the RD performance of Autoregressive and even also outperforming Autoregressive. Inside each patch, the decoding order must be carefully decided as a bad order negatively impacts performance; therefore, we also propose a decoding order optimization algorithm. Fangzheng Lin, Heming Sun, Jinming Liu 0001, Jiro Katto |
ICASSP | 3 |
| 2023 | Composable Image Coding for Machine via Task-oriented Internal Adaptor and External PriorabstractTraditional image coding standards are typically optimized with a focus on human perception, which conflicts with the fact that most of the images are now analyzed by machines. To enable a variety of downstream intelligent tasks, contemporary approaches either utilize traditional codecs for image compression which are then used for task analysis, or develop a unified feature compression paradigm with deep learning techniques. However, they might suffer from accumulative errors and poor compatibility/generalization due to the conflict between standardized codecs and diverse machine tasks. We argue that a favorable image coding for machine (ICM) framework should have highly efficient adaptation capability, and take the ultimate task goals into account. Oriented at this, we propose a composable ICM solution dubbed Com-ICM, which develops plug-and-play lightweight internal adaptors injected into the codec architecture for efficient task transfer, and leverages off-the-shelf (large) models to provide external prior information for further task-oriented semantics learning. The internal adaptors (from the architectural aspect) and external priors (from the precondition aspect) complement each other, resulting in a mutually beneficial effect. We evaluate Com-ICM on diverse vision benchmarks, including image classification, object detection, and semantic segmentation, demonstrating its effectiveness and superiority. We are also actively submitting Com-ICM as a technical proposal to the international organization for standardization. Jinming Liu 0001, Xin Jin 0014, Ruoyu Feng 0001, Zhibo Chen 0001, Wenjun Zeng 0001 |
VCIP | 1 |
| 2023 | PTS-LIC: Pruning Threshold Searching for Lightweight Learned Image CompressionabstractLearned Image Compression (LIC), which uses neural networks to compress images, has experienced significant growth in recent years. The hyperprior-module-based LIC model has achieved higher performance than classical codecs. However, the LIC models are too heavy (in calculation and parameter amounts) to apply to edge devices. To solve this problem, some former papers focus on structural pruning for LIC models. However, they either cause noticeable performance decrement or neglect the appropriate pruning threshold for each LIC model. These problems keep their pruning results sub-optimal. This paper proposes a Pruning Threshold Searching on the hyperprior module for different-quality LIC models. Our method removes most parameters and calculations while keeping the performance the same as the models before pruning. We removed at least 49.8% of parameters and 28.5% of calculations for the Channel-Wise-Context-Model-based models and 29.1% of parameters for the Cheng-2020 models. Ao Luo, Heming Sun, Jinming Liu 0001, Fangzheng Lin, Jiro Katto |
VCIP | 3 |
| 2022 | Memory-Efficient Learned Image Compression with Pruned Hyperprior ModuleabstractLearned Image Compression (LIC) gradually became more and more famous in these years. The hyperprior-module-based LIC models have achieved remarkable rate-distortion performance. However, the memory cost of these LIC models is too large to actually apply them to various devices, especially to portable or edge devices. The parameter scale is directly linked with memory cost. In our research, we found the hyperprior module is not only highly over-parameterized, but also its latent representation contains redundant information. Therefore, we propose a novel pruning method named ERHP in this paper to efficiently reduce the memory cost of hyperprior module, while improving the network performance. The experiments show our method is effective, reducing at least 22.6% parameters in the whole model while achieving better rate-distortion performance. Ao Luo, Heming Sun, Jinming Liu 0001, Jiro Katto |
ICIP | 3 |
| 2022 | Improving Multiple Machine Vision Tasks in the Compressed DomainabstractThere is a growing number of images that are analyzed by machines rather than just humans. Recently, most machine vision tasks are based on decoded images which require an image compression (encoding/decoding) framework. However, using the decoded images in the pixel-domain has two drawbacks: 1) the complexity is high for the decoder part, 2) the accuracy (e.g., mIoU, mean absolute error, and average precision) of machine vision tasks will be degraded since decoded images only aim to optimize the human perceived quality (e.g., PSNR) so that information required for machine vision tasks will be lost during the decoding process. In this paper, we improve the machine vision tasks in the compressed domain. 1) A gate module is utilized to effectively select some compressed-domain features. 2) Knowledge distillation is introduced to improve the accuracy. 3) A training strategy is explored to support multiple tasks including the image compression. The experimental results show that we can achieve better rate-accuracy/distortion and lower complexity compared with the state-of-the-art pixel-domain work that can take both machine and human vision tasks. Jinming Liu 0001, Heming Sun, Jiro Katto |
ICPR | 1 |
| 2022 | Semantic Segmentation In Learned Compressed DomainabstractMost machine vision tasks (e.g., semantic segmentation) are based on images encoded and decoded by image compression algorithms (e.g., JPEG). However, these decoded images in the pixel domain introduce distortion, and they are optimized for human perception, making the performance of machine vision tasks suboptimal. In this paper, we propose a method based on the compressed domain to improve segmentation tasks. i) A dynamic and a static channel selection method are proposed to reduce the redundancy of compressed representations that are obtained by encoding. ii) Two different transform modules are explored and analyzed to help the compressed representation be transformed as the features in the segmentation network. The experimental results show that we can save up to 15.8% bitrates compared with a state-of-the-art compressed domain-based work while saving up to about 83.6% bitrates and 44.8% inference time compared with the pixel domain-based method. Jinming Liu 0001, Heming Sun, Jiro Katto |
PCS | 1 |
| 2021 | Learning in Compressed Domain for Faster Machine Vision TasksabstractLearned image compression (LIC) has illustrated good ability for reconstruction quality driven tasks (e.g. PSNR, MS-SSIM) and machine vision tasks such as image understanding. However, most LIC frameworks are based on pixel domain, which requires the decoding process. In this paper, we develop a learned compressed domain framework for machine vision tasks. 1) By sending the compressed latent representation directly to the task network, the decoding computation can be eliminated to reduce the complexity. 2) By sorting the latent channels by entropy, only selective channels will be transmitted to the task network, which can reduce the bitrate. As a result, compared with the traditional pixel domain methods, we can reduce about 1/3 multiply-add operations (MACs) and 1/5 inference time while keeping the same accuracy. Moreover, proposed channel selection can contribute to at most 6.8% bitrate saving. Jinming Liu 0001, Heming Sun, Jiro Katto |
VCIP | 1 |