EDBT 2026 Demo / reviewers in the wild / expert
Chengtao Lv
dblp:343/4000
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play ToolkitabstractLarge Vision-Language Models (VLMs) exhibit impressive multi-modal capabilities but suffer from prohibitive computational and memory demands, due to their long visual token sequences and massive parameter sizes. To address these issues, recent works have proposed training-free compression methods. However, existing efforts often suffer from three major limitations: (1) Current approaches do not decompose techniques into comparable modules, hindering fair evaluation across spatial and temporal redundancy. (2) Evaluation confined to simple single-turn tasks, failing to reflect performance in realistic scenarios. (3) Isolated use of individual compression techniques, without exploring their joint potential. To overcome these gaps, we introduce LLMC+, a comprehensive VLM compression benchmark with a versatile, plug-and-play toolkit. LLMC+ supports over 20 algorithms across five representative VLM families and enables systematic study of token-level and model-level compression. Our benchmark reveals that: (1) Spatial and temporal redundancies demand distinct technical strategies. (2) Token reduction methods degrade significantly in multi-turn dialogue and detail-sensitive tasks. (3) Combining token and model compression achieves extreme compression with minimal performance loss. We believe LLMC+ will facilitate fair evaluation and inspire future research in efficient VLM. Chengtao Lv, Bilang Zhang, Yang Yong, Ruihao Gong, Yushi Huang, Shiqiao Gu, Jiajun Wu 0024, Yumeng Shi, Wenya Wang 0001 |
AAAI | 1 |
| 2026 | G2HFNet: GeoGran-Aware Hierarchical Feature Fusion Network for Salient Object Detection in Optical Remote Sensing ImagesabstractRemote sensing images captured from aerial perspectives often exhibit significant scale variations and complex backgrounds, posing challenges for salient object detection (SOD). Existing methods typically extract multi-level features at a single scale using uniform attention mechanisms, leading to suboptimal representations and incomplete detection results. To address these issues, we propose a GeoGran-Aware Hierarchical Feature Fusion Network (G2HFNet) that fully exploits geometric and granular cues in optical remote sensing images. Specifically, G2HFNet adopts Swin Transformer as the backbone to extract multi-level features and integrates three key modules: the multi-scale detail enhancement (MDE) module to handle object scale variations and enrich fine details, the dual-branch geo-gran complementary (DGC) module to jointly capture fine-grained details and positional information in mid-level features, and the deep semantic perception (DSP) module to refine high-level positional cues via self-attention. Additionally, a local-global guidance fusion (LGF) module is introduced to replace traditional convolutions for effective multi-level feature integration. Extensive experiments demonstrate that G2HFNet achieves high-quality saliency maps and significantly improves detection performance in challenging remote sensing scenarios. Bin Wan, Runmin Cong, Xiaofei Zhou 0003, Hao Fang 0010, Chengtao Lv, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | RSONet: Region-Guided Selective Optimization Network for RGB-T Salient Object DetectionabstractThis paper focuses on the inconsistency in salient regions between RGB and thermal images. To address this issue, we propose the Region-guided Selective Optimization Network for RGB-T Salient Object Detection, which consists of the region guidance stage and saliency generation stage. In the region guidance stage, three parallel branches with same encoder-decoder structure equipped with the context interaction (CI) module and spatial-aware fusion (SF) module are designed to generate the guidance maps which are leveraged to calculate similarity scores. Then, in the saliency generation stage, the selective optimization (SO) module fuses RGB and thermal features based on the previously obtained similarity values to mitigate the impact of inconsistent distribution of salient targets between the two modalities. After that, to generate high-quality detection result, the dense detail enhancement (DDE) module which adopts the multiple dense connections and visual state space blocks is applied to low-level features for optimizing the detail information. In addition, the mutual interaction semantic (MIS) module is placed in the high-level features to dig the location cues by the mutual fusion strategy. We conduct extensive experiments on the RGB-T dataset, and the results demonstrate that the proposed RSONet achieves competitive performance against 27 state-of-the-art SOD methods. Bin Wan, Runmin Cong, Xiaofei Zhou 0003, Hao Fang 0010, Chengtao Lv, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | A survey of low-bit large language models: Basics, systems, and algorithms
Ruihao Gong, Yifu Ding 0001, Chengtao Lv, Xingyu Zheng, Jinyang Du, Yang Yong, Shiqiao Gu, Haotong Qin, Jinyang Guo 0002, Dahua Lin, Michele Magno, Xianglong Liu 0001 |
Neural Networks | 4 |
| 2025 | Causality-Inspired Debiasing Learning for Open World Object DetectionabstractOpen world object detection (OWOD) aims to identify both known instances of trained classes and unknown ones. Despite recent advancements, existing methods exhibit a detection bias towards known classes, as detectors are exclusively trained under the supervision of known classes. To address this problem, we construct a causal graph to scrutinize OWOD from a causal perspective, revealing that the bias problem primarily arises due to the confounding effect of known classes, and the causality between unknown objects and their predictions learned by the detector is weak. Therefore, we propose a causality-inspired debiasing framework for OWOD, aiming to bolster the performance of OWOD models by eliminating confounders and encouraging appropriate features. Specifically, a semantic causal intervention module is proposed to remove the confounding effect from known classes to unknown features, which introduces the known semantics to interact fairly with all unknown features through backdoor adjustment. Moreover, an unknown causality enhancement module is employed to enhance the causality of unknown objects and their predictions acquired by the model, which imposes constraints for different unknown classes in feature space with the contrastive learning paradigm from the perspective of intervention effect. Extensive experiments conducted on the commonly-used OWOD benchmarks demonstrate that our framework consistently yields superior results on unknown classes compared with state-of-the-art methods by a large margin (+25.0% UD-Pre, +10.2% Recall on unknown classes) and even better on known classes (+1.4% mAP on known classes). Yuqing Ma, Chengtao Lv, Jiakai Wang, Xianglong Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | PTQ4SAM: Post-Training Quantization for Segment AnythingabstractSegment Anything Model (SAM) has achieved impressive performance in many computer vision tasks. However, as a large-scale model, the immense memory and computation costs hinder its practical deployment. In this paper, we pro-pose a post-training quantization (PTQ)frameworkfor Segment Anything Model, namely PTQ4SAM. First, we investigate the inherent bottleneck of SAM quantization attributed to the bimodal distribution in post-Key-Linear activations. We analyze its characteristics from both per-tensor and per-channel perspectives, and propose a Bimodal Integration strategy, which utilizes a mathematically equivalent sign operation to transform the bimodal distribution into a relatively easy-quantized normal distribution offline. Second, SAM encompasses diverse attention mechanisms (i.e., self-attention and two-way cross-attention), resulting in substantial variations in the post-Softmax distributions. Therefore, we introduce an Adaptive Granularity Quantization for Softmax through searching the optimal power-of-two base, which is hardware-friendly. Extensive experimen-tal results across various vision tasks (instance segmentation, semantic segmentation and object detection), datasets and model variants show the superiority of PTQ4SAM. For example, when quantizing SAM-L to 6-bit, we achieve loss-less accuracy for instance segmentation, about 0.5% drop with theoretical3.9x acceleration. The code is available at https://github.com/chengtao-lv/PTQ4SAM. Chengtao Lv, Hong Chen 0004, Jinyang Guo 0002, Yifu Ding 0001, Xianglong Liu 0001 |
CVPR | 1 |
| 2024 | QVD: Post-training Quantization for Video Diffusion ModelsabstractRecently, video diffusion models (VDMs) have garnered significant attention due to their notable advancements in generating coherent and realistic video content. However, processing multiple frame features concurrently, coupled with the considerable model size, results in high latency and extensive memory consumption, hindering their broader application. Post-training quantization (PTQ) is an effective technique to reduce memory footprint and improve computational efficiency. Unlike image diffusion, we observe that the temporal features, which are integrated into all frame features, exhibit pronounced skewness. Furthermore, we investigate significant inter-channel disparities and asymmetries in the activation of video diffusion models, resulting in low coverage of quantization levels by individual channels and increasing the challenge of quantization. To address these issues, we introduce the first PTQ strategy tailored for video diffusion models, dubbed QVD. Specifically, we propose the High Temporal Discriminability Quantization (HTDQ) method, designed for temporal features, which retains the high discriminability of quantized features, providing precise temporal guidance for all video frames. In addition, we present the Scattered Channel Range Integration (SCRI) method which aims to improve the coverage of quantization levels across individual channels. Experimental validations across various models, datasets, and bit-width settings demonstrate the effectiveness of our QVD in terms of diverse metrics. In particular, we achieve near-lossless performance degradation on W8A8, outperforming the current methods by 205.12 in FVD. Shilong Tian, Hong Chen 0014, Chengtao Lv, Yu Liu 0031, Jinyang Guo 0002, Xianglong Liu 0001, Shengxi Li, Hao Yang 0008 |
ACM Multimedia | 3 |
| 2024 | ADNet: Anti-noise dual-branch network for road defect detection
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | TMNet: Triple-modal interaction encoder and multi-scale fusion decoder network for V-D-T salient object detection
Bin Wan, Chengtao Lv, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Hongkui Wang, Chenggang Yan 0001 |
Pattern Recognit. | 2 |
| 2024 | A foreground-context dual-guided network for light-field salient object detectionabstractLight-field salient object detection (SOD) has become an emerging trend as it records comprehensive information about natural scenes that can benefit salient object detection in various ways. However, salient object detection models with light-field data as input have not been thoroughly explored. The existing methods cannot effectively suppress the noise, and it is difficult to distinguish the foreground and background under challenging conditions including self-similarity, complex backgrounds, large depth of field, and non-Lambertian scenarios. In order to extract the feature of light-field images effectively and suppress the noise in light-field, in this paper, we propose a foreground and context dual guided network. Specifically, we design a global context extraction module (GCEM) and a local foreground extraction module (LFEM). GCEM is used to suppress global noise and roughly predict saliency maps. GCEM also can extract global context information from deep-level features to guide decoding process. By extracting local information from shallow-level, LFEM refines the prediction obtained by GCEM. In addition, we use RGB images to enhance the light-field images before the input GCEM. Experimental results show that our proposed method is effective in suppressing global noise and achieves better results when dealing with transparent objects and complex backgrounds. The experimental results show that the proposed method outperforms several other state-of-the-art methods on three light-field datasets. Xin Zheng 0006, Deyang Liu, Chengtao Lv, Jiebin Yan |
Signal Process. Image Commun. | 4 |
| 2024 | MFFNet: Multi-Modal Feature Fusion Network for V-D-T Salient Object DetectionabstractThis article discusses the limitations of single- and two-modal salient object detection (SOD) methods and the emergence of multi-modal SOD techniques that integrate Visible, Depth, or Thermal information. However, current multi-modal methods often rely on simple fusion techniques such as addition, multiplication and concatenation, to combine the different modalities, which is ineffective for challenging scenes, such as low illumination and background messy. To address this issue, we propose a novel multi-modal feature fusion network (MFFNet) for V-D-T salient object detection, where the two key points are the triple-modal deep fusion encoder and the progressive feature enhancement decoder. The MFFNet's triple-modal deep fusion (TDF) module is designed to integrate the features of the three modalities and explore their complementarity by utilizing mutual optimization during the encoding phase. In addition, the progressive feature enhancement decoder consists of the weighted context-enhanced feature (WCF) module, region optimization (RO) module and boundary perception (BP) module to produce region-aware and contour-aware features. After that, a multi-scale fusion (MF) module is proposed to integrate these features and generate high-quality saliency maps. We conduct extensive experiments on the VDT-2048 dataset, and our results show that the proposed MFFNet outperforms 12 state-of-the-art multi-modal methods. Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | SRI-Net: Similarity retrieval-based inference network for light field salient object detection
Chengtao Lv, Xiaofei Zhou 0003, Deyang Liu, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 1 |