VLDB 2026 Research / reviewers in the wild / expert
Jiedong Zhuang
dblp:305/1138
· DBLP profile ↗
16ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-0551-5911ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Q Cache: Visual Attention Is Valuable in Less than Half of Decode Layers for Multimodal Large Language ModelabstractMultimodal large language models (MLLMs) are plagued by exorbitant inference costs attributable to the profusion of visual tokens within the vision encoder. The redundant visual tokens engenders a substantial computational load and key-value (KV) cache footprint bottleneck. Existing approaches focus on token-wise optimization, leveraging diverse intricate token pruning techniques to eliminate non-crucial visual tokens. Nevertheless, these methods often unavoidably undermine the integrity of the KV cache, resulting in failures in long-text generation tasks. To this end, we conduct an in-depth investigation towards the attention mechanism of the model from a new perspective, and discern that attention within more than half of all decode layers are semantic similar. Upon this finding, we contend that the attention in certain layers can be streamlined by inheriting the attention from their preceding layers. Consequently, we propose Lazy Attention, an efficient attention mechanism that enables cross-layer sharing of similar attention patterns. It ingeniously reduces layer-wise redundant computation in attention. In Lazy Attention, we develop a novel layer-shared cache, Q Cache, tailored for MLLMs, which facilitates the reuse of queries across adjacent layers. In particular, Q Cache is lightweight and fully compatible with existing inference frameworks, including Flash Attention and KV cache. Additionally, our method is highly flexible as it is orthogonal to existing token-wise techniques and can be deployed independently or combined with token pruning approaches. Empirical evaluations on multiple benchmarks demonstrate that our method can reduce KV cache usage by over 35% and achieve 1.5x throughput improvement, while sacrificing only approximately 1% of performance on various MLLMs. Compared with SOTA token-wise methods, our technique achieves superior accuracy preservation. Jiedong Zhuang, Haoji Hu |
AAAI | 1 |
| 2026 | PAGS-relight: Position-aware Gaussian Splatting scene relighting with multi-view diffusion models
Jiangnan Ye 0002, Jiedong Zhuang, Jianhong Bai, Lianrui Mu, Wuhao Tan, Syed Abdul Rahman Abu-Bakar, Haoji Hu |
Pattern Recognit. | 2 |
| 2026 | GC3VG: Generalized Multi-Task Visual Grounding With Coarse-to-Fine Consistency ConstraintsabstractIn this work, we propose an efficient and streamlined paradigm to address the challenge of consistency prediction in generalized multi-task visual grounding. While most existing approaches primarily focus on integrating multi-modal information and employing multi-task learning to enhance both visual and linguistic understanding, they often rely on joint supervision at the region and pixel levels to exploit task complementarities. In contrast, C3VG explores the relatively under-addressed problem ofconsistency across multi-task predictions. To this end, a multi-task visual grounding framework based on a coarse-to-fine architecture is introduced. Empirical studies demonstrate that the incorporation of both implicit and explicit consistency constraints substantially enhances the coherence between detection and segmentation outputs. However, C3VG is restricted to single-referent visual grounding scenarios and exhibits limited generalizability to real-world applications, which often involve multi-referents or even absent referent. To overcome these limitations, we proposeGC3VG, which incorporates three key advancements: (1) extension to generalized scenarios, including both multi-referent and non-referent cases; (2) aUnified Coherent Refinement Modulethat implicitly encodes region- and instance-level features while explicitly modeling their relational alignment through an IoUbased constraint; and (3) aGranularity-aware Hard-mining Alignmentstrategy that enforces prediction consistency in the feature space and simultaneously enhances the discriminative power of visual and linguistic representations. Extensive experiments on RefCOCO/+/g and gRefCOCO demonstrate the effectiveness and generalizability of the proposed framework. Kai Chen 0037, Wenxuan Cheng, Jiedong Zhuang, Zhenhua Feng 0001, Pengfei Zhu 0001, Wankou Yang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Multi-task Visual Grounding with Coarse-to-Fine Consistency ConstraintsabstractMulti-task visual grounding involves the simultaneous execution of localization and segmentation in images based on textual expressions. The majority of advanced methods predominantly focus on transformer-based multimodal fusion, aiming to extract robust multimodal representations. However, ambiguity between referring expression comprehension (REC) and referring image segmentation (RIS) is error-prone, leading to inconsistencies between multi-task predictions. Besides, insufficient multimodal understanding directly contributes to biased target perception. To overcome these challenges, we propose a Coarse-to-fine Consistency Constraints Visual Grounding architecture (C3VG), which integrates implicit and explicit modeling approaches within a two-stage framework. Initially, query and pixel decoders are employed to generate preliminary detection and segmentation outputs, a process referred to as the Rough Semantic Perception (RSP) stage. These coarse predictions are subsequently refined through the proposed Mask-guided Interaction Module (MIM) and a novel explicit bidirectional consistency constraint loss to ensure consistent representations across tasks, which we term the Refined Consistency Interaction (RCI) stage. Furthermore, to address the challenge of insufficient multimodal understanding, we leverage pre-trained models based on visual-linguistic fusion representations. Empirical evaluations on the RefCOCO, RefCOCO+, and RefCOCOg datasets demonstrate the efficacy and soundness of C3VG, which significantly outperforms state-of-the-art REC and RIS methods by a substantial margin. Jiedong Zhuang, Wankou Yang |
AAAI | 3 |
| 2025 | ST3: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token TrimmingabstractMultimodal large language models (MLLMs) enhance their perceptual capabilities by integrating visual and textual information. However, processing the massive number of visual tokens incurs a significant computational cost. Existing analysis of the MLLM attention mechanisms remains shallow, leading to coarse-grain token pruning strategies that fail to effectively balance speed and accuracy. In this paper, we conduct a comprehensive investigation of MLLM attention mechanisms with LLaVA. We find that numerous visual tokens and partial attention computations are redundant during the decoding process. Based on this insight, we propose Spatial-Temporal Visual Token Trimming (ST3), a framework designed to accelerate MLLM inference without retraining. ST3 consists of two primary components: 1) Progressive Visual Token Pruning (PVTP), which eliminates inattentive visual tokens across layers, and 2) Visual Token Annealing (VTA), which dynamically reduces the number of visual tokens in each layer as the generated tokens grow. Together, these techniques deliver around 2x faster inference with only about 30% KV cache memory compared to the original LLaVA, while maintaining consistent performance across various datasets. Crucially, ST3 can be seamlessly integrated into existing pre-trained MLLMs, providing a plug-and-play solution for efficient inference. Jiedong Zhuang, Haoji Hu |
AAAI | 1 |
| 2025 | PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity DiscriminationabstractRecent advances in visual grounding have largely shifted away from traditional proposal-based two-stage frameworks due to their inefficiency and high computational complexity, favoring end-to-end direct reference paradigms. However, these methods rely exclusively on the referred target for supervision, overlooking the potential benefits of prominent prospective targets. Moreover, existing approaches often fail to incorporate multi-granularity discrimination, which is crucial for robust object identification in complex scenarios. To address these limitations, we propose PropVG, an end-to-end proposal-based framework that, to the best of our knowledge, is the first to seamlessly integrate foreground object proposal generation with referential object comprehension without requiring additional detectors. Furthermore, we introduce a Contrastive-based Refer Scoring (CRS) module, which employs contrastive learning at both sentence and word levels to enhance the capability in understanding and distinguishing referred objects. Additionally, we design a Multi-granularity Target Discrimination (MTD) module that fuses object- and semantic-level information to improve the recognition of absent targets. Extensive experiments on gRefCOCO (GREC/GRES), Ref-ZOM, R-RefCOCO, and RefCOCO (REC/RES) benchmarks demonstrate the effectiveness of PropVG. The codes and models are available at https://github.com/Dmmm1997/PropVG. Wenxuan Cheng, Jiedong Zhuang, Jiang-jiang Liu, Hongshen Zhao, Zhenhua Feng 0001, Wankou Yang |
ICCV | 3 |
| 2025 | PHiD: Preserving human identity in pose-guided character animation
Wenjie Zheng 0003, Xingze Zou, Lianrui Mu, Jing Wang 0234, Jiaqi Hu 0005, Jiangnan Ye 0002, Jiedong Zhuang, Mudassar Ali 0002, Olumayowa Idowu, Haoji Hu |
Neurocomputing | 7 |
| 2024 | FALIP: Visual Prompt as Foveal Attention Boosts CLIP Zero-Shot Performance
Jiedong Zhuang, Jiaqi Hu 0005, Lianrui Mu, Jiangnan Ye 0002, Haoji Hu |
ECCV (10) | 1 |
| 2024 | FashionR2R: Texture-preserving Rendered-to-Real Image Translation with Diffusion ModelsabstractModeling and producing lifelike clothed human images has attracted researchers' attention from different areas for decades, with the complexity from highly articulated and structured content. Rendering algorithms decompose and simulate the imaging process of a camera, while are limited by the accuracy of modeled variables and the efficiency of computation. Generative models can produce impressively vivid human images, however still lacking in controllability and editability. This paper studies photorealism enhancement of rendered images, leveraging generative power from diffusion models on the controlled basis of rendering. We introduce a novel framework to translate rendered images into their realistic counterparts, which consists of two stages: Domain Knowledge Injection (DKI) and Realistic Image Generation (RIG). In DKI, we adopt positive (real) domain finetuning and negative (rendered) domain embedding to inject knowledge into a pretrained Text-to-image (T2I) diffusion model. In RIG, we generate the realistic image corresponding to the input rendered image, with a Texture-preserving Attention Control (TAC) to preserve fine-grained clothing textures, exploiting the decoupled features encoded in the UNet structure. Additionally, we introduce SynFashion dataset, featuring high-quality digital clothing images with diverse textures. Extensive experimental results demonstrate the superiority and effectiveness of our method in rendered-to-real image translation. Gaofeng He, Jiedong Zhuang |
NeurIPS | 4 |
| 2024 | Perceptual Image Compression with Text-Guided Multi-level Fusion
Jiaqi Hu 0005, Jiedong Zhuang, Lu Yu 0003, Haoji Hu |
PRCV (5) | 2 |
| 2024 | Mitigating Hallucination in Visual-Language Models via Re-balancing Contrastive Decoding
Jiayuan Yu, Lianrui Mu, Jiedong Zhuang, Jiaqi Hu 0005, Jiangnan Ye 0002, Haoji Hu |
PRCV (5) | 4 |
| 2024 | Trans-DONeRF for Transparent Object Rendering with Mixed Depth Prior
Jiangnan Ye 0002, Taoqi Bao, Lianrui Mu, Jiedong Zhuang, Haoji Hu |
PRCV (6) | 5 |
| 2024 | Zero-Shot Referring Image Segmentation with Hierarchical Prompts and Frequency Domain Fusion
Jiedong Zhuang, Jiaqi Hu 0005, Haoji Hu |
PRICAI (4) | 2 |
| 2024 | Data-Free Quantization of Vision Transformers Through Perturbation-Aware Image Synthesis
Lianrui Mu, Jiedong Zhuang, Jiangnan Ye 0002, Haoji Hu |
PRICAI (3) | 3 |
| 2024 | Vision-Based UAV Self-Positioning in Low-Altitude Urban EnvironmentsabstractUnmanned Aerial Vehicles (UAVs) rely on satellite systems for stable positioning. However, due to limited satellite coverage or communication disruptions, UAVs may lose signals for positioning. In such situations, vision-based techniques can serve as an alternative, ensuring the self-positioning capability of UAVs. However, most of the existing datasets are developed for the geo-localization task of the objects captured by UAVs, rather than UAV self-positioning. Furthermore, the existing UAV datasets apply discrete sampling to synthetic data, such as Google Maps, neglecting the crucial aspects of dense sampling and the uncertainties commonly experienced in practical scenarios. To address these issues, this paper presents a new dataset, DenseUAV, that is the first publicly available dataset tailored for the UAV self-positioning task. DenseUAV adopts dense sampling on UAV images obtained in low-altitude urban areas. In total, over 27K UAV- and satellite-view images of 14 university campuses are collected and annotated. In terms of methodology, we first verify the superiority of Transformers over CNNs for the proposed task. Then we incorporate metric learning into representation learning to enhance the model's discriminative capacity and to reduce the modality discrepancy. Besides, to facilitate joint learning from both the satellite and UAV views, we introduce a mutually supervised learning approach. Last, we enhance the Recall@K metric and introduce a new measurement, SDM@K, to evaluate both the retrieval and localization performance for the proposed task. As a result, the proposed baseline method achieves a remarkable Recall@1 score of 83.01% and an SDM@1 score of 86.50% on DenseUAV. The dataset and code have been made publicly available on https://github.com/Dmmm1997/DenseUAV. Enhui Zheng, Zhenhua Feng 0001, Lei Qi 0001, Jiedong Zhuang, Wankou Yang |
IEEE Trans. Image Process. | 5 |
| 2022 | A Transformer-Based Feature Segmentation and Region Alignment Method for UAV-View Geo-LocalizationabstractCross-view geo-localization is a task of matching the same geographic image from different views, e.g., unmanned aerial vehicle (UAV) and satellite. The most difficult challenges are the position shift and the uncertainty of distance and scale. Existing methods are mainly aimed at digging for more comprehensive fine-grained information. However, it underestimates the importance of extracting robust feature representation and the impact of feature alignment. The CNN-based methods have achieved great success in cross-view geo-localization. However it still has some limitations, e.g., it can only extract part of the information in the neighborhood and some scale reduction operations will make some fine-grained information lost. In particular, we introduce a simple and efficient transformer-based structure called Feature Segmentation and Region Alignment (FSRA) to enhance the model’s ability to understand contextual information as well as to understand the distribution of instances. Without using additional supervisory information, FSRA divides regions based on the heat distribution of the transformer’s feature map, and then aligns multiple specific regions in different views one on one. Finally, FSRA integrates each region into a set of feature representations. The difference is that FSRA does not divide regions manually, but automatically based on the heat distribution of the feature map. So that specific instances can still be divided and aligned when there are significant shifts and scale changes in the image. In addition, a multiple sampling strategy is proposed to overcome the disparity in the number of satellite images and that of images from other sources. Experiments show that the proposed method has superior performance and achieves the state-of-the-art in both tasks of drone view target localization and drone navigation. Jianhong Hu, Jiedong Zhuang, Enhui Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |