Zhe Li 0081

dblp:11/751-81 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0002-0230-0501ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
3D vision · 50% Robot manipulation · 44% Vision and language · 6%
Computer graphics and multimedia
2 papers
Image and video coding · 54% Multimedia analysis and retrieval · 46%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
0.912025
URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model · NeurIPS 2025
Computer vision › 3D vision › 3d reconstruction › object reconstruction
articulated object reconstruction
0.912025
URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model · NeurIPS 2025
Robotics › Robot manipulation
digital twin generation
0.912025
URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model · NeurIPS 2025
Robotics › Robot manipulation
robot simulation
0.912025
URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model · NeurIPS 2025
Image and video coding › image compression › lossy image compression
extreme low bit-rate compression
0.912025
Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image Compression · CVPR 2025
Image and video coding › image compression › learned image compression
generative image compression
0.912025
Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image Compression · CVPR 2025
Multimedia analysis and retrieval › deepfake detection
audio-visual deepfake detection
0.812024
MFMS: Learning Modality-Fused and Modality-Specific Features for Deepfake Detection and Localization Tasks · ACM Multimedia 2024
Multimedia analysis and retrieval
deepfake detection
0.812024
MFMS: Learning Modality-Fused and Modality-Specific Features for Deepfake Detection and Localization Tasks · ACM Multimedia 2024
Computer vision › 3D vision › 3d scene understanding
multimodal 3d understanding
0.312025
URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

vector-quantized encoder · 0.9point cloud segmentation · 0.9multimodal large language model · 0.9map-guided latent masking · 0.9diffusion model · 0.9autoregressive prediction · 0.9modality-specific feature learning · 0.8modality-fused feature learning · 0.8
YearPublicationVenuePosition
2025 Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image Compression
abstract
Leveraging the generative power of diffusion models, generative image compression has achieved impressive perceptual fidelity even at extremely low bitrates. However, current methods often neglect the non-uniform complexity of images, limiting their ability to balance global perceptual quality with local texture consistency and to allocate coding resources efficiently. To address this, we introduce the Map-guided Masking Realism Image Diffusion Codec (MRIDC), designed to optimize the trade- off between local distortion and global perceptual quality in extreme-low bitrate compression. MRIDC integrates a vector-quantized image encoder with a diffusion-based decoder. On the encoding side, we propose a Map-guided Latent Masking (MLM) module, which selectively masks elements in the latent space based on prior information, allowing adaptive resource allocation aligned with image complexity. On the decoding side, masked latents are completed using the Bidirectional Prediction Controllable Generation (BPCG) module, which guides the constrained generation process within the diffusion model to reconstruct the image. Experimental results show that MRIDC achieves state-of-the-art perceptual compression quality at extremely low bitrates, effectively preserving feature consistency in key regions and advancing the rate-distortion-perception performance curve, establishing new benchmarks in balancing compression efficiency with visual fidelity. Our code can be found at https://github.com/xjc97/mridc.
Jinchang Xu, Zhe Li 0081, Peidong Jia, Guoqing Xiang, Zhijian Hao, Shanghang Zhang
CVPR4
2025 URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model
abstract
Constructing accurate digital twins of articulated objects is essential for robotic simulation training and embodied AI world model building, yet historically requires painstaking manual modeling or multi-stage pipelines. In this work, we propose \textbf{URDF-Anything}, an end-to-end automatic reconstruction framework based on a 3D multimodal large language model (MLLM). URDF-Anything utilizes an autoregressive prediction framework based on point-cloud and text multimodal input to jointly optimize geometric segmentation and kinematic parameter prediction. It implements a specialized [SEG] token mechanism that interacts directly with point cloud features, enabling fine-grained part-level segmentation while maintaining consistency with the kinematic parameter predictions. Experiments on both simulated and real-world datasets demonstrate that our method significantly outperforms existing approaches regarding geometric segmentation (mIoU 17\% improvement), kinematic parameter prediction (average error reduction of 29\%), and physical executability (surpassing baselines by 50\%). Notably, our method exhibits excellent generalization ability, performing well even on objects outside the training set. This work provides an efficient solution for constructing digital twins for robotic simulation, significantly enhancing the sim-to-real transfer capability.
Zhe Li 0081, Xiang Bai, Jieyu Zhang 0001, Zhuangzhe Wu, Ying Li 0128, Chengkai Hou, Shanghang Zhang
NeurIPS1
2024 MFMS: Learning Modality-Fused and Modality-Specific Features for Deepfake Detection and Localization Tasks
abstract
This paper presents a summary of the proposed solution to the AV-Deepfake1M competition. Deepfake technology is developing fast, and realistic generation techniques of audio and videos have aroused public concerns. With this background, the AV-Deepfake1M competition aims to address the problem of audio-video Deepfake and provides a large-scale dataset named AV-Deepfake1M to boost the research in this area. In this paper, we present our solutions which have achieved top performance in this competition. We also provide more detailed experiments to prove the effectiveness of the modules used in our methods.
Changtao Miao, Jianshu Li, Wenzhong Deng, Weibin Yao, Zhe Li 0081, Bingyu Hu, Weiwei Feng, Qi Chu 0001
ACM Multimedia7