EDBT 2026 Demo / reviewers in the wild / expert
Zhe Li 0038
dblp:11/751-38
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0009-0008-3896-7694ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control FlowabstractGenerating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may cause conflict between the style and content, harming the integration. Differently, in this work we build a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration. Moreover, we extend the stylized motion generation from one modality, i.e. the style motion, to multiple modalities including texts and images through contrastive learning, leading to flexible style control on the motion generation. To further boost the performance, we advance the motion diffusion to motion-aligned temporal latent diffusion by developing a novel motion VAE. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, while also enabling multimodal signals control. The code of our method will be made publicly available. Zhe Li 0038, Yisheng He, Weichao Shen, Qi Zuo, Lingteng Qiu, Shenhao Zhu, Zilong Dong, Laurence T. Yang, Chang Xu 0002, Weihao Yuan 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | AniGS: Animatable Gaussian Avatar from a Single Image with Inconsistent Gaussian ReconstructionabstractGenerating animatable human avatars from a single image is essential for various digital human modeling applications. Existing 3D reconstruction methods often struggle to capture fine details in animatable models, while generative approaches for controllable animation, though avoiding explicit 3D modeling, suffer from viewpoint inconsistencies in extreme poses and computational inefficiencies. In this paper, we address these challenges by leveraging the power of generative models to produce detailed multi-view canonical pose images, which help resolve ambiguities in animatable human reconstruction. We then propose a robust method for 3D reconstruction of inconsistent images, enabling real-time rendering during inference. Specifically, we adapt a transformer-based video generation model to generate multi-view canonical pose images and normal maps, pretraining on a large-scale video dataset to improve generalization. To handle view inconsistencies, we recast the reconstruction problem as a 4D task and introduce an efficient 3D modeling approach using 4D Gaussian Splatting. Experiments demonstrate that our method achieves photorealistic, real-time animation of 3D human avatars from in-the-wild images, showcasing its effectiveness and generalization capability. Our code will be available on https://github.com/aigc3d/AniGS. Lingteng Qiu, Shenhao Zhu, Qi Zuo, Xiaodong Gu 0004, Zhe Li 0038, Weihao Yuan 0001, Liefeng Bo, Guanying Chen, Zilong Dong |
CVPR | 8 |
| 2025 | LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and CaptioningabstractLanguage plays a vital role in the realm of human motion. Existing methods have largely depended on CLIP text embeddings for motion generation, yet they fall short in effectively aligning language and motion due to CLIP’s pretraining on static image-text pairs. This work introduces LaMP, a novel Language-Motion Pretraining model, which transitions from a language-vision to a more suitable language-motion latent space. It addresses key limitations by generating motion-informative text embeddings, significantly enhancing the relevance and semantics of generated motion sequences. With LaMP, we advance three key tasks: text-to-motion generation, motion-text retrieval, and motion captioning through aligned language-motion representation learning. For generation, LaMP instead of CLIP provides the text condition, and an autoregressive masked prediction is designed to achieve mask modeling without rank collapse in transformers. For retrieval, motion features from LaMP’s motion transformer interact with query tokens to retrieve text features from the text transformer, and vice versa. For captioning, we finetune a large language model with the language-informative motion features to develop a strong motion captioning model. In addition, we introduce the LaMP-BertScore metric to assess the alignment of generated motions with textual descriptions. Extensive experimental results on multiple datasets demonstrate substantial improvements over previous methods across all three tasks. Project page: https://aigc3d.github.io/LaMP Zhe Li 0038, Weihao Yuan 0001, Yisheng He, Lingteng Qiu, Shenhao Zhu, Xiaodong Gu 0004, Weichao Shen, Zilong Dong, Laurence T. Yang |
ICLR | 1 |
| 2025 | Interpretable Multimodal Tucker Fusion Model With Information Filtering for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) integrates multiple sources of sentiment information for processing and has demonstrated superior performance compared to single-modal sentiment analysis, making it widely applicable in domains such as human–computer interaction and public opinion supervision. However, current MSA models heavily rely on black-box deep learning (DL) methods, which lack interpretability. Additionally, effectively integrating multimodal data, reducing noise and redundancy, as well as bridging the semantic gap between heterogeneous data remain challenging issues in multimodal DL. To address these challenges, we propose an interpretable multimodal Tucker fusion model with information filtering (IMTFMIF). We are the first to utilize the multimodal Tucker fusion model for MSA tasks. This approach maps multimodal data into a unified tensor space for fusion, effectively reducing modal heterogeneity and eliminating redundant information while maintaining interpretability. Furthermore, mutual information is employed to filter out task-irrelevant information and explain the association between input and output from an information flow perspective. We propose a novel approach to enhance the comprehension of multimodal data and optimize model performance in MSA tasks. Finally, extensive experiments conducted on three public multimodal datasets demonstrate that our proposed IMTFMIF achieves competitive performance compared to state-of-the-art methods. Laurence T. Yang, Zhe Li 0038, Xianjun Deng, Fulan Fan, Zecan Yang |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | Tensor-empowered Incomplete Multimodal Learning with Modality Reconstruction for Edge IntelligenceabstractThe distributed computing paradigm of edge computing effectively addresses the challenges of data transmission delay and data privacy security. With the increasing popularity of IoT devices and 5 G networks, edge computing has a broader range of applications. The advancement in AI technology enables the realization of edge intelligence, which conducts data processing and analysis on edge devices to avoid excessive data transmission to the cloud, enhance system response speed, and protect user data privacy. In various edge intelligent systems like smart homes and autonomous driving, multimodal data plays a crucial role. However, missing modalities in such systems may lead to model failure in real-world environments. To tackle this issue, we propose a tensor-empowered modality reconstruction network (TMRN) that utilizes an end-to-end variational autoencoder for reconstructing missing modal data. This approach effectively enhances model robustness while reducing model size and training complexity. Furthermore, we introduce a supervised method for feature reconstruction to better align with the true distribution of missing modal data by leveraging tensor feature fusion and label supervision techniques. Additionally, we design a task information disentanglement module to make multimodal representations more relevant to specific tasks by effectively separating task-relevant from task-irrelevant information. Extensive experiments demonstrate that TMRN achieves competitive performance compared to existing state-of-the-art methods. Laurence T. Yang, Zhe Li 0038, Fulan Fan, Zecan Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | General Point Model Pretraining with Autoencoding and AutoregressiveabstractThe pre-training architectures of large language models encompass various types, including autoencoding models, autoregressive models, and encoder-decoder models. We posit that any modality can potentially benefit from a large language model, as long as it undergoes vector quantization to become discrete tokens. Inspired by the General Language Model, we propose a General Point Model (GPM) that seamlessly integrates autoencoding and autoregressive tasks in a point cloud transformer. This model is versatile, allowing fine-tuning for downstream point cloud representation tasks, as well as unconditional and conditional generation tasks. GPM enhances masked prediction in autoencoding through various forms of mask padding tasks, leading to improved performance in point cloud understanding. Additionally, GPM demonstrates highly competitive results in unconditional point cloud generation tasks, even exhibiting the potential for conditional generation tasks by modifying the input's conditional information. Compared to models like Point-BERT, MaskPoint. and PointMAE, our GPM achieves superior performance in point cloud understanding tasks. Furthermore, the integration of autoregressive and autoencoding within the same transformer underscores its versatility across different downstream tasks. Codes are available at https://github.com/gentlefress/GPM Zhe Li 0038, Zhangyang Gao, Cheng Tan 0012, Bocheng Ren, Laurence T. Yang, Stan Z. Li |
CVPR | 1 |
| 2024 | MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive LearningabstractThe scarcity of annotated data has sparked signifi-cant interest in unsupervised pre-training methods that leverage medical reports as auxiliary signals for medi-cal visual representation learning. However, existing re-search overlooks the multi-granularity nature of medical visual representation and lacks suitable contrastive learning techniques to improve the models' generalizability across different granularities, leading to the underutilization of image-text information. To address this, we pro-pose MLIP, a novel framework leveraging domain-specific medical knowledge as guiding signals to integrate language information into the visual domain through image-text contrastive learning. Our model includes global contrastive learning with our designed divergence encoder, lo-cal token-knowledge-patch alignment contrastive learning, and knowledge-guided category-level contrastive learning with expert knowledge. Experimental evaluations reveal the efficacy of our model in enhancing transfer performance for tasks such as image classification, object detection, and semantic segmentation. Notably, MLIP surpasses state-of-the-art methods even with limited annotated data, highlighting the potential of multimodal pre-training in advancing medical representation learning.11Codes are available at https://github.com/gentlefress/MLIP Zhe Li 0038, Laurence T. Yang, Bocheng Ren, Zhangyang Gao, Cheng Tan 0012, Stan Z. Li |
CVPR | 1 |
| 2024 | Capturing Detail Variations for Lightweight Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) has recently overhauled novel view synthesis, but it requires extensive computations for training and captures variations in detail with difficulty. In this paper, we propose a novel framework, termed CD-TDRF, to mitigate these dilemmas. CD-TDRF factorizes a density voxel grid into a core tensor and three matrices via Tucker decomposition, reducing memory usage and accelerating training. To better capture variations in complex scenes, CD-TDRF uses a fully convolutional network to extract prior information from the training images. Moreover, three learnable appearance planes are constructed to preserve information about scene details, which enhances the rendering quality significantly. Our experimental results demonstrate that CD-TDRF has achieved competitive rendering quality on three popular datasets and speeds up training compared with traditional NeRF models. Laurence T. Yang, Bocheng Ren, Jinglin Zhao, Zhe Li 0038, Guolei Zeng |
ICASSP | 5 |
| 2024 | TD3D: Tensor-based Discrete Diffusion Process for 3D Shape GenerationabstractBased on the recent popularity of diffusion models, we have proposed a tensor-based diffusion model for 3D shape generation (TD3D). This generator is capable of tasks such as unconditional shape generation, shape completion and cross-modal shape generation. TD3D utilizes the Vector Quantized Variational Autoencoder (VQ-VAE) for encoding, compressing 3D shapes into compact latent representations, and then learns the discrete diffusion model based on it. To preserve high-dimensional feature information, we propose a tensor-based noise injection process. To capture 3D features in space faster and use them, a ResNet3D module is introduced during the denoising process. To fuse shape features obtained from ResNet3D and Self-Attention mechanisms, we employ a tensor-based Self-Attention mechanism (T-SA) fusion method. Lastly, a ResNet3D-assisted Multi-Frequency Fusion Module (R-MFM) is designed to aggregate high and low frequency features. Based on the aforementioned design, TD3D provides high fidelity, diverse generated samples, and have the ability to generate 3D shapes across modalities. Extensive experiments have demonstrated its superior performance in various 3D shape generation tasks. Jinglin Zhao, Debin Liu, Laurence T. Yang, Ruonan Zhao, Zhe Li 0038 |
ICME | 6 |
| 2023 | Enhancing Sentence Representation with Visually-supervised Multimodal Pre-trainingabstractLarge-scale pre-trained language models have garnered significant attention in recent years due to their effectiveness in extracting sentence representations. However, most pre-trained models currently use transformer-based encoder with a single modality and are primarily designed for specific tasks such as natural language inference and question-answering. Unfortunately, this approach neglects the complementary information provided by multimodal data, which can enhance the effectiveness of sentence representation. To address this issue, we propose a Visually-supervised Pre-trained Multimodal Model (ViP) for sentence representation. Our model leverages diverse label-free multimodal proxy tasks to embed visual information into language, facilitating effective modality alignment and complementarity exploration. Additionally, our model utilizes a novel approach to distinguish highly similar negative and positive samples. We conduct comprehensive downstream experiments on natural language understanding and sentiment classification, demonstrating that ViP outperforms both existing unimodal and multimodal pre-trained models. Our contributions include a novel approach to multimodal pre-training and a state-of-the-art model for sentence representation that incorporates visual information.1 Our code is available at https://github.com/gentlefress/ViP Zhe Li 0038, Laurence T. Yang, Bocheng Ren, Xianjun Deng |
ACM Multimedia | 1 |