Mengtian Li 0002

dblp:183/6276-2 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
16since 2021 · last 2026
0009-0001-7281-2812ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Spatio-Temporal Disentanglement and Constrained Self-Attention for Multi-Modal Deception Detection
abstract
Multi-modal deception detection is a challenging yet important task, having pivotal applications in many fields such as business credibility assessment and multimedia anti-frauds. Previous methods either rely solely on spatial features or overemphasize only temporal information within or across modalities, which may overlook potential critical clues. Motivated by these observations, we propose a Spatio-Temporal Representation Disentanglement (STRD) framework for multi-modal deception detection, which uses a dual-encoder structure to learn spatial and temporal representations for each modality. Specifically, we introduce a pre-trained foundation model to act as the spatial encoder and design a lightweight network as the temporal encoder, extracting spatial semantics and capturing dynamic temporal patterns. Then, we propose a Constrained Self-Attention Block (CSAB), in which self-attention distribution of each head is regarded as spatial distribution and is constrained to attend a certain facial local region. Furthermore, we present a Cross-Modal Correlation Fusion Block (CCFB) to achieve temporal synchronization across modalities by measuring the correlations between visual and audio features. Extensive experiments show that our STRD outperforms the state-of-the-art methods on challenging DOLOS, BOL, BgOL, and RLtrial benchmarks. Particularly, STRD improves by 2.12% and 1.88% over the previous best results in terms of ACC on the DOLOS and BOL datasets, respectively. Additionally, STRD outperforms previous methods in cross-dataset testing, highlighting its superior generalization ability.
Zhiwen Shao, Hancheng Zhu, Rui Yao 0006, Lixin Zou, Mengtian Li 0002, Bin Sheng 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2025 MDSAM: Intrinsic Cues Guided Segmentation for Mirror Detection
Lizhuang Ma, Mengtian Li 0002
CGI (1)4
2025 StageDesigner: Artistic Stage Generation for Scenography via Theater Scripts
abstract
In this work, we introduce StageDesigner, the first comprehensive framework for artistic stage generation using large language models combined with layout-controlled diffusion models. Given the professional requirements of stage scenography, StageDesigner simulates the workflows of seasoned artists to generate immersive 3D stage scenes. Specifically, our approach is divided into three primary modules: Script Analysis, which extracts thematic and spatial cues from input scripts; Foreground Generation, which constructs and arranges essential 3D objects; and Background Generation, which produces a harmonious background aligned with the narrative atmosphere and maintains spatial coherence by managing occlusions between foreground and background elements. Furthermore, we introduce the StagePro-V1 dataset, a dedicated dataset with 276 unique stage scenes spanning different historical styles and annotated with scripts, images, and detailed 3D layouts, specifically tailored for this task. Finally, evaluations using both standard and newly proposed metrics, along with extensive user studies, demonstrate the effectiveness of StageDesigner, showcasing its ability to produce visually and thematically cohesive stages that meet both artistic and spatial coherence standards. Project can be found at: https://deadsmither5.github.io/2025/01/03/StageDesigner/
Zhaoxing Gan, Mengtian Li 0002, Ruhua Chen, Zhongxia Ji, Sichen Guo, Huanling Hu, Guangnan Ye, Zuo Hu
CVPR2
2025 CustAny: Customizing Anything from A Single Example
abstract
Recent advances in diffusion-based text-to-image models have simplified creating high-fidelity images, but preserving the identity (ID) of specific elements, like a personal dog, is still challenging. Object customization, using reference images and textual descriptions, is key to addressing this issue. Current object customization methods are either object-specific, requiring extensive fine-tuning, or object-agnostic, offering zero-shot customization but limited to specialized domains. The primary issue of promoting zero-shot object customization from specific domains to the general domain is to establish a large-scale general ID dataset for model pre-training, which is time-consuming and labor-intensive. In this paper, we propose a novel pipeline to construct a large dataset of general objects and build the Multi-Category ID-Consistent (MC-IDC) dataset, featuring 315k text-image samples across 10k categories. With the help of MC-IDC, we introduce Customizing Anything (CustAny), a zero-shot framework that maintains ID fidelity and supports flexible text editing for general objects. CustAny features three key components: a general ID extraction module, a dual-level ID injection module, and an ID-aware decoupling module, allowing it to customize any object from a single reference image and text prompt. Experiments demonstrate that CustAny outperforms existing methods in both general object customization and specialized domains like human customization and virtual try-on. Our contributions include a large-scale dataset, the CustAny framework and novel ID processing to advance this field. The official project page is in https://lingjiekong-fdu.github.io.
Lingjie Kong, Chengming Xu 0001, Xiaobin Hu, Wenhui Han, Jinlong Peng, Donghao Luo 0001, Mengtian Li 0002, Jiangning Zhang, Chengjie Wang 0001, Yanwei Fu 0001
CVPR8
2025 FilmComposer: LLM-Driven Music Production for Silent Film Clips
abstract
In this work, we implement music production for silent film clips using LLM-driven method. Given the strong professional demands of film music production, we propose the FilmComposer, simulating the actual workflows of professional musicians. FilmComposer is the first to combine large generative models with a multi-agent approach, leveraging the advantages of both waveform music and symbolic music generation. Additionally, FilmComposer is the first to focus on the three core elements of music production for film—audio quality, musicality, and musical development—and introduces various controls, such as rhythm, semantics, and visuals, to enhance these key aspects. Specifically, FilmComposer consists of the visual processing module, rhythm-controllable MusicGen, and multi-agent assessment, arrangement and mix. In addition, our framework can seamlessly integrate into the actual music production pipeline and allows user intervention in every step, providing strong interactivity and a high degree of creative freedom. Furthermore, we propose MusicPro-7k which includes 7,418 film clips, music, description, rhythm spots and main melody, considering the lack of a professional and high-quality film music dataset. Finally, both the standard metrics and the new specialized metrics we propose demonstrate that the music generated by our model achieves state-of-the-art performance in terms of quality, consistency with video, diversity, musicality, and musical development. Project page: https://apple-jun.github.io/FilmComposer.github.io/
Qile He, Youjia Zhu, Qiwei He, Mengtian Li 0002
CVPR5
2025 Knowledge Transfer Across Modalities for Weakly Supervised Point Cloud Semantic Segmentation
abstract
Current weakly supervised point cloud semantic segmentation struggles with insufficient utilization of limited annotations in unimodal representation learning due to the sparse and textureless nature of point clouds. In this work, we leverage cross-modality information by transferring knowledge from image and text sources to the point cloud network. The intuition is that images contribute rich texture, color, and discriminative information, complementing point clouds to boost semantic segmentation performance. To reduce extensive computational resources for cross-modality fusion, we introduce the Multi-Scale Deformable Knowledge Transfer, an innovative training scheme that optimizes and extends the one-to-one mapping to flexible one-to-many relations between multi-modal data. Furthermore, we employ pre-trained image-text models to generate pseudo labels for point clouds and construct positive and negative samples for semantic contrastive regularization, facilitating the full exploitation of unlabeled data. The experimental results evaluated on SemanticKITTI and nuScenes demonstrate substantial improvements, achieving an average gain of 3.8% over the previous weakly supervised methods, and comparable performances to fully supervised approaches.
Yunhang Shen, Mengtian Li 0002, Ke Li 0015, Xing Sun 0001, Shaohui Lin, Lizhuang Ma
ICASSP3
2025 LMTalker: Sparse Landmark-guided Gaussian Splatting for High-fidelity Talking Head Synthesis
abstract
3D Gaussian splatting (3DGS) has demonstrated significant potential in audio-driven talking head synthesis. However, despite notable advancements in speed and fidelity, current methods still face challenges such as inaccurate lip movements and facial artifacts. To address these issues, we propose LMTalker, a sparse landmark-guided 3DGS method, applying facial landmarks for the first time in 3DGS-based talking head synthesis. Our method explicitly leverages sparse facial landmarks to guide the deformation of dense Gaussians, effectively reduces inconsistencies between the input audio and facial dynamics, leading to improved lip movement accuracy and facial fidelity. Furthermore, we utilize facial landmarks in a hierarchical way to achieve region-specific generation. By integrating audio information, we enhance the clarity and reduce artifacts in the inner mouth region. Experimental results demonstrate that our method surpasses existing methods in terms of fidelity and lip movements accuracy, while maintaining high rendering speed. Our project page is available at https://jiangzhiwen0520.github.io/LMTalker
Zhiwen Jiang, Xuemin Lei, Mengtian Li 0002
ICASSP4
2025 ArtNVG: Content-Style Separated Artistic Neighboring-View Gaussian Stylization
abstract
As demand from the film and gaming industries for 3D scenes with target styles grows, the importance of advanced 3D stylization techniques increases. However, recent methods often struggle to maintain local consistency in color and texture throughout stylized scenes, which is essential for maintaining aesthetic coherence. To solve this problem, this paper introduces ArtNVG, an innovative 3D stylization framework that efficiently generates stylized 3D scenes by leveraging reference style images. Built on 3D Gaussian Splatting (3DGS), ArtNVG achieves rapid optimization and rendering while upholding high reconstruction quality. Our framework realizes high-quality 3D stylization by incorporating two pivotal techniques: Content-Style Separated Control and Attention-based Neighboring-View Alignment. Content-Style Separated Control uses the CSGO model and the Tile ControlNet to decouple the content and style control, reducing risks of information leakage. Concurrently, Attention-based Neighboring-View Alignment ensures consistency of local colors and textures across neighboring views, significantly improving visual quality. Extensive experiments validate that ArtNVG surpasses existing SOTA methods, delivering superior results in content preservation, style alignment, and local consistency.
Zixiao Gu, Zhenye Zhang, Mengtian Li 0002, Zhongxia Ji, Ruhua Chen, Zuo Hu, Guangnan Ye
ICMR3
2025 EditMaster: Bridging Text instruction and Visual Example for Multimodal guided Image Editing
abstract
Recent advances in image editing systems reveal critical limitations in handling complex real-world scenarios requiring multimodal condition controls. While text instructions enable broad semantic guidance, visual examples provide precise visual reference in specific scenarios, existing unimodal approaches fail to synergize these complementary modalities effectively. We propose EditMaster, a unified framework that integrates text and visual controls through multimodal instruction learning, enabling precise image manipulation with bidirectional consistency. Our framework introduces three core innovations: A multimodal large language model enhanced with extended visual tokens replaces CLIP text encoders, generating pre-edited visual guidance that aligns textual commands with visual examples to guide diffusion model toward high-quality outputs; The Mask-Based Decoupled Residual Exemplar-Attention module preserves unedited regions through spatial masking while integrating visual details via residual pathways; A systematic data construction method converts unimodal editing datasets into a task-specific multimodal dataset, eliminating the need for de novo data construction. Experiments show that our approach outperforms unimodal baselines and excels in complex multimodal instruction editing, setting a new benchmark for this field.
Mengtian Li 0002, Jiewei Tang, Junyu Deng, Guangnan Ye, Yu-Gang Jiang 0001
ACM Multimedia2
2025 Domain-Incremental Learning Paradigm for scene understanding via Pseudo-Replay Generation
abstract
Scene understanding is a computer vision task that involves grasping the pixel-level distribution of objects. Unlike most research focuses on single-scene models, we consider a more versatile proposal: domain-incremental learning for scene understanding. This allows us to adapt well-studied single-scene models into multi-scene models, reducing data requirements and ensuring model flexibility. However, domain-incremental learning that leverages correlations between scene domains has yet to be explored. To address this challenge, we propose a Domain-Incremental Learning Paradigm (D-ILP) for scene understanding, along with a new strategy of Pseudo-Replay Generation (PRG) that does not require manual labeling. Specifically, D-ILP leverages pre-trained single-scene models and incremental images for supervised training to acquire new knowledge from other scenes. As a pre-trained generation model, PRG can controllably generate pseudo-replays resembling source images from incremental images and text prompts. These pseudo-replays are utilized to minimize catastrophic forgetting in the original scene. We perform experiments with three publicly accessible models: Mask2Former, Segformer, and DeepLabv3+. With successfully transforming these single-scene models into multi-scene models, we achieve high-quality parsing results for original and new scenes simultaneously. Meanwhile, the validity and rationality of our method are proved by the analysis of D-ILP.
Qile He, Mengtian Li 0002, Xin Tan 0002
Graph. Model.4
2024 Sonic VisionLM: Playing Sound with Vision Language Models
abstract
There has been a growing interest in the task of generating sound for silent videos, primarily because of its prac-ticality in streamlining video post-production. However, existing methods for video-sound generation attempt to di-rectly create sound from visual representations, which can be challenging due to the difficulty of aligning visual rep-resentations with audio representations. In this paper, we present Sonic VisionLM, a novel framework aimed at gen-erating a wide range of sound effects by leveraging vision-language models(VLMs). Instead of generating audio di-rectly from video, we use the capabilities of powerful VLMs. When provided with a silent video, our approach first iden-tifies events within the video using a VLM to suggest pos-sible sounds that match the video content. This shift in approach transforms the challenging task of aligning image and audio into more well-studied sub-problems of aligning image-to-text and text-to-audio through the popular diffusion models. To improve the quality of audio recommen-dations with LLMs, we have collected an extensive dataset that maps text descriptions to specific sound effects and de-veloped a time-controlled audio adapter. Our approach surpasses current state-of-the-art methods for converting video to audio, enhancing synchronization with the visuals, and improving alignment between audio and video components. Project page: https://yusiissy.github.io/SonicVisionLM.github.io/
Shengye Yu, Qile He, Mengtian Li 0002
CVPR4
2024 Class-imbalanced semi-supervised learning for large-scale point cloud semantic segmentation via decoupling optimization
abstract
Semi-supervised learning (SSL), thanks to the significant reduction of data annotation costs, has been an active research topic for large-scale 3D scene understanding. However, the existing SSL-based methods suffer from severe training bias, mainly due to class imbalance and long-tail distributions of the point cloud data. As a result, they lead to a biased prediction for the tail class segmentation. In this paper, we introduce a new decoupling optimization framework, which disentangles feature representation learning and classifier in an alternative optimization manner to shift the bias decision boundary effectively. In particular, we first employ two-round pseudo-label generation to select unlabeled points across head-to-tail classes. We further introduce multi-class imbalanced focus loss to adaptively pay more attention to feature learning across head-to-tail classes. We fix the backbone parameters after feature learning and retrain the classifier using ground-truth points to update its parameters. Extensive experiments demonstrate the effectiveness of our method outperforming previous state-of-the-art methods on both indoor and outdoor 3D point cloud datasets ( i.e. , S3DIS, ScanNet-V2, Semantic3D, and SemanticKITTI) using 1% and 1pt evaluation.
Mengtian Li 0002, Shaohui Lin, Yunhang Shen, Baochang Zhang 0001, Lizhuang Ma
Pattern Recognit.1
2023 A fine-grained vision and language representation framework with graph-based fashion semantic knowledge
Huiming Ding, Mengtian Li 0002, Lizhuang Ma
Comput. Graph.4
2022 HybridCR: Weakly-Supervised 3D Point Cloud Semantic Segmentation via Hybrid Contrastive Regularization
abstract
To address the huge labeling cost in large-scale point cloud semantic segmentation, we propose a novel hybrid contrastive regularization (HybridCR) framework in weakly-supervised setting, which obtains competitive performance compared to its fully-supervised counterpart. Specifically, HybridCR is the first framework to leverage both point consistency and employ contrastive regularization with pseudo labeling in an end-to-end manner. Fundamentally, HybridCR explicitly and effectively considers the semantic similarity between local neighboring points and global characteristics of 3D classes. We further design a dynamic point cloud augmentor to generate diversity and robust sample views, whose transformation parameter is jointly optimized with model training. Through extensive experiments, HybridCR achieves significant performance improvement against the SOTA methods on both indoor and outdoor datasets, e.g., S3DIS, ScanNet-V2, Semantic3D, and SemanticKITTI.
Mengtian Li 0002, Yuan Xie 0006, Yunhang Shen, Bo Ke, Ruizhi Qiao, Bo Ren 0002, Shaohui Lin, Lizhuang Ma
CVPR1
2022 Hyperspherical Learning in Multi-Label Classification
Bo Ke, Yunquan Zhu, Mengtian Li 0002, Xiujun Shu, Ruizhi Qiao, Bo Ren 0002
ECCV (25)3
2022 Paying attention for adjacent areas: Learning discriminative features for large-scale 3D scene segmentation
Mengtian Li 0002, Yuan Xie 0006, Lizhuang Ma
Pattern Recognit.1