Meng Shao

dblp:46/8774 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%
Artificial intelligence
2 papers
Vision and language · 100%

Topics — the 3 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › face editing
face swapping
0.912025
High Fidelity Face Swapping via Facial Texture and Structure Consistency Mining · IEEE Trans. Multim. 2025
Human-AI interaction
interactive content creation
0.912025
IterMeme: Expert-Guided Multimodal LLM for Interactive Meme Creation with Layout-Aware Generation · IJCAI 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
IterMeme: Expert-Guided Multimodal LLM for Interactive Meme Creation with Layout-Aware Generation · IJCAI 2025

Methods — techniques the papers use, named apart from their topics

parameter-shared dual-LLM · 2.6multimodal large language model · 2.6mixture of experts · 2.6structure consistency mining · 1.7facial texture consistency mining · 1.7
YearPublicationVenuePosition
2025 IterMeme: Expert-Guided Multimodal LLM for Interactive Meme Creation with Layout-Aware Generation
abstract
Meme creation is a creative process that blends images and text. However, existing methods lack critical components, failing to support intent-driven caption-layout generation and personalized generation, making it difficult to generate high-quality memes. To address this limitation, we propose IterMeme, an end-to-end interactive meme creation framework that utilizes a unified Multimodal Large Language Model (MLLM) to facilitate seamless collaboration among multiple components. To overcome the absence of a caption-layout generation component, we develop a robust layout representation method and construct a large-scale image-caption-layout dataset, MemeCap, which enhances the model’s ability to comprehend emotions and coordinate caption-layout generation effectively. To address the lack of a personalization component, we introduce a parameter-shared dual-LLM architecture that decouples the intricate representations of reference images and text. Furthermore, we incorporate the expert-guided M³OE for fine-grained identity properties (IP) feature extraction and cross-modal fusion. By dynamically injecting features into every layer of the model, we enable adaptive refinement of both visual and semantic information. Experimental results demonstrate that IterMeme significantly advances the field of meme creation by delivering consistently high-quality outcomes. The code, model, and dataset will be open-sourced to the community.
Yaqi Cai, Shancheng Fang, Yadong Qu, Meng Shao, Hongtao Xie 0001
IJCAI5
2025 TalkingAvatar: Learning 3D talking human avatar via NeRF
Lingyun Yu 0002, Chuanbin Liu 0001, Wu Liu 0005, Quanwei Yang, Meng Shao
Neurocomputing7
2025 High Fidelity Face Swapping via Facial Texture and Structure Consistency Mining
Lingyun Yu 0002, Quanwei Yang, Meng Shao, Hongtao Xie 0001
IEEE Trans. Multim.4
2024 Symmetrical Siamese Network for pose-guided person synthesis
Quanwei Yang, Lingyun Yu 0002, Yun Song, Meng Shao, Guoqing Jin, Hongtao Xie 0001
Comput. Vis. Image Underst.5