Jiji Tang

dblp:226/4763 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
0009-0001-9581-8790ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021
YearPublicationVenuePosition
2025 StoryWeaver: A Unified World Model for Knowledge-Enhanced Story Character Customization
abstract
Story visualization has gained increasing attention in artificial intelligence. However, existing methods still struggle with maintaining a balance between character identity preservation and text-semantics alignment, largely due to a lack of detailed semantic modeling of the story scene. To tackle this challenge, we propose a novel knowledge graph, namely Character-Graph (CG), which represents various story-related knowledge, including the characters, their attributes and the relationship. We then introduce StoryWeaver, an image generator that achieves Customization via Character-Graph (C-CG), capable of consistent story visualization with rich text semantics. To further improve the multi-character generation performance, we incorporate knowledge-enhanced spatial guidance (KE-SG) into StoryWeaver to precisely inject character semantics into generation. To validate the effectiveness of our proposed method, extensive experiments are conducted using a new benchmark called TBC-Bench. The experiments confirm that our StoryWeaver excels not only in creating vivid visual story plots but also in accurately conveying character identities across various scenarios with considerable storage efficiency, e.g., achieving an average increase of +9.03% DINO-I and +13.44% CLIP-T. Furthermore, ablation experiments are conducted to verify the superiority of each proposed module.
Jinlu Zhang 0002, Jiji Tang, Tangjie Lv, Xiaoshuai Sun
AAAI2
2025 InterID: Improving Multi-ID Interaction for Personalized Image Generation
abstract
Personalized image generation is an important topic in text-to-image generation, with multi-ID personalization drawing widespread attention. Recent advancements in multi-ID image generation have led to zero-shot generation. However, existing multi-ID personalized generation methods cannot satisfy the real-world requirements of storytelling image generation where individuals actively interact through various poses and expressions. In this paper, we propose InterID, a plug-and-play framework designed for interaction-improved multi-ID personalization. We introduce a semantic prior as the spatial control signal in decoding texts to images, alleviating the challenge of generating complex interactions. The prior is obtained by arranging the individual concepts parsed from the input text with the aid of large language models and it is then fused with corresponding ID features to guide the flow of ID-enhanced features toward the target area. These ID-enhanced features, combined with the original text features, condition the generation process. Moreover, our approach is highly adaptable and can be seamlessly integrated with other custom models fine-tuned from the same base model, functioning as a versatile plugin. Extensive experiments validate the effectiveness and robustness of our approach, showing its superiority in multi-ID personalization. Our Code is released at https://github.com/sisibeauty/InterID.
Siting Chen, Jiji Tang, Xiaoshuai Sun
ICME3
2025 Let storytelling tell vivid stories: A multi-modal-agent-based unified storytelling framework
Jiji Tang, Chuanqi Zang, Mingtao Pei, Wei Liang 0008, Zeng Zhao
Neurocomputing2
2024 Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-Modal Structured Representations
abstract
Large-scale vision-language pre-training has achieved significant performance in multi-modal understanding and generation tasks. However, existing methods often perform poorly on image-text matching tasks that require structured representations, i.e., representations of objects, attributes, and relations. The models cannot make a distinction between "An astronaut rides a horse" and "A horse rides an astronaut". This is because they fail to fully leverage structured knowledge when learning multi-modal representations. In this paper, we present an end-to-end framework Structure-CLIP, which integrates Scene Graph Knowledge (SGK) to enhance multi-modal structured representations. Firstly, we use scene graphs to guide the construction of semantic negative examples, which results in an increased emphasis on learning structured representations. Moreover, a Knowledge-Enhance Encoder (KEE) is proposed to leverage SGK as input to further enhance structured representations. To verify the effectiveness of the proposed framework, we pre-train our model with the aforementioned approaches and conduct experiments on downstream tasks. Experimental results demonstrate that Structure-CLIP achieves state-of-the-art (SOTA) performance on VG-Attribution and VG-Relation datasets, with 12.5% and 4.1% ahead of the multi-modal SOTA model respectively. Meanwhile, the results on MSCOCO indicate that Structure-CLIP significantly enhances the structured representations while maintaining the ability of general representations. Our code is available at https://github.com/zjukg/Structure-CLIP.
Jiji Tang, Zhuo Chen 0007, Zeng Zhao, Tangjie Lv, Zhipeng Hu, Wen Zhang 0015
AAAI2
2024 Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
abstract
Diffusion-based Image Editing (DIE) is an emerging research hot-spot, which often applies a semantic mask to control the target area for diffusion-based editing. However, most existing solutions obtain these masks via manual operations or off-line processing, greatly reducing their efficiency. In this paper, we propose a novel and efficient image editing method for Text-to-Image (T2I) diffusion models, termed Instant Diffusion Editing (InstDiffEdit). In particular, InstDiffEdit aims to employ the cross-modal attention ability of existing diffusion models to achieve instant mask guidance during the diffusion steps. To reduce the noise of attention maps and realize the full automatics, we equip InstDiffEdit with a training-free refinement scheme to adaptively aggregate the attention distributions for the automatic yet accurate mask generation. Meanwhile, to supplement the existing evaluations of DIE, we propose a new benchmark called Editing-Mask to examine the mask accuracy and local editing ability of existing methods. To validate InstDiffEdit, we also conduct extensive experiments on ImageNet and Imagen, and compare it with a bunch of the SOTA methods. The experimental results show that InstDiffEdit not only outperforms the SOTA methods in both image quality and editing results, but also has a much faster inference speed, i.e., +5 to +6 times. Our code available at https://anonymous.4open.science/r/InstDiffEdit-C306
Siyu Zou, Jiji Tang, Yiyi Zhou, Chaoyi Zhao, Zhipeng Hu, Xiaoshuai Sun
AAAI2
2023 Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D Representation
abstract
In recent years, 3D representation learning has turned to 2D vision-language pre-trained models to overcome data scarcity challenges. However, existing methods simply transfer 2D alignment strategies, aligning 3D representations with single-view 2D images and coarse-grained parent category text. These approaches introduce information degradation and insufficient synergy issues, leading to performance loss. Information degradation arises from overlooking the fact that a 3D representation should be equivalent to a series of multi-view images and more fine-grained subcategory text. Insufficient synergy neglects the idea that a robust 3D representation should align with the joint vision-language space, rather than independently aligning with each modality. In this paper, we propose a multi-view joint modality modeling approach, termed JM3D, to obtain a unified representation for point cloud, text, and image. Specifically, a novel Structured Multimodal Organizer (SMO) is proposed to address the information degradation issue, which introduces contiguous multi-view images and hierarchical text to enrich the representation of vision and language modalities. A Joint Multi-modal Alignment (JMA) is designed to tackle the insufficient synergy problem, which models the joint modality by incorporating language knowledge into the visual modality. Extensive experiments on ModelNet40 and ScanObjectNN demonstrate the effectiveness of our proposed method, JM3D, which achieves state-of-the-art performance in zero-shot 3D classification. JM3D outperforms ULIP by approximately 4.3% on PointMLP and achieves an improvement of up to 6.5% accuracy on PointNet++ in top-1 accuracy for zero-shot 3D classification on ModelNet40. The source code and trained models for all our experiments are publicly available at https://github.com/Mr-Neko/JM3D.
Haowei Wang 0001, Jiji Tang, Jiayi Ji, Xiaoshuai Sun, Minda Zhao, Lincheng Li, Zeng Zhao, Tangjie Lv, Rongrong Ji
ACM Multimedia2
2021 ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene Graphs
abstract
We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL tries to build the detailed semantic connections (objects, attributes of objects and relationships between objects) across vision and language, which are essential to vision-language cross-modal tasks. Utilizing scene graphs of visual scenes, ERNIE-ViL constructs Scene Graph Prediction tasks, i.e., Object Prediction, Attribute Prediction and Relationship Prediction tasks in the pre-training phase. Specifically, these prediction tasks are implemented by predicting nodes of different types in the scene graph parsed from the sentence. Thus, ERNIE-ViL can learn the joint representations characterizing the alignments of the detailed semantics across vision and language. After pre-training on large scale image-text aligned datasets, we validate the effectiveness of ERNIE-ViL on 5 cross-modal downstream tasks. ERNIE-ViL achieves state-of-the-art performances on all these tasks and ranks the first place on the VCR leaderboard with an absolute improvement of 3.7%.
Fei Yu 0010, Jiji Tang, Weichong Yin, Yu Sun 0029, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001
AAAI2