EDBT 2026 Demo / reviewers in the wild / expert
Sitong Su
dblp:299/4584
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0001-7406-116XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SeMv-3D: Toward Concurrency of Semantic and Multi-View Consistency in General Text-to-3D GenerationabstractGeneral Text-to-3D (GT23D) generation is crucial for creating diverse 3D content across objects and scenes, yet it faces two key challenges: 1) ensuring semantic consistency between input text and generated 3D models, and 2) maintaining multi-view consistency across different perspectives within 3D. Existing approaches typically address only one of these challenges, often leading to suboptimal results in semantic fidelity and structural coherence. To overcome these limitations, we propose SeMv-3D, a novel framework that jointly enhances semantic alignment and multi-view consistency in GT23D generation. At its core, we introduce Triplane Prior Learning (TPL), which effectively learns triplane priors by capturing spatial correspondences across three orthogonal planes using a dedicated Orthogonal Attention mechanism, thereby ensuring geometric consistency across viewpoints. Additionally, we present Prior-based Semantic Aligning in Triplanes (SAT), which enables consistent any-view synthesis by leveraging attention-based feature alignment to reinforce the correspondence between textual semantics and triplane representations. Extensive experiments demonstrate that our method sets a new state-of-the-art in multi-view consistency, while maintaining competitive performance in semantic consistency compared to methods focused solely on semantic alignment. These results emphasize the remarkable ability of our approach to effectively balance and excel in both dimensions, establishing a new benchmark in the field. Pengpeng Zeng, Lianli Gao, Sitong Su, Heng Tao Shen, Jingkuan Song |
IEEE Trans. Image Process. | 4 |
| 2024 | F³-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video SynthesisabstractRecently Text-to-Video (T2V) synthesis has undergone a breakthrough by training transformers or diffusion models on large-scale datasets. Nevertheless, inferring such large models incurs huge costs. Previous inference acceleration works either require costly retraining or are model-specific. To address this issue, instead of retraining we explore the inference process of two mainstream T2V models using transformers and diffusion models. The exploration reveals the redundancy in temporal attention modules of both models, which are commonly utilized to establish temporal relations among frames. Consequently, we propose a training-free and generalized pruning strategy called F3-Pruning to prune redundant temporal attention weights. Specifically, when aggregate temporal attention values are ranked below a certain ratio, corresponding weights will be pruned. Extensive experiments on three datasets using a classic transformer-based model CogVideo and a typical diffusion-based model Tune-A-Video verify the effectiveness of F3-Pruning in inference acceleration, quality assurance and broad applicability. Sitong Su, Jianzhi Liu, Lianli Gao, Jingkuan Song |
AAAI | 1 |
| 2024 | Training-Free Semantic Video Composition via Pre-trained Diffusion ModelabstractThe video composition task aims to integrate specified foregrounds and backgrounds from different videos into a harmonious composite. Current approaches, predominantly trained on videos with adjusted foreground color and lighting, struggle to address deep semantic disparities beyond superficial adjustments, such as domain gaps. Therefore, we propose a training-free pipeline employing a pre-trained diffusion model imbued with semantic prior knowledge, which can process composite videos with broader semantic disparities. Specifically, we process the video frames in a cascading manner and handle each frame in two processes with the diffusion model. In the inversion process, we propose Balanced Partial Inversion to obtain generation initial points that balance reversibility and modifiability. Then, in the generation process, we further propose Inter-Frame Augmented attention to augment foreground continuity across frames. Experimental results reveal that our pipeline successfully ensures the visual harmony and inter-frame coherence of the outputs, demonstrating efficacy in managing broader semantic disparities. Sitong Su, Junchen Zhu, Lianli Gao, Jingkuan Song |
ICME | 2 |
| 2024 | Allowing Supervision in Unsupervised Deformable- Instances Image-to-Image TranslationabstractReplacing objects in images is a practical functionality of Photoshop, e.g., clothes changing. This task is defined as Unsupervised Deformable-Instances Image-to-Image Translation (UDIT), which maps multiple foreground instances of a source domain to a target domain, involving significant changes in shape. Although previous works incorporate instance masks of source domain for instance shape indication, their translation still fails in shape because of inadequate utilization of shape information in masks. To mitigate this issue, we introduce an effective two-stage pipeline for UDIT called Mask-Guided Deformable-instances GAN++ (MGD-GAN++), which generates target masks in the first stage named Mask Morph and utilizes the masks to guide the synthesis of corresponding instances in the second stage named Mask-Guided Image Generation. To further provide sufficient supervision with existing unpaired datasets, an overall set of training schemes is proposed for the two stages of MGD-GAN++, coined as Aligned Supervision and Inpainting Supervision, respectively. Extensive experiments on four datasets demonstrate the significant advantages of our MGD-GAN++ over existing methods both quantitatively and qualitatively. Furthermore, our training time consumption is hugely reduced compared to the state-of-the-art. Yu Liu 0076, Sitong Su, Junchen Zhu, Feng Zheng 0001, Lianli Gao, Jingkuan Song |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Utilizing Greedy Nature for Multimodal Conditional Image Synthesis in TransformersabstractMultimodal Conditional Image Synthesis(MCIS) aims to generate images according to different modalities input and their combination, which allows users to describe their requirements in complementary ways, e.g. segmentation for shapes and text for attributes. Despite satisfying results in MCIS, a non-trivial issue is neglected. Some modalities are fully optimized and dominate the generation, while other modalities are sub-optimized and fail to contribute their complementary information. We coin this phenomenon as Modality Bias.Our analysis reveals that generative models own greedy nature. Specifically, the modality that shares less semantic gap with the synthesized modality will be greedily incorporated and thus takes a larger proportion in synthesis. The main idea of previous works in Modality Bias is to punish the greedy nature, which hurts the performance of dominant modalities and impedes their contribution to multimodal synthesis. Instead, we propose to utilize the greedy nature by setting dominant modalities as guidance for sub-optimized modalities through coordinated feature space, named Coordinated Knowledge Mining. Afterwards, improved uni-modalities are aggregated by fusing coordinated features to further boost the performance of multimodal image synthesis, called Coordinated Knowledge Fusion. Extensive experiments prove that our method not only increases uni-modal performance by a large margin, but also promotes multimodal image synthesis by fully utilizing complementary information from different modalities. Sitong Su, Junchen Zhu, Lianli Gao, Jingkuan Song |
IEEE Trans. Multim. | 1 |
| 2023 | From External to Internal: Structuring Image for Text-to-Image Attributes ManipulationabstractManipulating visual attributes of an image through a natural language description, known as text-to-image attributes manipulation (T2AM), is a challenging task. However, existing approaches tend to search the whole image to manipulate the target instance indicated by a description, thus they often fail to locate and manipulate the accurate text-relevant regions, and even disturb the text-irrelevant contents, e.g. texture and background. Meanwhile, the model efficiency needs to be improved. To tackle the above issues, we introduce a novel yet simple GAN-based approach, namelyStructuringImage forManipulating(SIMGAN), to narrow down the optimization areas from external to internal. It consists of two major components: 1)External Structuring(ExST), a pretrained segmentation network, for recognizing and separating the target instances and background from an image; and 2)Internal Structuring(InST) for seeking out and editing the text-relevant attributes of the target instances based on the given description and masked hierarchical image representations from ExST. Specifically, the InST structures target instances from outline to detail by firstly drawing the sketch and colors underpainting of instances with anOutline-Oriented Structuring(OuST), and then enhancing the text-relevant attributes and elaborating on details with aDetail-Oriented Structuring(DeST). Extensive experiments on benchmark datasets demonstrate that our framework significantly outperforms state-of-the-art both quantitatively and qualitatively. Compared with the state-of-the-art method ManiGAN, our approach reduces the training time by 88%, while the inferring time is three times faster. In addition, our approach is easily extended to solve the instance-level image-to-image translation problem, and the results exhibit the versatility and effectiveness of our approach. This code is released inhttps://github.com/qikizh/SIMGAN. Lianli Gao, Qike Zhao, Junchen Zhu, Sitong Su, Lechao Cheng, Lei Zhao 0017 |
IEEE Trans. Multim. | 4 |
| 2021 | Towards Unsupervised Deformable-Instances Image-to-Image TranslationabstractReplacing objects in images is a practical functionality of Photoshop, e.g., clothes changing. This task is defined as Unsupervised Deformable-Instances Image-to-Image Translation (UDIT), which maps multiple foreground instances of a source domain to a target domain, involving significant changes in shape. In this paper, we propose an effective pipeline named Mask-Guided Deformable-instances GAN (MGD-GAN) which first generates target masks in batch and then utilizes them to synthesize corresponding instances on the background image, with all instances efficiently translated and background well preserved. To promote the quality of synthesized images and stabilize the training, we design an elegant training procedure which transforms the unsupervised mask-to-instance process into a supervised way by creating paired examples. To objectively evaluate the performance of UDIT task, we design new evaluation metrics which are based on the object detection. Extensive experiments on four datasets demonstrate the significant advantages of our MGD-GAN over existing methods both quantitatively and qualitatively. Furthermore, our training time consumption is hugely reduced compared to the state-of-the-art. The code could be available at https://github.com/sitongsu/MGD_GAN. Sitong Su, Jingkuan Song, Lianli Gao, Junchen Zhu |
IJCAI | 1 |
| 2021 | Fully Functional Image Manipulation Using Scene Graphs in A Bounding-Box Free WayabstractRecently, performing semantic editing of an image by modifying a scene graph has been proposed to support high-level image manipulation, and plays an important role for image generation. However, existing methods are all based on bounding boxes, and they suffer from the bounding box constraint. First, a bounding box often involves other instances (e.g, objects or environments) which do not need to be modified, but existing methods manipulate all the contents included in the bounding box. Secondly, prior methods fail to support adding instances when the bounding box of the target instance cannot be provided. To address the two issues above, we propose a novel bounding box free approach, which consists of two parts: a Local Bounding Box Free (Local-BBox-Free) Mask Generation and a Global Bounding Box Free (Global-BBox-Free) Instance Generation. The first part relieves the model of reliance on bounding boxes by generating the mask of the target instance to be manipulated without using the target instance bounding box. This enables our method to be the first to support fully functional image manipulation using scene graphs, including adding, removing, replacing and repositing instances. The second part is designed to synthesize the target instance directly from the generated mask and then paste it back to the inpainted original image using the generated mask, which preserves the unchanged part to the largest extent and precisely controls the target instance generation. Extensive experiments on Visual Genome and COCO-Stuff demonstrate that our model significantly surpasses the state-of-the-art both quantitatively and qualitatively. Sitong Su, Lianli Gao, Junchen Zhu, Jie Shao 0001, Jingkuan Song |
ACM Multimedia | 1 |