EDBT 2026 Demo / reviewers in the wild / expert
Ming Tao 0002
dblp:112/1-2
· DBLP profile ↗
14ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-4662-7170ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ColView: Consistent Text-Guided Grayscale Scene Colorization From Multi-View ImagesabstractThe colorization of scenes from multi-view grayscale images plays a crucial role in applications such as augmented reality and virtual exhibitions. Existing methods combine NeRF with an automatic colorization model, averaging multiple colorized patches to reduce inconsistency. However, they still face three key limitations: (1) Current methods cannot produce diverse colorization results due to the lack of multimodal conditional inputs, (2) They struggle to maintain multi-view consistency caused by unreliable geometric correspondence and ineffective propagation mechanisms, and (3) Computational inefficiency from NeRF's dense ray sampling and numerical integration. In this paper, we propose ColView, a unified framework for text-guided grayscale scene colorization that achieves both automatic and controllable colorization of grayscale scenes from multi-view grayscale images. First, for flexible color control, we leverage text description as the input to guide the colorization process, which allows users to specify desired colors through natural language descriptions. Second, to ensure multi-view consistency, we introduce a multi-view consistent colorization module that explicitly models dependencies between different views. This module follows three key steps: cross-view attention mechanism for collaborative key-view colorization, feature matching for inter-view correspondence establishment, and correspondence-guided feature propagation. Third, to improve computational efficiency, we adopt 3D Gaussian Splatting as our underlying representation. This explicit point-based representation renders significantly faster than NeRF. Extensive experimental results demonstrate that our method achieves superior visual quality and computational efficiency. Our code and models are publicly available athttps://github.com/ChchNiu/ColView. Chaochao Niu, Ming Tao 0002, Bing-Kun Bao, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2026 | AdaEdit: Adaptive Diffusion Model for Invisible Target Oriented Text-Conditioned Image EditingabstractText-conditioned image editing aims to modify a source image into a target image according to a specified text description, tackling two core challenges: locating target editing regions and ensuring consistency in non-target editing areas. Existing approaches utilize manual selection or cross-modal attention to define editing regions and deploy diffusion models to generate edited images. Despite these recent advancements, two problems remain. First, current methods fail to locate editing areas described in the text but invisible in the image. Second, they struggle to ensure spatial consistency in non-targeted regions due to the global noise addition along with excessive denoising during the diffusion process. To overcome these limitations, we propose AdaEdit, which comprises an adaptive mask localization module and an adaptive denoising strategy for text-conditioned image editing. AdaEdit can accurately identify the editing area via the measurement of cross-modal semantic mismatch, even when the visual details are not explicitly described in the text inputs. The adaptive denoising strategy applies varying noise levels to differentiate between targeted and non-targeted regions, enhancing the stability and consistency of the non-edited areas. Extensive experiments demonstrate that our proposed method achieves excellent performance on MS-COCO, MagicBrush, and Laion. We also expand our application to iterative editing tasks, thereby extending its utility for generalized editing scenarios. Yefei Sheng, Jie Wang 0061, Ming Tao 0002, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | CLGC: Continuous Layout Guidance for Consistent Text-to-Video EditingabstractText-to-Video (T2V) editing aims to produce temporally consistent videos aligned with text prompts, simultaneously reconstructing the original spatial structure. Existing methods rely on cross-attention maps generated from fixed text prompts, which lack sufficient spatial information, leading to inaccurate object positioning across frames. Additionally, existing methods only rely on the first or former frame to synthesize the current frame, offering limited viewpoint information and causing flickering artifacts. To address these issues, we propose CLGC, a training-free framework for continuous layout-guided T2V editing. First, we introduce semantic masks for continuous object position layout guidance, refining cross-attention maps and ensuring accurate object positioning across frames. Second, we adaptively integrate extra reference frames into the self-attention for the current frame synthesis, enhancing temporal consistency in the edited video. Finally, we integrate parallel null-text inversion to improve DDIM sampling, achieving accurate reconstruction results. Extensive experiments demonstrate that CLGC excels at attribute editing and shape transformation, confirming its effectiveness in T2V editing. Xuancheng Xu, Ming Tao 0002, Bing-Kun Bao |
ICME | 2 |
| 2025 | D2Gaussian: Dynamic Control with Discretized 3D View Modeling for Text-Driven 3D Gaussian Splatting EditingabstractCurrent advances in text-driven 3D scene editing tasks typically render the 3D representations into multi-view images and modify the images with the text instructions. Context consistency across multiple views and cross-modal consistency in the single-view are the keys to effective 3D editing. Accordingly, existing methods introduce additional image constraints and apply pre-trained 2D editing models. However, they fix the same text instruction across all views and freeze the pre-trained 2D model for single-view editing, leading to deficient modeling of 3D scene views and results in inconsistent generations with visual artifacts. To address these limitations, we introduce a discretized 3D view modeling method and a diffusion-based multi-view consistent editing pipeline for text-driven 3D gaussian splatting editing, abbreviated as D2Gaussian. Specifically, our approach constructs a codebook that encodes continuous 3D view information into discrete token embeddings to model the spatial feature expressions. Then, the token embeddings are proposed to guide and finetune the diffusion-based image editing model with the dynamic addition of control conditions, yielding a multi-view consistent editing pipeline. Finally, we introduce a 3D editing dataset generation approach along with a 3D-CLIP-SIM metric to form a benchmark, 3D-MagicBrush, to provide more diverse evaluation scenarios for future 3D editing works. Experiments demonstrate that our method achieves better visual results and multi-view consistency than previous state-of-the-art methods. Yefei Sheng, Jie Wang 0061, Ming Tao 0002, Bing-Kun Bao |
ACM Multimedia | 3 |
| 2025 | Chain-of-Cooking: Cooking Process Visualization via Bidirectional Chain-of-Thought GuidanceabstractCooking process visualization is a promising task in the intersection of image generation and food analysis, which aims to generate an image for each cooking step of a recipe. However, most existing works focus on generating images of finished foods based on the given recipes, and face two challenges in visualizing the cooking process. First, the appearance of ingredients changes variously across cooking steps, it is difficult to generate the correct appearances of foods that match the textual description, leading to semantic inconsistency. Second, the current step might depend on the operations of previous step, it is crucial to maintain the contextual coherence of images in sequential order. In this work, we present a cooking process visualization model, called Chain-of-Cooking. Specifically, to generate correct appearances of ingredients, we present a Dynamic Patch Selection Module to retrieve previously generated image patches as references, which are most related to current textual contents. Furthermore, to enhance the coherence and keep the rational order of generated images, we propose a Semantic Evolution Module and a Bidirectional Chain-of-Thought (CoT) Guidance. To better utilize the semantics of previous texts, the Semantic Evolution Module establishes the semantical association between latent prompts and current cooking step, and merges it with the latent features. Then the CoT Guidance updates the merged features to guide the current cooking step remain coherent with the previous step. Moreover, we construct a dataset named CookViz, consisting of intermediate image-text pairs for the cooking process. Quantitative and qualitative experiments show that our method outperforms existing methods in generating coherent and semantic consistent cooking process. Mengling Xu, Ming Tao 0002, Bing-Kun Bao |
ACM Multimedia | 2 |
| 2025 | SEMACOL: Semantic-enhanced multi-scale approach for text-guided grayscale image colorization
Chaochao Niu, Ming Tao 0002, Bing-Kun Bao |
Pattern Recognit. | 2 |
| 2025 | CookGALIP: Recipe Controllable Generative Adversarial CLIPs With Sequential Ingredient Prompts for Food Image GenerationabstractGenerating food images from recipes is a challenging task in food analysis, as recipes contain lengthy texts far beyond the semantic information in food images, making it difficult to align the features of two modalities. Existing studies usually concatenate the representations of ingredients and cooking instructions directly, and use the concatenated representations to generate food images through generative adversarial networks (GANs). However, previous models generally ignore the sequential information contained in complicated procedural instructions, which leads to semantic inconsistency between recipes and generated food images. Furthermore, it is still difficult for current models to distinguish and control fine-grained features, causing the entangled ingredient features in food images. To this end, we propose CookGALIP, which strengthens semantic consistency and controllability for food image generation. Based on the recently proposed text-to-image framework GALIP, two modules are specially designed: 1) To incorporate the sequential relationships into the food image generation process, we propose a Recipe Fusion Module (RFM) to fuse the semantics of cooking instructions, so as to balance the semantic complexity between modalities and improve the semantic consistency of recipes and generated food images. 2) To distinguish and control the fine-grained ingredient features, we introduce the Ingredient Control Module (ICM) to generate sequential ingredient prompts, which enables more refined control over the recipe-to-food synthesis process. Experimental results on Recipe1M and Vireo Food-172 datasets show that the proposed model outperforms the state-of-the-art methods. Mengling Xu, Jie Wang 0061, Ming Tao 0002, Bing-Kun Bao, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2024 | StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
Ming Tao 0002, Bing-Kun Bao, Hao Tang 0005, Yaowei Wang 0001, Changsheng Xu |
ECCV (56) | 1 |
| 2024 | CoIn: A Lightweight and Effective Framework for Story Visualization and ContinuationabstractStory visualization aims to generate realistic and coherent images based on multi-sentence stories. However, current methods face challenges in achieving high-quality image generation while maintaining lightweight models and a fast generation speed. The main issue lies in the two existing frameworks. The independent framework prioritizes speed but sacrifices image quality with the non-collaborative image generation process and basic GAN-based learning. The autoregressive framework modifies the large pretrained text-to-image model in an auto-regressive manner with additional history modules, leading to large model size, resource-intensive requirements, and slow generation speed. To address these issues, we propose a lightweight and effective framework, namely CoIn. Specifically, we introduce a Context-aware Story Generator to predict shared context semantics for each image generator. Additionally, we propose an Intra-Story Interchange module that allows each image generator to exchange visual information with other image generators. Furthermore, we incorporate DINOv2 into the story and image discriminators to assess the story image quality more accurately. Extensive experiments show that our CoIn keeps the model size and generation speed of the independent framework, while achieving promising story image quality. Ming Tao 0002, Bing-Kun Bao, Hao Tang 0005, Yaowei Wang 0001, Changsheng Xu |
ACM Multimedia | 1 |
| 2024 | ISF-GAN: Imagine, Select, and Fuse with GPT-Based Text Enrichment for Text-to-Image SynthesisabstractText-to-Image synthesis aims to generate an accurate and semantically consistent image from a given text description. However, it is difficult for existing generative methods to generate semantically complete images from a single piece of text. Some works try to expand the input text to multiple captions via retrieving similar descriptions of the input text from the training set but still fail to fill in missing image semantics. In this article, we propose a GAN-based approach to Imagine, Select, and Fuse for Text-to-image synthesis, named ISF-GAN. The proposed ISF-GAN contains Imagine Stage and Select and Fuse Stage to solve the above problems. First, the Imagine Stage proposes a text completion and enrichment module. This module guides a GPT-based model to enrich the text expression beyond the original dataset. Second, the Select and Fuse Stage selects qualified text descriptions and then introduces a cross-modal attentional mechanism to interact these different sentence embeddings with the image features at different scales. In short, our proposed model enriches the input text information for completing missing semantics and introduces a cross-modal attentional mechanism to maximize the utilization of enriched text information to generate semantically consistent images. Experimental results on CUB, Oxford-102, and CelebA-HQ datasets prove the effectiveness and superiority of the proposed network. Code is available at https://github.com/Feilingg/ISF-GAN Yefei Sheng, Ming Tao 0002, Jie Wang 0061, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | DE-net: Dynamic Text-Guided Image Editing Adversarial NetworksabstractText-guided image editing models have shown remarkable results. However, there remain two problems. First, they employ fixed manipulation modules for various editing requirements (e.g., color changing, texture changing, content adding and removing), which results in over-editing or insufficient editing. Second, they do not clearly distinguish between text-required and text-irrelevant parts, which leads to inaccurate editing. To solve these limitations, we propose: (i) a Dynamic Editing Block (DEBlock) that composes different editing modules dynamically for various editing requirements. (ii) a Composition Predictor (Comp-Pred), which predicts the composition weights for DEBlock according to the inference on target texts and source images. (iii) a Dynamic text-adaptive Convolution Block (DCBlock) that queries source image features to distinguish text-required parts and text-irrelevant parts. Extensive experiments demonstrate that our DE-Net achieves excellent performance and manipulates source images more correctly and accurately. Ming Tao 0002, Bing-Kun Bao, Hao Tang 0005, Fei Wu 0004, Longhui Wei, Qi Tian 0001 |
AAAI | 1 |
| 2023 | GALIP: Generative Adversarial CLIPs for Text-to-Image SynthesisabstractSynthesizing high-fidelity complex images from text is challenging. Based on large pretraining, the autoregressive and diffusion models can synthesize photo-realistic images. Although these large models have shown notable progress, there remain three flaws. 1) These models require tremendous training data and parameters to achieve good performance. 2) The multi-step generation design slows the image synthesis process heavily. 3) The synthesized visual features are challenging to control and require delicately designed prompts. To enable high-quality, efficient, fast, and controllable text-to-image synthesis, we propose Generative Adversarial CLIPs, namely GALIP. GALIP leverages the powerful pretrained CLIP model both in the discriminator and generator. Specifically, we propose a CLIP-based discriminator. The complex scene understanding ability of CLIP enables the discriminator to accurately assess the image quality. Furthermore, we propose a CLIP-empowered generator that induces the visual concepts from CLIP through bridge features and prompts. The CLIP-integrated generator and discriminator boost training efficiency, and as a result, our model only requires about 3% training data and 6% learnable parameters, achieving comparable results to large pretrained autoregressive and diffusion models. Moreover, our model achieves ~120×faster synthesis speed and inherits the smooth latent space from GAN. The extensive experimental results demonstrate the excellent performance of our GALIP. Code is available at https://github.com/tobran/GALIP. Ming Tao 0002, Bing-Kun Bao, Hao Tang 0005, Changsheng Xu |
CVPR | 1 |
| 2022 | DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisabstractSynthesizing high-quality realistic images from text descriptions is a challenging task. Existing text-to-image Generative Adversarial Networks generally employ a stacked architecture as the backbone yet still remain three flaws. First, the stacked architecture introduces the entanglements between generators of different image scales. Second, existing studies prefer to apply and fix extra networks in adversarial learning for text-image semantic consistency, which limits the supervision capability of these networks. Third, the cross-modal attention-based text-image fusion that widely adopted by previous works is limited on several special image scales because of the computational cost. To these ends, we propose a simpler but more effective Deep Fusion Generative Adversarial Networks (DF-GAN). To be specific, we propose: (i) a novel one-stage text-to-image backbone that directly synthesizes high-resolution images without entanglements between different generators, (ii) a novel Target-Aware Discriminator composed of Matching-Aware Gradient Penalty and One-Way Output, which enhances the text-image semantic consistency without introducing extra networks, (iii) a novel deep text-image fusion block, which deepens the fusion process to make a full fusion between text and visual features. Compared with current state-of-the-art methods, our proposed DF-GAN is simpler but more efficient to synthesize realistic and text-matching images and achieves better performance on widely used datasets. Code is available at https://github.com/tobran/DF-GAN. Ming Tao 0002, Hao Tang 0005, Fei Wu 0004, Xiaoyuan Jing, Bing-Kun Bao, Changsheng Xu |
CVPR | 1 |
| 2020 | SiENet: Siamese Expansion Network for Image ExtrapolationabstractDifferent from image inpainting, image outpainting has relatively less context in the image center to capture and more content at the image border to predict. Therefore, classical encoder-decoder pipeline of existing methods may not predict the outstretched unknown content perfectly. In this paper, a novel two-stage siamese adversarial model for image extrapolation, named Siamese Expansion Network (SiENet) is proposed. Specifically, in two stages, a novel border sensitive convolution named adaptive filling convolution is designed for allowing encoder to predict the unknown content, alleviating the burden of decoder. Besides, to introduce prior knowledge to network and reinforce the inferring ability of encoder, siamese adversarial mechanism is designed to enable our network to model the distribution of covered long range feature as that of uncovered image feature. The results on four datasets has demonstrated that our method outperforms existing state-of-the-arts and could produce realistic results. Our code is released on https://github.com/nanjingxiaobawang/SieNet-Image-extrapolation. Feng Chen 0047, Cailing Wang, Ming Tao 0002, Guoping Jiang |
IEEE Signal Process. Lett. | 4 |