VLDB 2026 Research / reviewers in the wild / expert
Jie Wang 0061
dblp:29/5259-61
· DBLP profile ↗
10ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-8662-9488ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaEdit: Adaptive Diffusion Model for Invisible Target Oriented Text-Conditioned Image EditingabstractText-conditioned image editing aims to modify a source image into a target image according to a specified text description, tackling two core challenges: locating target editing regions and ensuring consistency in non-target editing areas. Existing approaches utilize manual selection or cross-modal attention to define editing regions and deploy diffusion models to generate edited images. Despite these recent advancements, two problems remain. First, current methods fail to locate editing areas described in the text but invisible in the image. Second, they struggle to ensure spatial consistency in non-targeted regions due to the global noise addition along with excessive denoising during the diffusion process. To overcome these limitations, we propose AdaEdit, which comprises an adaptive mask localization module and an adaptive denoising strategy for text-conditioned image editing. AdaEdit can accurately identify the editing area via the measurement of cross-modal semantic mismatch, even when the visual details are not explicitly described in the text inputs. The adaptive denoising strategy applies varying noise levels to differentiate between targeted and non-targeted regions, enhancing the stability and consistency of the non-edited areas. Extensive experiments demonstrate that our proposed method achieves excellent performance on MS-COCO, MagicBrush, and Laion. We also expand our application to iterative editing tasks, thereby extending its utility for generalized editing scenarios. Yefei Sheng, Jie Wang 0061, Ming Tao 0002, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | InstantPainting: Expanding GANs for Efficient Text-Conditioned Image Generation PlatformabstractText-conditioned image generation enables cross-modal comprehension. Recent emergence of many platforms have found applications in diverse domains like assisted designing and video gaming. However, there still exist challenges in existing platforms due to their expensive training and time-consuming generation processes. In this paper, we introduce an efficient text-conditioned image generation platform, termed InstantPainting. Unlike existing platforms based on large-scale pre-trained diffusion models, InstantPainting expands generative adversarial networks (GANs) to achieve efficient generation by using only about three percent pre-training data of other platforms. Compared to existing platforms, InstantPainting achieves the following functions at a very low deployment cost and approximately 4 to 5 times faster generation speeds: (1) Multi-category and multi-size image generation (2) Image stylization and controlled generation (3) Creative generation, including the generation of poetry pictures and counterfactual images. The proposed platform provides web application implementations for PC and mobile, users can create high-quality images directly through the user interface. Bing-Kun Bao, Yefei Sheng, Jie Wang 0061, Sisi You |
AAAI | 3 |
| 2025 | D2Gaussian: Dynamic Control with Discretized 3D View Modeling for Text-Driven 3D Gaussian Splatting EditingabstractCurrent advances in text-driven 3D scene editing tasks typically render the 3D representations into multi-view images and modify the images with the text instructions. Context consistency across multiple views and cross-modal consistency in the single-view are the keys to effective 3D editing. Accordingly, existing methods introduce additional image constraints and apply pre-trained 2D editing models. However, they fix the same text instruction across all views and freeze the pre-trained 2D model for single-view editing, leading to deficient modeling of 3D scene views and results in inconsistent generations with visual artifacts. To address these limitations, we introduce a discretized 3D view modeling method and a diffusion-based multi-view consistent editing pipeline for text-driven 3D gaussian splatting editing, abbreviated as D2Gaussian. Specifically, our approach constructs a codebook that encodes continuous 3D view information into discrete token embeddings to model the spatial feature expressions. Then, the token embeddings are proposed to guide and finetune the diffusion-based image editing model with the dynamic addition of control conditions, yielding a multi-view consistent editing pipeline. Finally, we introduce a 3D editing dataset generation approach along with a 3D-CLIP-SIM metric to form a benchmark, 3D-MagicBrush, to provide more diverse evaluation scenarios for future 3D editing works. Experiments demonstrate that our method achieves better visual results and multi-view consistency than previous state-of-the-art methods. Yefei Sheng, Jie Wang 0061, Ming Tao 0002, Bing-Kun Bao |
ACM Multimedia | 2 |
| 2025 | CookGALIP: Recipe Controllable Generative Adversarial CLIPs With Sequential Ingredient Prompts for Food Image GenerationabstractGenerating food images from recipes is a challenging task in food analysis, as recipes contain lengthy texts far beyond the semantic information in food images, making it difficult to align the features of two modalities. Existing studies usually concatenate the representations of ingredients and cooking instructions directly, and use the concatenated representations to generate food images through generative adversarial networks (GANs). However, previous models generally ignore the sequential information contained in complicated procedural instructions, which leads to semantic inconsistency between recipes and generated food images. Furthermore, it is still difficult for current models to distinguish and control fine-grained features, causing the entangled ingredient features in food images. To this end, we propose CookGALIP, which strengthens semantic consistency and controllability for food image generation. Based on the recently proposed text-to-image framework GALIP, two modules are specially designed: 1) To incorporate the sequential relationships into the food image generation process, we propose a Recipe Fusion Module (RFM) to fuse the semantics of cooking instructions, so as to balance the semantic complexity between modalities and improve the semantic consistency of recipes and generated food images. 2) To distinguish and control the fine-grained ingredient features, we introduce the Ingredient Control Module (ICM) to generate sequential ingredient prompts, which enables more refined control over the recipe-to-food synthesis process. Experimental results on Recipe1M and Vireo Food-172 datasets show that the proposed model outperforms the state-of-the-art methods. Mengling Xu, Jie Wang 0061, Ming Tao 0002, Bing-Kun Bao, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2024 | ISF-GAN: Imagine, Select, and Fuse with GPT-Based Text Enrichment for Text-to-Image SynthesisabstractText-to-Image synthesis aims to generate an accurate and semantically consistent image from a given text description. However, it is difficult for existing generative methods to generate semantically complete images from a single piece of text. Some works try to expand the input text to multiple captions via retrieving similar descriptions of the input text from the training set but still fail to fill in missing image semantics. In this article, we propose a GAN-based approach to Imagine, Select, and Fuse for Text-to-image synthesis, named ISF-GAN. The proposed ISF-GAN contains Imagine Stage and Select and Fuse Stage to solve the above problems. First, the Imagine Stage proposes a text completion and enrichment module. This module guides a GPT-based model to enrich the text expression beyond the original dataset. Second, the Select and Fuse Stage selects qualified text descriptions and then introduces a cross-modal attentional mechanism to interact these different sentence embeddings with the image features at different scales. In short, our proposed model enriches the input text information for completing missing semantics and introduces a cross-modal attentional mechanism to maximize the utilization of enriched text information to generate semantically consistent images. Experimental results on CUB, Oxford-102, and CelebA-HQ datasets prove the effectiveness and superiority of the proposed network. Code is available at https://github.com/Feilingg/ISF-GAN Yefei Sheng, Ming Tao 0002, Jie Wang 0061, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | 3D GAN image synthesis and dataset quality assessment for bacterial biofilmabstractMOTIVATION: Data-driven deep learning techniques usually require a large quantity of labeled training data to achieve reliable solutions in bioimage analysis. However, noisy image conditions and high cell density in bacterial biofilm images make 3D cell annotations difficult to obtain. Alternatively, data augmentation via synthetic data generation is attempted, but current methods fail to produce realistic images. RESULTS: This article presents a bioimage synthesis and assessment workflow with application to augment bacterial biofilm images. 3D cyclic generative adversarial networks (GAN) with unbalanced cycle consistency loss functions are exploited in order to synthesize 3D biofilm images from binary cell labels. Then, a stochastic synthetic dataset quality assessment (SSQA) measure that compares statistical appearance similarity between random patches from random images in two datasets is proposed. Both SSQA scores and other existing image quality measures indicate that the proposed 3D Cyclic GAN, along with the unbalanced loss function, provides a reliably realistic (as measured by mean opinion score) 3D synthetic biofilm image. In 3D cell segmentation experiments, a GAN-augmented training model also presents more realistic signal-to-background intensity ratio and improved cell counting accuracy. AVAILABILITY AND IMPLEMENTATION: https://github.com/jwang-c/DeepBiofilm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jie Wang 0061, Nazia Tabassum, Tanjin Taher Toma, Yibo Wang 0013, Andreas Gahlmann, Scott T. Acton |
Bioinform. | 1 |
| 2021 | Graph-Theoretic Post-Processing of Segmentation With Application to Dense BiofilmsabstractRecent deep learning methods have provided successful initial segmentation results for generalized cell segmentation in microscopy. However, for dense arrangements of small cells with limited ground truth for training, the deep learning methods produce both over-segmentation and under-segmentation errors. Post-processing attempts to balance the trade-off between the global goal of cell counting for instance segmentation, and local fidelity to the morphology of identified cells. The need for post-processing is especially evident for segmenting 3D bacterial cells in densely-packed communities called biofilms. A graph-based recursive clustering approach, m-LCuts, is proposed to automatically detect collinearly structured clusters and applied to post-process unsolved cells in 3D bacterial biofilm segmentation. Construction of outlier-removed graphs to extract the collinearity feature in the data adds additional novelty to m-LCuts. The superiority of m-LCuts is observed by the evaluation in cell counting with over 90% of cells correctly identified, while a lower bound of 0.8 in terms of average single-cell segmentation accuracy is maintained. This proposed method does not need manual specification of the number of cells to be segmented. Furthermore, the broad adaptation for working on various applications, with the presence of data collinearity, also makes m-LCuts stand out from the other approaches. Jie Wang 0061, Ji Zhang 0023, Yibo Wang 0013, Andreas Gahlmann, Scott T. Acton |
IEEE Trans. Image Process. | 1 |
| 2019 | Lcuts: Linear Clustering of Bacteria Using Recursive Graph CutsabstractBacterial biofilm segmentation poses significant challenges due to lack of apparent structure, poor imaging resolution, limited contrast between conterminous cells and high density of cells that overlap. Although there exist bacterial segmentation algorithms in the existing art, they fail to delineate cells in dense biofilms, especially in 3D imaging scenarios in which the cells are growing and subdividing in a complex manner. A graph-based data clustering method, £Cuts, is presented with the application on bacterial cell segmentation. By constructing a weighted graph with node features in locations and principal orientations, the proposed method can automatically classify and detect differently oriented aggregations of linear structures (represent by bacteria in the application). The method assists in the assessment of several facets, such as bacterium tracking, cluster growth, and mapping of migration patterns of bacterial biofilms. Quantitative and qualitative measures for 2D data demonstrate the superiority of proposed method over the state of the art. Preliminary 3D results exhibit reliable classification of the cells with 97% accuracy. Jie Wang 0061, Tamal Batabyal, Ji Zhang 0023, Arslan Aziz, Andreas Gahlmann, Scott T. Acton |
ICIP | 1 |
| 2018 | Nonlinear Shape Regression for Filtering Segmentation Results from Calcium ImagingabstractA shape filter is presented to repair segmentation results obtained in calcium imaging of neurons in vivo. This post-segmentation algorithm can automatically smooth the shapes obtained from a preliminary segmentation, while precluding the cases where two neurons are counted as one combined component. The shape filter is realized using a square-root velocity to project the shapes on a shape manifold in which distances between shapes are based on elastic changes. Two data-driven weighting methods are proposed to achieve a trade-off between shape smoothness and consistency with the data. Intuitive comparisons of proposed methods via projection onto Cartesian maps demonstrate the smoothing ability of the shape filter. Quantitative measures also prove the superiority of our methods over models that do not employ any weighting criterion. Jie Wang 0061, Zhongxiao Fu, Nasrin Sadeghzadehyazdi, Jonathan Kipnis, Scott T. Acton |
ICIP | 1 |
| 2017 | Bact-3D: A level set segmentation approach for dense multi-layered 3D bacterial biofilmsabstractIn microscopy, new super-resolution methods are emerging that produce three-dimensional images at resolutions ten times smaller than that provided by traditional light microscopy. Such technology is enabling the exploration of structure and function in living tissues such as bacterial biofilms that have mysterious interconnections and organization. Unfortunately, the standard tools used in the image analysis community to perform segmentation and other higher-level analyses cannot be applied naïvely to these data. This paper presents Bact-3D, a 3D method for segmenting super-resolution images of multi-leveled, living bacteria cultured in vitro. The method incorporates a novel initialization approach that exploits the geometry of the bacterial cells as well an iterative local level set evolution that is tailored to the biological application. In experiments where segmentation is used as precursor to cell detection, the Bact-3D matches or improves upon the Dice score and mean-squared error of two existing methods, while yielding a substantial improvement in cell detection accuracy. In addition to providing improvements in performance over the state-of-the-art, this report also characterizes the tradeoff between imaging resolution and segmentation quality. Jie Wang 0061, Rituparna Sarkar, Arslan Aziz, Andrea Vaccari, Andreas Gahlmann, Scott T. Acton |
ICIP | 1 |