Zhida Feng

dblp:297/8802 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
0009-0004-6348-9621ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Simplifying Control Mechanism in Text-to-Image Diffusion Models
abstract
ControlNet has significantly advanced controllable image generation by integrating dense conditions (such as depth and canny edges) with text-to-image diffusion models. However, ControlNet's integration requires an additional amount nearly equal to half of the base diffusion model's parameters, making it inefficient. To address this, we introduce Simple-ControlNet, an efficient and streamlined network for controllable text-to-image generation. It employs a single-scale projection layer to incorporate condition information into the denoising U-Net. It is supplemented by Low-Rank Adapter (LoRA) parameters to facilitate condition learning. Impressively, Simple-ControlNet requires fewer than 3 million parameters for the control mechanism, substantially less than the 300 million needed by ControlNet. Our extensive experiments confirm that Simple-ControlNet matches and surpasses ControlNet's performance across a broad range of tasks and base diffusion models, showcasing its utility and efficiency.
Zhida Feng, Li Chen 0011, Yuenan Sun, Jiaxiang Liu 0004, Shikun Feng
AAAI1
2025 Fine-tuned Multimodal Large Language Models are Zero-shot Learners in Image Quality Assessment
abstract
Image quality assessment (IQA) has traditionally relied on task-specific models, often limiting their adaptability and generalization to diverse image content. To this end, this paper introduces LV-IQA, a fine-tuned Multimodal Large Language Model (MLLM) via Visual Grounding, demonstrating zero-shot learning in IQA. The key contributions of this paper involve the proposal of a cross-modal chain of thought rooted in hierarchical semantics and quality levels. Additionally, To enhance its adaptability, the paper incorporates a visual grounding prompt generated through segmentation masks and bounding boxes to learn the correspondence between original images and their hierarchical semantics. By observing the original images and their corresponding visual grounding information, LV-IQA is empowered to employ a sophisticated chain of thought, involving the perception and comprehension of local details and area structures, to infer the quality of the images. Through experiments, LV-IQA demonstrates reliable zero-shot capabilities in IQA.
Zhida Feng, Shikun Feng
ICME3
2025 Extending Pretrained Diffusion Models for Medical Image-Label Generation
abstract
Medical image segmentation often suffers from the limitation of small-scale annotated datasets. To address this challenge, we propose Paired Diffusion Models (PDM), a framework that adapts large-scale pre-trained diffusion models (e.g., Stable Diffusion) to generate paired image-label data. By modifying the neural network’s architecture to accept and output both images and their pixel-level annotations simultaneously, PDM enables a single model to generate an image and its corresponding binary segmentation mask. Furthermore, we introduce a mask decoder that directly decodes binary masks from the latent space, bypassing the need for an intermediate RGB-to-binary conversion step. By carefully initializing additional parameters, PDM stabilizes training and enhances the accuracy of the generated segmentation masks. Experimental results demonstrate that pretraining with images generated by PDM significantly boosts the performance of medical image segmentation models, especially in scenarios with limited annotated data.
Yuenan Sun, Li Chen 0011, Zhida Feng
MMAsia3
2024 Named Entity Driven Zero-Shot Image Manipulation
abstract
We introduced StyleEntity, a zero-shot image manipulation model that utilizes named entities as proxies during its training phase. This strategy enables our model to manipulate images using unseen textual descriptions during inference, all within a single training phase. Additionally, we proposed an inference technique termed Prompt Ensemble Latent Averaging (PELA). PELA averages the manipulation directions derived from various named entities during inference, effectively eliminating the noise directions, thus achieving stable manipulation. In our experiments, StyleEntity exhibited superior performance in a zero-shot setting compared to other methods. The code, model weights, and datasets are available at https://github.com/feng-zhida/StyleEntity.
Zhida Feng, Li Chen 0011, Jing Tian 0002, Jiaxiang Liu 0004, Shikun Feng
CVPR1
2024 Unleashing Fine-Coarse Curve Perception Via Trunk-Branch Perturbation
abstract
Segmenting intricate curve structures like retinal blood vessels, encompassing both fine and coarse details, remains a significant challenge. This work proposes a novel module that divides these complex curve structures into trunks and branches, fuses the input as auxiliary information, and optimizes the breakpoints of different curve parts through losses. In order to balance the redundant information that may lead to model overfitting, a unique feature perturbation strategy is introduced after the backbone decoding process to enhance the model’s robustness to complex curve structure segmentation tasks. Experiments show that this method can effectively distinguish different topological structures of blood vessels and maintain high segmentation accuracy even at blood vessel intersections or breakpoints, which holds immense potential for diverse future applications in image segmentation.
Yunxiang Cao, Li Chen 0011, Zhida Feng, Xiaoming Liu 0004
ICIP4
2024 B-Walk: Bernoulli Principle Guided Biased Random Walk for Curve Connection
abstract
In the segmentation of curve structures, the discontinuity may lead to an incomplete topological representation of the structure. The current methods for reconnecting curve structures lack physical explanations and are limited in their effectiveness in image processing. In order to address these constraints, a new algorithm for reconnecting curve structures in segmentation is proposed by combining fluid mechanics principles, especially Bernoulli’s principle and random walks. This algorithm calculates the similarity of fracture curves to find the fracture curve.Redefined the energy calculation method during the reconnection process and calculated the probability of energy transfer. It adopts a biased random walk guided by the energy transfer probability in the graph until the walker reaches the target. This innovative approach provides a more comprehensive physical explanation and improves the effectiveness of image processing.
Zhuang Sun, Li Chen 0011, Zhida Feng, Xiaoming Liu 0004
ICIP3
2024 DD-Net: Dynamic Network Architecture for Optimized Curve Segmentation and Reduce Computational Redundancy
Yunxiang Cao, Zhida Feng
ICPR (24)4
2024 Cross-Modal Ship Grounding: Towards Large Model for Enhanced Few-Shot Learning
Quan Hu, Zhida Feng, Yaojie Chen
ICPR (30)3
2023 ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-Experts
abstract
Recent progress in diffusion models has revolutionized the popular technology of text-to-image generation. While existing approaches could produce photorealistic high-resolution images with text conditions, there are still several open problems to be solved, which limits the further improvement of image fidelity and text relevancy. In this paper, we propose ERNIE-ViLG 2.0, a large-scale Chinese text-to-image diffusion model, to progressively upgrade the quality of generated images by: (1) incorporating fine-grained textual and visual knowledge of key elements in the scene, and (2) utilizing different denoising experts at different denoising stages. With the proposed mechanisms, ERNIE-ViLG 2.01not only achieves a new state-of-the-art on MS-COCO with zero-shot FID-30k score of 6.75, but also significantly outperforms recent models in terms of image fidelity and image-text alignment, with side-by-side human evaluation on the bilingual prompt set ViLG-300.
Zhida Feng, Zhenyu Zhang 0006, Yewei Fang, Lanxin Li, Xuyi Chen, Jiaxiang Liu 0004, Weichong Yin, Shikun Feng, Yu Sun 0004, Li Chen 0011, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001
CVPR1
2023 Joint Skeleton and Boundary Features Networks for Curvilinear Structure Segmentation
Zhida Feng, Yunxiang Cao
ICIC (5)3