Zheng Ding

dblp:08/9587 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
14since 2021 · last 2026
0009-0004-2115-8357ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 A Low-phase-error FD-fNIRS Readout Circuit with Sub-1V Transimpedance Amplifier and LC-ADC-based Amplitude Control Loop
Zheng Ding, Nan Zeng, Jian Zhao 0004, Mohamad Sawan, Guoxing Wang, Cheng Chen 0054
ISCAS3
2026 Live Demonstration:A Reconfigurable and Self-Regulating Wearable NIRS Platform for Multi-Scenario Monitoring
Qianke Zeng, Nan Zeng, Zheng Ding, Yanyu Lu, Jian Zhao 0004, Mohamad Sawan, Shan Fu, Guoxing Wang, Cheng Chen 0054
ISCAS5
2026 DMDGRN: A data augmentation-based multilayer directed graph convolutional network for gene regulatory network inference
Pi-Jing Wei, Mingzhu Sun, Zheng Ding, Chun-Hou Zheng 0001
J. Biomed. Informatics3
2025 UTD-SCnet: Underwater Target Detection in Sonar Image based on Spatial and Channel Attention Net
abstract
The underwater sonar image target detection is a significant research direction in marine exploration tasks. The noise interference and blurred feature details of sonar images pose challenges to the extraction of effective feature information. And the similarity between the target shadow and the real target in the sonar image also increases the difficulty of the model to accurately identify the target. The previous sonar image target detection models have the issue of being sensitive to the threshold in Non-Maximum Suppression (NMS). Given the existing issues, this paper proposes UTD-SCnet. Firstly, the proposed Dual Multi-scale Spatial Attention module (DMSA) enhances the spatial relationship between features and increases the model’s distinguishability between real targets and shadows. Subsequently, the sensitivity of the threshold of NMS in sonar image target detection is analyzed. A transformer decoder is introduced to eliminate this influence and improve the detection accuracy of the model. Finally the Cross stage partial Orthogonal channel attention Network (CON) strengthens the model’s ability to extract effective feature information of sonar images. Experiments demonstrate that the proposed model achieved an mAP50:95 of 0.588 on the URPC2022 (Underwater Robot Picking Contest), which is 3.6% higher than the baseline model. The detection performance (mAP50:95 reached 0.580) on the URPC2021 and the noise addition experiment both demonstrated the superiority of our proposed UTD-SCnet.
Zheng Ding, Zhichao Tang, Kaitao Wu, Xiangdang Huang
IJCNN2
2024 Restoration by Generation with Constrained Priors
abstract
The inherent generative power of denoising diffusion mod-els makes them well-suited for image restoration tasks where the objective is to find the optimal high-quality image within the generative space that closely resembles the input im-age. We propose a method to adapt a pretrained diffusion model for image restoration by simply adding noise to the input image to be restored and then denoise. Our method is based on the observation that the space of a generative model needs to be constrained. We impose this constraint by finetuning the generative model with a set of anchor images that capture the characteristics of the input image. With the constrained space, we can then leverage the sampling strat-egy used for generation to do image restoration. We evaluate against previous methods and show superior performances on multiple real-world restoration datasets in preserving identity and image quality. We also demonstrate an important and practical application on personalized restoration, where we use a personal album as the anchor images to constrain the generative space. This approach allows us to produce results that accurately preserve high-frequency details, which previous works are unable to do. Project webpage: https://gen2res.github.io.
Zheng Ding, Xuaner Cecilia Zhang, Zhuowen Tu, Zhihao Xia
CVPR1
2024 TokenCompose: Text-to-Image Diffusion with Token-Level Supervision
abstract
We present TokenCompose, a Latent Diffusion Model for text-to-image generation that achieves enhanced consistency between user-specified text prompts and model-generated images. Despite its tremendous success, the standard denoising process in the Latent Diffusion Model takes text prompts as conditions only, absent explicit constraint for the consistency between the text prompts and the image contents, leading to unsatisfactory results for composing multiple object categories. Our proposed TokenCompose aims to improve multi-category instance composition by introducing the token-wise consistency terms between the image content and object segmentation maps in the finetuning stage. TokenCompose can be applied directly to the existing training pipeline of text-conditioned diffusion models without extra human labeling information. By finetuning Stable Diffusion with our approach, the model exhibits significant improvements in multi-category instance composition and enhanced photorealism for its generated images.11Project done while Zirui WangZhizhou Sha and Yilin Wang interned at UC San Diego.
Zhizhou Sha, Zheng Ding, Yilin Wang 0025, Zhuowen Tu
CVPR3
2024 HOIDiffusion: Generating Realistic 3D Hand-Object Interaction Data
abstract
3D hand-object interaction data is scarce due to the hardware constraints in scaling up the data collection pro-cess. In this paper, we propose HOIDiffusion for generating realistic and diverse 3D hand-object interaction data. Our model is a conditional diffusion model that takes both the 3D hand-object geometric structure and text description as inputs for image synthesis. This offers a more control-lable and realistic synthesis as we can specify the structure and style inputs in a disentangled manner. HOIDiffusion is trained by leveraging a diffusion model pre-trained on large-scale natural images and a few 3D human demonstrations. Beyond controllable image synthesis, we adopt the generated 3D data for learning 6D object pose estimation and show its effectiveness in improving perception systems. Project page: https://mq-zhang1.github.io/HOIDiffusion.
Zheng Ding, Sifei Liu, Zhuowen Tu, Xiaolong Wang 0004
CVPR3
2024 Explorative Inbetweening of Time and Space
Haiwen Feng, Zheng Ding, Zhihao Xia, Simon Niklaus, Victoria Fernández Abrevaya, Michael J. Black, Xuaner Cecilia Zhang
ECCV (78)2
2024 Dolfin: Diffusion Layout Transformers Without Autoencoder
Yilin Wang 0025, Liangjun Zhong, Zheng Ding, Zhuowen Tu
ECCV (51)4
2024 Patched Denoising Diffusion Models For High-Resolution Image Synthesis
abstract
We propose an effective denoising diffusion model for generating high-resolution images (e.g., 1024$\times$512), trained on small-size image patches (e.g., 64$\times$64). We name our algorithm Patch-DM, in which a new feature collage strategy is designed to avoid the boundary artifact when synthesizing large-size images. Feature collage systematically crops and combines partial features of the neighboring patches to predict the features of a shifted image patch, allowing the seamless generation of the entire image due to the overlap in the patch feature space. Patch-DM produces high-quality image synthesis results on our newly collected dataset of nature images (1024$\times$512), as well as on standard benchmarks of LHQ(1024$\times$ 1024), FFHQ(1024$\times$ 1024) and on other datasets with smaller sizes (256$\times$256), including LSUN-Bedroom, LSUN-Church, and FFHQ. We compare our method with previous patch-based generation methods and achieve state-of-the-art FID scores on all six datasets. Further, Patch-DM also reduces memory complexity compared to the classic diffusion models. Project page: https://patchdm.github.io.
Zheng Ding, Jiajun Wu 0001, Zhuowen Tu
ICLR1
2024 Inference of gene regulatory networks based on directed graph convolutional networks
abstract
Inferring gene regulatory network (GRN) is one of the important challenges in systems biology, and many outstanding computational methods have been proposed; however there remains some challenges especially in real datasets. In this study, we propose Directed Graph Convolutional neural network-based method for GRN inference (DGCGRN). To better understand and process the directed graph structure data of GRN, a directed graph convolutional neural network is conducted which retains the structural information of the directed graph while also making full use of neighbor node features. The local augmentation strategy is adopted in graph neural network to solve the problem of poor prediction accuracy caused by a large number of low-degree nodes in GRN. In addition, for real data such as E.coli, sequence features are obtained by extracting hidden features using Bi-GRU and calculating the statistical physicochemical characteristics of gene sequence. At the training stage, a dynamic update strategy is used to convert the obtained edge prediction scores into edge weights to guide the subsequent training process of the model. The results on synthetic benchmark datasets and real datasets show that the prediction performance of DGCGRN is significantly better than existing models. Furthermore, the case studies on bladder uroepithelial carcinoma and lung cancer cells also illustrate the performance of the proposed model.
Pi-Jing Wei, Ziqiang Guo, Zheng Ding, Yansen Su, Chun-Hou Zheng 0001
Briefings Bioinform.4
2023 DiffusionRig: Learning Personalized Priors for Facial Appearance Editing
abstract
We address the problem of learning person-specific facial priors from a small number (e.g., 20) of portrait photos of the same person. This enables us to edit this specific person's facial appearance, such as expression and lighting, while preserving their identity and high-frequency facial details. Key to our approach, which we dub DiffusionRig, is a diffusion model conditioned on, or “rigged by,“ crude 3D face models estimated from single in-the-wild images by an off-the-shelf estimator. On a high level, DiffusionRig learns to map simplistic renderings of 3D face models to realistic photos of a given person. Specifically, DiffusionRig is trained in two stages: It first learns generic facial priors from a large-scale face dataset and then person-specific priors from a small portrait photo collection of the person of interest. By learning the CGI-to-photo mapping with such personalized priors,DiffusionRig can “rig“ the lighting, facial expression, head pose, etc. of a portrait photo, conditioned only on coarse 3D models while preserving this person's identity and other high-frequency characteristics. Qualitative and quantitative experiments show that DiffusionRig outperforms existing approaches in both identity preservation and photorealism. Please see the project website: https://diffusionrig.github.io for the supplemental material, video, code, and data.
Zheng Ding, Xuaner Cecilia Zhang, Zhihao Xia, Lars Jebe, Zhuowen Tu, Xiuming Zhang
CVPR1
2023 MasQCLIP for Open-Vocabulary Universal Image Segmentation
abstract
We present a new method for open-vocabulary universal image segmentation, which is capable of performing instance, semantic, and panoptic segmentation under a unified framework. Our approach, called MasQCLIP, seamlessly integrates with a pre-trained CLIP model by utilizing its dense features, thereby circumventing the need for extensive parameter training. MasQCLIP emphasizes two new aspects when building an image segmentation method with a CLIP model: 1) a student-teacher module to deal with masks of the novel (unseen) classes by distilling information from the base (seen) classes; 2) a fine-tuning process to update model parameters for the queries Q within the CLIP model. Thanks to these two simple and intuitive designs, MasQCLIP is able to achieve state-of-the-art performances with a substantial gain over the competing methods by a large margin across all three tasks, including open-vocabulary instance, semantic, and panoptic segmentation. Project page is at https://masqclip.github.io/.
Tianyi Xiong, Zheng Ding, Zhuowen Tu
ICCV3
2023 Open-Vocabulary Universal Image Segmentation with MaskCLIP
abstract
In this paper, we tackle an emerging computer vision task, open-vocabulary universal image segmentation, that aims to perform semantic/instance/panoptic segmentation (background semantic labeling + foreground instance segmentation) for arbitrary categories of text-based descriptions in inference time. We first build a baseline method by directly adopting pre-trained CLIP models without finetuning or distillation. We then develop MaskCLIP, a Transformer-based approach with a MaskCLIP Visual Encoder, which is an encoder-only module that seamlessly integrates mask tokens with a pre-trained ViT CLIP model for semantic/instance segmentation and class prediction. MaskCLIP learns to efficiently and effectively utilize pre-trained partial/dense CLIP features within the MaskCLIP Visual Encoder that avoids the time-consuming student-teacher training process. MaskCLIP outperforms previous methods for semantic/instance/panoptic segmentation on ADE20K and PASCAL datasets. We show qualitative illustrations for MaskCLIP with online custom categories. Project website: https://maskclip.github.io.
Zheng Ding, Jieke Wang, Zhuowen Tu
ICML1
2020 Guided Variational Autoencoder for Disentanglement Learning
abstract
We propose an algorithm, guided variational autoencoder (Guided-VAE), that is able to learn a controllable generative model by performing latent representation disentanglement learning. The learning objective is achieved by providing signal to the latent encoding/embedding in VAE without changing its main backbone architecture, hence retaining the desirable properties of the VAE. We design an unsupervised and a supervised strategy in Guided-VAE and observe enhanced modeling and controlling capability over the vanilla VAE. In the unsupervised strategy, we guide the VAE learning by introducing a lightweight decoder that learns latent geometric transformation and principal components; in the supervised strategy, we use an adversarial excitation and inhibition mechanism to encourage the disentanglement of the latent variables. Guided-VAE enjoys its transparency and simplicity for the general representation learning task, as well as disentanglement learning. On a number of experiments for representation learning, improved synthesis/sampling, better disentanglement for classification, and reduced classification errors in meta learning have been observed.
Zheng Ding, Yifan Xu 0009, Weijian Xu, Gaurav Parmar, Yang Yang 0010, Max Welling, Zhuowen Tu
CVPR1
2011 Hecto-Scale Frame Rate Face Detection System for SVGA Source on FPGA Board
abstract
This paper proposes techniques for face detection and gives the implementation details for an FPGA development board. We analyze and discuss the relation between the system computation cost and selection of the image scaling factor. We give a new method to select the stop threshold for the image reduction process, which reduces the total computation by half. We also provide a color image output mode to let our system enjoy more human-oriented design. Test results show that the system achieves real-time face detection speed (100 fps) and a high face detection rate (87.2%) for an SVGA (600 × 800) video source. The low power consumption (3.5W) is another advantage over previous work.
Zheng Ding, Tinghui Wang, Wei Shu, Min-You Wu
FCCM1