Shengyuan Zhang

dblp:58/2695 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Security and privacy · 3Databases, data management, data science and information retrieval · 3Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Diffusion Distillation with Direct Preference Optimization for Efficient 3D LiDAR Scene Completion
abstract
The slow sampling speed of diffusion models hinders their application in 3D LiDAR scene completion. To address this, we propose Distillation-DPO, a novel framework that accelerates sampling through score distillation while simultaneously enhancing generation quality via preference alignment. Distillation-DPO follows a three-step procedure. First, the student model generates paired completion scenes with different initial noises. Second, using LiDAR scene evaluation metrics as preference, we construct winning and losing sample pairs. Third, as our core innovation, Distillation-DPO optimizes the student model by exploiting the difference in score functions between the teacher and student models on the paired completion scenes. This operation performs variational score distillation of the student model but simultaneously encourages the distilled student to prefer the winning samples over the losing ones. Extensive experiments demonstrate that Distillation-DPO achieves higher-quality scene completion than state-of-the-art diffusion models, while accelerating sampling by over 5-fold. To our knowledge, our work is the first to integrate the preference learning principle of DPO into the distillation of diffusion models, offering a new framework of preference-aligned distillation.
Shengyuan Zhang, Zejian Li, Ling Yang 0006, Pei Chen 0005, Anyang Wei, Perry Pengyun Gu, Lingyun Sun
AAAI2
2026 Aniso-GS: Anisotropic appearance field for complex highlight modeling in 3D Gaussian splatting
abstract
3D Gaussian splatting (3DGS) has significantly advanced the development of novel-view synthesis in both academic research and industrial applications due to its real-time rendering capability and excellent rendering quality. However, the visual quality of 3DGS is constrained when modeling complex scenes with specular highlights and anisotropic effects due to the limitations of low-order spherical harmonics in capturing high-frequency details. To address this issue, we propose a new method, Aniso-GS, which models the view-dependent appearance of each Gaussian by constructing an anisotropic appearance field. This introduces spatially adaptive spherical harmonic features that map low-order spherical harmonics into higher dimensions, significantly enhancing the model’s ability to capture high-frequency details. Additionally, we introduce a multi-view information smoothing strategy that improves generalization to novel viewpoints by smoothing model parameters across different training views. Furthermore, our adaptive densification strategy that suppresses Gaussian growth in low-frequency observation regions while promoting refinement in high-frequency regions, thus improving visual quality without increasing the total number of Gaussians. Quantitative and qualitative results demonstrate that our method achieves state-of-the-art visual quality on five public benchmark datasets. On the challenging Anisotropic Synthetic dataset, the PSNR improves by 5.86 dB compared to vanilla 3DGS, which clearly demonstrates the superior performance of our method in modeling specular highlights and anisotropic effects.
Gang Li 0047, Zhongliang Fu 0001, Zhao Liu 0004, Shengyuan Zhang
Comput. Graph.4
2026 SCoRE: Standardized Human Evaluation Provides a Reliable Measure for Semantic Consistency of Text-to-Image Generation
Zejian Li, Qi Liu 0076, Jiaman Pan, Lefan Hou, Xiangfei Hu, Jiarui Ma, Shengyuan Zhang, Jiesi Zhang, Xuetao Tian, Xiaoming Deng 0001
Int. J. Comput. Vis.8
2026 RS-prompt: Prompt-driven domain-incremental learning framework for remote sensing farmland segmentation
Zhaoxiang Cao, Yuchun Huang, Shengyuan Zhang
Neurocomputing3
2026 From Sketch to Reality: Enabling High-Quality, Cross-Category 3D Model Generation From Free-Hand Sketches With Minimal Data
abstract
This paper presents a novel approach for generating high-quality, cross-category 3D models from free-hand sketches with limited training data. We propose the first semi-supervised learning method to our knowledge for sketch-to-3D model conversion. Innovatively, we design a coarse-to-fine pipeline to perform the semi-supervised learning in the coarse stage and train a diffusion-based refiner to get a high-resolution 3D model. We designed a sketch-augmentation method for semi-supervised learning and integrated priors such as CLIP loss, shape prototypes, and adversarial loss to help generate high-quality results even with abstract and imprecise sketches. We also introduce an innovative procedural 3D generation method based on CAD code, which helps pre-train part of the network before fine-tuning with limited real data. Our approach, coupled with a specifically designed curriculum learning, allows us to generate high-quality 3D models across multiple categories with as few as 300 sketch-3D model pairs, marking a significant advancement over previous single-category approaches. In addition, we introduce the KO2D dataset, the largest collection of hand-drawn sketch-3D pairs to support further research in this area. As sketches are a far more intuitive and detailed way for users to express their unique ideas, we believe that this paper can move us closer to democratizing 3D content creation, enabling anyone to transform their ideas into 3D models effortlessly.
Ying Zang, Chunan Yu, Jing Li 0145, Shengyuan Zhang, Lanyun Zhu, Chaotao Ding, Renjun Xu, Tianrun Chen
IEEE Trans. Vis. Comput. Graph.5
2025 Distilling Diffusion Models to Efficient 3D LiDAR Scene Completion
abstract
Diffusion models have been applied to 3D LiDAR scene completion due to their strong training stability and high completion quality. However, the slow sampling speed limits the practical application of diffusion-based scene completion models since autonomous vehicles require an efficient perception of surrounding environments. This paper proposes a novel distillation method tailored for 3D Li- DAR scene completion models, dubbed ScoreLiDAR, which achieves efficient yet high-quality scene completion. Score- LiDAR enables the distilled model to sample in significantly fewer steps after distillation. To improve completion quality, we also introduce a novel Structural Loss, which encourages the distilled model to capture the geometric structure of the 3D LiDAR scene. The loss contains a scene-wise term constraining the holistic structure and a point-wise term constraining the key landmark points and their relative configuration. Extensive experiments demonstrate that ScoreLiDAR significantly accelerates the completion time from 30.55 to 5.37 seconds per frame (>5x) on SemanticKITTI and achieves superior performance compared to state-of-the-art 3D LiDAR scene completion models. Our model and code are publicly available on https://github.com/happyw1nd/ScoreLiDAR.
Shengyuan Zhang, Ling Yang 0006, Zejian Li, Chenye Meng, Tianrun Chen, Anyang Wei, Perry Pengyun Gu, Lingyun Sun
ICCV1
2025 Distribution Backtracking Builds A Faster Convergence Trajectory for Diffusion Distillation
abstract
Accelerating the sampling speed of diffusion models remains a significant challenge. Recent score distillation methods distill a heavy teacher model into a student generator to achieve one-step generation, which is optimized by calculating the difference between two score functions on the samples generated by the student model. However, there is a score mismatch issue in the early stage of the score distillation process, since existing methods mainly focus on using the endpoint of pre-trained diffusion models as teacher models, overlooking the importance of the convergence trajectory between the student generator and the teacher model. To address this issue, we extend the score distillation process by introducing the entire convergence trajectory of the teacher model and propose $\textbf{Dis}$tribution $\textbf{Back}$tracking Distillation ($\textbf{DisBack}$). DisBask is composed of two stages: $\textit{Degradation Recording}$ and $\textit{Distribution Backtracking}$. $\textit{Degradation Recording}$ is designed to obtain the convergence trajectory by recording the degradation path from the pre-trained teacher model to the untrained student generator. The degradation path implicitly represents the intermediate distributions between the teacher and the student, and its reverse can be viewed as the convergence trajectory from the student generator to the teacher model. Then $\textit{Distribution Backtracking}$ trains the student generator to backtrack the intermediate distributions along the path to approximate the convergence trajectory of the teacher model. Extensive experiments show that DisBack achieves faster and better convergence than the existing distillation method and achieves comparable or better generation performance, with an FID score of 1.38 on the ImageNet 64$\times$64 dataset. DisBack is easy to implement and can be generalized to existing distillation methods to boost performance.
Shengyuan Zhang, Ling Yang 0006, Zejian Li, Chenye Meng, Chang-yuan Yang, Guang Yang 0022, Lingyun Sun
ICLR1
2025 Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models
abstract
Recent advancements in diffusion models (DMs) have been propelled by alignment methods that post-train models to better conform to human preferences. However, these approaches typically require computation-intensive training of a base model and a reward model, which not only incurs substantial computational overhead but may also compromise model accuracy and training efficiency. To address these limitations, we propose Inversion-DPO, a novel alignment framework that circumvents reward modeling by reformulating Direct Preference Optimization (DPO) with DDIM inversion for DMs. Our method conducts intractable posterior sampling in Diffusion-DPO with the deterministic inversion from winning and losing samples to noise and thus derive a new post-training paradigm. This paradigm eliminates the need for auxiliary reward models or inaccurate appromixation, significantly enhancing both precision and efficiency of training. We apply Inversion-DPO to a basic task of text-to-image generation and a challenging task of compositional image generation. Extensive experiments show substantial performance improvements achieved by Inversion-DPO compared to existing post-training methods and highlight the ability of the trained generative models to generate high-fidelity compositionally coherent images. For the post-training of compostitional image geneation, we curate a paired dataset consisting of 11,140 images with complex structural annotations and comprehensive scores, designed to enhance the compositional capabilities of generative models. Inversion-DPO explores a new avenue for efficient, high-precision alignment in diffusion models, advancing their applicability to complex realistic generation tasks. Our code is available at https://github.com/MIGHTYEZ/Inversion-DPO
Zejian Li, Yize Li 0001, Chenye Meng, Zhongni Liu, Ling Yang 0006, Shengyuan Zhang, Guang Yang 0022, Chang-yuan Yang, Lingyun Sun
ACM Multimedia6
2024 Reducing Spatial Fitting Error in Distillation of Denoising Diffusion Models
abstract
Denoising Diffusion models have exhibited remarkable capabilities in image generation. However, generating high-quality samples requires a large number of iterations. Knowledge distillation for diffusion models is an effective method to address this limitation with a shortened sampling process but causes degraded generative quality. Based on our analysis with bias-variance decomposition and experimental observations, we attribute the degradation to the spatial fitting error occurring in the training of both the teacher and student model in the distillation. Accordingly, we propose Spatial Fitting-Error Reduction Distillation model (SFERD). SFERD utilizes attention guidance from the teacher model and a designed semantic gradient predictor to reduce the student's fitting error. Empirically, our proposed model facilitates high-quality sample generation in a few function evaluations. We achieve an FID of 5.31 on CIFAR-10 and 9.39 on ImageNet 64x64 with only one step, outperforming existing diffusion methods. Our study provides a new perspective on diffusion distillation by highlighting the intrinsic denoising ability of models.
Shengzhe Zhou, Zejian Li, Shengyuan Zhang, Lefan Hou, Chang-yuan Yang, Guang Yang 0022, Lingyun Sun
AAAI3
2024 A novel multi-model 3D object detection framework with adaptive voxel-image feature fusion
abstract
Abstract The multifaceted nature of sensor data has long been a hurdle for those seeking to harness its full potential in the field of 3D object detection. Although the utilisation of point clouds as input has yielded exceptional results, the challenge of effectively combining the complementary properties of multi‐sensor data looms large. This work presents a new approach to multi‐model 3D object detection, called adaptive voxel‐image feature fusion (AVIFF). Adaptive voxel‐image feature fusion is an end‐to‐end single‐shot framework that can dynamically and adaptively fuse point cloud and image features, resulting in a more comprehensive and integrated analysis of the camera sensor and the LiDar sensor data. With the aid of the adaptive feature fusion module, spatialised image features can be adroitly fused with voxel‐based point cloud features, while the Dense Fusion module ensures the preservation of the distinctive characteristics of 3D point cloud data through the use of a heterogeneous architecture. Notably, the authors’ framework features a novel generalised intersection over union loss function that enhances the perceptibility of object localsation and rotation in 3D space. Comprehensive experimentation has validated the efficacy of the authors’ proposed modules, firmly establishing AVIFF as a novel framework in the field of 3D object detection.
Zhao Liu 0004, Zhongliang Fu 0001, Gang Li 0047, Shengyuan Zhang
IET Comput. Vis.4
2022 Deep Multimodal Fusion Network for Semantic Segmentation Using Remote Sensing Image and LiDAR Data
abstract
Extracting semantic information from very-high-resolution (VHR) aerial images is a prominent topic in the Earth observation research. An increasing number of different sensor platforms are appearing in remote sensing, each of which can provide corresponding multimodal supplemental or enhanced information, such as optical images, light detection and ranging (LiDAR) point clouds, infrared images, or inertial measurement unit (IMU) data. However, these current deep networks for LiDAR and VHR images have not fully utilized the complete potential of multimodal data. The stacked multimodal fusion network (MFNet) ignores the structural differences between the modalities and the manual statistical characteristics within the modalities. For multimodal remote sensing data and its corresponding carefully designed handcrafted features, we designed a novel deep MFNet that can use multimodal VHR aerial images and LiDAR data and the corresponding intramodal features, such as LiDAR-derived features [slope and normalized digital surface model (NDSM)] and imagery-derived features [infrared–red–green (IRRG), normalized difference vegetation index (NDVI), and difference of Gaussian (DoG)]. Technically, we introduce the attention mechanism and multimodal learning to adaptively fuse intermodal and intramodal features. Specifically, we designed a multimodal fusion mechanism, pyramid dilation blocks, and a multilevel feature fusion module. Through these modules, our network realized the adaptive fusion of multimodal features, improved the receptive field, and enhanced the global-to-local contextual fusion effect. Moreover, we used a multiscale supervision training scheme to optimize the network. Extensive experimental results and ablation studies on the ISPRS semantic dataset and IEEE GRSS DFC Zeebrugge dataset show the effectiveness of our proposed MFNet.
Yangjie Sun, Zhongliang Fu 0001, Chuanxia Sun, Yinglei Hu, Shengyuan Zhang
IEEE Trans. Geosci. Remote. Sens.5
2014 Secure universal designated verifier identity-based signcryption
abstract
ABSTRACT In 2003, Steinfeld et al. introduced the notion of universal designated verifier signature (UDVS), which allows a signature holder, who receives a signature from the signer, to convince a designated verifier whether he is possession of a signer's signature; at the same time, the verifier cannot transfer such conviction to anyone else. These signatures devote to protect the receiver's privacy, that is, the receiver may want to prove to any designated verifier who he is in possession of such signature signed by the known signer but reluctant to disclose it. Moreover, the receiver also does not want the verifier to be able to convince anyone that he is in possession of such signature. In the existing UDVS schemes, a secure channel is required between the signer and the signature holder to transfer the signature. This paper, for the first time, proposes the notion of universal designated verifier signcryption without this secure channel by combining the notions of UDVS and signcryption. We give the formal definitions and a concrete construction of universal designated verifier identity‐based signcryption scheme. We also give the formal security proofs for our scheme under the random oracle model. Copyright © 2013 John Wiley & Sons, Ltd.
Changlu Lin, Pinhui Ke, Lein Harn, Shengyuan Zhang
Secur. Commun. Networks5
2013 On the linear complexity and the autocorrelation of generalized cyclotomic binary sequences of length 2p m
Pinhui Ke, Jie Zhang 0004, Shengyuan Zhang
Des. Codes Cryptogr.3
2012 New classes of quaternary cyclotomic sequence of length 2pm with high linear complexity
Pinhui Ke, Shengyuan Zhang
Inf. Process. Lett.2
2012 Constructions of binary array set with zero-correlation zone
Pinhui Ke, Shengyuan Zhang, Fuchun Lin
Inf. Sci.2
2012 Deterministic Construction of Compressed Sensing Matrices via Algebraic Curves
abstract
Compressed sensing is a sampling technique which provides a fundamentally new approach to data acquisition. Comparing with traditional methods, compressed sensing makes full use of sparsity so that a sparse signal can be reconstructed from very few measurements. A central problem in compressed sensing is the construction of sensing matrices. While random sensing matrices have been studied intensively, only a few deterministic constructions are known. Inspired by algebraic geometry codes, we introduce a new deterministic construction via algebraic curves over finite fields, which is a natural generalization of DeVore's construction using polynomials over finite fields. The diversity of algebraic curves provides numerous choices for sensing matrices. By choosing appropriate curves, we are able to construct binary sensing matrices which are superior to Devore's ones. We hope this connection between algebraic geometry and compressed sensing will provide a new point of view and stimulate further research in both areas.
Shuxing Li, Gennian Ge, Shengyuan Zhang
IEEE Trans. Inf. Theory4
2011 Identity-Based Strong Designated Verifier Signature Scheme with Full Non-Delegatability
abstract
Jakobsson et al. first proposed the notions of designated verifier signature (DVS) and an enhanced version, strong designated verifier signature (SDVS) that only the designated verifier can verify the signature's validity. Since then, many improved schemes and variants have been proposed. In 2005, Lipmaa et al. introduced a delegation-attack and proposed a corresponding security concept, namely non-delegatability, which means that both the signer and the designated verifier can not delegate signing right to a third party to generate a valid signature. In this paper, we present a stronger security notion for the SDVS schemes, full non-delegatability, which not only needs the non-delegatability of signing, but also requires that non-delegatability of verifying. And we also analyze some previous SDVS schemes and find that all of them are not secure under the new security notion. At last, we propose an improved SDVS scheme based on identity (ID-SDVS). The proposed scheme is provably secure under the new security notion in the random oracle model.
Changlu Lin, Yong Li 0002, Shengyuan Zhang
TrustCom4
2009 Improved lower bound on the number of balanced symmetric functions over GF
Pinhui Ke, Liuling Huang, Shengyuan Zhang
Inf. Sci.3
2002 Fuzzy weights of evidence method implemented in GeoDAS GIS for information extraction and integration for prediction of point events
abstract
Fuzzy weights of evidence (FWofE) is a data integration method for predictive purpose in supporting decision-making. It is an extension of the ordinary weights of evidence method, which as an artificial intelligent method that has been commonly used in geographic information systems for information extraction and integration for prediction of point events on the basis of diverse spatial evidential layers. The FWofE method has been implemented in a newly developed GIS system - GeoData Analysis System (GeoDAS). FWofE as an artificial intelligent spatial decision support method has broad applications in planning, site selection and prediction of point events such as for mineral resources assessment in geology, prediction of spring and flowing wells in hydrology and planning, and prediction of distribution of diseases in health policy study. In the current paper, a case study of mineral potential prediction in southern Nova Scotia, Canada, is used to demonstrate the application of the GeoDAS and FWofE.
Qiuming Cheng, Shengyuan Zhang
IGARSS2