EDBT 2026 Demo / reviewers in the wild / expert
Sanyuan Zhang
dblp:29/855
· DBLP profile ↗
40ranked-venue papers
1as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
Junan Lin, Daizong Liu, Xianke Chen, Xiaoye Qu, Xun Yang 0001, Jixiang Zhu, Sanyuan Zhang, Jianfeng Dong |
ACM Multimedia | 7 |
| 2025 | Improving generative trajectory prediction via collision-free modeling and goal scene reconstruction
Zhaoxin Su, Gang Huang 0004, Zhou Zhou 0003, Yongfu Li 0001, Sanyuan Zhang, Wei Hua 0002 |
Pattern Recognit. Lett. | 5 |
| 2025 | Advancing neural aesthetic assessment of artistic images based on bundle features integration
Simin Yan, Shuchang Xu, Aiping Lei, Sanyuan Zhang |
Vis. Comput. | 4 |
| 2024 | TICondition: Expanding Control Capabilities for Text-to-Image Generation with Multi-Modal Conditions
Sanyuan Zhang |
MMM (1) | 3 |
| 2024 | Advances in retinal microaneurysms detection, segmentation and datasets for the diagnosis of diabetic retinopathy: a systematic literature review
Muhammad Zeeshan Tahir, Sanyuan Zhang |
Multim. Tools Appl. | 3 |
| 2024 | Correction to: Advances in retinal microaneurysms detection, segmentation and datasets for the diagnosis of diabetic retinopathy: a systematic literature review
Muhammad Zeeshan Tahir, Sanyuan Zhang |
Multim. Tools Appl. | 3 |
| 2024 | Self-Driven Dual-Path Learning for Reference-Based Line Art Colorization Under Limited DataabstractSynthesizing color images based on line arts while considering the styles of reference photos is a flexible form of artistic creation that has recently attracted public attention. Previous approaches usually require large datasets at training, causing great inconvenience to the application. Besides, the sparsity of line art pictures often leads to a failure in learning valid mappings. To this end, we present SDL, a self-driven dual-path framework for reference-based line art colorization under limited data. Given small training sets containing sketch-image pairs, SDL first utilizes a novel Dynamic Pseudo Sample Generator (DPSG) to produce quantities of fake samples. Then, we introduce a dual-path network to achieve better visual effects, in which the Content-Generation Path reconstructs reliable content features to help establish multi-level correspondence in the Content-Color Aggregation Module (CCAM) of the Color-Transfer Path. Furthermore, we develop a Region-aware Contrastive Scheme (RCS) to focus on fine-grained details and a Style-augmented Contrastive Scheme (SCS) to encourage style consistency. Experiments verify the superiority of our model compared with existing works. We also demonstrate SDL outperforms state-of-the-art self-driven methods even though they adopt much more data than us ($30\times $on CelebA-HQ Dataset and$17\times $on ASCP Dataset). Shukai Wu, Weiming Liu 0005, Shuchang Xu, Sanyuan Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | An art-oriented pixelation method for cartoon images
Shuchang Xu, Sanyuan Zhang |
Vis. Comput. | 3 |
| 2023 | FlexIcon: Flexible Icon Colorization via Guided Images and PalettesabstractAutomatic icon colorization systems show great potential value as they can serve as a source of inspiration for designers. Despite yielding promising results, previous reference-guided approaches ignore how to effectively fuse icon structure and style, leading to unpleasant color effects. Meanwhile, they cannot take free-style palettes as inputs, which is less user-friendly. To this end, we present FlexIcon, a Flexible Icon colorization model based on guided images and palettes. To promote visual quality, our model first leverages a Hybrid Multi-expert Module to aggregate better structural features, followed by dynamically integrating the global style with each individual pixel of the structure map via the Pixel-Style Aggregation Layer. We also introduce an efficient learning scheme for free-style palette-based colorization, editing, interpolation, and diverse generation. Extensive experiments demonstrate the superiority of our framework compared with state-of-the-art approaches. In addition, we contribute a Mandala dataset to the multimedia community and further validate the application value of the proposed model. Shukai Wu, Shuchang Xu, Weiming Liu 0005, Sanyuan Zhang |
ACM Multimedia | 6 |
| 2023 | Image recoloring based on fast and flexible palette extraction
Simin Yan, Shuchang Xu, Wenzhen Yang, Sanyuan Zhang |
Multim. Tools Appl. | 4 |
| 2022 | Style Image Harmonization via Global-Local Style Mutual Guided
Juncheng Shuai, Sanyuan Zhang |
ACCV (7) | 4 |
| 2022 | Improving Reference-Based Image Colorization For Line Arts Via Feature Aggregation And Contrastive LearningabstractThe tremendous semantic discrepancy between the line art drawings without texture and the reference pictures containing rich color challenges current image-to-image translation models. Previous works attempt to establish cross-domain correspondence. However, they fail to capture more detailed features. A Reference-based Line art Translation Network (RLTN) is introduced with a Multi-level Feature Aggregation Module (MFAM) to improve the performance. The MFAM concentrates on more meaningful information for feature matching by utilizing the Multi-stream High Frequency Block (MHFB) and the Pixel-wise Correlation Block (PCB). We also employ the Channel-level Attention Block (CAB) and the Spatial-level Attention Block (SAB) for a better fusion of features. Moreover, a Style-based Contrastive Loss (SCL) is proposed to maintain the style similarity between the synthesized images and the reference examples. Experiments conducted on three datasets demonstrate the effectiveness of our model in producing more pleasing visual effects compared with state-of-the-art approaches. Shukai Wu, Qingqin Wang, Shuchang Xu, Sanyuan Zhang |
ICASSP | 4 |
| 2022 | Crossmodal Transformer Based Generative Framework for Pedestrian Trajectory PredictionabstractProviding guidance about collision avoidance, pedestrian trajectory prediction is an important task for autonomous driving. In this paper, to produce plausible trajectory predictions in the first-person view circumstance, we propose a crossmodal transformer based generative framework which could leverage sequences of cues from multiple modalities as well as pedestrian attributes. For the encoder, crossmodal transformers are exploited during the past stage to explore the cross-relation features of four modality-modality pairs, which are then fused with the help of a branch assigning operation and a modality attention module. For the decoder, we employ a bézier curve interpolation based method to project encoder features into trajectory results. Our training process not only considers the pedestrian's intention of crossing road but also optimizes our model to achieve more accurate predictions at the terminal time steps. Experimental results demonstrate that our framework outperforms state-of-the-art methods on both JAAD and PIE datasets. Especially, compared with the best baseline, our method could achieve 15.1%/14.3% and 14.3%/22.2% improvement for deterministic/multimodal prediction in the metric of box center final displacement error on JAAD and PIE, respectively. Zhaoxin Su, Gang Huang 0004, Sanyuan Zhang, Wei Hua 0002 |
ICRA | 3 |
| 2022 | RefFaceNet: Reference-based Face Image Generation from Line Art Drawings
Shukai Wu, Weiming Liu 0005, Qingqin Wang, Sanyuan Zhang, Zhenjie Hong, Shuchang Xu |
Neurocomputing | 4 |
| 2021 | CR-LSTM: Collision-prior Guided Social Refinement for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a challenge because of the complex social interactions in context and the elusive intention of each pedestrian. Collision avoidance is one of the most common social interactions in real world, while existing data-driven works have not handled it well yet. In order to address this issue, we propose a framework that considers the theory about the minimum distance between each pedestrian-pedestrian pair and the corresponding time as the collision related prior knowledge. With the prior, our social refinement module, called Collision-prior Guided Refinement, can be guided to understand the collision situations of a crowd through a message passing mechanism. To focus on more useful information from context, we also introduce pedestrian-wise attention and collision gate to jointly judge collision potential for all pedestrian-pedestrian pairs. Experimental results demonstrate that our framework can achieve competitive results on ETH and UCY datasets by comparing with existing works. In addition, it indicates the superiority of our framework in the aspect of collision avoidance. Zhaoxin Su, Sanyuan Zhang, Wei Hua 0002 |
IROS | 2 |
| 2021 | Characterizing Network Anomaly Traffic with Euclidean Distance-Based Multiscale Fuzzy EntropyabstractThe prosperity of mobile networks and social networks brings revolutionary conveniences to our daily lives. However, due to the complexity and fragility of the network environment, network attacks are becoming more and more serious. Characterization of network traffic is commonly used to model and detect network anomalies and finally to raise the cybersecurity awareness capability of network administrators. As a tool to characterize system running status, entropy-based time-series complexity measurement methods such as Multiscale Entropy (MSE), Composite Multiscale Entropy (CMSE), and Fuzzy Approximate Entropy (FuzzyEn) have been widely used in anomaly detection. However, the existing methods calculate the distance between vectors solely using the two most different elements of the two vectors. Furthermore, the similarity of vectors is calculated using the Heaviside function, which has a problem of bouncing between 0 and 1. The Euclidean Distance-Based Multiscale Fuzzy Entropy (EDM-Fuzzy) algorithm was proposed to avoid the two disadvantages and to measure entropy values of system signals more precisely, accurately, and stably. In this paper, the EDM-Fuzzy is applied to analyze the characteristics of abnormal network traffic such as botnet network traffic and Distributed Denial of Service (DDoS) attack traffic. The experimental analysis shows that the EDM-Fuzzy entropy technology is able to characterize the differences between normal traffic and abnormal traffic. The EDM-Fuzzy entropy characteristics of ARP traffic discovered in this paper can be used to detect various types of network traffic anomalies including botnet and DDoS attacks. Xiao Wang 0047, Wei Zhang 0138, Sanyuan Zhang |
Secur. Commun. Networks | 5 |
| 2020 | Saliency Guided Subdivision for Single-View Mesh ReconstructionabstractIn this paper, we present a novel deep architecture to recover a 3D shape in triangular mesh from a single image based on mesh deformation. Most existing deformation-based methods produce uniform mesh predictions by repeatedly applying global subdivision but fail to require the highlighted details due to the memory limits. To address this problem, we propose a novel saliency guided subdivision method to achieve the trade-off between detail generation and memory consumption. Instead of using local geometric cues such as curvature, we introduce a global point-based saliency voting operation to guide the adaptive mesh subdivision and deformation explicitly. Moreover, we propose the oriented chamfer loss to mitigate the mesh self-intersection problem in subdivision. We further make our network configurable and explore the best structure combination. Extensive experiments show that our method can both produce visually pleasing results with fine details and achieve better performance compared to other state-of-the-art methods. Weicai Ye, Guofeng Zhang 0001, Sanyuan Zhang, Hujun Bao |
3DV | 4 |
| 2020 | An End-To-End Network For Detecting Multi-Domain Fractures On X-Ray ImagesabstractAutomated fracture detection on medical images is a crucial prerequisite for orthopedic diagnosis. However, due to the considerable variation of bone structures, it is challenging to detect fractures on images filmed from various body parts utilizing a single model. In this paper, we treat each body part as a domain and propose a novel Multi-domain Fracture Detection Network (MFDN), which is composed of two sub-networks, namely, a domain classification network for predicting the domain type of an image and a fracture detection network for detecting fractures on X-ray images of different domains. By constructing Feature Enhancement Modules and Multi-Feature-Enhanced R-CNN, the proposed MFDN extracts better feature representations for each domain. Experimental results on real-world datasets show the effectiveness of our model which has been used in clinical diagnosis with the best performance on all the domains. Shukai Wu, Lifeng Yan, Yizhou Yu, Sanyuan Zhang |
ICIP | 5 |
| 2020 | Maximum spatial-temporal isometric cluster for dynamic surface correspondence
Zhihao Cheng, Fuchang Liu, Sanyuan Zhang |
Vis. Comput. | 4 |
| 2018 | Image-based 3D model retrieval using manifold learningabstractWe propose a new framework for image-based three-dimensional (3D) model retrieval. We first model the query image as a Euclidean point. Then we model all projected views of a 3D model as a symmetric positive definite (SPD) matrix, which is a point on a Riemannian manifold. Thus, the image-based 3D model retrieval is reduced to a problem of Euclid-to-Riemann metric learning. To solve this heterogeneous matching problem, we map the Euclidean space and SPD Riemannian manifold to the same high-dimensional Hilbert space, thus shrinking the great gap between them. Finally, we design an optimization algorithm to learn a metric in this Hilbert space using a kernel trick. Any new image descriptors, such as the features from deep learning, can be easily embedded in our framework. Experimental results show the advantages of our approach over the state-of-the-art methods for image-based 3D model retrieval. Pan-pan Mu, Sanyuan Zhang, Yin Zhang 0006, Xiuzi Ye |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2017 | Using Transferred Deep Model in Combination with Prior Features to Localize Multi-style Ship License Numbers in Nature ScenesabstractShip License Numbers (SLNs) localization is an important part of waterway intelligent transportation systems. Unfortunately, this issue has been neglected for a long time. In this paper, we present an effective approach for localizing multi-style SLNs in nature scenes. The problem of locating SLNs is posed as the detection of character sequences which possess SLNs prior features. First, faced with the difficulty of no training data, a transfer learning-based deep convolutional neural network is designed to detect character sequences in SLNs. In the second step, to accurately locate SLNs from the detected character sequences, the prior features of SLNs are considered. Three SLNs prior features are summarized. An SLNs region generating algorithm and a low-level similarity-based fake SLNs filtering algorithm are presented, respectively. The accurate positions of the SLNs in the input image are obtained in this stage. The proposed approach is finally tested on ZJUSHIPS950 dataset. The approach achieves a FPPI of 0.42 and a F-measure of 0.614 on 1374 labeled SLNs, surpassing several related methods by a large margin. Controlled experiment results also prove the impressive performances of the proposed SLNs region generating and fake SLNs filtering algorithms. Xingzheng Lyu, Sanyuan Zhang, Zhenjie Hong, Xiuzi Ye |
ICTAI | 4 |
| 2017 | Co-clustering with Manifold and Double Sparse Representation
Sanyuan Zhang |
IDEAL | 2 |
| 2017 | Laplacian sparse dictionary learning for image classification based on sparse representationabstractSparse representation is a mathematical model for data representation that has proved to be a powerful tool for solving problems in various fields such as pattern recognition, machine learning, and computer vision. As one of the building blocks of the sparse representation method, dictionary learning plays an important role in the minimization of the reconstruction error between the original signal and its sparse representation in the space of the learned dictionary. Although using training samples directly as dictionary bases can achieve good performance, the main drawback of this method is that it may result in a very large and inefficient dictionary due to noisy training instances. To obtain a smaller and more representative dictionary, in this paper, we propose an approach called Laplacian sparse dictionary (LSD) learning. Our method is based on manifold learning and double sparsity. We incorporate the Laplacian weighted graph in the sparse representation model and impose the l 1 -norm sparsity on the dictionary. An LSD is a sparse overcomplete dictionary that can preserve the intrinsic structure of the data and learn a smaller dictionary for each class. The learned LSD can be easily integrated into a classification framework based on sparse representation. We compare the proposed method with other methods using three benchmark-controlled face image databases, Extended Yale B, ORL, and AR, and one uncontrolled person image dataset, i-LIDS-MA. Results show the advantages of the proposed LSD algorithm over state-of-the-art sparse representation based classification methods. Jia Sheng, Sanyuan Zhang |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2017 | Penalty-based haptic rendering technique on medicinal healthy dental detection
Yi Li 0013, Sanyuan Zhang, Xiuzi Ye |
Multim. Tools Appl. | 2 |
| 2017 | Healthy human sitting posture estimation in RGB-D scenes using object context
Yi Li 0013, Sanyuan Zhang, Xiuzi Ye |
Multim. Tools Appl. | 3 |
| 2016 | Construction of retinal vascular trees via curvature orientation priorabstractConstructing vascular trees is a prerequisite for graph-based methods of arterial/venous classification, vessel abnormities measurement and various diseases evaluation. Previous works mainly focused on geometrical and topological properties of vessel segments and applied fixed rules or constraint optimization to building vessel trees. In this paper, we propose a novel vascular tree generation method with vessel curvature orientation as a prior. We firstly build a rough graph from vessel centerline image and then modify graph misrepresentations (false edge and missing edge) based on vessel landmarks, which can be extracted from multi-scale curvature orientation histogram. Next we separate different trees at vessel junctions by curvature orientation clustering. Two graph-based applications, vessel classification and vessel diameter ratio measurement, are tested on two different image sets to validate our vascular tree construction approach. Experiments demonstrate that our approach generates a more reliable vascular graph and achieves a comparatively high vessel classification accuracy of 83.21% and 85.06% within entire image in RITE and INSPIRE database respectively. Xingzheng Lyu, Shunren Xia, Sanyuan Zhang |
BIBM | 4 |
| 2016 | Human articulated body recognition method in high-resolution monitoring images
Yi Li 0013, Sanyuan Zhang, Xiuzi Ye |
Neurocomputing | 3 |
| 2016 | Data analysis on virtual stiffness in 6DoFs haptic rendering system
Yi Li 0013, Sanyuan Zhang, Xiuzi Ye |
Neurocomputing | 2 |
| 2016 | Mining location-aware discriminative blocklets for recognizing landmark architectures
Yi Li 0013, Sanyuan Zhang |
Multim. Syst. | 2 |
| 2016 | Haptic rendering method based on generalized penetration depth computation
Yi Li 0013, Yin Zhang 0006, Xiuzi Ye, Sanyuan Zhang |
Signal Process. | 4 |
| 2015 | Detection of engineering vehicles in high-resolution monitoring imagesabstractThis paper presents a novel formulation for detecting objects with articulated rigid bodies from high-resolution monitoring images, particularly engineering vehicles. There are many pixels in high-resolution monitoring images, and most of them represent the background. Our method first detects object patches from monitoring images using a coarse detection process. In this phase, we build a descriptor based on histograms of oriented gradient, which contain color frequency information. Then we use a linear support vector machine to rapidly detect many image patches that may contain object parts, with a low false negative rate and a high false positive rate. In the second phase, we apply a refinement classification to determine the patches that actually contain objects. In this stage, we increase the size of the image patches so that they include the complete object using models of the object parts. Then an accelerated and improved salient mask is used to improve the performance of the dense scale-invariant feature transform descriptor. The detection process returns the absolute position of positive objects in the original images. We have applied our methods to three datasets to demonstrate their effectiveness. Yin Zhang 0006, Sanyuan Zhang, Zhong-yan Liang, Xiuzi Ye |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2015 | An optimization method for penalty-based six-degrees-of-freedom haptic rendering system
Yi Li 0013, Yin Zhang 0006, Xiuzi Ye, Sanyuan Zhang |
Signal Process. Image Commun. | 4 |
| 2013 | Six-degree-of-freedom haptic rendering using translational and generalized penetration depth computationabstractWe present six-degree-of-freedom (6DoF) haptic rendering algorithms using translational (PDt) and generalized penetration depth (PDg). Our rendering algorithm can handle any type of object/object haptic interaction using penalty-based response and makes no assumption about the underlying geometry and topology. Moreover, our rendering algorithm can effectively deal with multiple contacts. Our penetration depth algorithms for PDtand PDgare based on a contact-space projection technique combined with iterative, local optimization on the contact-space. We circumvent the local minima problem, imposed by the local optimization, using motion coherence present in the haptic simulation. Our experimental results show that our methods can produce high-fidelity force feedback for general polygonal models consisting of tens of thousands of triangles at near-haptic rates, and are successfully integrated into an off-the-shelf 6DoF haptic device. We also discuss the benefits of using different formulations of penetration depth in the context of 6DoF haptics. Yi Li 0013, Min Tang 0004, Sanyuan Zhang, Young J. Kim |
World Haptics | 3 |
| 2008 | Reverse innovative design - an integrated product design methodology
Xiuzi Ye, Hongzheng Liu, Sanyuan Zhang |
Comput. Aided Des. | 6 |
| 2007 | Approximate the swept volume of revolutions along curved trajectoriesabstractSwept volume has been applied to many applications areas such as NC machining simulation and verification, robot workspace analysis, collision detection, and CAD. The numerical computation of swept volume remains to be a very challenging problem in terms of simplicity, efficiency and accuracy. The paper presents a novel swept volume approximation method for revolution generator solids sweeping along curved trajectories. This algorithms consists of the following main steps: (1) Discretization of the trajectory curve and establishments of reference frames; (2) Extraction and approximation of envelop profile at each discretized position based on the velocity vector calculated; and (3) Generation of envelop surfaces from the extracted envelop profile curves, and attachment of the ingress and egress surfaces to the envelop surface to form the SV Solid. Examples show that our algorithm is simple, efficient and accurate. Zhiqi Xu, Xiuzi Ye, Sanyuan Zhang |
Symposium on Solid and Physical Modeling | 4 |
| 2005 | Multi-level access control for collaborative CADabstractAccess control is one of the key steps to ensuring the availability, confidentiality and integrity of system resources. While it has been fully applied in various network systems, it's still relatively new for collaborative CAD. This paper provides a new framework of multi-level access control for collaborative CAD based on hierarchical product modeling. A layered privilege model (LPM) has been developed to provide multi-granularity access permissions in the part level and feature level. An extended hierarchy role-based access control (EHRBAC) is implemented, integrated with LPM and common Hierarchy RBAC, to provide hierarchical access control of collaborative feature modeling in the component/part level and design feature level. Mesh simplification techniques are used to create variable level-of-detail for individual features. Cuihao Fang, Xiuzi Ye, Sanyuan Zhang |
CSCWD (1) | 4 |
| 2005 | Uniform color transferabstractWe present in this paper a general algorithm called uniform color transfer for transferring dominant colors in one image to another. Our algorithm frees users from interactions required for colorizing grayscale image or transferring color between images. Different from existing algorithms, we cluster the source color image into regions using the GMM-EM method, and the target image using the K-means algorithm. We then impose chromatic mean values on corresponding source image regions in the target image. Experiments showed that our algorithm can give better results. Shuchang Xu, Yin Zhang 0006, Sanyuan Zhang, Xiuzi Ye |
ICIP (3) | 3 |
| 2005 | Radius-Normal Histogram and Hybrid Strategy for 3D Shape RetrievalabstractRecent development in computer hardware and modeling technologies has led to a fast increasing of 3D models. To help user find expected models accurately and efficiently, technologies on content-based 3D model retrievals have become one of the most challenging research topic recently. In this paper, radius-normal histogram(RNH) is proposed to describe shape contents and used for shape retrievals. The RNH shape descriptor first uses a series of concentric spheres to capture the point distribution information of the given model. Then for points in each concentric sphere, a radian normal angle is computed to extract the local geometry features. Finally, the radius-normal histogram is constructed by using the extracted shape signatures. The proposed shape representation remains invariant under rotations. It can be generated from the given 3D model efficiently and easily as well. Performance comparisons for the shape benchmark database have proven that the proposed algorithm can achieve better retrieving performance than other similar histogram-based shape representations. To further improve the retrieving performance, different shape features are combined based on dynamic weight selections and a hierarchical architecture is used to speed up the matching process. Yin Zhang 0006, Sanyuan Zhang, Xiuzi Ye |
SMI | 3 |
| 2005 | A hybrid method for robust car plate character recognition
Xiuzi Ye, Sanyuan Zhang |
Eng. Appl. Artif. Intell. | 3 |
| 2001 | Cubic algebraic curves based on geometric constraints
Sanyuan Zhang, Hujun Bao, Baogang Wei |
Comput. Aided Geom. Des. | 1 |