Haiyan Sun

dblp:94/6808 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
24since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Oracle Bone Inscriptions Recognition Based on Spatial Transformer Network and Few-Shot Learning
Xingquan Cai, Lixin Ding, Haiyan Sun
ICIC (2)5
2025 Dynamic Convolution and Dimensional Joint Attention Based Denoising of Oracle Topography Images
Xingquan Cai, Mengrui Dai, Haiyan Sun
ICIC (3)5
2025 Interpretable Action Quality Assessment with Temporal Parsing
Xingquan Cai, Haiyan Sun
ICIC (5)5
2025 MH-CL: Self-Supervised Action Recognition Method Based on Multi-view Hypergraphic Contrastive Learning for Intangible Cultural Heritage Dance
abstract
Compared to daily movements, dance movements of the intangible cultural heritage exhibit complexities and intricate coordination between distant joints, such as the coordinated movements required in the three bends of the Dai dance involving the feet, knees, hips, and arms. These characteristics pose challenges for existing recognition methods, resulting in low accuracy. To address this, we propose a self-supervised action recognition method for intangible cultural heritage dance based on multi-view hypergraph contrastive learning. We begin by inputting the 3D human pose sequence of the dance video, enhancing it with data through multi-view and random frame methods. Subsequently, an action encoder network is constructed to extract action features using joint dynamic and static hypergraph convolution. A multilayer projection head is then designed to map the feature representations to a one-dimensional space, optimizing the encoder loss by comparing the loss functions. The experimental results demonstrate the superior performance of our approach over the QDPuzzle [1] and AimCLR [2] algorithms under two evaluation protocols in the NTU RGB + D 60 [3], NTU RGB+D 120 [4], and PKUMMD [5] datasets. In addition, ablation experiments show the ability of the method to recognize more accurate 3D human pose sequences, particularly in intangible cultural heritage dance videos featuring distant joint coordination and addressing the issue of missing segments.
Xingquan Cai, Pengyan Cheng, Kaijie Qu, Haiyan Sun
IJCNN5
2025 A multilevel feature-based method for mapping sparse point clouds to CAD models
Youlong Zeng, Haiyan Sun, Zhuoyi Chen
Comput. Aided Geom. Des.2
2025 3D human pose estimation using spatiotemporal hypergraphs and its public benchmark on opera videos
Xingquan Cai, LiZhe Chen, YiJie Wu, Haiyan Sun
Vis. Comput.5
2024 Retinex-Based Low-Light Mural Image Enhancement with Color Correction
Xingquan Cai, Haiyan Ma, Haiyan Sun
CGI (1)5
2024 Decoupled Estimation of Human Pose and Shape for ICH Performance Video Based on L-C-HRNet
Xingquan Cai, Shike Liu, Haiyan Sun
CGI (1)4
2024 A Novel Auxiliary Task Framework in 3D Human Pose Estimation for Opera Videos
abstract
Influenced by the costume and limb occlusion of dance movements in opera videos, the 2D human pose estimation methods struggle to accurately locate the 2D joint coordinates of the occluded parts. This inaccuracy leads to lower precision in estimating 3D human poses from 2D joint coordinates in opera videos. To enhance the learning of more effective spatio-temporal dependencies from 2D joint coordinates and improve the accuracy of 3D human pose estimation, this paper proposes a novel auxiliary task framework. We first designed three auxiliary tasks to mask some of the 2D joint coordinates, disorder the spatial position of the joints and the video frame order, with the goal of recovering the corrupted 2D joint coordinates. Then to address these auxiliary tasks, we propose a multi-feature representation Transformer network to capture 2D joint coordinates spatio-temporal features from local to global perspective by constructing local adaptive graph convolution network, segmented time-aware network and global spatio-temporal self-attention module respectively. Finally, an adaptive weight allocation module is utilized to integrate local and global features to output the 3D joint coordinates. Extensive comparative and ablation experiments on the Human3.6M, MPI-INF-3DHP and opera datasets demonstrate that our method surpasses all comparative methods in MPJPE accuracy. Furthermore, the auxiliary task framework designed in this paper effectively captures comprehensive and efficient spatio-temporal dependencies in 2D joint coordinates from opera videos.
Xingquan Cai, Shanshan He, Haoyu Song 0005, Haiyan Sun
ICMR5
2023 An Image Extraction Method for Traditional Dress Pattern Line Drawings Based on Improved CycleGAN
Xingquan Cai, Sichen Jia, Jiali Yao, Haiyan Sun
CGI5
2023 Abnormal Behavior Detection Method based on Spatio-temporal Dual-flow Network for Surveillance Videos
abstract
To address the issues that current surveillance videos abnormal behavior detection methods are affected by complex environments such as video blurring and long distances, resulting in low detection accuracy and slow speed, we propose an abnormal behavior detection method based on spatio-temporal dual-flow network for surveillance videos. Firstly, the input surveillance videos are sampled with different frame rates and a spatio-temporal dual-flow partial convolution network is constructed to extract spatial and motion information features from the spatial and temporal flow video sequences, respectively. Then, a cross-modal dual-attention fusion mechanism is introduced after each feature extraction of the dual-flow partial convolution network to enhance the exchange of feature information between spatial and temporal flows. Finally, the extracted motion and spatial features are fused to output detection results. Experiments show that our method reduces the number of floating-point operations and model parameters on the Kinetics-600 dataset compared to the MViT-L algorithm with guaranteed accuracy. Compared to lightweight networks such as X3D-XL, it improves the accuracy in Top-1 and Top-5 by 4.3% and 3.6%, respectively. The experiments prove that the method can reduce the model’s computational complexity while obtaining more accurate abnormal behavior detection results when the surveillance videos is blurred or the abnormal behavior occurs far away.
Xingquan Cai, Haiyan Sun
ICTAI5
2023 A Method for 3D Human Pose Estimation and Similarity Calculation in Tai Chi Videos
abstract
Human pose estimation from video sequences has become a hot research topic in the domain of robotics and computer vision. However, existing three-dimensional (3D) pose estimation methods usually analyze individual frames, which have a low accuracy due to various human movement speed, limiting its practical application. In this paper, we propose a method for estimating 3D pose and calculating similarity from Tai Chi video sequences based on Seq2Seq network. Specifically, using 2D joint point coordinate sequence of the original image as input, our method constructs an encoder and a decoder to build a Seq2Seq network. Our method introduces an attention mechanism for weighing the input data to obtain an intermediate vector and decode it to estimate the 3D joint point sequence. Afterwards, using a template video and a target video as input, our method calculates the cost of passing through each point within the constraints to construct a cost matrix for video similarity. With the cost matrix, our method can determine the optimal path and use the correspondence of the video sequence to calculate the image similarity of the corresponding frame. The experimental data show that the proposed method can effectively improve the accuracy of 3D pose estimation, and increase the speed for video similarity calculation.
Xingquan Cai, Yuqing Huo, Haiyan Sun, Jiaqi Ji
Int. J. Pattern Recognit. Artif. Intell.5
2023 An Automatic Music-Driven Folk Dance Movements Generation Method Based on Sequence-To-Sequence Network
abstract
Music-driven automatic dance movement generation has become a hot research topic in the field of computer vision and internet of things in the recent past. To address the problems of increasing loss of Chinese folk dance culture, high cost of manual choreography methods and requirements for professional background, this paper proposes an automatic generation method for folk dance movements. Firstly, the proposed method collects paired folk music and dance videos to construct a synchronized folk music–dance dataset, extracting music and dance features using a feature extraction tool and a multi-scale fusion high-resolution network, respectively. Afterward, a sequence-to-sequence network model is constructed and then trained based on music features and dance features to synthesize rhythmically matched dance sequences for new music clips. Finally, an easy-to-use and effective automatic folk dance choreography method is implemented. Experimental data show that the proposed method performs well in automatic folk dance generation and the generated dances have folk characteristics and match the rhythm of the given music.
Xingquan Cai, Mengyao Xi, Sichen Jia, Haiyan Sun
Int. J. Pattern Recognit. Artif. Intell.6
2023 A Social Distance Monitoring Method Based on Improved YOLOv4 for Surveillance Videos
abstract
Social distance monitoring is of great significance for public health in the era of COVID-19 pandemic. However, existing monitoring methods cannot effectively detect social distance in terms of efficiency, accuracy, and robustness. In this paper, we proposed a social distance monitoring method based on an improved YOLOv4 algorithm. Specifically, our method constructs and pre-processes a dataset. Afterwards, our method screens the valid samples and improves the K-means clustering algorithm based on the IoU distance. Then, our method detects the target pedestrians using a trained improved YOLOv4 algorithm and gets the pedestrian target detection frame location information. Finally, our method defines the observation depth parameters, generates the 3D feature space, and clusters the offending aggregation groups based on the L2 parametric distance to finally realize the pedestrian social distance monitoring of 2D video. Experiments show that the proposed social distance monitoring method based on improved YOLOv4 can accurately detect pedestrian target locations in video images, where the pre-processing operation and improved K-means algorithm can improve the pedestrian target detection accuracy. Our method can cluster the offending groups without going through calibration mapping transformation to realize the pedestrian social distance monitoring of 2D videos.
Xingquan Cai, Pengyan Cheng, Dingwei Feng, Haiyan Sun, Jiaqi Ji
Int. J. Pattern Recognit. Artif. Intell.5
2023 A Multi-Object Detection Method Based on Adaptive Feature Adjustment of 3D Point Cloud in Indoor Scenes
abstract
Due to the complexity and diversity of indoor environment objects and interference occlusions, the accuracy of multi-object target detection based on 3D point cloud is limited. To address this issue, we present a multi-target detection method based on adaptive feature adjustment (AFA) of 3D point cloud. First, our method preprocesses the dataset and constructs a backbone module. Afterwards, our method uses an improved PointNet[Formula: see text] network for feature adaptive learning, where an AFA module is added to learn the influence relationship between point pairs. The proposed method then establishes the relationship between contexts in the local point set area and extracts the feature of point cloud. Using the idea of Hough voting, our method can generate some votes close to the particle. Using these votes to generate proposal, the proposed method adds CBAM attention mechanisms to both modules of voting and proposal, which can fuse the feature information of the channel and expand the receptive field in space. Our method can enhance the important features and weaken the unimportant features, making the extracted features more directional and enhancing the expressiveness of the network. Finally, the generated results are visualized to complete the multi-target detection of 3D point cloud. To verify the effectiveness of our proposed method, two large datasets with real 3D scanning, scanNet2 and SunRGB-D, are used for training the network. The experimental results show that the proposed method can improve the effectiveness of point cloud target detection in indoor scenes, getting a higher detection accuracy.
Haiyan Sun, Keng Chen, Sichen Jia, Xingquan Cai
Int. J. Pattern Recognit. Artif. Intell.1
2023 An Improved PoinTr Point Cloud Completion Method Based on Feature Enhancement
abstract
To address the issue that point cloud data is often incomplete and difficult to obtain, we propose a point cloud completion method to improve the PoinTr method based on feature enhancement. In dataset preprocessing, the farthest point of the original point cloud is sampled to obtain the central point coordinates. Our method constructs an MLP network, where the local information of these central points is obtained and the location embedding is performed. Combining network and SENet network, the local features of the point cloud are extracted and enhanced, and the location embedding and local features are added to obtain the point proxies of the original point cloud. Afterward, our method predicts the missing part of the point cloud by using an Encoder to model the relationship between the point cloud structure information and points, and then using a Decoder to learn the relationship between the missing and existing parts of the point cloud and reconstruct the missing point cloud. Our method also modifies the attention mechanism to make the features more global and enhance the network expression. Finally, the point cloud is refined, and is realized by predicting multiple points around each point of the coarse point cloud through the FoldingNet network, and the final output is the complete point cloud. Experimental results show that the proposed method can not only reduce the performance overhead, but also improve the effects of point cloud completion.
Haiyan Sun, Zaichao Lin, Qingtao Lu, Sichen Jia, Xingquan Cai
Int. J. Pattern Recognit. Artif. Intell.1
2023 A Digital Simulation and Re-Editing Method for Clothing Patterns Based on Deep Learning and Somatosensory Interaction
abstract
To address the issues in clothing pattern style migration, this paper proposes a digital simulation and re-editing method for clothing patterns based on deep learning and somatosensory interaction. First, the proposed method encodes the black-and-white line drawing image, generating random noise images through a diffusion process, introducing color information for synthesis, and using a decoder to reconstruct a colored image. Afterwards, an improved VGG19 model is used to reconstruct content features and perform linear color transformation on style images, enabling pattern style migration through the construction of a Gram matrix and resulting in colored clothing texture patterns. Finally, a KinectV2 is utilized for fabric simulation, overlaying colorful clothing texture patterns to achieve 3D virtual dressing. The experimental results show that the proposed method improves the structural similarity index measure (SSIM) by 9–11% and the peak signal-to-noise ratio (PSNR) by 3–8% when compared to existing algorithms. The experiments provide evidence that the proposed method effectively mitigates color overflow, delivers precise image coloring, and accomplishes realistic restoration of clothing texture. Furthermore, the method offers an improved garment fit to fulfill the user’s interaction requirements.
Haiyan Sun, Jiali Yao, Xingquan Cai
Int. J. Pattern Recognit. Artif. Intell.1
2023 Automatic generation of Labanotation based on human pose estimation in folk dance videos
Xingquan Cai, Sichen Jia, Haiyan Sun
Neural Comput. Appl.5
2022 MT-3000: a heterogeneous multi-zone processor for HPC
Kai Lu 0001, Yang Guo 0003, Chun Huang 0006, Sheng Liu 0001, Ruibo Wang, Jianbin Fang, Tao Tang 0001, Zhaoyun Chen, Biwei Liu, Zhong Liu 0003, Yuanwu Lei, Haiyan Sun
CCF Trans. High Perform. Comput.13
2022 A Low Distortion Mesh Parameterization Mapping Method Based on Proxy Function and Combined Newton
abstract
To address the issues of low efficiencies and serious mapping distortions in current mesh parameterization methods, we present a low distortion mesh parameterization mapping method based on proxy function and combined Newton’s method in this paper. First, the proposed method calculates visual blind areas and distortion prone areas of a 3D mesh model, and generates a model slit. Afterwards, the method performs the Tutte mapping on the cut three-dimensional mesh model, measures the mapping distortion of the model, and outputs a distortion metric function and distortion values. Finally, the method sets iteration parameters, establishes a reference mesh, and finds the optimal coordinate points to get a convergent mesh model. When calculating mapping distortions, Dirichlet energy function is used to measure the isometric mapping distortion, and MIPS energy function is used to measure the conformal mapping distortion. To find the minimum value of the mapping distortion metric function, we use an optimal solution method combining proxy functions and combined Newton’s method. The experimental data show that the proposed method has high execution efficiency, fast descending speed of mapping distortion energy and stable optimal value convergence quality. When a texture mapping is performed, the texture is evenly colored, close laid and uniformly lined, which meets the standards in practical applications.
Xingquan Cai, Dingwei Feng, Mohan Cai, Haiyan Sun
Int. J. Pattern Recognit. Artif. Intell.5
2022 Real-Time Leaf Recognition Method Based on Image Segmentation and Feature Extraction
abstract
Leaf recognition has been an important research field of image recognition in the recent past. However, traditional leaf recognition methods can be easily affected by environments and cannot realize multi-leaf recognition under a complex background in real time. In this work, we present a real-time leaf recognition method based on image segmentation and feature recognition. First, we denoise the input of a leaf image, performing a leaf segmentation with an improved FCN network model, and then optimize the contour edge with a CRF algorithm to get a leaf segmentation image. Second, we extract the content features of the segmented leaf image with an Inception-V2 network model to get a feature map of the leaf image. Third, we input the feature map into an RPN network to obtain a set of regional candidate frames and then integrate the feature map and the information of candidate frames in a RoI Pooling layer, which can extract the feature map of a candidate frame area and scale it to a fixed-size feature map. Finally, we send the feature map to a fully connected layer to classify each preselection box content through the calculation of preselection feature maps, and then obtain the final accurate position of the prediction box by utilizing a bounding box regression. The experimental results show that the proposed method can achieve multi-leaf recognitions with high accuracy and fast speed under complex environments in real time.
Xingquan Cai, Yuqing Huo, Yunbo Chen, Mengyao Xi, Yuxin Tu, Haiyan Sun
Int. J. Pattern Recognit. Artif. Intell.7
2022 Image Attribute Migration Based on Decoupling and Adaptive Layer Instance Normalization
abstract
The issue of image attribute migration is one of the hot research topics in the field of computer vision, which has received extensive research interest. However, current unsupervised image attribute migration models using symmetric generative adversarial network structure do not work well on datasets with large geometric variations, where the results lack diversity and are of low quality. To address these problems, we present an image attribute migration model based on decoupling and adaptive layer instance normalization. First, a codec structure based on a decoupled representation is constructed as the generator, and an adaptive layer instance normalization operation is used in the decoder. Then, the iterations of the model are constrained by various improved loss functions. We conducted controlled experiments and compared the results of our method with other methods using several datasets with large geometric variations. The experimental results demonstrate that the proposed method can achieve high quality and diverse image attribute migration.
Xingquan Cai, Fajian Li, Keng Chen, Yuechao Wei, Haiyan Sun
Int. J. Pattern Recognit. Artif. Intell.5
2022 POGT: A Peking Opera Gesture Training System Using Infrared Sensors
abstract
Peking opera is one of the national cultural heritages in China. However, it is difficult for people to learn the gestures in Peking opera performance, which limits the spread of this traditional culture. To address this issue, we propose a Peking opera gesture training system using infrared sensors. Specifically, we build a character avatar for demonstrating the gestures in Peking opera in the proposed system. Based on the data collected by infrared sensors, a method for calculating gesture similarity is proposed and is applied for the training of Peking opera gestures, which allows natural interactions and provides interactive feedback for user gestures. We conducted multiple experiments to verify the feasibility and effectiveness of the training system. The experimental results showed that the proposed system can overcome the difficulties in the traditional learning process of Peking opera gestures, which helps users to achieve the goal of learning standard Peking opera gestures. The proposed training system greatly eases the learning of Peking opera gestures, adding vitality into the culture of traditional Peking opera.
Xingquan Cai, Zaichao Lin, Yakun Ge, Haiyan Sun
Int. J. Pattern Recognit. Artif. Intell.6
2021 Terrain Elevation Map Synthesis Method based on Single Sample and User Sketch
abstract
Terrain synthesis has been a hot topic in the field of computer graphics and image processing. However, there are still issues in terrain synthesis where synthesis results are difficult to control and not realistic enough. To address these problems, this paper proposes an interactive terrain elevation map generation method based on the synthesis of a single sample terrain elevation map. First, we propose a method to extract the skeleton from a terrain elevation map and a user sketch. Second, we construct a skeleton sample feature map based on the terrain elevation map and the user sketch. Finally, we propose a matching cost function to match image patches of the terrain sample and the user sketch. The proposed method can obtain a synthesis result containing the features of both the terrain sample and the user sketch, and then generates a synthetic terrain elevation map. The experimental results demonstrate the effectiveness of the proposed method, where the synthesized results can meet the needs of users.
Xingquan Cai, Haiyan Sun, Amanda Gozho, Yakun Ge, Runbo Cai
Int. J. Pattern Recognit. Artif. Intell.3
2016 A strongly secure pairing-free certificateless authenticated key agreement protocol under the CDH assumption
Haiyan Sun, Qiaoyan Wen, Wenmin Li 0001
Sci. China Inf. Sci.1
2016 Haze removal based on multiple scattering model with superpixel algorithm
Rui Wang 0039, Haiyan Sun
Signal Process.3
2015 A strongly secure identity-based authenticated key agreement protocol without pairings under the GDH assumption
abstract
Among the existing identity-based authenticated key agreement ID-AKA protocols, there are only a few of them that can resist to leakage of ephemeral secret keys, which is about the protection of the session secret key after the ephemeral secret keys of users are compromised. However, all these ID-AKA protocols with leakage of ephemeral secret keys resistance require expensive bilinear pairing operations. In this paper, we present a pairing-free ID-AKA protocol with ephemeral secrets leakage resistance. We also provide a full proof of its security in the extended Canetti-Krawczyk model, which not only can capture resistance to leakage of ephemeral secret keys but also can capture other basic security properties such as master key forward security and key compromise impersonation resistance. Compared with the existing ID-AKA protocols, our scheme is a good trade-off between security and efficiency. Copyright © 2015 John Wiley & Sons, Ltd.
Haiyan Sun, Qiaoyan Wen, Hua Zhang 0001, Zhengping Jin
Secur. Commun. Networks1
2014 Lifetime holes aware register allocation for clustered VLIW processors
abstract
This paper presents an on-the-fly register allocator which dynamically detects and utilises lifetime holes for clustered VLIW processors. A lifetime hole is an interval in which a variable does not contain a valid value. A register holding a lifetime hole can be allocated to another variable whose live range fits in the lifetime hole, leading to more efficient utilisation of registers. We propose efficient techniques for dynamically utilising lifetime holes and incorporate these techniques into our on-the-fly register allocator. We have simulated our register allocator and a linear scan register allocator without considering lifetime holes by using the MediaBench II benchmark suite. Our simulation results show that our register allocator reduces the number of spills by 12.5%, 11.7%, 12.7%, for three different processor models, respectively.
Xuemeng Zhang, Hui Wu 0001, Haiyan Sun, Jingling Xue
DATE3
2013 A novel pairing-free certificateless authenticated key agreement protocol with provable security
Haiyan Sun, Qiaoyan Wen, Hua Zhang 0001, Zhengping Jin
Frontiers Comput. Sci.1