Xingquan Cai

dblp:22/4065 · DBLP profile ↗
← Back
44ranked-venue papers
34as first author
40since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 16 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 12 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 CRAFT: Calibrated Robust Adaptive Framework for Texture-Heterogeneity
Xingquan Cai, Yaoyao Xing, Dong Lin, Ying Li 0039
ICIC (12)1
2026 Robust Vehicle and Pedestrian Detection in Adverse Weather via a Weather-Aware Transformer
abstract
Pedestrian and vehicle in adverse weather detection technology based on deep learning has attracted increasing attention. However, YOLO-based detectors degrade under low contrast and noise, causing missed small or occluded targets, while DETR-based models suffer from slow convergence and limited multi-scale sensitivity. To address these challenges, a robust weather-aware transformer for vehicle and pedestrian detection, RWT-DETR, is proposed. Firstly, a Strip-Oriented Feature Gating (SOFG) module is designed to improve the capture of strip-like and directional features of traffic targets. Secondly, a Signed Attention Decomposition (SAD) mechanism is designed to enable the encoder to model both positive and negative correlations, thereby enhancing the model's robustness against noise interference and feature blurring. Finally, an Adaptive Multi-scale Fusion with Gating (AMFG) module is proposed to achieve dynamic calibration and denoising of cross-scale features, significantly improving semantic alignment. Experimental results show that, compared with the baseline, the proposed method improves [email protected] and [email protected]:0.95 by 2.4% and 3.0%, respectively. Meanwhile, it outperforms existing mainstream models, confirming its effectiveness under adverse weather conditions.
Xingquan Cai, Yuanxiang Zhao, Ying Li 0039
IEEE Signal Process. Lett.1
2026 CAS-ODE: Jointly Learning Adaptive Structures and Continuous Dynamics for Emotion Recognition in Conversation
abstract
Multimodal emotion recognition in conversation (MERC) plays a crucial role in applications such as empathetic dialogue systems, intelligent tutoring, and mental health monitoring. Existing graph neural network (GNN)-based approaches often rely on fixed conversational topologies, which may introduce redundant message passing and fail to capture evolving high-order dependencies. In addition, conventional GNNs model temporal dynamics in discrete steps, leading to oversmoothing and a limited ability to represent continuous emotional changes. To address these limitations, we propose CAS-ODE (Contrastive Adaptive Spatiotemporal ODE Network), a novel framework that jointly learns adaptive conversational structures and continuous temporal dynamics. CAS-ODE employs a variational hypergraph autoencoder with contrastive regularization to infer robust high-order relations, while a graph ordinary differential equation models the smooth evolution of emotional states. Extensive experiments on the IEMOCAP and MELD benchmarks validate the superiority of CAS-ODE, which achieves state-of-the-art performance and highlights the effectiveness of synergistically learning conversational structures and their continuous temporal evolution.
Guangzi Zhang, Xingquan Cai
IEEE Signal Process. Lett.5
2026 JC-STNet: a physics-inspired joint-centric spatiotemporal network for real-time 2D pose estimation
Xingquan Cai, Kaijie Qu, Chuansheng Liu, Xinzhu Pu
Vis. Comput.1
2025 MatLayerNet: A Multi-agent-Based Method for Text-to-PBR Material Generation
Xingquan Cai
CGI (3)1
2025 Dynamic Prompting and Cross-Modal Attention for Context-Aware Multimodal Emotion Recognition
Guangzi Zhang, Shanshan He, Xingquan Cai
CGI (2)5
2025 Oracle Bone Inscriptions Recognition Based on Spatial Transformer Network and Few-Shot Learning
Xingquan Cai, Lixin Ding, Haiyan Sun
ICIC (2)1
2025 Dynamic Convolution and Dimensional Joint Attention Based Denoising of Oracle Topography Images
Xingquan Cai, Mengrui Dai, Haiyan Sun
ICIC (3)1
2025 Interpretable Action Quality Assessment with Temporal Parsing
Xingquan Cai, Haiyan Sun
ICIC (5)1
2025 HDF-YOLO: A High-Precision Ship Detection Method in SAR Images Based on Improved YOLOv11
Xingquan Cai, Lixin Ding
ICIC (5)1
2025 Abnormal Behavior Detection Method in Surveillance Videos Based on a Lightweight Cross-Modal Attention Network
Xingquan Cai, Shanshan He
IJCNN1
2025 MH-CL: Self-Supervised Action Recognition Method Based on Multi-view Hypergraphic Contrastive Learning for Intangible Cultural Heritage Dance
abstract
Compared to daily movements, dance movements of the intangible cultural heritage exhibit complexities and intricate coordination between distant joints, such as the coordinated movements required in the three bends of the Dai dance involving the feet, knees, hips, and arms. These characteristics pose challenges for existing recognition methods, resulting in low accuracy. To address this, we propose a self-supervised action recognition method for intangible cultural heritage dance based on multi-view hypergraph contrastive learning. We begin by inputting the 3D human pose sequence of the dance video, enhancing it with data through multi-view and random frame methods. Subsequently, an action encoder network is constructed to extract action features using joint dynamic and static hypergraph convolution. A multilayer projection head is then designed to map the feature representations to a one-dimensional space, optimizing the encoder loss by comparing the loss functions. The experimental results demonstrate the superior performance of our approach over the QDPuzzle [1] and AimCLR [2] algorithms under two evaluation protocols in the NTU RGB + D 60 [3], NTU RGB+D 120 [4], and PKUMMD [5] datasets. In addition, ablation experiments show the ability of the method to recognize more accurate 3D human pose sequences, particularly in intangible cultural heritage dance videos featuring distant joint coordination and addressing the issue of missing segments.
Xingquan Cai, Pengyan Cheng, Kaijie Qu, Haiyan Sun
IJCNN1
2025 A Multi-Person Motion Prediction Method Based on Multi-Angle Coding of Joint-Relation for Intangible Cultural Heritage Dance Videos
abstract
3D multi-person pose prediction is a key computer vision task. The task involves not only the motion trajectories of individuals, but also the complex interactions between individuals, which is especially important in intangible cultural heritage dance scenes. Although existing methods have improved performance, they still face challenges such as insufficient capture of individuals and interaction details. Therefore, we propose a multi-person motion prediction method based on multi-angle coding of joint-relation for intangible cultural heritage dance videos. Firstly, in order to effectively capture joint and relation features, we adapt a multi-angle coding strategy, which integrates displacement and time information in the joint branch, and uses connectivity matrix and interaction magnitude matrix in the relation branch; then, in order to realize effective information transfer between joint and relation, we design the Joint-Relation Fusion Module (JRFM), which allows the joint branch to update the joint features by querying the information of the relation branch using an attention mechanism; at the same time, the relation branch further optimizes the relation features by local updating, thus realizing the effective fusion of the joint and relation features; finally, we decode the fused joint-relation features to output the predicted 3D pose sequence. Experiments show that compared to the UnityGraph [1] method, our method reduces the MPJPE metrics by 0.7mm, 0.8mm, 0.3mm, and 0.4mm at 3s for the CMU-Mocap, MuPoTS-3D, and Mix1 & Mix2(9~15 persons) datasets, respectively, and the VIM metrics on the 3DPW dataset by 0.5 mm, the results of these evaluations clearly demonstrate the superiority of our method.
Xingquan Cai, Kaijie Qu, Mengrui Dai
IJCNN1
2025 Adaptive Hypergraph-Based 3D Multi-Person Pose Estimation Method for Intangible Cultural Heritage Dance Videos
abstract
Despite recent advancements, 3D multi-person pose estimation from monocular videos remains challenging due to common issues such as occlusions caused by clothing and limbs, as well as inaccuracies in person detection. Current 3D multi-person pose estimation methods typically treat individuals as independent entities for estimation. This methodology has significant limitations, particularly in its failure to fully account for the rich interactive information between individuals and within limbs. To address these challenges, we propose a novel method for monocular 3D multi-person pose estimation. Firstly, it constructs Intra-Hypergraph and Inter-Hypergraph to represent inter-limb and inter-individual interaction information. Subsequently, we design adaptive hypergraph convolutional networks to extract spatial features from 2D human pose sequences. Finally, after the temporal attention module outputs the coordinates of the predicted 3D joint points. Quantitative and qualitative evaluations demonstrate the effectiveness of the proposed method.
Xingquan Cai, Kaijie Qu, Mengrui Dai, Ying Li 0039
ICMR1
2025 Multi-Agent Learning With Hierarchical Biomechanical Priors for Efficient 3D Human Pose Estimation in Virtual Reality
abstract
ABSTRACT In virtual reality (VR) applications, real‐time and robust 3D human pose estimation is paramount to enhance user experience, yet existing methodologies often encounter challenges such as high computational burden, occlusion sensitivity, and inadequate adaptation to complex actions. To mitigate these issues, we propose a novel 3D human pose estimation method based on a multi‐agent hierarchical biomechanical priors architecture. This method achieves efficient heatmap prediction in the local feature space through a parallel agents architecture, while simultaneously integrating a hierarchical loss function and dynamic context modeling. It incorporates virtual avatars' geometric constraints into network training, thereby enhancing pose plausibility and effectively addressing occlusion and intricate actions. Moreover, it substantially improves cross‐frame stability and estimation accuracy in occlusion scenarios through multiview spatiotemporal consistency optimization. Compared to existing methods, the proposed framework provides adaptability to the unique demands of virtual environments with reduced computational cost. We experimentally validate our approach on the widely used Human3.6M and MPI‐INF‐3DHP datasets, and further demonstrate through ablation experiments that the dynamic occlusion compensation module, which fuses multimodal perception with a spatio‐temporal diffusion mechanism, significantly enhances the robustness of pose estimation under occlusion scenarios with virtual costumes.
Xingquan Cai, Kaijie Qu, Shanshan He
Comput. Animat. Virtual Worlds1
2025 3D human pose estimation using spatiotemporal hypergraphs and its public benchmark on opera videos
Xingquan Cai, LiZhe Chen, YiJie Wu, Haiyan Sun
Vis. Comput.1
2024 Retinex-Based Low-Light Mural Image Enhancement with Color Correction
Xingquan Cai, Haiyan Ma, Haiyan Sun
CGI (1)1
2024 Decoupled Estimation of Human Pose and Shape for ICH Performance Video Based on L-C-HRNet
Xingquan Cai, Shike Liu, Haiyan Sun
CGI (1)1
2024 A Novel Auxiliary Task Framework in 3D Human Pose Estimation for Opera Videos
abstract
Influenced by the costume and limb occlusion of dance movements in opera videos, the 2D human pose estimation methods struggle to accurately locate the 2D joint coordinates of the occluded parts. This inaccuracy leads to lower precision in estimating 3D human poses from 2D joint coordinates in opera videos. To enhance the learning of more effective spatio-temporal dependencies from 2D joint coordinates and improve the accuracy of 3D human pose estimation, this paper proposes a novel auxiliary task framework. We first designed three auxiliary tasks to mask some of the 2D joint coordinates, disorder the spatial position of the joints and the video frame order, with the goal of recovering the corrupted 2D joint coordinates. Then to address these auxiliary tasks, we propose a multi-feature representation Transformer network to capture 2D joint coordinates spatio-temporal features from local to global perspective by constructing local adaptive graph convolution network, segmented time-aware network and global spatio-temporal self-attention module respectively. Finally, an adaptive weight allocation module is utilized to integrate local and global features to output the 3D joint coordinates. Extensive comparative and ablation experiments on the Human3.6M, MPI-INF-3DHP and opera datasets demonstrate that our method surpasses all comparative methods in MPJPE accuracy. Furthermore, the auxiliary task framework designed in this paper effectively captures comprehensive and efficient spatio-temporal dependencies in 2D joint coordinates from opera videos.
Xingquan Cai, Shanshan He, Haoyu Song 0005, Haiyan Sun
ICMR1
2024 Frequency-importance gaussian splatting for real-time lightweight radiance field rendering
Lizhe Chen, Yuyao Ge, Xingquan Cai
Multim. Tools Appl.6
2023 An Image Extraction Method for Traditional Dress Pattern Line Drawings Based on Improved CycleGAN
Xingquan Cai, Sichen Jia, Jiali Yao, Haiyan Sun
CGI1
2023 Hand Movement Recognition and Analysis Based on Deep Learning in Classical Hand Dance Videos
Xingquan Cai, Qingtao Lu, Fajian Li, Shike Liu
CGI (3)1
2023 An Ancient Murals Inpainting Method Based on Bidirectional Feature Adaptation and Adversarial Generative Networks
Xingquan Cai, Qingtao Lu, Jiali Yao
CGI1
2023 GFENet: Group-Free Enhancement Network for Indoor Scene 3D Object Detection
Feng Zhou 0007, Ju Dai, JunJun Pan, Mengxiao Zhu 0004, Xingquan Cai, Chen Wang 0043
CGI (3)5
2023 A Driver Abnormal Behavior Detection Method Based on Improved YOLOv7 and OpenPose
Xingquan Cai, Jiali Yao, Pengyan Cheng
ICIC (5)1
2023 Abnormal Behavior Detection Method based on Spatio-temporal Dual-flow Network for Surveillance Videos
abstract
To address the issues that current surveillance videos abnormal behavior detection methods are affected by complex environments such as video blurring and long distances, resulting in low detection accuracy and slow speed, we propose an abnormal behavior detection method based on spatio-temporal dual-flow network for surveillance videos. Firstly, the input surveillance videos are sampled with different frame rates and a spatio-temporal dual-flow partial convolution network is constructed to extract spatial and motion information features from the spatial and temporal flow video sequences, respectively. Then, a cross-modal dual-attention fusion mechanism is introduced after each feature extraction of the dual-flow partial convolution network to enhance the exchange of feature information between spatial and temporal flows. Finally, the extracted motion and spatial features are fused to output detection results. Experiments show that our method reduces the number of floating-point operations and model parameters on the Kinetics-600 dataset compared to the MViT-L algorithm with guaranteed accuracy. Compared to lightweight networks such as X3D-XL, it improves the accuracy in Top-1 and Top-5 by 4.3% and 3.6%, respectively. The experiments prove that the method can reduce the model’s computational complexity while obtaining more accurate abnormal behavior detection results when the surveillance videos is blurred or the abnormal behavior occurs far away.
Xingquan Cai, Haiyan Sun
ICTAI1
2023 An Extended Labanotation Generation Method Based on 3D Human Pose Estimation for Intangible Cultural Heritage Dance Videos
abstract
To address the issues of low accuracy in existing 3D human pose estimation (HPE) methods and the limited level of details in Labanotation, we propose an extended Labanotation generation method for intangible cultural heritage dance videos based on 3D HPE. First, a 2D human pose sequence of the performer is inputted along with spatial location embeddings, where multiple spatial transformer modules are employed to extract spatial features of human joints and generate cross-joint multiple hypotheses. Afterward, temporal features are extracted by a self-attentive module and the correlation between different hypotheses is learned using bilinear pooling. Finally, the 3D joint coordinates of the performer are predicted, which are matched with the corresponding extended Labanotation symbols using the Laban template matching method to generate extended Labanotation. Experimental results show that, compared with VideoPose and CrossFormer algorithms, the Mean Per Joint Position Error (MPJPE) of the proposed method is reduced by 3.7[Formula: see text]mm and 0.6[Formula: see text]mm, respectively on Human3.6M dataset, and the generated extended Labanotation can better describe the movement details compared with the basic Labanotation.
Xingquan Cai, Pengyan Cheng, Jiali Yao
Int. J. Pattern Recognit. Artif. Intell.1
2023 A Method for 3D Human Pose Estimation and Similarity Calculation in Tai Chi Videos
abstract
Human pose estimation from video sequences has become a hot research topic in the domain of robotics and computer vision. However, existing three-dimensional (3D) pose estimation methods usually analyze individual frames, which have a low accuracy due to various human movement speed, limiting its practical application. In this paper, we propose a method for estimating 3D pose and calculating similarity from Tai Chi video sequences based on Seq2Seq network. Specifically, using 2D joint point coordinate sequence of the original image as input, our method constructs an encoder and a decoder to build a Seq2Seq network. Our method introduces an attention mechanism for weighing the input data to obtain an intermediate vector and decode it to estimate the 3D joint point sequence. Afterwards, using a template video and a target video as input, our method calculates the cost of passing through each point within the constraints to construct a cost matrix for video similarity. With the cost matrix, our method can determine the optimal path and use the correspondence of the video sequence to calculate the image similarity of the corresponding frame. The experimental data show that the proposed method can effectively improve the accuracy of 3D pose estimation, and increase the speed for video similarity calculation.
Xingquan Cai, Yuqing Huo, Haiyan Sun, Jiaqi Ji
Int. J. Pattern Recognit. Artif. Intell.1
2023 An Automatic Music-Driven Folk Dance Movements Generation Method Based on Sequence-To-Sequence Network
abstract
Music-driven automatic dance movement generation has become a hot research topic in the field of computer vision and internet of things in the recent past. To address the problems of increasing loss of Chinese folk dance culture, high cost of manual choreography methods and requirements for professional background, this paper proposes an automatic generation method for folk dance movements. Firstly, the proposed method collects paired folk music and dance videos to construct a synchronized folk music–dance dataset, extracting music and dance features using a feature extraction tool and a multi-scale fusion high-resolution network, respectively. Afterward, a sequence-to-sequence network model is constructed and then trained based on music features and dance features to synthesize rhythmically matched dance sequences for new music clips. Finally, an easy-to-use and effective automatic folk dance choreography method is implemented. Experimental data show that the proposed method performs well in automatic folk dance generation and the generated dances have folk characteristics and match the rhythm of the given music.
Xingquan Cai, Mengyao Xi, Sichen Jia, Haiyan Sun
Int. J. Pattern Recognit. Artif. Intell.1
2023 A Social Distance Monitoring Method Based on Improved YOLOv4 for Surveillance Videos
abstract
Social distance monitoring is of great significance for public health in the era of COVID-19 pandemic. However, existing monitoring methods cannot effectively detect social distance in terms of efficiency, accuracy, and robustness. In this paper, we proposed a social distance monitoring method based on an improved YOLOv4 algorithm. Specifically, our method constructs and pre-processes a dataset. Afterwards, our method screens the valid samples and improves the K-means clustering algorithm based on the IoU distance. Then, our method detects the target pedestrians using a trained improved YOLOv4 algorithm and gets the pedestrian target detection frame location information. Finally, our method defines the observation depth parameters, generates the 3D feature space, and clusters the offending aggregation groups based on the L2 parametric distance to finally realize the pedestrian social distance monitoring of 2D video. Experiments show that the proposed social distance monitoring method based on improved YOLOv4 can accurately detect pedestrian target locations in video images, where the pre-processing operation and improved K-means algorithm can improve the pedestrian target detection accuracy. Our method can cluster the offending groups without going through calibration mapping transformation to realize the pedestrian social distance monitoring of 2D videos.
Xingquan Cai, Pengyan Cheng, Dingwei Feng, Haiyan Sun, Jiaqi Ji
Int. J. Pattern Recognit. Artif. Intell.1
2023 A Multi-Object Detection Method Based on Adaptive Feature Adjustment of 3D Point Cloud in Indoor Scenes
abstract
Due to the complexity and diversity of indoor environment objects and interference occlusions, the accuracy of multi-object target detection based on 3D point cloud is limited. To address this issue, we present a multi-target detection method based on adaptive feature adjustment (AFA) of 3D point cloud. First, our method preprocesses the dataset and constructs a backbone module. Afterwards, our method uses an improved PointNet[Formula: see text] network for feature adaptive learning, where an AFA module is added to learn the influence relationship between point pairs. The proposed method then establishes the relationship between contexts in the local point set area and extracts the feature of point cloud. Using the idea of Hough voting, our method can generate some votes close to the particle. Using these votes to generate proposal, the proposed method adds CBAM attention mechanisms to both modules of voting and proposal, which can fuse the feature information of the channel and expand the receptive field in space. Our method can enhance the important features and weaken the unimportant features, making the extracted features more directional and enhancing the expressiveness of the network. Finally, the generated results are visualized to complete the multi-target detection of 3D point cloud. To verify the effectiveness of our proposed method, two large datasets with real 3D scanning, scanNet2 and SunRGB-D, are used for training the network. The experimental results show that the proposed method can improve the effectiveness of point cloud target detection in indoor scenes, getting a higher detection accuracy.
Haiyan Sun, Keng Chen, Sichen Jia, Xingquan Cai
Int. J. Pattern Recognit. Artif. Intell.5
2023 An Improved PoinTr Point Cloud Completion Method Based on Feature Enhancement
abstract
To address the issue that point cloud data is often incomplete and difficult to obtain, we propose a point cloud completion method to improve the PoinTr method based on feature enhancement. In dataset preprocessing, the farthest point of the original point cloud is sampled to obtain the central point coordinates. Our method constructs an MLP network, where the local information of these central points is obtained and the location embedding is performed. Combining network and SENet network, the local features of the point cloud are extracted and enhanced, and the location embedding and local features are added to obtain the point proxies of the original point cloud. Afterward, our method predicts the missing part of the point cloud by using an Encoder to model the relationship between the point cloud structure information and points, and then using a Decoder to learn the relationship between the missing and existing parts of the point cloud and reconstruct the missing point cloud. Our method also modifies the attention mechanism to make the features more global and enhance the network expression. Finally, the point cloud is refined, and is realized by predicting multiple points around each point of the coarse point cloud through the FoldingNet network, and the final output is the complete point cloud. Experimental results show that the proposed method can not only reduce the performance overhead, but also improve the effects of point cloud completion.
Haiyan Sun, Zaichao Lin, Qingtao Lu, Sichen Jia, Xingquan Cai
Int. J. Pattern Recognit. Artif. Intell.5
2023 A Digital Simulation and Re-Editing Method for Clothing Patterns Based on Deep Learning and Somatosensory Interaction
abstract
To address the issues in clothing pattern style migration, this paper proposes a digital simulation and re-editing method for clothing patterns based on deep learning and somatosensory interaction. First, the proposed method encodes the black-and-white line drawing image, generating random noise images through a diffusion process, introducing color information for synthesis, and using a decoder to reconstruct a colored image. Afterwards, an improved VGG19 model is used to reconstruct content features and perform linear color transformation on style images, enabling pattern style migration through the construction of a Gram matrix and resulting in colored clothing texture patterns. Finally, a KinectV2 is utilized for fabric simulation, overlaying colorful clothing texture patterns to achieve 3D virtual dressing. The experimental results show that the proposed method improves the structural similarity index measure (SSIM) by 9–11% and the peak signal-to-noise ratio (PSNR) by 3–8% when compared to existing algorithms. The experiments provide evidence that the proposed method effectively mitigates color overflow, delivers precise image coloring, and accomplishes realistic restoration of clothing texture. Furthermore, the method offers an improved garment fit to fulfill the user’s interaction requirements.
Haiyan Sun, Jiali Yao, Xingquan Cai
Int. J. Pattern Recognit. Artif. Intell.5
2023 Automatic generation of Labanotation based on human pose estimation in folk dance videos
Xingquan Cai, Sichen Jia, Haiyan Sun
Neural Comput. Appl.1
2022 MCGNet: Multi-Level Context-aware and Geometric-aware Network for 3D Object Detection
abstract
Hough voting based on PointNet++ [1] is effective against 3D object detection, which has been verified by VoteNet [2], H3DNet [3], etc. However, we find there is still room for improvements in two aspects. The first is that most existing methods ignores the particular significance of different format inputs and geometric primitives for predicting object proposals. The second is that the feature extracted by PointNet++ overlooks contextual information about each object. In this paper, to tackle the above issues, we introduce MCGNet to learn multi-level geometric-aware and scale-aware contextual information for 3D object detection. Specifically, our network mainly consists of the baseline module based on H3DNet, geometric-aware module, and context-aware module. The baseline module feeding with four-types inputs (Point, Edge, Surface, and Line) concentrates on extracting diversified geometric primitives, i.e., BB centers, BB face centers, and BB edge centers. The geometric-aware module is proposed to learn the different contributions among the four-types feature maps and the three geometric primitives. The context-aware module aims to establish long-range dependencies features for either four-types feature maps or three geometric primitives. Extensive experiments on two large datasets with real 3D scans, SUN RGB-D and ScanNet datasets, demonstrate that our method is effective against 3D object detection.
Keng Chen, Feng Zhou 0007, Ju Dai, Pei Shen, Xingquan Cai, Fengquan Zhang
ICIP5
2022 A Low Distortion Mesh Parameterization Mapping Method Based on Proxy Function and Combined Newton
abstract
To address the issues of low efficiencies and serious mapping distortions in current mesh parameterization methods, we present a low distortion mesh parameterization mapping method based on proxy function and combined Newton’s method in this paper. First, the proposed method calculates visual blind areas and distortion prone areas of a 3D mesh model, and generates a model slit. Afterwards, the method performs the Tutte mapping on the cut three-dimensional mesh model, measures the mapping distortion of the model, and outputs a distortion metric function and distortion values. Finally, the method sets iteration parameters, establishes a reference mesh, and finds the optimal coordinate points to get a convergent mesh model. When calculating mapping distortions, Dirichlet energy function is used to measure the isometric mapping distortion, and MIPS energy function is used to measure the conformal mapping distortion. To find the minimum value of the mapping distortion metric function, we use an optimal solution method combining proxy functions and combined Newton’s method. The experimental data show that the proposed method has high execution efficiency, fast descending speed of mapping distortion energy and stable optimal value convergence quality. When a texture mapping is performed, the texture is evenly colored, close laid and uniformly lined, which meets the standards in practical applications.
Xingquan Cai, Dingwei Feng, Mohan Cai, Haiyan Sun
Int. J. Pattern Recognit. Artif. Intell.1
2022 Real-Time Leaf Recognition Method Based on Image Segmentation and Feature Extraction
abstract
Leaf recognition has been an important research field of image recognition in the recent past. However, traditional leaf recognition methods can be easily affected by environments and cannot realize multi-leaf recognition under a complex background in real time. In this work, we present a real-time leaf recognition method based on image segmentation and feature recognition. First, we denoise the input of a leaf image, performing a leaf segmentation with an improved FCN network model, and then optimize the contour edge with a CRF algorithm to get a leaf segmentation image. Second, we extract the content features of the segmented leaf image with an Inception-V2 network model to get a feature map of the leaf image. Third, we input the feature map into an RPN network to obtain a set of regional candidate frames and then integrate the feature map and the information of candidate frames in a RoI Pooling layer, which can extract the feature map of a candidate frame area and scale it to a fixed-size feature map. Finally, we send the feature map to a fully connected layer to classify each preselection box content through the calculation of preselection feature maps, and then obtain the final accurate position of the prediction box by utilizing a bounding box regression. The experimental results show that the proposed method can achieve multi-leaf recognitions with high accuracy and fast speed under complex environments in real time.
Xingquan Cai, Yuqing Huo, Yunbo Chen, Mengyao Xi, Yuxin Tu, Haiyan Sun
Int. J. Pattern Recognit. Artif. Intell.1
2022 Image Attribute Migration Based on Decoupling and Adaptive Layer Instance Normalization
abstract
The issue of image attribute migration is one of the hot research topics in the field of computer vision, which has received extensive research interest. However, current unsupervised image attribute migration models using symmetric generative adversarial network structure do not work well on datasets with large geometric variations, where the results lack diversity and are of low quality. To address these problems, we present an image attribute migration model based on decoupling and adaptive layer instance normalization. First, a codec structure based on a decoupled representation is constructed as the generator, and an adaptive layer instance normalization operation is used in the decoder. Then, the iterations of the model are constrained by various improved loss functions. We conducted controlled experiments and compared the results of our method with other methods using several datasets with large geometric variations. The experimental results demonstrate that the proposed method can achieve high quality and diverse image attribute migration.
Xingquan Cai, Fajian Li, Keng Chen, Yuechao Wei, Haiyan Sun
Int. J. Pattern Recognit. Artif. Intell.1
2022 POGT: A Peking Opera Gesture Training System Using Infrared Sensors
abstract
Peking opera is one of the national cultural heritages in China. However, it is difficult for people to learn the gestures in Peking opera performance, which limits the spread of this traditional culture. To address this issue, we propose a Peking opera gesture training system using infrared sensors. Specifically, we build a character avatar for demonstrating the gestures in Peking opera in the proposed system. Based on the data collected by infrared sensors, a method for calculating gesture similarity is proposed and is applied for the training of Peking opera gestures, which allows natural interactions and provides interactive feedback for user gestures. We conducted multiple experiments to verify the feasibility and effectiveness of the training system. The experimental results showed that the proposed system can overcome the difficulties in the traditional learning process of Peking opera gestures, which helps users to achieve the goal of learning standard Peking opera gestures. The proposed training system greatly eases the learning of Peking opera gestures, adding vitality into the culture of traditional Peking opera.
Xingquan Cai, Zaichao Lin, Yakun Ge, Haiyan Sun
Int. J. Pattern Recognit. Artif. Intell.1
2021 Terrain Elevation Map Synthesis Method based on Single Sample and User Sketch
abstract
Terrain synthesis has been a hot topic in the field of computer graphics and image processing. However, there are still issues in terrain synthesis where synthesis results are difficult to control and not realistic enough. To address these problems, this paper proposes an interactive terrain elevation map generation method based on the synthesis of a single sample terrain elevation map. First, we propose a method to extract the skeleton from a terrain elevation map and a user sketch. Second, we construct a skeleton sample feature map based on the terrain elevation map and the user sketch. Finally, we propose a matching cost function to match image patches of the terrain sample and the user sketch. The proposed method can obtain a synthesis result containing the features of both the terrain sample and the user sketch, and then generates a synthetic terrain elevation map. The experimental results demonstrate the effectiveness of the proposed method, where the synthesized results can meet the needs of users.
Xingquan Cai, Haiyan Sun, Amanda Gozho, Yakun Ge, Runbo Cai
Int. J. Pattern Recognit. Artif. Intell.1
2020 Immersive Interactive Virtual Fish Swarm Simulation Based on Infrared Sensors
abstract
Virtual simulation and 3D interaction have shown great potentials in a variety of domains for our future life. For a virtual fish swarm simulation system, the simulation of cohesion behaviors of fish swarm and the interaction between human and fish swarm are two key components to create immersive interactive experiences. However, it is a huge challenge to create a realistic fish swarm simulation system while providing a natural and comfortable interaction. In this paper, we propose a method for immersive virtual fish swarm simulation based on infrared sensors. Based on dynamic weight constraints, we propose a particle swarm optimization method for fish swarm cohesion simulation, which separates a particle swarm by the state of each particle and dynamically controls the particle swarm, making the movement behavior of virtual fish more realistic. In addition, an interactive fast skinning method is proposed for cartoon fishes, which leverages image segmentation, Optical Character Recognition (OCR) and bone skinning are used to generate cartoon fishes based on user-created colors. With infrared sensors, we propose a method for virtual fish swarm interaction, where the positions of human skeleton are processed by an action analyzer, achieving real-time user interactions with fish swarms. With all the proposed techniques integrated in a system, the experimental results show that our method is feasible and effective.
Xingquan Cai, Yakun Ge, Honghao Buni
Int. J. Pattern Recognit. Artif. Intell.1
2018 Real-Time Calibration and Registration Method for Indoor Scene with Joint Depth and Color Camera
abstract
Traditional vision registration technologies require the design of precise markers or rich texture information captured from the video scenes, and the vision-based methods have high computational complexity while the hardware-based registration technologies lack accuracy. Therefore, in this paper, we propose a novel registration method that takes advantages of RGB-D camera to obtain the depth information in real-time, and a binocular system using the Time of Flight (ToF) camera and a commercial color camera is constructed to realize the three-dimensional registration technique. First, we calibrate the binocular system to get their position relationships. The systematic errors are fitted and corrected by the method of B-spline curve. In order to reduce the anomaly and random noise, an elimination algorithm and an improved bilateral filtering algorithm are proposed to optimize the depth map. For the real-time requirement of the system, it is further accelerated by parallel computing with CUDA. Then, the Camshift-based tracking algorithm is applied to capture the real object registered in the video stream. In addition, the position and orientation of the object are tracked according to the correspondence between the color image and the 3D data. Finally, some experiments are implemented and compared using our binocular system. Experimental results are shown to demonstrate the feasibility and effectiveness of our method.
Fengquan Zhang, Tingsheng Lei, Xingquan Cai, Xuqiang Shao, Jian Chang 0001, Feng Tian 0009
Int. J. Pattern Recognit. Artif. Intell.4
2017 Real-time single camera natural user interface engine development
Wei Song 0004, Xingquan Cai, Yulong Xi, Seoungjae Cho, Kyungeun Cho
Multim. Tools Appl.2
2008 Complex Effects Simulation Based Large Particles System on GPU
Xingquan Cai, Zhitong Su
ISNN (2)1