VLDB 2026 Research / reviewers in the wild / expert
Xiangyang Wang 0003
dblp:54/6596-3
· DBLP profile ↗
27ranked-venue papers
10as first author
15since 2021 · last 2025
0000-0003-1394-6068ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 11 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tic action recognition for children tic disorder with end-to-end video semi-supervised learning
Xiangyang Wang 0003, Rui Wang 0034, Jinhua Sun |
Vis. Comput. | 1 |
| 2024 | DFSTrack: Dual-stream fusion Siamese network for human pose tracking in videos
Xiangyang Wang 0003, Yuhui Tian, Fudi Geng, Rui Wang 0034 |
Image Vis. Comput. | 1 |
| 2024 | TQRFormer: Tubelet query recollection transformer for action detection
Xiangyang Wang 0003, Rui Wang 0034, Jinhua Sun |
Image Vis. Comput. | 1 |
| 2024 | Emotion recognition based on brain-like multimodal hierarchical perception
Xianxun Zhu, Xiangyang Wang 0003, Rui Wang 0034 |
Multim. Tools Appl. | 3 |
| 2024 | Enhanced keypoint information and pose-weighted re-ID features for multi-person pose estimation and tracking
Xiangyang Wang 0003, Tao Pei, Rui Wang 0034 |
Mach. Vis. Appl. | 1 |
| 2024 | Self-supervised Siamese keypoint inference network for human pose estimation and tracking
Xiangyang Wang 0003, Yuhui Tian, Rui Wang 0034 |
Mach. Vis. Appl. | 1 |
| 2024 | DPR-GAN: Dual-Stream Progressive Refinement for Adversarial 3D Point Cloud GenerationabstractAbstract Point cloud generation aims to transfer a latent code to realistic 3D shapes through generative models. However, most of progressive generative methods ignore the spatial relationship among different stages and suffer from the loss of contextual information. To address this issue, we propose dual-stream progressive refinement adversarial network (DPR-GAN), which utilizes a dual-stream structure to establish the relationship between two adjacent stages. Such a mechanism can learn the spatial context and preserve more spatial details of point clouds at different stages. In addition, DPR-GAN adopts 3D gridding transformation to guide the shape deformation. In this way, 3D gridding transformation can learn a reasonable correspondence between the local regions of 3D shapes and latent codes. Benefiting from the uniformity and adaptability of the 3D grids, our proposed DPR-GAN can improve the quality and consistency of generated point clouds. We conduct comprehensive experiments to demonstrate that the proposed DPR-GAN is capable of generating pluralistic point clouds, as compared with state-of-the-art generation methods in terms of both visual and quantitative evaluations. Xiangyang Wang 0003, Rui Wang 0034 |
Neural Process. Lett. | 1 |
| 2023 | Human pose estimation based on lightweight basicblock
Xiangyang Wang 0003, Rui Wang 0034 |
Mach. Vis. Appl. | 3 |
| 2023 | Enhancing multi-scale information exchange and feature fusion for human pose estimation
Rui Wang 0034, Wanyu Wu, Xiangyang Wang 0003 |
Vis. Comput. | 3 |
| 2022 | MTPose: Human Pose Estimation with High-Resolution Multi-scale Transformers
Rui Wang 0034, Fudi Geng, Xiangyang Wang 0003 |
Neural Process. Lett. | 3 |
| 2022 | Learning Enriched Global Context Information for Human Pose Estimation
Rui Wang 0034, Xiangyang Wang 0003 |
Neural Process. Lett. | 4 |
| 2022 | GA-CNN: Convolutional Neural Network Based on Geometric Algebra for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) have achieved state-of-the-art performance in hyperspectral images (HSIs) classification, which is widely used for the analysis of remotely sensed images. HSI includes spectral and spatial information from several hundreds of spectral data channels. Recent CNN models deal with various bands of HSIs as independent channels, which may lead to the loss of dependencies between different channels or the loss of associated information between each channel and the global. This article proposes a novel CNN model based on geometric algebra (GA), dubbed GA-CNN, to process the HSIs in a holistic way without losing the interrelationship among channels. Specifically, taking advantage of GA, different band images are represented as GA multivectors to capture the inherent structures and preserve the correlation of those channels. In particular, all the basic modules of our model, such as convolutional layers and the backpropagation algorithm, are extended to the GA domain. We evaluate the performance of the proposed GA-CNN model in classification tasks on four well-known HSI datasets. The experimental results indicate that our GA-CNN model outperforms traditional and state-of-the-art real-valued CNNs with higher classification accuracy and fewer model parameters. Rui Wang 0034, Yi Wang 0063, Xiangyang Wang 0003, Wenming Cao 0001, Wei Xiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | RGA-CNNs: convolutional neural networks based on reduced geometric algebra
Rui Wang 0034, Miaomiao Shen, Xiangyang Wang 0003, Wenming Cao 0001 |
Sci. China Inf. Sci. | 3 |
| 2021 | Attention Refined Network for Human Pose Estimation
Xiangyang Wang 0003, Jiangwei Tong, Rui Wang 0034 |
Neural Process. Lett. | 1 |
| 2021 | A Novel Approach of Intelligent Computing for Multiperson Pose Estimation with Deep High Spatial Resolution and Multiscale FeaturesabstractCurrently, human pose estimation (HPE) methods mainly rely on the design framework of Convolutional Neural Networks (CNNs). These CNNs typically consist of high‐to‐low‐resolution subnetworks (encoder) to learn semantic information and low‐to‐high subnetworks (decoder) to raise the resolution for keypoint localization. Because too low‐resolution feature maps in encoder will inevitably lose some spatial information, which cannot be recovered in the upsampling stages, keeping high spatial resolution features is critical for human pose estimation. On the other hand, due to scale variation of human body parts, multiscale features are also very important for human pose estimation. In this paper, a novel backbone network is proposed specifically for HPE, named High Spatial Resolution and Multiscale Networks (HSR‐MSNet), which maintain high spatial resolution features in deeper layers of the encoder and meanwhile construct multiscale features within one single residual block via subgroup splitting and fusion of feature maps. Experiments show that our approach outperforms other state‐of‐the‐art methods with more accurate keypoint locations on COCO dataset. Xiangyang Wang 0003, Yijie Shi, Chunhua Qian, Rui Wang 0034 |
Wirel. Commun. Mob. Comput. | 2 |
| 2020 | Barycentric convolution surfaces based on general planar polygon skeletons
Xiaoqiang Zhu, Chenze Song, Xiangyang Wang 0003, Lihua You, Xiaogang Jin 0001 |
Graph. Model. | 4 |
| 2020 | Enhancing feature fusion for human pose estimation
Rui Wang 0034, Jiangwei Tong, Xiangyang Wang 0003 |
Mach. Vis. Appl. | 3 |
| 2019 | Screwing assembly oriented interactive model segmentation in HMD VR environmentabstractAbstract Although different approaches of segmenting and assembling geometric models for 3D printing have been proposed, it is difficult to find any research studies, which investigate model segmentation and assembly in head‐mounted display (HMD) virtual reality (VR) environments for 3D printing. In this work, we propose a novel and interactive segmentation method for screwing assembly in the environments to tackle this problem. Our approach divides a large model into semantic parts with a screwing interface for repeated tight assembly. Specifically, after a user places the cutting interface, our algorithm computes the bounding box of the current part automatically for subsequent multicomponent semantic Boolean segmentations. Afterwards, the bolt is positioned with an improved K3M image thinning algorithm and is used for merging paired components with union and subtraction Boolean operations respectively. Moreover, we introduce a swept Boolean‐based rotation collision detection and location method to guarantee a collision‐free screwing assembly. Experiments show that our approach provides a new interactive multicomponent semantic segmentation tool that supports not only repeated installation and disassembly but also tight and aligned assembly. Xiaoqiang Zhu, Shenshuai Chen, Xiangyang Wang 0003, Lihua You, Zhigang Deng 0001, Xiaogang Jin 0001 |
Comput. Animat. Virtual Worlds | 6 |
| 2019 | Depth-aware saliency detection using convolutional neural networks
Zhi Liu 0003, Mengke Huang, Xiangyang Wang 0003 |
J. Vis. Commun. Image Represent. | 5 |
| 2017 | Brush2Model: Convolution surface-based brushes for 3D modelling in head-mounted display-based virtual environmentsabstractAbstract Easy and efficient 3D modelling in virtual environments is an important and unsolved topic. This paper proposes a new modelling approach to tackle this issue. It invents convolution surface‐based brushes to directly draw 3D models in head‐mounted display‐based virtual environments. In order to maximize the efficiency, flexibility, and capacity of our proposed modelling approach, we propose three different skeleton‐based convolution surfaces to tackle different modelling tasks: point skeleton‐based convolution surfaces for metaball shapes, line skeleton‐based convolution surfaces for cylindrical shapes, and polygon skeleton‐based convolution surfaces for planar surfaces. Their combination makes 3D modelling more flexible and powerful. The high efficiency is further raised by our developed closed‐form solutions for point skeletons, ends of line skeletons, and edges of polygonal skeletons. Different user‐friendly sweeping schemes are provided to facilitate intuitive inputs for various complex shape generation. Unlike Google's Tilt Brush, which is used to create disconnected sheet‐like surfaces only, our proposed convolution surface‐based brushes can produce smoothly blended manifold surfaces, and novice users can easily learn and use them to create various interesting 3D models efficiently. Xiaoqiang Zhu, Lihua You, Xiangyang Wang 0003, Xiaogang Jin 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2017 | Adaptive saliency fusion based on quality assessment
Xiaofei Zhou 0003, Zhi Liu 0003, Guangling Sun, Xiangyang Wang 0003 |
Multim. Tools Appl. | 4 |
| 2016 | Improving Saliency Detection Via Multiple Kernel Boosting and Adaptive FusionabstractThis letter proposes a novel framework to improve the saliency detection performance of an existing saliency model, which is used to generate the initial saliency map. First, a novel regional descriptor consisting of regional self-information, regional variance, and regional contrast on a number of features with local, global, and border context is proposed to describe the segmented regions at multiple scales. Then, regarding saliency computation as a regression problem, a multiple kernel boosting method based on support vector regression (MKB-SVR) is proposed to generate the complementary saliency map. Finally, an adaptive fusion method via learning a quality prediction model for saliency maps is proposed to effectively fuse the initial saliency map with the complementary saliency map and obtain the final saliency map with improvement on saliency detection performance. Experimental results on two public datasets with the state-of-the-art saliency models validate that the proposed method consistently improves the saliency detection performance of various saliency models. Xiaofei Zhou 0003, Zhi Liu 0003, Guangling Sun, Linwei Ye, Xiangyang Wang 0003 |
IEEE Signal Process. Lett. | 5 |
| 2015 | Classification of MRI under the Presence of Disease Heterogeneity using Multi-Task Learning: Application to Bipolar Disorder
Xiangyang Wang 0003, Tiffany M. Chaim, Marcus V. Zanetti, Christos Davatzikos |
MICCAI (1) | 1 |
| 2013 | The LogitBoost Based on Joint Feature for Face DetectionabstractIn this paper, joint Haar-like feature is used for detecting faces in images. Our method is based on the co-occurrence Haar-like features which can capture the structural characteristic of the face, make it possible to construct more effective weak classifier. As the Haar-like, the joint Haar-like feature can be calculated very effective and has the robustness to addition of noise and change in illumination. The face detector is learned by stage wise selection which is different with Viola and Jones Detector is that we use LogitBoost. We perform two experiments: In the Experiment 1, we show that the LogitBoost [9] obtain higher performance than AdaBoost [2] [4] [11]. We have confirmed that our method based on LogitBoost yielded higher performance than AdaBoost. In the Experiment 2, the performance increase according with the num of the combined features. However, the time spent increase at huge growth rate, too. Therefore, we will research that how we can choose the optimal F at the acceptable training time in the latter work. Shishi Duan, Xiangyang Wang 0003 |
ICIG | 2 |
| 2013 | Object Tracking with Sparse Representation and Annealed Particle FilterabstractIn this paper, we propose a new visual tracking algorithm, SRAPF, for object tracking, which is based on sparse representation and annealed particle filter. To find the tracking target at a new frame, each target candidate is sparsely represented by target templates and trivial templates. The sparsity is achieved by solving a l1-regularized least squares problem. After that, Instead of tracking objects in the common particle filter framework, we solve the sparse representation problem in an annealed particle filter framework. Then the candidate with the largest likelihood is taken as the tracking target. In the APF framework, the sampling covariance and annealing factor items are incorporated into the tracking process. The annealing strategy can achieve "Smart sampling" to avoid generating invalid particles corresponding to impossible target object. Both qualitative and quantitative evaluations on challenging image sequences demonstrate that the proposed tracking algorithm performs better in comparison with the L1 tracking algorithm. Xiangyang Wang 0003 |
ICIG | 2 |
| 2012 | Multi-Task low-rank and sparse matrix recovery for human motion segmentationabstractThis paper proposes a new algorithm, named Multi-Task Robust Principal Component Analysis (MTRPCA), to collaboratively integrate multiple visual features and motion priors for human motion segmentation. Given the video data described by multiple features, the human motion part is obtained by jointly decomposing multiple feature matrices into pairs of low-rank and sparse matrices. The inference process is formulated as a convex optimization problem that minimizes a constrained combination of nuclear norm and ℓ2,1-norm, which can be solved efficiently with Augmented Lagrange Multiplier (ALM) method. Compared to previous methods, which usually make use of individual features, the proposed method seamlessly integrates multiple features and priors within a single inference step, and thus produces more accurate and reliable results. Experiments on the HumanEva human motion dataset show that the proposed MTRPCA is novel and promising. Xiangyang Wang 0003, Guangcan Liu |
ICIP | 1 |
| 2007 | Feature selection based on rough sets and particle swarm optimization
Xiangyang Wang 0003, Jie Yang 0002, Xiaolong Teng, Weijun Xia, Richard Jensen |
Pattern Recognit. Lett. | 1 |