VLDB 2026 Research / reviewers in the wild / expert
Xuna Wang
dblp:267/7777
· DBLP profile ↗
12ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From static to dynamic: a hybrid directed graph neural network approach for entity recognition in information propagation networks
Xuna Wang, Qingmei Tan |
Appl. Intell. | 1 |
| 2026 | Advancing Automated Reporting in Manufacturing With Dual-Channel GCN and CNN Architectures
Xuna Wang, Hong Lyu |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2025 | Exploring Complexity-Calibrated morphological distribution for whole slide image classification and difficulty-grading
Xuna Wang, Weiming Fan, Yuping Guo, Junfen Fu, Yingke Xu |
Medical Image Anal. | 2 |
| 2025 | Recognizing Video Activities in the Wild via View-to-Scene Joint LearningabstractRecognizing video actions in the wild is challenging for visual control systems. In-the-wild videos show actions not seen in training data, recorded from various angles and scenes with the same labels. Most existing methods address this challenge by developing complex frameworks to extract spatiotemporal features. To achieve view robustness and scene generalization cost-effectively, we explore view consistency and scene joint understanding. Based on this, we propose a neural network (called Wild-VAR) to learn view and scene information jointly without any 3D pose ground truth labels, a new approach to recognizing video actions in the wild. Unlike most existing methods, first, we propose a Cubing module to self-learn body consistency between views instead of comprehensive image features, boosting the generalization performance of across-view settings. Specifically, we map 3D representations to multiple 2D features and then adopt a self-adaptive scheme to constrain 2D features from different perspectives. Moreover, we propose temporal neural networks (called T-Scene) to develop a recognizing framework, enabling Wild-VAR to flexibly learn scenes across time, including key interactors and context, in video sequences. Extensive experiments show that Wild-VAR consistently outperforms state-of-the-art methods on four benchmarks. Notably, with only half the computation costs, Wild-VAR improves accuracy by 2.2% and 1.3% on the Kinetics-400 and the Something-Somthing V2 datasets, respectively. Note to Practitioners—In human-robot interaction tasks, video action recognition technology is a prerequisite for visual control. In real applications, humans move freely in 3D space, which results in significant changes in the view of video capture and constantly changing scenes. Deep Neural Networks are limited by the perspectives and scenarios contained in the training data, resulting in most existing methods are only effective for identifying actions from 2–4 fixed views, and the background is single. Therefore, existing models are often difficult to generalize to unconstrained application environments. Human view and video scene understanding are often treated separately. Inspired by the human visual system, this paper proposes a view-to-scene video processing method in a cost-efficient way. In real-world applications, this lightweight method can be integrated into robots to help identify human behavior in complex environments. Fewer parameters indicate that the method can be easily migrated to different types of behaviors, and the reduced computational costs represent the ability to achieve real-time performance under limited hardware conditions. Xuna Wang, Xu Cheng 0003, Zhaojie Ju, Yingke Xu |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Spatiotemporal View-Reset Deep Learning With Attentional GRUs for Skeleton-Based Human Action RecognitionabstractSkeleton-based human action recognition has attracted significant attention. However, the skeleton spatial invariance and temporal context modeling are recent challenges for most existing methods. This work proposes a view-reset network, which integrates two important branches: spatial view-reset module (SVRM) and temporal attention module (TAM). In the SVRM, the skeleton model from different perspectives is reset in a unified coordinate system, eliminating the influence of viewpoint changes. In the TAM, the perception of temporal features is jointly enhanced by weighting each frame’s importance in the context. Furthermore, the pretrained residual network (ResNet) is used for prediction. The sample size is increased through data augmentation to improve the robustness of the model. The SVRM, TAM, and ResNet form an end-to-end learning network. The ablation study proved that the model could record the key skeletons and frames in the sequence and then reset the human body to a new position, making it easy for learning. The proposed model is evaluated on four challenging benchmarks based on the performance of the cross-view evaluation metrics. Experiments prove that the proposed model has superior performance and surpasses many state-of-the-art algorithms, with an increase of 1.91% over the top ten on the NTU RGB+D 60 dataset. Xuna Wang, Hongwei Gao 0002, Zide Liu, Zhaojie Ju |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2025 | Robust Compact Human Pose Learning Against Open-World Visual PerturbationsabstractSkeleton-based action recognition has achieved remarkable progress. However, in open-world scenarios, limited human visual labels, drifting skeletal structures, and novel action categories introduce complex visual disturbances that severely limit the robustness of pose representations. Herein, we propose UnicornPose, a Universal Compact Human Pose Representation, that learns robust skeleton correlations and recognizes action across various open-world scenarios. The core advantages include: 1) Continuously modeling human skeletal structures along the action timeline to construct a rich feature volume of human poses, ensuring sufficient information for universal representation. 2) Utilizing a multiview decoupling method to compress visual information further, obtaining robust pose representations that facilitate easier generalization across different open-world scenarios. 3) Coherence training and regularization constraint methods should be employed to enhance the generalization capability of noise-containing pose representations. These contributions enable UnicornPose to effectively counter noise interference and surpass the existing top results by 3-4%. Xuna Wang, Yuping Guo, Weiming Fan, Zhiyong Wang 0009 |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Video Object Detection Considering Dynamic Neighborhood Feature MultiplexingabstractVideo object detection is essential for human-interaction applications, including bimanual manipulation sensing (BMS). The effects of video detection in practical applications still need to be improved, as they are restricted by long-range spatiotemporal dependency analysis. How do humans sense bimanual manipulation in videos, especially for deteriorated clips? We argue that humans analyze the current clips based on earlier memory, namely, long-term spatial and temporal dependencies (LTSTD). However, most existing methods have yet to report significant results, as the limited exploration of these dependencies limits them. Developing an easy-to-integrate module is generally preferred for future applications rather than designing a complex end-to-end framework. Therefore, we propose a dynamic neighborhood feature multiplexing mechanism for online video object detection in this article, which is better at learning LTSTD in flexible and robust ways, boosting existing detection results, called DNFM. Specifically, we develop dynamic memory enhancement neural networks for better long-term feature aggregation with negligible additional computation costs. We multiplex each frame feature to aggregate key enhanced representations under the guidance of dynamic memory recall. The DNFM contributes to various famous detectors in BMS and other challenging detection tasks, and particular attention has been devoted to “low-quality” frame detection. Experimental results show that, while achieving state-of-the-art detection performance, DNFM clearly illustrates the easy-to-integrate operation for boosting the video object detection results. Xuna Wang, Dalin Zhou, Yingke Xu, Zhaojie Ju |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Patch-Slide Discriminative Joint Learning for Weakly-Supervised Whole Slide Image Representation and Classification
Xuna Wang, Yingke Xu |
MICCAI (3) | 2 |
| 2021 | Multi-Attribute Preferences Mining Method for Group Users with the Process of Noise Reduction
Qingmei Tan, Xuna Wang |
J. Comput. Sci. Technol. | 2 |
| 2020 | Attention-based deep neural network for Internet platform group users' dynamic identification and recommendation
Xuna Wang, Qingmei Tan, Mark Goh 0001 |
Expert Syst. Appl. | 1 |
| 2020 | A deep neural network of multi-form alliances for personalized recommendations
Xuna Wang, Qingmei Tan, Lifan Zhang |
Inf. Sci. | 1 |
| 2020 | DAN: a deep association neural network approach for personalization recommendationabstractThe collaborative filtering technology used in traditional recommendation systems has a problem of data sparsity. The traditional matrix decomposition algorithm simply decomposes users and items into a linear model of potential factors. These limitations have led to the low accuracy in traditional recommendation algorithms, thus leading to the emergence of recommendation systems based on deep learning. At present, deep learning recommendations mostly use deep neural networks to model some of the auxiliary information, and in the process of modeling, multiple mapping paths are adopted to map the original input data to the potential vector space. However, these deep neural network recommendation algorithms ignore the combined effects of different categories of data, which can have a potential impact on the effectiveness of the recommendation. Aimed at this problem, in this paper we propose a feedforward deep neural network recommendation method, called the deep association neural network (DAN), which is based on the joint action of multiple categories of information, for implicit feedback recommendation. Specifically, the underlying input of the model includes not only users and items, but also more auxiliary information. In addition, the impact of the joint action of different types of information on the recommendation is considered. Experiments on an open data set show the significant improvements made by our proposed method over the other methods. Empirical evidence shows that deep, joint recommendations can provide better recommendation performance. Xuna Wang, Qingmei Tan |
Frontiers Inf. Technol. Electron. Eng. | 1 |