Yingke Xu

dblp:89/7528 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
11since 2021 · last 2027
0000-0002-8317-0608ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2027 Detecting rare whole slide image via cross-space multi-prompt generation and collaboration
Tianfeng Wu, Sakirin Tam, Yingke Xu
Expert Syst. Appl.9
2025 Segmenting Vessels Encapsulating Tumor Clusters via Fine-Grained Visual Prompt
Shenjian Gu, Yuping Guo, Yingke Xu
MICCAI (16)7
2025 Exploring Complexity-Calibrated morphological distribution for whole slide image classification and difficulty-grading
Xuna Wang, Weiming Fan, Yuping Guo, Junfen Fu, Yingke Xu
Medical Image Anal.6
2025 Recognizing Video Activities in the Wild via View-to-Scene Joint Learning
abstract
Recognizing video actions in the wild is challenging for visual control systems. In-the-wild videos show actions not seen in training data, recorded from various angles and scenes with the same labels. Most existing methods address this challenge by developing complex frameworks to extract spatiotemporal features. To achieve view robustness and scene generalization cost-effectively, we explore view consistency and scene joint understanding. Based on this, we propose a neural network (called Wild-VAR) to learn view and scene information jointly without any 3D pose ground truth labels, a new approach to recognizing video actions in the wild. Unlike most existing methods, first, we propose a Cubing module to self-learn body consistency between views instead of comprehensive image features, boosting the generalization performance of across-view settings. Specifically, we map 3D representations to multiple 2D features and then adopt a self-adaptive scheme to constrain 2D features from different perspectives. Moreover, we propose temporal neural networks (called T-Scene) to develop a recognizing framework, enabling Wild-VAR to flexibly learn scenes across time, including key interactors and context, in video sequences. Extensive experiments show that Wild-VAR consistently outperforms state-of-the-art methods on four benchmarks. Notably, with only half the computation costs, Wild-VAR improves accuracy by 2.2% and 1.3% on the Kinetics-400 and the Something-Somthing V2 datasets, respectively. Note to Practitioners—In human-robot interaction tasks, video action recognition technology is a prerequisite for visual control. In real applications, humans move freely in 3D space, which results in significant changes in the view of video capture and constantly changing scenes. Deep Neural Networks are limited by the perspectives and scenarios contained in the training data, resulting in most existing methods are only effective for identifying actions from 2–4 fixed views, and the background is single. Therefore, existing models are often difficult to generalize to unconstrained application environments. Human view and video scene understanding are often treated separately. Inspired by the human visual system, this paper proposes a view-to-scene video processing method in a cost-efficient way. In real-world applications, this lightweight method can be integrated into robots to help identify human behavior in complex environments. Fewer parameters indicate that the method can be easily migrated to different types of behaviors, and the reduced computational costs represent the ability to achieve real-time performance under limited hardware conditions.
Xuna Wang, Xu Cheng 0003, Zhaojie Ju, Yingke Xu
IEEE Trans Autom. Sci. Eng.6
2025 Prior Visual-Guided Self-Supervised Learning Enables Color Vignetting Correction for High-Throughput Microscopic Imaging
abstract
Vignetting constitutes a prevalent optical degradation that significantly compromises the quality of biomedical microscopic imaging. However, a robust and efficient vignetting correction methodology in multi-channel microscopic images remains absent at present. In this paper, we take advantage of a prior knowledge about the homogeneity of microscopic images and radial attenuation property of vignetting to develop a self-supervised deep learning algorithm that achieves complex vignetting removal in color microscopic images. Our proposed method, vignetting correction lookup table (VCLUT), is trainable on both single and multiple images, which employs adversarial learning to effectively transfer good imaging conditions from the user visually defined central region of its own light field to the entire image. To illustrate its effectiveness, we performed individual correction experiments on data from five distinct biological specimens. The results demonstrate that VCLUT exhibits enhanced performance compared to classical methods. We further examined its performance as a multi-image-based approach on a pathological dataset, revealing its advantage over other state-of-the-art approaches in both qualitative and quantitative measurements. Moreover, it uniquely possesses the capacity for generalization across various levels of vignetting intensity and an ultra-fast model computation capability, rendering it well-suited for integration into high-throughput imaging pipelines of digital microscopy.
Jianhang Wang, Luhong Jin, Yunqi Zhu, Shujun Fu, Yingke Xu
IEEE J. Biomed. Health Informatics8
2025 Semi-Supervised Instance Segmentation in Whole Slide Images via Dense Spatial Variability Enhancing
abstract
Current whole slide image (WSI) segmentation aims at extracting tumor regions from the background. Unlike this, segmenting distinct tumor areas (instances) within a WSI driven by limited annotated data remains under-explored. In this paper, we formally propose semi-supervised instance segmentation (Semi-IS) in WSIs. We address a key challenge: learning intra-class similarity and inter-class dissimilarity driven by unlabeled data. Specifically, we generally perceive the patch as composed of tokens (together), not the patch alone. We employ contrastive learning to develop a segmentation framework. In the Semi-IS, we find that the boundaries of segmented instances are usually disturbed by noise. We jointly eliminate and preserve noise features to address this problem. We conduct extensive experiments to evaluate the effectiveness and generalizability of Semi-IS, including histopathology and cellular pathology. The results show that in clinical multi instance segmentation tasks, Semi-IS achieves almost full-supervised state-of-the-art results with only 30% annotated data. Semi-IS can improve segmentation accuracy by about 2% on public cell pathology datasets.
Dong Hua, Junfen Fu, Yingke Xu
IEEE J. Biomed. Health Informatics6
2025 Video Object Detection Considering Dynamic Neighborhood Feature Multiplexing
abstract
Video object detection is essential for human-interaction applications, including bimanual manipulation sensing (BMS). The effects of video detection in practical applications still need to be improved, as they are restricted by long-range spatiotemporal dependency analysis. How do humans sense bimanual manipulation in videos, especially for deteriorated clips? We argue that humans analyze the current clips based on earlier memory, namely, long-term spatial and temporal dependencies (LTSTD). However, most existing methods have yet to report significant results, as the limited exploration of these dependencies limits them. Developing an easy-to-integrate module is generally preferred for future applications rather than designing a complex end-to-end framework. Therefore, we propose a dynamic neighborhood feature multiplexing mechanism for online video object detection in this article, which is better at learning LTSTD in flexible and robust ways, boosting existing detection results, called DNFM. Specifically, we develop dynamic memory enhancement neural networks for better long-term feature aggregation with negligible additional computation costs. We multiplex each frame feature to aggregate key enhanced representations under the guidance of dynamic memory recall. The DNFM contributes to various famous detectors in BMS and other challenging detection tasks, and particular attention has been devoted to “low-quality” frame detection. Experimental results show that, while achieving state-of-the-art detection performance, DNFM clearly illustrates the easy-to-integrate operation for boosting the video object detection results.
Xuna Wang, Dalin Zhou, Yingke Xu, Zhaojie Ju
IEEE Trans. Syst. Man Cybern. Syst.7
2024 Patch-Slide Discriminative Joint Learning for Weakly-Supervised Whole Slide Image Representation and Classification
Xuna Wang, Yingke Xu
MICCAI (3)5
2024 Versatile Graph Neural Networks Toward Intuitive Human Activity Understanding
abstract
Benefiting from the advanced human visual system, humans naturally classify activities and predict motions in a short time. However, most existing computer vision studies consider those two tasks separately, resulting in an insufficient understanding of human actions. Moreover, the effects of view variations remain challenging for most existing skeleton-based methods, and the existing graph operators cannot fully explore multiscale relationship. In this article, a versatile graph-based model (Vers-GNN) is proposed to deal with those two tasks simultaneously. First, a skeleton representation self-regulated scheme is proposed. It is among the first trials that successfully integrate the idea of view adaptation into a graph-based human activity analysis system. Next, several novel graph operators are proposed to model the positional relationships and learn the abstract dynamics between different human joints and parts. Finally, a practical multitask learning framework and a multiobjective self-supervised learning scheme are proposed to promote both the tasks. The comparative experimental results show that Vers-GNN outperforms the recent state-of-the-art methods for both the tasks, with the to date highest recognition accuracies on the datasets of NTU RGB + D (CV: 97.2%), UWA3D (88.7%), and CMU (1000 ms: 1.13).
Yingke Xu, Zhaojie Ju
IEEE Trans. Neural Networks Learn. Syst.2
2023 Marrying Global-Local Spatial Context for Image Patches in Computer-Aided Assessment
abstract
Computer-aided assessment using whole slide images (WSIs) is one of the critical steps in clinical procedures. How do doctors recognize cancer in a WSI? A quick answer is that they consider the spatial structure of a WSI rather than only considering single patches. We argue that two clues are essential for computer-aided deep learning: 1) global spatial context and 2) local semantic information. This is because local, semi-local, and global tissue observing are the principal assessment means of pathologists, perfectly corresponding with both clues. However, most existing methods only consider local spatial information learning within each patch rather than developing an effective local-to-global reaction, leading to an incapable of capturing robust and enriched representation. Toward a new area for computer-aided assessment, we propose novel neural networks to learn the global–local spatial context in WSIs, called GLSCL. The GLSCL is among the first trials that understand both clues for WSI understanding. Furthermore, the proposed novel operators enable the GLSCL to learn spatial semantic representation sufficiently. We evaluate the GLSCL using renal cell carcinoma (RCC) samples with synthetic ambiguity collected from the public benchmark and clinical procedures. Enhanced by global and local spatial information, the GLSCL achieves state-of-the-art performance, including classification accuracy, survival prediction index, and cancer tissue attention rate.
Maode Lai, Zhaojie Ju, Yingke Xu
IEEE Trans. Syst. Man Cybern. Syst.6
2022 View-Robust Neural Networks for Unseen Human Action Recognition in Videos
abstract
Data-driven deep learning achieved excellent performance for human action recognition. However, unseen action recognition remains a challenge for most existing neural networks. Because the action categories, collection perspectives, and scenarios considered during data collection are limited. Compared with class-unseen action recognition, view-unseen action recognition in videos is under-explored. This paper proposes view-robust neural networks (VR-Net) to recognize unseen actions in videos. The VR-Net consists of a 3D pose estimation module, skeleton adaptive transformation neural networks, and classification modules. We first extract 3D skeleton models from the video sequence based on existing pose estimation methods. Next, we propose a skeleton representation transformation scheme and achieve it based on Convolutional Neural Networks (VR-CNN) and Graph Neural Networks (VR-GCN), resulting in the optimal skeleton representations. Futhermore, we explore an associate optimization scheme and a fused output method. We evaluate the proposed neural networks on three challenging benchmarks, i.e., NTU RGB-D dataset (NTU), Kinetics-400 dataset, and Human3.6M dataset (H3.6M). The experimental results show that view robust neural networks achieve the top performance compared to state-of-the-art RGB-based and skeleton-based works, such as 93.6% on the NTU (CV) and 94.6% on the Kinetics-400 dataset (Top-5). The proposed neural networks significantly improve the recognition performance for unseen action recognition, such as 86.8% on the H3.6M (View 2).
Zhaojie Ju, Yingke Xu
SMC5