Yong Su 0003

dblp:29/201-3 · DBLP profile ↗
← Back
31ranked-venue papers
13as first author
22since 2021 · last 2025
0000-0002-6851-4142ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 11 since 2021Computer networks · 6 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 From Gaze to Masks: Gaze-Based Weakly-Supervised Medical Image Segmentation
abstract
Weakly supervised medical image segmentation seeks to reduce reliance on exhaustive pixel-wise annotations by leveraging coarse or indirect supervision, thereby lowering annotation costs while maintaining segmentation quality. Existing methods rely on coarse heuristics, lacking spatial fidelity and semantic localization in complex medical imagery. To address this limitation, we propose a Gaze-based Weakly Supervised Medical Image Segmentation (GWMIS) framework that exploits gaze annotations as rich, low-cost semantic priors. GWMIS consists of two streams: a pseudo-label generation stream and a cross-modal fusion segmentation stream. The former fuses gaze-derived Gaussian priors with Segment Anything Model (SAM) for structure-consistent pseudo masks; the latter aligns multi-scale visual and gaze features via dual-encoder crossattention for anatomically faithful segmentation. Evaluations on colonic polyp (3,784 images, five public datasets) and brain tumor segmentation (3,929 MRI slices, LGG dataset) demonstrate superior performance over state-of-the-art weakly supervised approaches.
Meng Xing, Wanlong Zhang, Yong Su 0003, Qifa Peng, Shuang Zhu
BIBM3
2025 Gaze-Guided Learning: Avoiding Shortcut Bias in Visual Classification
Jiahang Li 0004, Shibo Xue, Yong Su 0003
CogSci3
2025 Gaze-and-Machine Dual-Driven Attention Fusion Network for Medical Image Classification
Qifa Peng, Shuang Zhu, Yong Su 0003, Meng Xing
ICIC (28)3
2025 Coming Out of the Dark: Human Pose Estimation in Low-light Conditions
abstract
Human pose estimation in low-light conditions is vital for applications such as surveillance and autonomous systems, yet the severe visual distortions hinder both manual annotation and estimation precision. Existing approaches typically rely on additional reference information to mitigate these issues, however, customized data collection equipment poses limitations on their scalability. To alleviate the issue, we construct a Low-Light Images and Poses (LLIP) dataset, which includes only paired low-light images and pose annotations obtained using off-the-shelf motion capture devices. Furthermore, we propose a Multi-grained High-frequency Feature Consistency Learning framework (MHFCL), which does not rely on additional reference information. MHFCL employs a Retinex-inspired restoration stream to recover high-frequency details and integrates them into pose estimation using a multi-grained consistency mechanism. Experiments demonstrate that our approach achieves a new benchmark in low-light pose estimation, while maintaining competitive performance in well-lit conditions.
Yong Su 0003, Meng Xing, Changjae Oh, Xuewei Liu, Jieyang Li
IJCAI1
2025 MeshPAD: Payload-aware mesh distortion for 3D steganography based on geometric deep learning
Weilong Peng, Keke Tang, Weixuan Tang 0002, Yong Su 0003, Meie Fang, Ping Li 0016
Expert Syst. Appl.4
2025 Spatio-temporal graph-based self-labeling for video anomaly detection
Meng Xing, Zhiyong Feng 0002, Yong Su 0003, Changjae Oh, Valeriya V. Gribova, Vladimir Fedorovich Filaretoy, De-Shuang Huang
Neurocomputing3
2025 Semantic-driven dual consistency learning for weakly supervised video anomaly detection
Yong Su 0003, Yuyu Tan, Simin An, Meng Xing, Zhiyong Feng 0002
Pattern Recognit.1
2025 Dual-Detector Reoptimization for Federated Weakly Supervised Video Anomaly Detection via Adaptive Dynamic Recursive Mapping
abstract
Federated weakly supervised video anomaly detection represents a significant advancement in privacy-preserving collaborative learning, enabling distributed clients to train anomaly detectors using only video-level annotations. However, the inherent challenges of optimizing noisy representation with coarse-grained labels often result in substantial local model errors, which are exacerbated during federated aggregation, particularly in heterogeneous scenarios. To address these limitations, we propose a novel dual-detector framework incorporating adaptive dynamic recursive mapping, which significantly enhances local model accuracy and robustness against representation noise. Our framework integrates two complementary components: a channel-averaged anomaly detector and a channel-statistical anomaly detector, which interact through cross-detector adaptive decision parameters to enable iterative optimization and stable anomaly scoring across all instances. Furthermore, we introduce the scene-similarity adaptive local aggregation algorithm, which dynamically aggregates and learns private models based on scene similarity, thereby enhancing generalization capabilities across diverse scenarios. Extensive experiments conducted on the NVIDIA Jetson AGX Xavier platform using the ShanghaiTech and UBnormal datasets demonstrate the superior performance of our approach in both centralized and federated settings. Notably, in federated environments, our method achieves remarkable improvements of 6.2% and 12.3% in AUC compared to state-of-the-art methods, underscoring its effectiveness in resource-constrained scenarios and its potential for real-world applications in distributed video surveillance systems.
Yong Su 0003, Jiahang Li 0004, Simin An, Hengpeng Xu, Weilong Peng
IEEE Trans. Ind. Informatics1
2024 Learning by Erasing: Conditional Entropy Based Transferable Out-of-Distribution Detection
abstract
Detecting OOD inputs is crucial to deploy machine learning models to the real world safely. However, existing OOD detection methods require an in-distribution (ID) dataset to retrain the models. In this paper, we propose a Deep Generative Models (DGMs) based transferable OOD detection that does not require retraining on the new ID dataset. We first establish and substantiate two hypotheses on DGMs: DGMs exhibit a predisposition towards acquiring low-level features, in preference to semantic information; the lower bound of DGM's log-likelihoods is tied to the conditional entropy between the model input and target output. Drawing on the aforementioned hypotheses, we present an innovative image-erasing strategy, which is designed to create distinct conditional entropy distributions for each individual ID dataset. By training a DGM on a complex dataset with the proposed image-erasing strategy, the DGM could capture the discrepancy of conditional entropy distribution for varying ID datasets, without re-training. We validate the proposed method on the five datasets and show that, without retraining, our method achieves comparable performance to the state-of-the-art group-based OOD detection methods. The project codes will be open-sourced on our project website.
Meng Xing, Zhiyong Feng 0002, Yong Su 0003, Changjae Oh
AAAI3
2024 Anomalies cannot materialize or vanish out of thin air: A hierarchical multiple instance learning with position-scale awareness for video anomaly detection
Yong Su 0003, Yuyu Tan, Simin An, Meng Xing
Expert Syst. Appl.1
2024 VPE-WSVAD: Visual prompt exemplars for weakly-supervised video anomaly detection
Yong Su 0003, Yuyu Tan, Meng Xing, Simin An
Knowl. Based Syst.1
2024 M2AST:MLP-mixer-based adaptive spatial-temporal graph learning for human motion prediction
Junyi Tang, Simin An, Yuanwei Liu, Yong Su 0003, Jin Chen 0002
Multim. Syst.4
2024 Spatio-Temporal Articulation & Coordination Co-attention Graph Network for human motion prediction
Shuang Zhu, Jin Chen 0002, Yong Su 0003
Signal Process.3
2023 Prime: Privacy-preserving video anomaly detection via Motion Exemplar guidance
Yong Su 0003, Haohao Zhu, Yuyu Tan, Simin An, Meng Xing
Knowl. Based Syst.1
2023 Energy-Based Temporal Summarized Attentive Network for Zero-Shot Action Recognition
abstract
Recently, Action Recognition (AR) is facing the scalability problem, since collecting and annotating data for the ever-growing action categories is exhausting and inappropriate. As an alternative to AR, Zero-Shot Action Recognition (ZSAR) is getting more and more attention in the community, as they could utilize a shared semantic/attribute space to recognize novel categories without annotated data. Different from the AR focuses on learning the correlation between actions, ZSAR needs to consider the correlation of action-action, label-label and action-label at the same time. However, as far as we know, there is no work to provide structural guidance for the framework design of ZSAR according to its task characteristics. In this paper, we demonstrate the rationality of using the Energy-Based Model (EBM) to guide the framework design of ZSAR based on their inference mechanism. Furthermore, under the guidance of EBM, we propose an Energy-based Temporal Summarized Attentive Network (ETSAN) to achieve ZSAR. Specifically, to ensure the effectiveness of cross-modal matching, EBM needs to capture the correlations of input-input, output-output and input-output, based on discriminative and focused input and output space. To this end, we first design the Temporal Summarized Attentive Mechanism (TSAM) to capture the correlation of action-action by constructing discriminative and focused input space. Then, a Label Semantic Adaptive Mechanism (LSAM) is proposed to learn the correlation of label-label by adjusting the semantic structure according to the target task. Finally, we devise an Energy Score Estimation Mechanism (ESEM) to measure the compatibility (i.e. energy score) between video representation and label semantic embedding. With end-to-end training, our framework can capture all three of the correlations mentioned above simultaneously by minimizing the energy score of the correct action-label pair. Experiments on the HMDB51 and UCF101 datasets show that the proposed architecture achieves comparable results among methods based on the spatial-temporal visual feature of sequence-level, which demonstrates the efficiency of the EBM in guiding the framework design of ZSAR. In addition, our code is available athttps://github.com/oOHCIOo/ETSAN.
Cheng Qi, Zhiyong Feng 0002, Meng Xing, Yong Su 0003, Jinqing Zheng
IEEE Trans. Multim.4
2022 VDSSA: Ventral & Dorsal Sequential Self-attention AutoEncoder for Cognitive-Consistency Disentanglement
Yi Yang 0074, Yong Su 0003, Simin An
PRCV (2)2
2021 DVAMN: Dual Visual Attention Matching Network for Zero-Shot Action Recognition
Cheng Qi, Zhiyong Feng 0002, Meng Xing, Yong Su 0003
ICANN (5)4
2021 Spatio-temporal multi-factor model for individual identification from biological motion
Yong Su 0003, Weilong Peng, Meng Xing, Zhiyong Feng 0002
Ad Hoc Networks1
2021 VDARN: Video Disentangling Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition
Yong Su 0003, Meng Xing, Simin An, Weilong Peng, Zhiyong Feng 0002
Ad Hoc Networks1
2021 Disentangling style on dynamic aligned poses for individual identification
Yong Su 0003, Meng Xing, Weilong Peng, Zhiyong Feng 0002
Ad Hoc Networks1
2021 Ventral & Dorsal Stream Theory based Zero-Shot Action Recognition
Meng Xing, Zhiyong Feng 0002, Yong Su 0003, Weilong Peng
Pattern Recognit.3
2021 Bayesian Covariance Representation with Global Informative Prior for 3D Action Recognition
Zhiyong Feng 0002, Yong Su 0003, Meng Xing
ACM Trans. Multim. Comput. Commun. Appl.3
2020 Spatio-temporal metric learning for individual recognition from locomotion
Yong Su 0003, Simin An, Zhiyong Feng 0002, Meng Xing
J. Vis. Commun. Image Represent.1
2020 Cross-covariance matrix: Time-shifted correlations for 3D action recognition
Zhiyong Feng 0002, Yong Su 0003, Meng Xing
Signal Process.3
2020 An Image Cues Coding Approach for 3D Human Pose Estimation
abstract
Although Deep Convolutional Neural Networks (DCNNs) facilitate the evolution of 3D human pose estimation, ambiguity remains the most challenging problem in such tasks. Inspired by the Human Perception Mechanism (HPM), we propose an image-to-pose coding method to fill the gap between image cues and 3D poses, thereby alleviating the ambiguity of 3D human pose estimation. First, in 3D pose space, we divide the whole 3D pose space into multiple subregions named pose codes , turning a disambiguation problem into a classification problem. The proposed coding mechanism covers multiple camera views and provides a complete description for 3D pose space. Second, it is noteworthy that the articulated structure of the human body lies on a sophisticated product manifold and the error accumulation in the chain structure will undoubtedly affect the coding performance. Therefore, in image space, we extract the image cues from independent local image patches rather than the whole image. The mapping relationship between image cues and 3D pose codes is established by a set of DCNNs. The image-to-pose coding method transforms the implicit image cues into explicit constraints. Finally, the image-to-pose coding method is integrated into a linear matching mechanism to construct a 3D pose estimation method that effectively alleviates the ambiguity. We conduct extensive experiments on widely used public benchmarks. The experimental results show that our method effectively alleviates the ambiguity in 3D pose recovery and is robust to the variations of view.
Meng Xing, Zhiyong Feng 0002, Yong Su 0003
ACM Trans. Multim. Comput. Commun. Appl.3
2019 Discriminative Saliency-pose-attention Covariance for Action Recognition
abstract
Most covariance-based representations of actions are focused on the statistical features of poses by empirical averaging weighting. Note that these poses have a variety of saliency levels for different actions. Neglecting pose saliency could degrade the discriminative power of the covariance features, and further reduce the performance of action recognition. In this paper, we propose a novel saliency weighting covariance feature representation, Saliency-Pose-Attention Covariance(SPA-Cov), which reduces the negative effects from the ambiguous pose samples. Specifically, we utilize a discriminative approach to derive probability distribution of action categories for each pose, which is modeled by the uncertainty of information entropy to obtain the salient weighting. Experimental results show that our proposed method efficiently improves the discriminative power of the generated covariance. In some databases, the proposed SPA-Cov outperforms the state-of-the-art variant methods which are based on kernel matrix, Bayesian posterior features, temporal hierarchical features, etc.
Zhiyong Feng 0002, Yong Su 0003, Meng Xing
ICASSP3
2019 Spatio-Temporal Multi-Factor Discriminant Analysis for Individual Identification
abstract
Individual identification from skeleton sequence is a challenging task because various covariate factors may produce drastic changes in human poses, thus causing large intra-class variability. In this paper, we present a novel spatio-temporal multi-factor discriminant analysis (ST-MFDA) to reduce the impact of covariate factors on identification performance. In our approach, a set of paired factor-specific spatio-temporal projections are learned to project motion features from different factors into a common discriminant subspace. In the common subspace, features from the same individual are united and those from different individuals are separated by Fisher discriminant criterion. We show that spatio-temporal projections of various factors can be jointly learned by solving an iterative generalized eigenvalue problem. According to the experimental results from three public databases, we confirm that spatio-temporal projections trained in this way lead to significant improvements in identification accuracy under the different covariate factors.
Yong Su 0003, Zhiyong Feng 0002
ICME1
2019 Dynamic hand gesture recognition using motion pattern and shape descriptors
Meng Xing, Jing Hu 0007, Zhiyong Feng 0002, Yong Su 0003, Weilong Peng, Jinqing Zheng
Multim. Tools Appl.4
2018 Spatio-Temporal Large Margin Nearest Neighbor (ST-LMNN) Based on Riemannian Features for Individual Identification
abstract
The major challenge for individual identification from periodical locomotion is to explore the unique spatio-temporal motion characteristic of each individual. In this paper, we present a novel spatio-temporal metric learning approach to activate the discriminating components of Riemannian features. Specifically, to deal with articulated motion embedded in a high dimensional space, we extend the Active Appearance Model to Riemannian manifold to measure the intrinsic variations of poses. Then we extract high order geometric features of each joint, which naturally suggests the biometric signatures of enrolled individuals. To model the most discriminant features in a linear subspace, we propose a Spatio-Temporal Large Margin Nearest Neighbor (ST-LMNN) algorithm to learn the low-dimensional linear embedding in the spatial and temporal domain, respectively. According to the experimental results from two public databases and a new built database, the proposed approach can achieve more accurate identification results in walking and running.
Yong Su 0003, Zhiyong Feng 0002, Meng Xing
ICME1
2018 Sequential Articulated Motion Reconstruction from a Monocular Image Sequence
abstract
In this article, we present a sequential approach for articulated motion estimation from a 2D skeleton sequence. This is a challenging task due to the complexity of human movements and the inherent depth ambiguities. The proposed approach models the human movement on a kinematic manifold with the tangent bundle, which is a natural geometrical representation of articulated motion. Combined with a second-order stochastic dynamic model based on the Markov hypothesis, we generalize the Extended Rauch Tung Striebel smoother to a Riemannian manifold to simulate the process of human movement. The human motor system might violate the Markov hypothesis when the human body is subject to external forces, and therefore a refinement stage is introduced to correct the estimation error. Specifically, the current estimation is refined in a feasible solution region consisting of a set of local estimations. This region is called a simplex, in which each element can be represented by a convex hull of all ingredients. We have proved that the refinement problem can be converted into a convex optimization problem with the simplicial constraint. Since the proposed formulation conforms to the principles of kinematic and spatio-temporal continuity of articulated motion, the reconstruction ambiguity can be alleviated essentially. The performance of the proposed algorithm is conducted on multiple synthetic sequences from the CMU and the HDM05 MoCap databases. The results show that, without requiring any training data, the proposed approach achieves greater accuracy over state-of-the-art baselines. Furthermore, the proposed approach outperforms two baselines on real sequences from the Human3.6m MoCap database.
Yong Su 0003, Zhiyong Feng 0002, Weilong Peng, Meng Xing
ACM Trans. Multim. Comput. Commun. Appl.1
2017 Parametric T-Spline Face Morphable Model for Detailed Fitting in Shape Subspace
abstract
Pre-learnt subspace methods, e.g., 3DMMs, are significant exploration for the synthesis of 3D faces by assuming that faces are in a linear class. However, the human face is in a nonlinear manifold, and a new test are always not in the pre-learnt subspace accurately because of the disparity brought by ethnicity, age, gender, etc. In the paper, we propose a parametric T-spline morphable model (T-splineMM) for 3D face representation, which has great advantages of fitting data from an unknown source accurately. In the model, we describe a face by C^2 T-spline surface, and divide the face surface into several shape units (SUs), according to facial action coding system (FACS), on T-mesh instead of on the surface directly. A fitting algorithm is proposed to optimize coefficients of T-spline control point components along pre-learnt identity and expression subspaces, as well as to optimize the details in refinement progress. As any pre-learnt subspace is not complete to handle the variety and details of faces and expressions, it covers a limited span of morphing. SUs division and detail refinement make the model fitting the facial muscle deformation in a larger span of morphing subspace. We conduct experiments on face scan data, kinect data as well as the space-time data to test the performance of detail fitting, robustness to missing data and noise, and to demonstrate the effectiveness of our model. Convincing results are illustrated to demonstrate the effectiveness of our model compared with the popular methods.
Weilong Peng, Zhiyong Feng 0002, Chao Xu 0003, Yong Su 0003
CVPR4