Songsen Yu

dblp:95/8453 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
23since 2021 · last 2026
0000-0003-0347-0932ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Pixel-wise perception-distortion trade-off for single image super-resolution
Hai Su, Yanghui Wei, ZhenWen Jian, Songsen Yu
Signal Process. Image Commun.4
2025 ProFiMed: Prompt-Based Federated Medical Image Classification via Pruned Optimal Transport and Residual Cross-Correlation Self-Attention
abstract
Federated learning (FL) enables collaborative modelling of medical images without sharing raw data, yet the severe feature- and label-shift across hospitals compromises the performance of a single global model. Recent promptbased FL frameworks such as FedOTP mitigate this issue by splitting learnable prompts into a shared and a client-specific part, but still (i) rely on entropy-regularised unbalanced optimal transport (UOT) that oversmoothes tiny lesions, and (ii) inherit CLIP's location-invariant self-attention, which blurs pathological details. To overcome these limitations we introduce ProFiMed, a communication-efficient framework that combines pruned UOT with a Residual Correlative Self-Attention (RCSA) block. Pruned UOT clips the OT cost matrix and reformulates the transport as a KL projection, yielding noise-resilient vision-text alignment at$\mathcal{O}(V)$complexity. RCSA fuses cross-layer cross-correlation attention with correlative self-attention, re-injecting position sensitivity into the final Transformer block without updating any backbone weights. Under a 16-shot protocol on five public medical benchmarks (CT, MRI, histology and dermoscopy) our method surpasses the state-of-the-art FedOTP backbone by 5.67 pp on average and outperforms six strong plug-in variants, while keeping the same privacy and communication budget. Extensive ablations reveal that pUOT and RCSA are complementary, robust to threshold and hyper-parameter choices, and scalable to heterogeneous hospital distributions.
Shuo Dai, Songsen Yu
BIBM3
2025 FedEM-AC: Enhancing Federated Learning for Medical Image Classification Through EM Aggregation and Adaptive Multi-Layer Attention Fusion
Shuo Dai, Songsen Yu
ICIC (19)4
2025 PowerOpsNet: Integrating YOLOv5-A and GCN for Real-Time Safety Behavior Recognition in Power Grid Environments
abstract
Ensuring operational safety in power grids is critical, particularly in substations where occlusion and small-object challenges are prevalent. We propose PowerOpsNet, a unified end-to-end framework for real-time unsafe behavior recognition. It combines an optimized YOLOv5-A detector — enhanced via Mosaic-8 augmentation, multi-scale features, and a dynamic focal loss — with a CNN-LSTM-GCN-based behavior recognition module. A benchmark dataset comprising 11 object types and 18 actions was collected from real substations. Experiments demonstrate 91.3% detection precision and 87.6% recognition accuracy, with inference speeds of 370ms and 200ms, respectively. PowerOpsNet supports real-time safety monitoring and is extensible to other high-risk domains.
Shenyu Yan, Songsen Yu
SMC2
2025 Classification and regression Task Integration in distillation for object detectors
Hai Su, ZhenWen Jian, Yanghui Wei, Songsen Yu
Neurocomputing4
2025 DenseJIN: Dense Depth Image Steganography Model With Joint Invertible and Noninvertible Mechanisms
abstract
Image steganography discreetly embeds secret information within a carrier, allowing covert communication and enabling the receiver to extract the concealed data when needed. Previous techniques for image steganography had limitations in achieving imperceptibility and security when dealing with images containing intricate textures. In this paper, we introduce DenseJIN, an innovative model for dense depth image steganography. DenseJIN joins invertible and noninvertible mechanisms to achieve effective and secure information hiding. The invertible component of DenseJIN ensures that the stego image maintains high imperceptibility and security, while the noninvertible component enables high-quality recovery of the secret image. In the invertible component, we employ a dense connection for each invertible block in the forward process and a straightforward series connection during the reverse process. In the forward process of the network, the secret image is embedded, while the backward process is responsible for extracting the embedded secret image. To perform the noninvertible step, we incorporate a modified Unet architecture, enabling deep fine-grained feature extraction from cover images and secret images. Our experimental results indicate that DenseJIN surpasses other contemporary image steganography methods. On average, DenseJIN achieves a remarkable improvement of over 1.75 dB in PSNR for secret image recovery across DIV2K, COCO and ImageNet.
Delin Duan, Shuyuan Shen, Songsen Yu, Yibo Yuan, Qidong Zhou, Haojie Lv, Huanjie Lin
IEEE Trans. Circuits Syst. Video Technol.3
2024 Value Function Evaluation with Data Augmentation for Offline Reinforcement Learning
Xianwei Zhou, Chulue Zhang, Songsen Yu
ICIC (2)4
2024 Context Embedding Similarity based Semi-Supervised Active Learning for Time Series
abstract
Time series data, a crucial form of structured data across various domains, often necessitates a substantial quantity of labels for supervised learning. The complexity of annotating time series data, however, results in a high annotation cost. To address this challenge, active learning emerges as a strategic approach, selecting key samples and interacting with experts for data annotation. Its primary objective is to acquire high-quality training data at the lowest possible cost. This paper proposes a novel Semi-supervised ACTive learning (SACT) framework tailored for time series data. Initially, we establish a criterion to assess the importance of unlabeled samples, considering the unique characteristics of time series. We employ context embedding similarity scores to segment the time series data, prioritizing samples with high scores for labeling. Furthermore, to enhance the utilization of a single query, we propagate labels to adjacent data points. By experimenting on different combinations of query strategies and label propagation methods, our approach consistently outperforms other semi-supervised active learning frameworks for time series data in most scenarios. This superior performance underscores the efficacy of our approach.
Xianwei Zhou, Songsen Yu, Shiqi Wu, Wencong Zhang, Chulue Zhang
IJCNN3
2024 Unpaired Multi-scenario Sketch Synthesis via Texture Enhancement
Songsen Yu, Shiqi Wu, Jun Liang 0002
PRCV (4)1
2024 Successive model-agnostic meta-learning for few-shot fault time series prognosis
Hai Su, Jiajun Hu, Songsen Yu, Juhua Liu, Xiangyang Qin
Neurocomputing3
2024 A Lightweight Object Detection Model for Low-End UAVs
abstract
With the rapid growth and widespread use of low-end commercial unmanned aerial vehicles (UAVs), it is critical to develop an object detection system that works well with these devices. This paper designs an efficient and lightweight object detection model specifically designed for low-end UAVs. Through excellent information interaction and the refined use of some techniques, our model can obtain multi-scale features and better focus on the details of different parts. Extensive experiments show that our model still maintains comparable accuracy while it consumes fewer parameters and FLOPs. In addition, our model has been applied to the tracking system of low-end UAVs, greatly enhancing the tracking performance.
Jun Liang 0002, Nuo Zhou, Jiahao Long, Songsen Yu, Muhammad Faizan Khan
Int. J. Pattern Recognit. Artif. Intell.4
2024 SRPSGAN: Super-resolution with pose and expression robust spatial-aware generative adversarial network for makeup transfer
Yiming Ru, Jingzeng Huang, Songsen Yu, Yihua Chen 0002
Multim. Tools Appl.4
2024 Contrast-based unsupervised hashing method with margin limit
Hai Su, Zhenyu Ke, Songsen Yu, Jianwei Fang, Yuchen Zhong
Multim. Tools Appl.3
2024 Dual contrast discriminator with sharing attention for video anomaly detection
Yiwenhao Zeng, Yihua Chen 0002, Songsen Yu, Mingzhang Yang, Rongrong Chen 0002
Mach. Vis. Appl.3
2024 Fair and Privacy-Preserved Data Trading Protocol by Exploiting Blockchain
abstract
With the popularity of the mobile Internet, data is increasingly becoming a new resource. Therefore, the trading of such data resources has become an increasing demand. In this paper, we propose a fair privacy-preserving data trading protocol based on blockchain. Firstly, our data trading protocol achieves fairness by carefully combining the probabilistic approaches and the fully homomorphic encryption techniques. Moreover, our protocol allows online arbitration when misbehavior occurs in the trading process is detected. Note that previous data trading protocols need a Trusted Third Party (TTP) or an offline arbitrator to solve disputes, weakening the trust of those protocols. Secondly, the data validity verification process of our protocol is more flexible. Most Importantly, different from all previous designs which only achieve privacy against communication channel eavesdroppers, our protocol achieves privacy against any eavesdropper and the passive arbitrator. The above-distinguishing properties of our protocol are mainly benefited from the homomorphic encryption and double encryption techniques. In addition, our data trading protocol can be instantiated with post-quantum primitives and thus achieves post-quantum security. To demonstrate the feasibility of the proposed protocol, we conduct a comprehensive evaluation with the instantiated cryptographic primitives based on the Ethereum test network.
Parhat Abla, Taotao Li, Debiao He, Huawei Huang, Songsen Yu, Yan Zhang 0002
IEEE Trans. Inf. Forensics Secur.5
2024 A novel robust image watermarking algorithm based on polar decomposition and image geometric correction
Qidong Zhou, Shuyuan Shen, Songsen Yu, Delin Duan, Yibo Yuan, Haojie Lv, Huanjie Lin
Vis. Comput.3
2023 Tennis Action Recognition Based on Multi-Branch Mixed Attention
Xianwei Zhou, Zhenfeng Li, Jiale Lei, Songsen Yu
KSEM (2)6
2023 Deep Reinforcement Learning for Group-Aware Robot Navigation in Crowds
Xianwei Zhou, Songsen Yu
KSEM (4)4
2023 Seam carving based on dynamic energy regulation
Hai Su, Zigui Ye, Songsen Yu
Multim. Tools Appl.4
2023 A deep hashing method of likelihood function adaptive mapping
Hai Su, Jianwei Fang, Songsen Yu
Neural Comput. Appl.4
2023 Deep supervised hashing with hard example pairs optimization for image retrieval
Hai Su, Meiyin Han, Junle Liang, Jun Liang 0002, Songsen Yu
Vis. Comput.5
2022 A multi-scale semantic attention representation for multi-label image recognition with graph networks
abstract
Multi-label image recognition is a basic and challenging task in computer vision and multimedia fields. Graph Convolutional Networks (GCNs) are often used to learn the multi-label semantic features and multi-label dependency. Although the label semantic features in GCNs can learn the global image visual representation well, they are rarely used on the local image regions. Therefore, we try to use GCNs to learn global and local features at the same time, and make a balance between them. In this paper, we give a multi-scale semantic attention model MS-SGA-GCN including three main modules (i.e., MS, SGA and GCN) for multi-label image recognition. The Multi-Scale module (MS) utilizes feature maps of different sizes to obtain global features and have strong generalization capabilities. Semantic Guide Attention module (SGA) applies the label embeddings learned by GCNs to guide the generation of the cross-modality class-specific attention maps, which can discover the locations of semantically related regions for each label. Experiments show that our model on two datasets MS-COCO and PASCAL VOC2007 separately achieves the classification accuracy by 83.4% and 94.2%, which has a competitive advantage over other mainstream models.
Jun Liang 0002, Feiteng Xu, Songsen Yu
Neurocomputing3
2021 Sketch works ranking based on improved transfer learning model
Songsen Yu, ZeSheng Lin, Jun Liang 0002, GangXu Shu, JiaLin Yu, Ao Zhu
Multim. Tools Appl.1
2018 A novel adaptive data hiding based on improved EMD and interpolation
Shu-Yuan Shen, Songsen Yu
Multim. Tools Appl.3