Yulong Hu

dblp:203/3759 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Enhancing Colorectal Lesion Segmentation Through Internal Feature Extraction and Computational Modeling Insights
abstract
Colorectal cancer (CRC) is a prevalent malignancy with significant social and healthcare implications, ranking among the top causes of cancer-related mortality worldwide. In this computational social systems context, we address the challenge of accurately segmenting CRC lesions from colonoscopy images, which is pivotal for early cancer detection and treatment. The complexity of intestinal environments and variations in medical expertise contribute to the high rate of undetected or misdiagnosed lesions, underscoring the need for advanced computational models. This study presents ColoSegNet, a novel self-supervised deep learning framework designed to enhance the accuracy and effectiveness of CRC diagnosis. By leveraging a comprehensive and annotated colorectal lesion segmentation dataset (CLSD), ColoSegNet incorporates a temporal correlation module to extract critical features from colonoscopy video frames, significantly improving the segmentation of colorectal lesions. Furthermore, ColoSegNet employs a masked autoencoder (MAE) module for self-supervised image reconstruction, preserving the original image integrity and facilitating precise segmentation. Comparative assessments against established models such as UNet, PraNet, and Deeplab V3 demonstrate ColoSegNet’s superior performance in detailed feature representation and overall segmentation accuracy. This research not only contributes to the field of medical imaging but also to computational social systems, by capturing inherent data patterns and integrating specialized modules for feature representation in a healthcare context. Our findings provide valuable insights into the early and accurate detection of CRC, a critical issue given the disease’s high incidence and mortality rates, and its impact on social systems.
Yulong Hu, Dehui Qiu, Rui Li 0115, Liguo Deng, Tinghui Ye, Shengtao Zhu, Xiujing Sun, Weilong Yao, Fa Zhang 0001
IEEE Trans. Comput. Soc. Syst.2
2025 Hierarchical Augmentation and Region-Aware Contrastive Learning for Semi-Supervised Semantic Segmentation of Remote Sensing Images
abstract
Semi-supervised semantic segmentation has gained significant attention as a method to reduce the substantial expense associated with pixel-level labeling. The existing methods primarily rely on consistency regularization or self-training. Recent consistency regularization methods augment the input images with weak or strong augmentation (SA) to improve the performance. However, such simple augmentations are not sufficient to simulate the variations in remote sensing images. The self-training methods exclude the noisy pseudo labels by some selection from the unsupervised training process to obtain better performance. The selection may lead to semantic information loss and bias of latent distribution. To solve the above two problems, we propose hierarchical augmentation (HA) and region-aware contrastive (RC) learning, namely HARC, for remote sensing images. The HA strategy simulates three levels of remote sensing image variations, i.e., spatial variations, uniform spectral variations, and uneven spectral variations. It can significantly enhance the model’s capability to handle more intricate variations. The RC learning learns a class-wise feature distribution of all unlabeled samples instead of some screened unlabeled samples. It can eliminate semantic information loss and enhance the model’s resistance to noise from pseudo labels. Our method is evaluated on three public remote sensing datasets, and the experimental results demonstrate its superiority over state-of-the-art (SOTA) semi-supervised methods.
Bin Sun 0001, Shutao Li 0001, Yulong Hu
IEEE Trans. Geosci. Remote. Sens.4
2025 BMG-Q: Localized Bipartite Match Graph Attention Q-Learning for Ride-Pooling Order Dispatch
abstract
This paper introduces Localized Bipartite Match Graph Attention Q-Learning (BMG-Q), a novel Multi-Agent Reinforcement Learning (MARL) algorithm framework tailored for ride-pooling order dispatch. BMG-Q advances ride-pooling decision-making process with the localized bipartite match graph underlying the Markov Decision Process, enabling the development of novel Graph Attention Double Deep Q Network (GATDDQN) as the MARL backbone to capture the dynamic interactions among ride-pooling vehicles in fleet. Our approach enriches the state information for each agent with GATDDQN by leveraging a localized bipartite interdependence graph and enables a centralized global coordinator to optimize order matching and agent behavior using Integer Linear Programming (ILP). Enhanced by gradient clipping and localized graph sampling, our GATDDQN improves scalability and robustness. Furthermore, the inclusion of a posterior score function in the ILP captures the online exploration-exploitation trade-off and reduces the potential overestimation bias of agents, thereby elevating the quality of the derived solutions. Through extensive experiments and validation, BMG-Q has demonstrated superior performance in both training and operations for thousands of vehicle agents, outperforming benchmark reinforcement learning frameworks by around 10% in accumulative rewards and showing a significant reduction in overestimation bias by over 50%. Additionally, it maintains robustness amidst task variations and fleet size changes, establishing BMG-Q as an effective, scalable, and robust framework for advancing ride-pooling order dispatch operations.
Yulong Hu, Siyuan Feng 0006
IEEE Trans. Intell. Transp. Syst.1
2024 SS-SwinUnet: A Distillation Method of Swin Transformer for Superior Ocular Image Segmentation
abstract
Precise segmentation of the pupil, iris, and sclera is critical for diagnosing and treating ocular diseases such as glaucoma, strabismus, and retinal disorders. However, the fine structural differences within the eye and the interference of complex backgrounds, especially with VR devices prone to reflections, tilts, distortions, and occlusions, present significant challenges. In this paper, we introduce SS-SwinUnet, a novel segmentation method that integrates Swin Transformer and knowledge distillation to achieve superior performance. Specifically, SS-SwinUnet balances feature transfer between the encoder and decoder, reducing redundancy and enhancing representation. Additionally, we incorporate a Boundary Difference over Union Loss to improve boundary segmentation accuracy. We also propose an eye modeling method that parameterizes segmentation results to optimize the semantic segmentation of ocular structures. We constructed the TongRenD dataset, comprising 400 VR-captured videos and 4,100 images, which, along with the TEyeD dataset, was used in our experiments. Results demonstrate that SS-SwinUnet significantly outperforms existing medical image segmentation methods across multiple datasets.
Bowei Ma, Dehui Qiu, Ze Xiong, Yulong Hu, Liguo Deng, Huimei Yuan, Fa Zhang 0001
BIBM4
2023 A Tongue Feature Extraction Method Based on a Sublingual Vein Segmentation
abstract
Sublingual vein features including swelling, varicose and cyanosis are essential for the symptoms differentiation and treatment selection in Traditional Chinese Medicine (TCM) tongue diagnosis, especially reflecting the state of human blood circulation. However, automatic and accurate extraction of sublingual vein features remains a great challenge, limited by both the lack of datasets for sublingual images and the influence of noise from non-tongue and non-sublingual vein components. In this paper, we propose a novel tongue features extraction method based on segmenting the sublingual vein instead of the whole tongue bottom, in which a sublingual vein segmentation framework based on a Polyp-PVT network is developed to eliminate the noise from the surrounding part of the sublingual vein. Meanwhile, we first adopt a transformer-based method such as Swin-Transformer network to extract sublingual vein features by virtue of the awesome capability of the transformer network. In addition, we construct a large dataset including 4018 sublingual vein images for the segmentation and classification of sublingual veins. Experimental results have shown that the tongue feature extraction method combined with a sublingual vein segmentation can greatly outperform the existing tongue feature extracting methods.
Yulong Hu, Dehui Qiu, Fa Zhang 0001, Bin Hu 0001
BIBM1
2017 Watch, attend and parse: An end-to-end neural network based approach to handwritten mathematical expression recognition
Jianshu Zhang 0001, Jun Du 0002, Shiliang Zhang, Dan Liu 0008, Yulong Hu, Jin-Shui Hu, Si Wei, Li-Rong Dai 0001
Pattern Recognit.5