Hongyuan Yu

dblp:232/2265 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 HyCTAS: Multi-objective hybrid convolution-transformer architecture search for real-time image segmentation
Hongyuan Yu, Cheng Wan 0006, Xiyang Dai, Mengchen Liu, Dongdong Chen 0001, Bin Xiao 0004, Yan Huang 0008, Liang Wang 0001
Neurocomputing1
2026 DCNMST: A Deep Contrastive Network With Multiple Self-Supervised Tasks for Diabetic Retinopathy Grading Classification in Internet of Medical Things
Yanfei Sun, Xiangjun Han, Hongyuan Yu, Yifan Zhong, Dongyong Zhang
IEEE Internet Things J.3
2025 SPENet: Self-guided Prototype Enhancement Network for Few-Shot Medical Image Segmentation
Chao Fan 0001, Xibin Jia, Anqi Xiao, Hongyuan Yu, Zhenghan Yang, Yan Huang 0008, Liang Wang 0001
MICCAI (5)4
2025 Sample-Aware RandAugment: Search-Free Automatic Data Augmentation for Effective Image Recognition
Anqi Xiao, Weichen Yu, Hongyuan Yu
Int. J. Comput. Vis.3
2025 DRBRN: A Deep Reconstruction Bottleneck Representation Network for CT-Image-Based Pancreatic Cancer Diagnosis in Internet of Medical Things
abstract
Automatic classification and survival prediction of pancreatic cancer based on computed tomography (CT) images have become a key research direction in the field of intelligent healthcare. Various deep learning-based diagnostic models have been developed for classifying pancreatic cancer from CT images. However, CT images often contain noise and a significant amount of irrelevant information unrelated to pancreatic cancer diagnosis. This irrelevant information weakens the discriminative power of the extracted features, leading to the problem of information redundancy. In addition, existing models generally suffer from poor feature generalization. To address these challenges, we propose a Deep Reconstruction Bottleneck Representation Network (DRBRN) for pancreatic cancer diagnosis, based on Internet of Medical Things (IoMT) and CT imaging. In essence, the DRBRN leverages the information bottleneck principle and a reconstruction task to learn more compact and generalizable features from CT images. It effectively filters out redundant information and noise, preserving only the critical features relevant to the diagnosis of pancreatic cancer. The experimental results on real-world pancreatic cancer CT dataset demonstrate that our proposed DRBRN model achieved the best performance across all three metrics, with a Precision of 0.996, Accuracy of 0.995, and F1-Score of 0.995.
Hongyuan Yu, Xiangjun Han, Yangyang Yu, Caiyu Yi, Ling Ren 0011, Ruoshi Liu, Dongyong Zhang
IEEE Internet Things J.2
2025 SliceMamba With Neural Architecture Search for Medical Image Segmentation
abstract
Despite the progress made in Mamba-based medical image segmentation models, existing methods utilizing unidirectional or multi-directional feature scanning mechanisms struggle to effectively capture dependencies between neighboring positions, limiting the discriminant representation learning of local features. These local features are crucial for medical image segmentation as they provide critical structural information about lesions and organs. To address this limitation, we propose SliceMamba, a simple yet effective locally sensitive Mamba-based medical image segmentation model. SliceMamba features an efficient Bidirectional Slicing and Scanning (BSS) module, which performs bidirectional feature slicing and employs varied scanning mechanisms for sliced features with distinct shapes. This design keeps spatially adjacent features close in the scan sequence, preserving the local structure of the image and enhancing segmentation performance. Additionally, to fit the varying sizes and shapes of lesions and organs, we introduce an Adaptive Slicing Search method that automatically identifies the optimal feature slicing method based on the characteristics of the target data. Extensive experiments on two skin lesion datasets (ISIC2017 and ISIC2018), two polyp segmentation datasets (Kvasir and ClinicDB), one ultra-wide field retinal hemorrhage segmentation dataset (UWF-RHS), and one multi-organ segmentation dataset (Synapse) demonstrate the effectiveness of our method.
Chao Fan 0001, Hongyuan Yu, Yan Huang 0008, Liang Wang 0001, Zhenghan Yang, Xibin Jia
IEEE J. Biomed. Health Informatics2
2024 Consecutive knowledge meta-adaptation learning for unsupervised medical diagnosis
Hongliu Li, Yawen Hou, Xiuyi Chen, Hongyuan Yu
Knowl. Based Syst.5
2023 Cyclic Differentiable Architecture Search
abstract
Differentiable ARchiTecture Search, i.e., DARTS, has drawn great attention in neural architecture search. It tries to find the optimal architecture in a shallow search network and then measures its performance in a deep evaluation network. The independent optimization of the search and evaluation networks, however, leaves a room for potential improvement by allowing interaction between the two networks. To address the problematic optimization issue, we propose new joint optimization objectives and a novel Cyclic Differentiable ARchiTecture Search framework, dubbed CDARTS. Considering the structure difference, CDARTS builds a cyclic feedback mechanism between the search and evaluation networks with introspective distillation. First, the search network generates an initial architecture for evaluation, and the weights of the evaluation network are optimized. Second, the architecture weights in the search network are further optimized by the label supervision in classification, as well as the regularization from the evaluation network through feature distillation. Repeating the above cycle results in a joint optimization of the search and evaluation networks and thus enables the evolution of the architecture to fit the final evaluation network. The experiments and analysis on CIFAR, ImageNet and NATS-Bench [95] demonstrate the effectiveness of the proposed approach over the state-of-the-art ones. Specifically, in the DARTS search space, we achieve 97.52% top-1 accuracy on CIFAR10 and 76.3% top-1 accuracy on ImageNet. In the chain-structured search space, we achieve 78.2% top-1 accuracy on ImageNet, which is 1.1% higher than EfficientNet-B0. Our code and models are publicly available at https://github.com/microsoft/Cream.
Hongyuan Yu, Houwen Peng, Yan Huang 0008, Jianlong Fu, Hao Du 0006, Liang Wang 0001, Haibin Ling
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 SiamON: Siamese Occlusion-Aware Network for Visual Tracking
abstract
Occlusion has been proven to be one of the most challenging factors faced by most visual trackers. There are mainly two difficulties, the first one is that the number of occlusion samples are very limited even though collecting a large-scale training data set, and another one is how to correctly learn the features of the target when comes to occlusion situations. In this paper, we tried to solve these two problems together in our proposed model. To this end, we propose a novel Siamese Occlusion-aware Network (SiamON) for high-performance visual tracking. In particular, we predefine some soft-masks to solve the problem of fewer occlusion samples, which perceive patterns of occlusion contents at different locations and take these masks as the conditions to guide occlusion-aware feature learning. Meanwhile, we propose a target-aware attention mechanism allows the model to pay more attention to the target and further weaken the impact of occlusion. Extensive experiments on several popular benchmarks show that our tracking method exceeds many state-of-the-art trackers especially in the presence of occlusion and meets the requirements of real-time.
Chao Fan 0001, Hongyuan Yu, Yan Huang 0008, Caifeng Shan, Liang Wang 0001, Chenglong Li 0002
IEEE Trans. Circuits Syst. Video Technol.2
2022 Regularized Graph Structure Learning with Semantic Knowledge for Multi-variates Time-Series Forecasting
abstract
Multivariate time-series forecasting is a critical task for many applications, and graph time-series network is widely studied due to its capability to capture the spatial-temporal correlation simultaneously. However, most existing works focus more on learning with the explicit prior graph structure, while ignoring potential information from the implicit graph structure, yielding incomplete structure modeling. Some recent works attempts to learn the intrinsic or implicit graph structure directly, while lacking a way to combine explicit prior structure with implicit structure together. In this paper, we propose Regularized Graph Structure Learning (RGSL) model to incorporate both explicit prior structure and implicit structure together, and learn the forecasting deep networks along with the graph structure. RGSL consists of two innovative modules. First, we derive an implicit dense similarity matrix through node embedding, and learn the sparse graph structure using the Regularized Graph Generation (RGG) based on the Gumbel Softmax trick. Second, we propose a Laplacian Matrix Mixed-up Module (LM3) to fuse the explicit graph and implicit graph together. We conduct experiments on three real-word datasets. Results show that the proposed RGSL model outperforms existing graph forecasting algorithms with a notable margin, while learning meaningful graph structure simultaneously. Our code and models are made publicly available at https://github.com/alipay/RGSL.git.
Hongyuan Yu, Weichen Yu, Yan Huang 0008, Liang Wang 0001, Alex X. Liu
IJCAI1
2022 Generalized Inter-class Loss for Gait Recognition
abstract
Gait recognition is a unique biometric technique that can be performed at a long distance non-cooperatively and has broad applications in public safety and intelligent traffic systems. The previous gait works focus more on minimizing the intra-class variance while ignoring the significance of constraining inter-class variance. To this end, we propose a generalized inter-class loss that resolves the inter-class variance from both sample-level feature distribution and class-level feature distribution. Instead of equal penalty strength on pair scores, the proposed loss optimizes sample-level inter-class feature distribution by dynamically adjusting the pairwise weight. Further, in class-level distribution, the proposed loss adds a constraint on the uniformity of inter-class feature distribution, which forces the feature representations to approximate a hypersphere and keep maximal inter-class variance. In addition, the proposed method automatically adjusts the margin between classes which enables the inter-class feature distribution to be more flexible. The proposed method can be generalized to different gait recognition networks and achieves significant improvements. We conduct a series of experiments on CASIA-B and OUMVLP, and the experimental results show that the proposed loss can significantly improve the performance and achieves the state-of-the-art performances.
Weichen Yu, Hongyuan Yu, Yan Huang 0008, Liang Wang 0001
ACM Multimedia2
2022 Learning a Robust Part-Aware Monocular 3D Human Pose Estimator via Neural Architecture Search
Zerui Chen, Yan Huang 0008, Hongyuan Yu, Liang Wang 0001
Int. J. Comput. Vis.3
2021 Joint Learning Appearance and Motion Models for Visual Tracking
Wenmei Xu, Hongyuan Yu, Wei Wang 0115, Chenglong Li 0002, Liang Wang 0001
PRCV (1)2
2021 End-to-end video text detection with online tracking
Hongyuan Yu, Yan Huang 0008, Lihong Pi, Chengquan Zhang, Liang Wang 0001
Pattern Recognit.1
2020 Towards Part-Aware Monocular 3D Human Pose Estimation: An Architecture Search Approach
Zerui Chen, Yan Huang 0008, Hongyuan Yu, Yiru Guo, Liang Wang 0001
ECCV (3)3
2020 Cream of the Crop: Distilling Prioritized Paths For One-Shot Neural Architecture Search
abstract
One-shot weight sharing methods have recently drawn great attention in neural architecture search due to high efficiency and competitive performance. However, weight sharing across models has an inherent deficiency, i.e., insufficient training of subnetworks in the hypernetwork. To alleviate this problem, we present a simple yet effective architecture distillation method. The central idea is that subnetworks can learn collaboratively and teach each other throughout the training process, aiming to boost the convergence of individual models. We introduce the concept of prioritized path, which refers to the architecture candidates exhibiting superior performance during training. Distilling knowledge from the prioritized paths is able to boost the training of subnetworks. Since the prioritized paths are changed on the fly depending on their performance and complexity, the final obtained paths are the cream of the crop. We directly select the most promising one from the prioritized paths as the final architecture, without using other complex search methods, such as reinforcement learning or evolution algorithms. The experiments on ImageNet verify such path distillation method can improve the convergence ratio and performance of the hypernetwork, as well as boosting the training of subnetworks. The discovered architectures achieve superior performance compared to the recent MobileNetV3 and EfficientNet families under aligned settings. Moreover, the experiments on object detection and more challenging search space show the generality and robustness of the proposed method. Code and models are available at \url{https://github.com/neurips-20/cream.git}.
Houwen Peng, Hao Du 0006, Hongyuan Yu, Jing Liao 0001, Jianlong Fu
NeurIPS3
2019 An End-to-End Video Text Detector with Online Tracking
abstract
Video text detection is considered as one of the most difficult tasks in document analysis due to the following two challenges: 1) the difficulties caused by video scenes, i.e., motion blur, illumination changes, and occlusion; 2) the properties of text including variants of fonts, languages, orientations, and shapes. Most existing methods attempt to enhance the performance of video text detection by cooperating with video text tracking, but treat these two tasks separately. In this work, we propose an end-to-end video text detection model with online tracking to address these two challenges. Specifically, in the detection branch, we adopt ConvLSTM to capture spatial structure information and motion memory. In the tracking branch, we convert the tracking problem to text instance association, and an appearance-geometry descriptor with memory mechanism is proposed to generate robust representation of text instances. By integrating these two branches into one trainable framework, they can promote each other and the computational cost is significantly reduced. Experiments on existing video text benchmarks including ICDAR2013 Video, Minetto and YVT demonstrate that the proposed method significantly outperforms state-of-the-art methods. Our method improves F-score by about 2% on all datasets and it can run realtime with 24.36 fps on TITAN Xp.
Hongyuan Yu, Chengquan Zhang, Junyu Han, Errui Ding, Liang Wang 0001
ICDAR1
2019 Recurrent Deconvolutional Generative Adversarial Networks with Application to Video Generation
Hongyuan Yu, Yan Huang 0008, Lihong Pi, Liang Wang 0001
PRCV (2)1
2018 Identity-Enhanced Network for Facial Expression Recognition
Xingang Wang 0003, Shilei Zhang, Lingxi Xie, Hongyuan Yu
ACCV (4)6