Yiyao Liu

dblp:195/6650 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-source multi-task meta-learning with task-oriented distribution alignment for gastric cancer analysis in CT images
Ning Yuan, Yiyao Liu, Yingpeng Xie, Jixin Luan, Kuan Lv, Tianfu Wang 0001, Harry Qin, LinLin Shen, Guolin Ma, Bai Ying Lei
Expert Syst. Appl.4
2026 Hierarchical feature-guided dynamic collaborative learning transformer model for ventricular septal defect identification
Cheng Zhao 0003, Peng Yang 0011, Zhuo Xiang, Yiyao Liu, Bei Xia, Harry Qin, Tianfu Wang 0001, Bai Ying Lei, Luyao Zhou
Neurocomputing5
2026 Diff-magnifier: Utilizing state space model for diffusion processes in breast tumor pathological image super-resolution and classification
Yiyao Liu, Tianfu Wang 0001, Shimao Zhu, Bai Ying Lei
Pattern Recognit.1
2025 Overlapping Free: Anchorless UWB-Assisted Relative Pose Estimation for Multi-Robot Systems
abstract
Accurate Relative Pose Estimation (RPE) is critical for effective collaboration of multi-robot systems. Traditional methods using cameras or LiDARs heavily rely on overlapping Fields of View (FoV) between robots, which is highly demanding in practical applications and may hinder collaboration efficiency. To accommodate this issue, we propose Anchorless UWB-Assisted Relative Pose Estimation (AURPE), a novel approach that leverages ultra-wideband (UWB) technology in an anchorless setup to achieve multi-robot RPE without requiring overlapping FoVs or external infrastructure. AURPE first estimates the initial relative poses between robots using inter-robot UWB ranging combined with a Bayesian framework and constrained optimization. During robot operation, AURPE continuously refines the relative poses by integrating UWB measurements with LiDAR-inertial odometry (LIO) and employs a consensus voting mechanism to identify the most reliable pose estimates. Additionally, a pose graph-based backend optimization is incorporated to enhance the accuracy of both initial and real-time relative pose. Extensive simulations and real-world experiments demonstrate that AURPE achieves accurate RPE even in non-overlapping scenarios where traditional methods fail. Compared to state-of-the-art point cloud registration methods, AURPE shows superior performance in both accuracy and robustness, highlighting its potential to significantly enhance cooperative tasks in multi-robot systems operating in complex environments.
Yanpu Yun, Guohao Peng, Jun Zhang 0042, Yiyao Liu, Kaimin Mao, Danwei Wang
ICRA5
2025 LCSPose: Efficient, Accurate and Scalable Markerless 6-DoF Pose Estimation of a Quay Crane Spreader Based on LiDAR and Camera
abstract
Accurate Six Degrees of Freedom (6-DoF) pose estimation of Ship-To-Shore (STS) quay crane spreaders is crucial for ensuring safe and efficient container handling in port automation. However, existing pose estimation techniques face significant challenges, as camera-based systems either rely on markers, which are prone to damage, or struggle with depth estimation inaccuracies. Additionally, 3D sensor-based approaches, particularly point cloud registration (PCR), face challenges such as initial pose errors, high-latency inference, and difficulties in object identification based purely on geometric features. To address these limitations, we propose LCSPose, a LiDAR-camera fusion-based 6-DoF pose estimation method that is marker-free, accurate, efficient, and scalable. Our approach integrates three key modules: (1) a semantic-geometric segmentation module for spreader segmentation and outlier removal, (2) a spatial consistency template sampling module based on Spatial Consistency Score (SC-Score) for reliable template selection across varying distances, and (3) a multi-view coarse-to-fine pose refinement module which incorporates multi-view PCA alignment for robust initial posture prior estimation and iterative pose refinement strategy for long-range registration. Our method demonstrates a 60% improvement in registration recall over state-of-the-art (SOTA) PCR methods, achieving up to 6 cm in translation error and 0.19 degrees in rotation error, while maintaining real-time processing at 20Hz.
Jun Zhang 0042, Guohao Peng, Yanpu Yun, Yiyao Liu, Yuanzhe Wang, Danwei Wang
ICRA5
2025 Feature knowledge distillation-based model lightweight for prohibited item detection in X-ray security inspection images
Yiyao Liu, Jinfeng Yang, Haigang Zhang, Bai Ying Lei
Adv. Eng. Informatics4
2025 Label-guided graph learning network via two-stage cross-modal fusion for multi-label skin disease diagnosis
Cheng Zhao 0003, Chunlun Xiao, Feifei Jin, Zhuo Xiang, Yiyao Liu, Lehang Guo, Tianfu Wang 0001, Bai Ying Lei
Eng. Appl. Artif. Intell.6
2025 ABVS breast tumour segmentation via integrating CNN with dilated sampling self-attention and feature interaction Transformer
Yiyao Liu, Jinyao Li, Yi Yang 0001, Cheng Zhao 0003, Peng Yang 0011, Xiaofei Deng, Tianfu Wang 0001, Bai Ying Lei
Neural Networks1
2025 Federated learning via multi-attention guided UNet for thyroid nodule segmentation of ultrasound images
Zhuo Xiang, Xiaoyu Tian, Yiyao Liu, Minsi Chen, Cheng Zhao 0003, Li-Na Tang, En-Sheng Xue, Hong-Yuan Xue, Ying-Jia Li, Quan-Shui Li, Chang-Jun Wu, Tian-Tian Ren, Jin-Yu Wu, Tianfu Wang 0001, Wen-Ying Liu, Bo-Ji Liu, Li-Ping Sun, Chong-Ke Zhao, Hui-Xiong Xu, Bai Ying Lei
Neural Networks3
2025 Dual-Scale Swin Transformer via Feature Alignment and Adversarial Discrimination for Retinopathy of Prematurity Diagnosis
abstract
Retinopathy of prematurity (ROP) is a retinal vascular disease that primarily affects premature infants with low birth weight. It is a leading cause of childhood blindness worldwide, but it can often be effectively managed with appropriate and timely diagnosis and treatment. To address the impact of image style on model classification performance, this paper proposes a dual-scale Swin Transformer (DS-Swin-T) network for ROP. The network comprises three components: image synthesis (IS), feature alignment, and advanced adversarial learning. The IS module generates synthesis style images as an intermediate latent space between source and target styles, reducing style difference. The DS-Swin-T serves as the primary framework for image feature extraction. Detail and style encoders extract features in the shallow feature space, with detail and style losses aligning these features to ensure consistency across styles. To extract rich style-invariant features and ensure consistent classification within the same category, adversarial learning is applied in the advanced feature space. Finally, feature fusion units process dual-scale classification representations. Our method achieves an average accuracy of 97.91% on the source style dataset. When transferred to other target style datasets, our method effectively mitigates the performance degradation caused by style difference, reaching a maximum average accuracy of 93.66%. Extensive experiments demonstrate the effectiveness of our method.
Shaobin Chen, Yiyao Liu, Hai Xie, Zhenquan Wu, Yingpeng Xie, Cheng Zhao 0003, Tianfu Wang 0001, Bai Ying Lei
IEEE J. Biomed. Health Informatics3
2025 FAMF-Net: Feature Alignment Mutual Attention Fusion With Region Awareness for Breast Cancer Diagnosis via Imbalanced Data
abstract
Automatic and accurate classification of breast cancer in multimodal ultrasound images is crucial to improve patients' diagnosis and treatment effect and save medical resources. Methodologically, the fusion of multimodal ultrasound images often encounters challenges such as misalignment, limited utilization of complementary information, poor interpretability in feature fusion, and imbalances in sample categories. To solve these problems, we propose a feature alignment mutual attention fusion method (FAMF-Net), which consists of a region awareness alignment (RAA) block, a mutual attention fusion (MAF) block, and a reinforcement learning-based dynamic optimization strategy(RDO). Specifically, RAA achieves region awareness through class activation mapping and performs translation transformation to achieve feature alignment. When MAF utilizes a mutual attention mechanism for feature interaction fusion, it mines edge and color features separately in B-mode and shear wave elastography images, enhancing the complementarity of features and improving interpretability. Finally, RDO uses the distribution of samples and prediction probabilities during training as the state of reinforcement learning to dynamically optimize the weights of the loss function, thereby solving the problem of class imbalance. The experimental results based on our clinically obtained dataset demonstrate the effectiveness of the proposed method. Our code will be available at: https://github.com/Magnety/Multi_modal_Image.
Yiyao Liu, Jinyao Li, Cheng Zhao 0003, Harry Qin, Tianfu Wang 0001, Bai Ying Lei
IEEE Trans. Medical Imaging1
2024 Cross-View Detection of Crowded Objects Based on Multi-Sensor Fusion
abstract
Traditional object detection methods are limited by single-sensor constraints, high computational requirements, and poor real-time performance. In addition, occlusion often occurs under the condition of restricted single-view. In this paper, we introduce a camera and LiDAR fusion-based object detection method, which achieves excellent detection performance under limited computational resources. We also explores a fusion detection method deployed with multi-view, which can effectively solve the occlusion issue encountered by single view. The proposed method is valuable for single view as well as multi-view in various application scenarios. Our fusion method significantly improves detection accuracy and reliability, and solves the problems of data discrepancy, interference between sensors, and occlusion due to restricted view. Simulations and extensive experiments show that our proposed object detection method exhibited high accuracy and relatively low computational time.
Zhipeng Gu, Guohao Peng, Yanpu Yun, Yiyao Liu, Zhenyu Wu 0001, Jun Zhang 0042, Xudong Suo, Danwei Wang
ICARCV4
2024 LB-R2R-Calib: Accurate and Robust Extrinsic Calibration of Multiple Long Baseline 4D Imaging Radars for V2X
abstract
As a new sensor, 4D radar (x, y, z, velocity) has great potential for V2X, due to its 3D point cloud, direct doppler velocity output, long distance ranging, low-cost, and more importantly, robust perception in all weathers. However, the extrinsic calibration of multiple long baseline 4D radars is rarely researched in V2X, which is the key to fuse multi-radars. The main reasons are three-folds: (1) New sensor. Thus, it is not surprising that little related work can be found. (2) Long baseline and large viewpoint-difference. Current works are mainly focused on unmanned vehicles, which is short baseline and small viewpoint-difference. (3) Sparse, noisy, and very cluttered 4D radar point cloud. Thus, it is challenging to rapidly and accurately locate the target and extract the feature. In this paper, LB-R2R-Calib (Long Baseline Radar to Radar extrinsic Calibration) is proposed to address these problems. The novelties are: (1) A new target is introduced: an eight-quadrant corner reflector enclosed by a foam sphere. The benefit is the target center is a viewpoint-invariant feature. Thus, it is ideal for large viewpoint-difference calibration. (2) A new feature extraction algorithm is proposed to rapidly locate the target and extract the target center from a very cluttered point cloud, as we observed some important characteristics of 4D radar. Experiments with two 4D radars in real environments with four configurations demonstrate our method is highly accurate and robust.
Jun Zhang 0042, Fangwei Zhang, Zhenyu Wu 0001, Guohao Peng, Yiyao Liu, Qiyang Lyu, Mingxing Wen, Danwei Wang
ICRA6
2023 Privacy-Preserving and Verifiable Outsourcing Inference Against Malicious Servers
abstract
Outsourcing inference enables users to outsource neural network inference tasks to a service provider (e.g., a remote server). This paradigm has brought enormous convenience and effectively solved the resource-limited issues of users, especially for mobile devices. However, it still suffers from two challenges: (1) The user's input data and inference results contain a large amount of private information, which should not be disclosed. (2) The server in this setting may be malicious and hence violate the inference procedure. For example, the server may use a low-quality model to reduce costs or return wrong inference results. While several privacy-preserving inference works have been proposed, they cannot solve the above two problems at the same time. In this work, we propose PPVI, a secure and verifiable outsourcing inference scheme against malicious service providers. PPVI designs a hybrid check technique for inference integrity verification and employs leveled homomorphic encryption to protect users' privacy. These ingredients together make it possible to protect users' privacy and verify the inference correctness in outsourcing inference simultaneously. Extensive experiment results demonstrate our scheme has an excellent performance in terms of verification accuracy and communication and computational overhead.
Yiyao Liu, Hongwei Li 0001, Meng Hao 0001, Guiqiang Hu
GLOBECOM1
2023 4DRadarSLAM: A 4D Imaging Radar SLAM System for Large-scale Environments based on Pose Graph Optimization
abstract
LiDAR-based SLAM may easily fail in adverse weathers (e.g., rain, snow, smoke, fog), while mmWave Radar remains unaffected. However, current researches are primarily focused on 2D$(x,y)$or 3D ($x, y$, doppler) Radar and 3D LiDAR, while limited work can be found for 4D Radar ($x, y, z$, doppler). As a new entrant to the market with unique characteristics, 4D Radar outputs 3D point cloud with added elevation information, rather than 2D point cloud; compared with 3D LiDAR, 4D Radar has noisier and sparser point cloud, making it more challenging to extract geometric features (edge and plane). In this paper, we propose a full system for 4D Radar SLAM consisting of three modules: 1) Front-end module performs scan-to-scan matching to calculate the odometry based on GICP, considering the probability distribution of each point; 2) Loop detection utilizes multiple rule-based loop pre-filtering steps, followed by an intensity scan context step to identify loop candidates, and odometry check to reject false loop; 3) Back-end builds a pose graph using front-end odometry, loop closure, and optional GPS data. Optimal pose is achieved through$\mathrm{g}2\mathrm{o}$. We conducted real experiments on two platforms and five datasets (ranging from 240m to 4.8km) and will make the code open-source to promote further research at: https://github.com/zhuge2333/4DRadarSLAM
Jun Zhang 0042, Huayang Zhuge, Zhenyu Wu 0001, Guohao Peng, Mingxing Wen, Yiyao Liu, Danwei Wang
ICRA6
2023 L2V2T2Calib: Automatic and Unified Extrinsic Calibration Toolbox for Different 3D LiDAR, Visual Camera and Thermal Camera
abstract
Extrinsic calibration between LiDAR-Camera and LiDAR-LiDAR has been researched extensively, because it is the foundation for sensor fusion. Meanwhile, many projects are open-sourced and significantly promote related research. However, limited solutions can unify the calibration between repetitive scanning and non-repetitive scanning 3D LiDAR, sparse and dense 3D LiDAR, visual and thermal camera. Currently, to achieve that, we normally need to use different targets and extract different features for different sensor combinations. Sometimes, human intervention is required to locate the target. It is inconvenient and time-consuming. In this paper, L2V2T2Calib is introduced and open-sourced as a trial to unify the calibration. 1). A four-circular-holes board is adopted for all sensors. The four circle centers can be detected by all the sensors, thus are ideal common features. Previous works also use this target, but the algorithms don’t consider non-repetitive scanning LiDARs, thus cannot be directly applied. 2). To unify the process, an important step is to automatically and robustly detect the target from different types of LiDARs. However, this does not receive enough attention. We propose a method based on template matching. It is simple, but effective and general to different depth sensors. 3). We provide two types of output, minimizing 2D re-projection error (Min2D) and minimizing 3D matching error (Min3D), for different users. And their performance is compared. Extensive experiments conducted in both simulation and real environment demonstrate L2V2T2Calib is accurate, robust, more importantly, unified. The code will be open-sourced to promote related research at: https://github.com/Clothooo/lvt2calib
Jun Zhang 0042, Yiyao Liu, Mingxing Wen, Yufeng Yue, Danwei Wang
IV2
2021 Semi-supervised Attention-Guided VNet for Breast Cancer Detection via Multi-task Learning
Yiyao Liu, Yi Yang 0001, Tianfu Wang 0001, Bai Ying Lei
ICIG (2)1
2018 Pedestrian Motion Learning Based Indoor WLAN Localization via Spatial Clustering
abstract
Applications on Location Based Services (LBSs) have driven the increasing demand for indoor localization technology. The conventional location fingerprinting based localization involves heavy time and labor cost for database construction, while the well‐known Simultaneous Localization and Mapping (SLAM) technique requires assistant motion sensors as well as complicated data fusion algorithms. To solve the above problems, a new pedestrian motion learning based indoor Wireless Local Area Network (WLAN) localization approach is proposed in this paper to achieve satisfactory LBS without the demand for location calibration or motion sensors. First of all, the concept of pedestrian motion learning is adopted to construct users’ motion paths in the target environment. Second, based on the timestamp relation of the collected Received Signal Strength (RSS) sequences, the RSS segments are constructed to obtain the signal clusters with the newly defined high‐dimensional linear distance. Third, the PageRank algorithm is performed to establish the hotspot mapping relations between the physical and signal spaces which are then used to localize the target. Finally, the experimental results show that the proposed approach can effectively estimate the target’s locations and analyze users’ motion preference in indoor environment.
Yanmeng Wang, Mu Zhou, Yiyao Liu
Wirel. Commun. Mob. Comput.4
2017 Simultaneous pathway mapping and behavior understanding with crowdsourced sensing in WLAN environment
Mu Zhou, Zengshan Tian, Yiyao Liu
Ad Hoc Networks4