VLDB 2026 Research / reviewers in the wild / expert
Yuxuan Liu 0013
dblp:42/7844-13
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-2431-2653ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Simultaneous surgical stereo depth and motion estimation via brightness-aware self-supervised learning
Yuxuan Liu 0013, Xinyao Zhou, Yating Luo, Yunfei Luan, Zhennan Xiao, Yao Guo 0002, Guang-Zhong Yang |
Pattern Recognit. | 1 |
| 2025 | Towards Accurate Brain Electrode Implantation via Cross-modality Fusion of White-light and Photoacoustic MicroscopyabstractInvasive flexible neural electrodes are becoming increasingly prevalent in monitoring and modulating brain neural activity, necessitating the precise and minimally invasive implantation of these electrodes to a depth of a few millimeters beneath the cerebral surface. Although Neuralink has pioneered robot-assisted neural electrode implantation guided by microscopy, it currently lacks the ability to detect non-cerebral surface microvessels that are invisible under the white-light microscope, leading to inaccurate implantation planning and a high risk of trauma. To address this limitation, we introduce a vascular-enhanced strategy that fuses intraoperative white-light microscopy and preoperative photoacoustic microscopy and applies the fusion results to our established microsurgical robotic system for brain electrode implantation. Specifically, a multi-modality data preprocessing pipeline is devised to extract representative features, and a 2.5D fusion network that incorporates a depth encoding mechanism is proposed to predict cross-modality correspondence. The enhanced fusion results are utilized for implantation planning and intraoperative guidance during in vivo surgical procedures. Both quantitative and qualitative results are presented to demonstrate the effectiveness of our proposed cross-modality fusion methods. Furthermore, in vivo surgical implementations on mice underscore the potential of the proposed approach for achieving more precise and minimally invasive brain electrode implantation. Yuxuan Liu 0013, Yating Luo, Yunfei Luan, Xinyao Zhou, Jianxin Yang, Yao Guo 0002, Guang-Zhong Yang |
IROS | 1 |
| 2025 | Deep Coarse-to-Fine Networks for Robust Segmentation and Pose Estimation of Surgical Suturing ThreadsabstractAutonomous suturing is a critical challenge in robot-assisted surgery, where accurate segmentation and pose estimation of suturing threads are essential prerequisites. However, suturing threads are easily occluded by moving instruments and embedded in deformable tissues which make the task much more challenging. To address this, we propose a coarse-to-fine network for detailed segmentation and pose estimation of suturing threads. The coarse stage aims to capture global thread structure, while the fine stage refines the detailed structure through error residual correction. A spatial context fusion module is incorporated to improve the perception of occluded regions, and weighted balanced cross entropy loss as well as hard sample mining strategy is implemented to enhance small target segmentation performance. To deal with severe occlusions, topological constraints are utilized to effectively identify and reconstruct invisible thread segments. Experiments have been conducted on three datasets collected from different surgical scenes including phantom, endoscopy, and microsurgery. Both quantitative and qualitative results have demonstrated that our proposed framework outperforms baseline methods on segmentation and pose estimation of suturing threads, particularly in detecting occluded threads. Our proposed framework generalizes well across different surgical scenarios, showing its potential for automatic suturing. Xinyao Zhou, Yuxuan Liu 0013, Musen Zhang, Yao Guo 0002, Guang-Zhong Yang |
IROS | 2 |
| 2025 | FPM-R2Net: Fused Photoacoustic and operating Microscopic imaging with cross-modality Representation and Registration Network
Yuxuan Liu 0013, Yating Luo, Sung-Liang Chen, Yao Guo 0002, Guang-Zhong Yang |
Medical Image Anal. | 1 |
| 2025 | PoseSDF++: Point Cloud-Based 3-D Human Pose Estimation via Implicit Neural RepresentationabstractPredicting accurate human pose from 3-D visual observation presents a formidable challenge in computer vision, with numerous applications across various industries. However, most existing studies tackled this issue by regressing the 3-D pose from depth maps via 2-D convolutional neural networks or parametric human models, with limited development in point cloud-based methods. To this end, we propose PoseSDF++, i.e., a point cloud-based encoder–decoder network utilizing implicit neural representation to perform 3-D human pose estimation (HPE) and nonparametric shape reconstruction simultaneously. Leveraging the representative capacity of the signed distance function (SDF), we conceptualize the 3-D HPE as a multiple-shape reconstruction task and propose a distance-aware regression method to accurately estimate the 3-D joint positions. In specific, our PoseSDF++ consists of three modules: first,a hierarchical encoderwith vector neuron layers extracts the multiscale rotation equivariant features from the point clouds captured from an arbitrary viewpoint, addressing the degradation issue caused by viewpoint variation of implicit representation; second,a shape decodermaps the extracted feature and the query to its corresponding shape SDF; third,a pose decodercomputes the distance between the query and the target keypoints, namely, the pose SDF. Extensive experiments on four publicly available datasets demonstrate that our PoseSDF++ achieves competitive performance against the state-of-the-art point cloud-based methods and covering the human hand (HANDS 2019), lower limbs (ICL-Gait), and full body (DFAUST, LiDARHuman2.6M) pose estimation. Jianxin Yang, Yuxuan Liu 0013, Xiao Gu 0003, Guang-Zhong Yang, Yao Guo 0002 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Fast Photoacoustic Microscopy with Robot Controlled Microtrajectory OptimizationabstractPhotoacoustic Microscopy (PAM) is a relatively new imaging modality in biomedicine. However, point-by-point raster scanning in PAM suffers from low imaging speed. Sparse sampling has been studied in recent years and with the development of deep learning algorithms, extensive efforts have been devoted to sparse image reconstruction while little attention has been paid to sparse sampling trajectory design required for actual implementation. The use of real-time adaptive robotically controlled sampling with micro-scale accuracy with due consideration of physical constraints can pave the way for using PAM for robot-assisted microsurgery. This work proposes a fast PAM scheme with robot-controlled microtrajectory optimization. The proposed method is adaptive to imaging details of different regions of interest (ROI) and detailed experiments have been conducted on both simulation and in-vivo settings. Results show that our proposed method can achieve faster scanning speed than traditional raster scanning and improved image quality in ROI than the standard spiral trajectory, which demonstrates the effectiveness of our proposed method and its potential to be deployed in other point-by-point scanning systems. Yating Luo, Yuxuan Liu 0013, Sung-Liang Chen, Yao Guo 0002, Guang-Zhong Yang |
ICRA | 2 |
| 2023 | EgoHMR: Egocentric Human Mesh Recovery via Hierarchical Latent Diffusion ModelabstractEgocentric vision has gained increasing popularity in social robotics, demonstrating great potentials for personal assistance and human-centric behavior analysis. Holistic per-ception of human body itself is a prerequisite for downstream applications, including action recognition and anticipation. Extensive research has been performed for human mesh recovery from the exocentric images captured from a third-person view, but limited studies are conducted for heavily distorted yet occluded egocentric images. In this paper, we propose Egocentric Human Mesh Recovery (EgoHMR), a novel hierarchical network based on latent diffusion models. Our method takes a single egocentric frame as the input and it can be trained in an end-to-end manner without supervision of 2D pose. The network is built upon the latent diffusion model by incorporating both global and local features in a hierarchical structure. To train the proposed network, we generate weak labels from synchronized exocentric images. The proposed method can perform human mesh recovery directly from egocentric images and detailed quantitative and qualitative experiments have been conducted to demonstrate the effectiveness of the proposed EgoHMR method. Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang |
ICRA | 1 |
| 2023 | EasyGaze3D: Towards Effective and Flexible 3D Gaze Estimation from a Single RGB CameraabstractEye gaze can convey rich information of human intentions, which enables the social robots to comprehend the cognition and behavior of human targets. However, the existing 3D gaze estimation methods generally have high requirements either on the dedicated hardware or the quantity and quality of training databases, which largely limits their practical application values. This paper proposes EasyGaze3D, an effective 3D gaze estimation framework using a single RGB camera. First, the framework detects the 2D facial landmarks and recovers the 3D facial shape from the input image, and derives the required camera parameters with these features. Then, without loss of generality, the gaze direction can be regarded as the vector pointing from the eyeball center to the pupil center, which are derived respectively from the detected facial landmarks and the spherical fitting performed on the recovered 3D facial shape. Besides, we propose a flexible yet efficient calibration module, namely Easy-Cali, for deriving the subject-specific 3D facial shape and eyeball centers. The features calibrated by Easy-Cali can further boost the performance of EasyGaze3D. Experimental results show that our proposed method, being plug-and-play and without the need of training on large-scale dataset, can achieve superior performance against the existing methods based on deep models. Jianxin Yang, Yuxuan Liu 0013, Zhen Li 0026, Guang-Zhong Yang, Yao Guo 0002 |
IROS | 3 |
| 2023 | EgoFish3D: Egocentric 3D Pose Estimation From a Fisheye Camera via Self-Supervised LearningabstractEgocentric vision has gained increasing popularity recently, opening new avenues for human-centric applications. However, the use of the egocentric fisheye cameras allows wide angle coverage but image distortion is introduced along with strong human body self-occlusion imposing significant challenges in data processing and model reconstruction. Unlike previous work only leveraging synthetic data for model training, this paper presents a new real-world EgoCentric Human Pose (ECHP) dataset. To tackle the difficulty of collecting 3D ground truth using motion capture systems, we simultaneously collect images from a head-mounted egocentric fisheye camera as well as from two third-person-view cameras, circumventing the environmental restrictions. By using self-supervised learning under multi-view constraints, we propose a simple yet effective framework, namely EgoFish3D, for egocentric 3D pose estimation from a single image in different real-world scenarios. The proposed EgoFish3D incorporates three main modules. 1)The third-person-view moduletakes two exocentric images as input and estimates the 3D pose represented in the third-person camera frame; 2)the egocentric modulepredicts the 3D pose in the egocentric camera frame; and 3)the interactive moduleestimates the rotation matrix between the third-person and the egocentric views. Experimental results on our ECHP dataset and existing benchmark datasets demonstrate the effectiveness of the proposed EgoFish3D, which can achieve superior performance to existing methods. Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang |
IEEE Trans. Multim. | 1 |
| 2022 | Tackling Long-Tailed Category Distribution Under Domain Shifts
Xiao Gu 0003, Yao Guo 0002, Zeju Li, Jianing Qiu, Qi Dou 0001, Yuxuan Liu 0013, Benny P. L. Lo, Guang-Zhong Yang |
ECCV (23) | 6 |
| 2022 | PoseSDF: Simultaneous 3D Human Shape Reconstruction and Gait Pose Estimation Using Signed Distance FunctionsabstractVision-based 3D human pose estimation and shape reconstruction play important roles in robot-assisted healthcare monitoring and personal assistance. However, 3D data captured from a single viewpoint always encounter occlusions and exhibit substantial heterogeneity across different views, resulting in significant challenges for both tasks. Extensive approaches have been proposed to perform each task separately, but few of them present a unified solution. In this paper, we propose a novel network based on signed distance functions, namely PoseSDF, to simultaneously reconstruct 3D lower limb shape and estimate gait pose by two dedicated branches. To promote multi-task learning, several strategies are developed to ensure that these two branches leverage the same latent shape code while exchanging information between them. More importantly, an auxiliary RotNet is incorporated into the inference phase, overcoming the inherent limitations of implicit neural functions under cross-view scenarios. Experimental results demonstrate that our proposed PoseSDF can achieve both high-quality shape reconstruction and precise pose estimation, generalizing well on the data from novel views, gait patterns, as well as real-world. Jianxin Yang, Yuxuan Liu 0013, Xiao Gu 0003, Guang-Zhong Yang, Yao Guo 0002 |
ICRA | 2 |
| 2022 | Ego+X: An Egocentric Vision System for Global 3D Human Pose Estimation and Social Interaction CharacterizationabstractEgocentric vision is an emerging topic, which has demonstrated great potential in assistive healthcare scenarios, ranging from human-centric behavior analysis to personal social assistance. Within this field, due to the heterogeneity of visual perception from first-person views, egocentric pose estimation is one of the most significant prerequisites for enabling various downstream applications. However, existing methods for egocentric pose estimation mainly focus on predicting the pose represented in the camera coordinates from a single image, which ignores the latent cues in the temporal domain and results in less accuracy. In this paper, we propose Ego+X, an egocentric vision based system for 3D canonical pose estimation and human-centric social interaction characterization. Our system is composed of two head-mounted egocentric cameras, where one is faced downwards and the other looks outwards. By leveraging the global context provided by visual SLAM, we first propose Ego-Glo for spatial-accurate and temporal-consistent egocentric 3D pose estimation in the canonical coordinate system. With the help of an egocentric camera looking outwards, we then propose Ego-Soc by extending Ego-Glo to various social interaction tasks, e.g., object detection and human-human interaction. Quantitative and qualitative experiments have been conducted to demonstrate the effectiveness of our proposed Ego+X. Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang |
IROS | 1 |