VLDB 2026 Research / reviewers in the wild / expert
Long Shao
dblp:30/10281
· DBLP profile ↗
14ranked-venue papers
0as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ensemble empirical mode decomposition and sample entropy-based adaptive boosting model for solar radiation forecasting for enhanced hydrogen production and carbon dioxide mitigation
Enguang Liu, Xiquan Wang, Shuaizhi Wang, Zezhong Feng, Hanqing Gu, Pan Ding, Long Shao |
Eng. Appl. Artif. Intell. | 8 |
| 2026 | Research on Memory Algorithm Based on Time Series Model and Reinforcement LearningabstractABSTRACT Spaced repetition is a highly effective method of memorization that helps learners to remember large amounts of content efficiently. This paper presents a spaced repetition framework integrating time‐series modelling with reinforcement learning. We propose the GLD‐HLR model, which utilizes Discrete Cosine Transform (DCT) to decouple multi‐scale temporal features in the frequency domain and a Legendre Projection Unit (LPU) to represent continuous memory trajectories via orthogonal basis functions. This architecture significantly reduces computational complexity while enhancing responsiveness to non‐linear memory changes. Furthermore, a PPO‐MMC algorithm is developed to optimize review intervals within a continuous state space. By achieving joint learning of memory prediction and policy scheduling, the framework effectively minimizes review costs while maximizing long‐term retention. This paper validated through ablation and comparative experiments that the mean absolute error (MAE) of the GLD‐HLR model's recall probability predictions remained below 0.03, achieving at least a 4% reduction compared to the LSTM‐HLR model. The mean absolute percentage error (MAPE) for half‐life predictions was below 0.2, which is smaller than the prediction errors of all other models. The PPO‐MMC algorithm achieved a cumulative number of words learned (WTL) exceeding 8000 within 1000 days, with the number of words memorized at the target half‐life (THR) surpassing 7000. This indicates that the algorithm can efficiently help learners master a large number of vocabulary words within a limited time frame and achieve long‐term retention. Long Shao, Yaxiu Qiao, Yuran Yang |
Expert Syst. J. Knowl. Eng. | 2 |
| 2026 | SG-3DGS: Sequential Growing 3D Gaussian Splatting for Scene Reconstruction of Monocular Endoscope VideoabstractThe reconstruction of monocular endoscope video scenes is essential for enhancing the application and analysis of surgical endoscopic images. However, restricted by the narrow space of endoscopic movement and the obstruction of vision within cavities, it is difficult for most conventional methods to perform high-quality reconstruction. To address these challenges, a novel dynamic growing 3D Gaussian splatting architecture is proposed to construct the 3D model of endoscopic scene without precomputed camera poses or Structure from Motion. Firstly, to establish spatial feature associations between interframes, a 2D-3D displacement fields are designed by utilizing dense feature matches and depth prediction. On this basis, a novel displacement field variational optimization is developed to obtain relative poses by minimizing the energy functional associated with field transformation. Secondly, to address the constraint of the endoscopic view, by Gaussian sequential transformation and differential gradient field optimization, a novel Sequential Gaussian Growing Module is proposed to grow the local Gaussian model sequentially. Finally, a novel Forward-Reconstruction&Backward-Optimization architecture is proposed to generate the global Gaussian model. The evaluation is conducted on two public endoscopic datasets: Scared and C3VD. The experimental results demonstrate that the proposed method outperforms state-of-the-art methods in both quantitative metrics (PSNR, SSIM, LPIPS, ATE, RMSE, MAE) and qualitative comparisons. The project page is https://iheckzza.github.io/ DG-3DGS/. Hong Song 0003, Jingfan Fan, Long Shao, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 4 |
| 2026 | Pelvic Fracture Reduction Planning via Joint Shape-Intensity ReferenceabstractPelvic fracture reduction planning is clinically critical yet technically demanding due to the complex anatomical structure of pelvis and the topological discontinuities introduced by fractures. Existing computer-assisted planning approaches dominantly rely on shape-based models, overlooking the rich CT intensity information that is essential for accurate and patient-specific planning. To address this limitation, we propose SIRDiff, a novel framework that incorporates anatomical shape and CT intensity information to generate biomechanically plausible reference models for pelvic fracture reduction planning. SIRDiff comprises three key components: 1) the structure-aware diffusion model to reconstruct the global anatomical structure, 2) the topology-adaptive structural conditioning strategy that maps fracture landmarks into a healthy anatomical graph domain for robust structure guidance, and 3) the detail-preserved autoencoder to ensure the fine-grained image reconstruction from latent representations. Additionally, SIRDiff adopts a multi-task learning approach to jointly predict the reference CT image and corresponding bone segmentation map, which enhances its potential for clinical application and ensures better anatomical consistency. Despite being trained exclusively on synthetic fracture data, SIRDiff shows the strong generalizability to real clinical cases and consistently outperforms existing methods across multiple clinically relevant evaluation metrics, demonstrating its potential as a robust and deployable solution for pelvic fracture reduction planning. Xirui Zhao, Deqiang Xiao, Long Shao, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Yucong Lin, Hong Song 0003, Junqiang Wang, Jian Yang 0009 |
IEEE Trans. Medical Imaging | 4 |
| 2026 | Enhanced CT-CBCT image registration for orthopedic surgery: Integrating rigid-elastic motion modelsabstractComputed tomography (CT) and cone-beam computed tomography (CBCT) image registration play pivotal roles in computer-assisted navigation for orthopedic surgery. Traditional methods often apply uniform deformation models, neglecting the biomechanical differences between rigid structures and soft tissues, which compromises registration accuracy, especially during significant bone displacements. To address this issue, we introduce RE-Reg, a rigid-elastic CT-CBCT image registration framework that jointly learns rigid bone motion and soft tissue deformation. RE-Reg incorporates a rigid alignment (RA) module to estimate global bone motion and an elastic deformation (ED) module to model soft tissue deformation, preserving bony structures through bone shape preservation (BSP) loss. Our comprehensive evaluation on publicly available datasets demonstrates that RE-Reg significantly outperforms existing methods in terms of registration accuracy and rigid bone structure preservation, achieving a 1.3% improvement in Dice similarity coefficient (DSC) and a 23% reduction in rigid bone deformation ( % Δ vol ) compared with the best baseline. This framework not only enhances anatomical fidelity but also ensures biomechanical plausibility and provides a valuable tool for image-guided orthopedic surgery. This code is available at https://github.com/Zq-Huang/RE-Reg. Deqiang Xiao, Hongxun Liu, Long Shao, Danni Ai, Jingfan Fan, Tianyu Fu 0003, Yucong Lin, Hong Song 0003, Jian Yang 0009 |
Virtual Real. Intell. Hardw. | 4 |
| 2026 | Augmented reality surgical navigation: Clinical applications, key technologies, and future directionsabstractSurgical navigation has evolved significantly through advances in augmented reality, virtual reality, and mixed reality, improving precision and safety across many clinical applications, including neurosurgery, maxillofacial, spinal, and arthroplasty procedures. By integrating preoperative imaging with real-time intraoperative data, these systems provide dynamic guidance, reduce radiation exposure, and minimize tissue damage. Key challenges persist, including intraoperative registration accuracy, flexible tissue deformation, respiratory compensation, and real-time imaging quality. Emerging solutions include artificial intelligence-driven segmentation, deformation-field modeling, and hybrid registration techniques. Future developments will include lightweight, portable systems, improved non-rigid registration algorithms, and greater clinical adoption. Despite advances in rigid-tissue applications, soft-tissue navigation requires additional innovation to address motion variability and registration reliability, ultimately advancing minimally invasive surgery and precision medicine. Jingfan Fan, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Long Shao, Tao Chen 0022, Hong Song 0003, Yongtian Wang, Jian Yang 0009 |
Virtual Real. Intell. Hardw. | 8 |
| 2025 | Multi-Task Learning for Optical Flow-Guided Self-Supervised Depth Estimation and Semantic Segmentation in Endoscopic SurgeryabstractSemantic segmentation in surgical scenes requires precise differentiation of organs and tissues, as well as strong generalization capabilities in complex environments, while the training of the model demands a large number of labeled images. To address these challenges and improve efficiency, we propose a joint-learning framework for self-supervised depth estimation and semantic segmentation, enhanced by optical flow. Our approach effectively leverages high-dimensional input features, using depth estimation to guide and improve semantic segmentation performance at a lower labeling cost. Meanwhile, an optical flow module is introduced and jointly trained with the segmentation network, with its outputs fused with RGB images as multimodal inputs to enhance dynamic feature perception and exploit semantic cues from motion. We evaluate our method on the CholecSeg8k and the experimental results demonstrate the effectiveness and robustness of our proposed approach. Xuexin Jiang, Sifan Cao, Long Shao, Jingfan Fan, Jian Yang 0009 |
BIBM | 4 |
| 2025 | SPPReg: Structure-Aware Partial-to-Complete Point Cloud Registration in Computer-Assisted Orthopedic SurgeryabstractAccurate alignment between partial intraoperative and complete preoperative bone surfaces is essential for navigation in computer-assisted orthopedic surgery. However, this task remains challenging due to low surface overlap, significant initial pose discrepancies, and noise inherent in intraoperative data, which often compromise the effectiveness of existing registration methods. To address these challenges, we propose a structure-aware partial-to-complete point cloud registration framework, named SPPReg, for accurate intraoperative-to-preoperative alignment, featuring a two-stage coarse-to-fine design. In the coarse alignment stage, a point completion network reconstructs missing structures in partial scans and leverages global geometric features to facilitate initial alignment under large pose variations. For the fine registration stage, we adopt a self-attention-based feature matching strategy that constructs a feature similarity matrix to establish accurate point correspondences. To reduce uncertainty interference, we design an overlap estimation block that learns point-wise overlap scores to select representative and reliable correspondences within overlapping regions, thereby improving the accuracy of fine registration. Comparative and ablation studies on a public bone point cloud dataset demonstrate that our method outperforms existing approaches in both accuracy and robustness, highlighting its effectiveness and potential for clinical application. Deqiang Xiao, Jingyi Bian, Long Shao, Hong Song 0003, Jian Yang 0009 |
BIBM | 4 |
| 2025 | Double-Shot 3D Shape Measurement With a Dual-Branch Network for Structured Light Projection ProfilometryabstractThe structured light (SL)-based three-dimensional (3D) measurement techniques with deep learning have been widely studied to improve measurement efficiency, among which fringe projection profilometry (FPP) and speckle projection profilometry (SPP) are two popular methods. However, they generally use a single projection pattern for reconstruction, resulting in fringe order ambiguity or poor reconstruction accuracy. To alleviate these problems, we propose a parallel dual-branch Convolutional Neural Network (CNN)-Transformer network (PDCNet), to take advantage of convolutional operations and self-attention mechanisms for processing different SL modalities. Within PDCNet, a Transformer branch is used to capture global perception in the fringe images, while a CNN branch is designed to collect local details in the speckle images. To fully integrate complementary features, we design a double-stream attention aggregation module (DAAM) that consists of a parallel attention subnetwork for aggregating multi-scale spatial structure information. This module can dynamically retain local and global representations to the maximum extent. Moreover, an adaptive mixture density head with bimodal Gaussian distribution is proposed for learning a representation that is precise near discontinuities. Compared to the standard disparity regression strategy, this adaptive mixture head can effectively improve performance at object boundaries. Extensive experiments demonstrate that our method can reduce fringe order ambiguity while producing high-accuracy results on self-made datasets. Mingyang Lei, Jingfan Fan, Long Shao, Hong Song 0003, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Structured Light Image Planar-Topography Feature Decomposition for Generalizable 3D Shape MeasurementabstractThe application of structured light (SL) techniques has achieved remarkable success in three-dimensional (3D) measurements. Traditional methods generally calculate SL information pixel by pixel to obtain the measurement results. Recently, the rise of deep learning (DL) has led to significant developments in this task. However, existing DL-based methods generally learn all features within the image in an end-to-end manner, ignoring the distinction between SL and non-SL information. Therefore, these methods may encounter difficulties in focusing on subtle variations in SL patterns across different scenes, thereby degrading measurement precision. To overcome this challenge, we propose a novel SL Image Planar-Topography Feature Decomposition Network (SIDNet). To fully utilize the information from different SL modality images (fringe and speckle), we decompose different modalities into topography features (modality-specific) and planar features (modality-shared). A physics-driven decomposition loss is proposed to make the topography/planar features dissimilar/similar, which guides the network to distinguish between SL and non-SL information. Moreover, to obtain modality-fused features with global overview and local detail information, we propose a wrapped phase-driven feature fusion module. Specifically, a novel Tri-modality Mamba block is designed to integrate different sources with the guidance of the wrapped phase features. Extensive experiments demonstrate the superiority of our SIDNet in multiple simulated 3D measurement scenes. Moreover, our method shows better generalization ability than other DL models and can be directly applicable to unseen real-world scenes. Mingyang Lei, Jingfan Fan, Long Shao, Hong Song 0003, Deqiang Xiao, Danni Ai, Tianyu Fu 0003, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Landmark and Pose Prediction in Occluded Facial Point Cloud via Explicit Joint Feature Fusion NetworkabstractFacial point clouds collected in practical applications often suffer from pose variations and occlusion. Existing studies typically focus on either pose estimation or landmarks localization, neglecting to fully utilize the effective information from various facial features, thus limiting the improvement of prediction accuracy. Therefore, we propose an innovative 3D facial multi-task prediction network. The proposed network embeds the output of related tasks into feature extraction from the point level to the global level based on the physical dependencies between tasks. This facilitates explicit multi-task knowledge transfer, enabling the simultaneous prediction of facial landmarks, occlusion, and head pose. We introduce a training strategy based on posterior knowledge correction to iteratively refine and improve multi-task prediction results. Moreover, no single dataset provides annotations for all these tasks at once, so we synthesized a 3D landmarks, occlusion and pose (3D-LOP) dataset, which includes annotations for landmarks coordinates, occlusion probability, and head pose. The proposed method was compared with state-of-the-art methods on two public datasets and 3D-LOP. The landmarks localization accuracy improved by 7.1% on the two public datasets, and the pose estimation accuracy and stability on 3D-LOP improved by 28.5% and 32.7%, respectively. The performance on wild data also shows its potential in practical applications. Jingfan Fan, Long Shao, Mingyang Lei, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Hong Song 0003, Yucong Lin, Jian Yang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | PrixMatch: Semi-supervised Network for Multi-modal Medical Image Segmentation with Cross-modal Data Augmentation and Adaptive Prior Knowledge ThresholdingabstractSemi-supervised medical image segmentation has made significant strides, yet most existing methods are confined to single-modality data, limiting both the volume of data and the generalizability of the models. Multi-modal data can provide richer information, expand the dataset and enhance model robustness. However, integrating multi-modal learning into semi-supervised medical image segmentation presents challenges, primarily in how to deal with the scarcity of labels and alignment across different modalities simultaneously. In this paper, we propose PrixMatch, a multi-modal semi-supervised model with a teacher-student strategy for medical image segmentation. Initially, we propose a cross-modal data augmentation strategy, which randomly exchanges image blocks of the same location between different modalities, to guide the student model to learn cross-modal consistency without the need for additional network modules. Secondly, we design a cross-modal adaptive pseudo-label threshold setting strategy, which can align the prior anatomical knowledge of different modalities, and combine the modal-aligned prior knowledge and model learning state to filter the pseudo-labels at the pixel-level, flexibly alleviating the confirmation bias that occurs during semi-supervised training. Experiments demonstrate that PrixMatch achieves a Dice Similarity Coefficient (DSC) of 87.2% on the BTCV (CT) and CHAOS (MR) multi-modal datasets with only 10% labeling ratio, bringing nearly 5.5% improvement over the latest state-of-the-art method. Hong Song 0003, Yucong Lin, Long Shao, Jingfan Fan, Tianyu Fu 0003, Danni Ai, Deqiang Xiao, Jian Yang 0009 |
BIBM | 4 |
| 2021 | Novel Augmented Reality System for Oral and Maxillofacial Surgery
Lele Ding, Long Shao, Zehua Zhao, Tao Zhang 0152, Danni Ai, Jian Yang 0009, Yongtian Wang |
ICIG (2) | 2 |
| 2016 | Single event upset rate modeling for ultra-deep submicron complementary metal-oxide-semiconductor devices
Xiaofei Jia, Chongguang Dai, Jing Liu 0006, Long Shao, Zhaoqing Liu |
Sci. China Inf. Sci. | 7 |