EDBT 2026 Demo / reviewers in the wild / expert
Songyue Yang
dblp:290/6443
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Synchronizing with Attentional Sampling: Modality-Specific Flicker Guidance for Gaze and Hand-Eye Interaction in VRabstractCurrent attentional guidance systems in Virtual Reality (VR) typically treat the user as a static receiver, employing fixed visual cues that fail to account for the fluctuating cognitive state. We address this limitation by positing that effective guidance requires modulating external stimuli to correspond with the brain’s intrinsic attentional sampling mechanism. Using EEG in an immersive environment, we provide direct evidence that this sampling periodicity is not fixed; specifically, the intensified sensorimotor integration load of hand-eye coordination drives the endogenous rhythm to decelerate from the alpha band (∼8 Hz) to the theta band (∼4 Hz). This neural adaptation dictates the temporal requirements for external guidance: identifying optimal parameters via Pareto optimization, we demonstrate that while a 4 Hz cue suffices for gaze, the computationally demanding coordination task requires a higher-frequency (7 Hz) cue to ensure sufficient temporal signal density. Validated in ecological VR scenarios, our adaptive strategy significantly enhanced interaction efficiency without increasing cognitive load. Beyond the specific implementation of flicker, this work establishes a critical design principle for next-generation attention-aware interfaces: maximizing performance by synchronizing information presentation with the user’s task-induced sensorimotor state. Songyue Yang, Kang Yue, Haolin Gao, Mei Guo, Zhonghao Zhu, Fanlu Zeng, Yu Liu 0081 |
VR | 1 |
| 2026 | Multi-modal vehicle trajectory prediction via hierarchical attention and raster-vector maps encoding in unstructured road environments
Zhifa Chen, Peng Chen 0021, Songyue Yang, Rentao Sun, Guizhen Yu |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Focus-guided feature fusion network for lightweight image super-resolution
Tongtai Cao, Songyue Yang, Yue Liu 0005 |
Neural Networks | 4 |
| 2026 | ResiDet: Robust point cloud object detection framework under degraded visual environments with state space model
Shengdi Sun, Songyue Yang, Runsen Liu, Guizhen Yu |
Pattern Recognit. | 2 |
| 2026 | 3DRailNet: A Multifocal Cameras Fusion Network for Long-Range 3-D Rail-Track DetectionabstractAccurate 3-D rail-track detection is vital to the perception of the railway environments for autonomous trains. However, existing methods based on monocular image cannot capture 3-D spatial features, facing challenges in detecting 3-D rail-track in turnouts and distant scenarios. This study introduces 3DRailNet, a long-range 3-D rail-track detection network using multifocal cameras. 3DRailNet consists of two modules: disparity-based feature extraction (DFE) module and long-short rail-track detection (LSRD) module. Specifically, the DFE module utilizes multifocal images to generate a disparity image and depth image, acquiring 3-D depth features to enhance the spatial information of rail-track. Based on the 3-D depth features, the LSRD module designs a detection head for long and short focal cameras to predict the 3-D position of rail-track. Experimental results demonstrate that the mean F1 score (mF1) of our proposed 3DRailNet is 83.2%, establishing it as the state-of-the-art method in this field. All these results indicate that 3DRailNet has the potential to be readily applicable in 3-D rail-track detection in railway environments. Guizhen Yu, Bin Zhou 0007, Songyue Yang |
IEEE Trans. Ind. Informatics | 5 |
| 2026 | Investigating the Effects of Egocentric FPV vs. Exocentric TPV Perspectives and Collaboration Modes on Sense of Agency in VR TeleoperationabstractWith the advancement of virtual reality (VR), human-robot interaction within remote collaboration contexts has emerged as a pivotal research area. However, the nuanced interplay between the user's perspective and the robot's level of automation, and its subsequent impact on the operator's sense of agency (SoA), remains underexplored. This study investigates how perspective and collaboration mode-differentiated by the robot's level of automation-affect the SoA in VR-based human-robot collaboration. We implemented a 2 (perspective: first-person view[FPV] vs. third-person view[TPV]) × 3 (collaboration mode: Direct Remote Control, Supervisory Cooperation, Spectator Mode) experimental design. SoA was assessed via a multimodal approach, measuring both the explicit judgement of agency (JoA) and the implicit feeling of agency (FoA). The results revealed a dissociation between these two components: explicit JoA was dictated by the collaboration mode, with active modes yielding a higher explicit JoA than the passive Spectator Mode. In contrast, implicit FoA exhibited a significant interaction effect. Under FPV, the influence of collaboration mode was pronounced, where Direct Remote Control elicited the strongest FoA; however, under TPV, the differences among collaboration modes disappeared. The relationship between cybersickness and SoA was also explored. These findings suggest that SoA is not a unitary concept but a multifaceted experience, co-shaped by the user's perspective and their role as either a decision-maker or an executor. Furthermore, we investigate the relationship between cybersickness and SoA. Our study offers novel insights for the design of VR-based human-robot systems. Zhonghao Zhu, Kang Yueh, Songyue Yang, Fanlu Zeng, Yue Liu 0005 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Improving Pointing Accuracy for 3D Target Selection in Virtual Reality Through Depth Perception Biases CorrectionabstractAccurate 3D target selection in virtual reality (VR) is fundamentally impeded by pointing uncertainty along the depth axis, a challenge that existing 2D pointing models fail to address due to the complexities of depth perception. Near-eye interactions in VR are influenced by binocular depth cues and vergence-accommodation conflicts (VAC), which introduce significant depth perception biases that impair predictive performance. To address this issue, we first investigate these factors and derive a Gaussian distribution to model near-field depth biases within a 2.5m range. Second, to analyze pointing performance across this extended depth range, we classify 3D target motions into three distinct types: motion-indepth, motion-in-plane, and combined motion. Our analysis identifies that interaction depth and motion amplitude are the two most critical factors influencing pointing accuracy. Accordingly, by incorporating these factors alongside our perceptual bias Gaussian into the Ternary-Gaussian framework, we demonstrate significantly improved predictive performance across diverse 3D motion scenarios. These findings enhance the understanding of user perception in virtual environments and support the development of precise, context-aware interaction cues. Future research can extend these models to design real-time adaptive interfaces, thereby elevating user experiences in VR. Songyue Yang, Kang Yue, Haolin Gao, Yiyi Yang, Mei Guo, Yu Liu 0081, Zhonghao Zhu, Yue Liu 0005 |
ISMAR | 1 |
| 2025 | Lightweight Local-Global Dual-Path Feature Fusion Network for Infrared Small Target Image Super-Resolution and EnhancementabstractInfrared imaging is widely used in remote sensing and military target recognition due to its strong resistance to interference in complex environments. However, imaging mechanisms and hardware limitations cause infrared images to have low-resolution, sparse textures, and significant background noise, which severely restrict the detection of small targets such as low-altitude drones and weak thermal emitters. To overcome these limitations, we propose a lightweight Local–Global Dual-Path Feature Fusion Network (LDFF-Net) that enhances the resolution and quality of infrared images, providing high-quality inputs for subsequent detection tasks. The network includes a Small Target Feature Recognition Module (STFRM) composed of three key components. The Enhanced High Frequency Perception Module (EHFPM) strengthens high-frequency details of small targets while suppressing noise, enabling robust local feature extraction. The State-Space Model (SSM) captures long-range dependencies with linear complexity and models semantic relationships between targets and background to compensate for the limited receptive field of local features. The Adaptive Feature Fusion Unit (AFFU) combines local and global features adaptively to improve the saliency of small targets. During training, we introduce a realistic degradation process based on visible-light images to generate training samples that include complex degradation patterns and noise, which enhances the model’s robustness and generalization. Evaluation on the ARCHIVE and SIRST datasets demonstrates that LDFF-Net outperforms existing state-of-the-art methods across eight widely used full-reference and no-reference metrics, including PSNR, LPIPS, FID, and NIQE. This result confirms the model’s effectiveness in enhancing both the super-resolution and detection performance of infrared small target images. The code and pretrained model weights are publicly available at https://github.com/98Hao/LDFF-Net. Songyue Yang, Tongtai Cao, Yue Liu 0005 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Dynamic Changes of Latency Perception Threshold in Virtual Reality: Behavioral and EEG EvidenceabstractVirtual Reality (VR) technologies in fields such as telehealth, teleconferencing, and virtual education are significantly affected by end-to-end latency, which notably impacts users' interactive experience and performance. Previous research suggests that a perceptual threshold may exist-once latency is reduced below a certain level, users no longer perceive it, and their interactive performance remains largely unaffected. However, there is no consensus on the exact value of this absolute latency perception threshold. In this study, we employed an experimental design based on Fitts' law to investigate whether interaction strategies and task difficulty can alter the latency perception threshold (LPT), and how variations in this threshold influence users' interactive performance. The results show that the LPT is approximately 130-170 ms, and that when interaction strategies prioritize speed or when tasks become more challenging, users exhibit heightened sensitivity to latency. Due to the presence of the LPT, the effect of latency on interactive performance follows a nonlinear pattern, and building on this finding, we refined a Fitts' law model to incorporate the influence of latency. Notably, electroencephalogram (EEG) signals can still capture users' perception of latency when they are unaware of minor latency, demonstrating a level of sensitivity that exceeds conscious awareness. Our findings provide insights into latency effects on performance and perception, guiding the design of more responsive VR interaction systems. Songyue Yang, Kang Yue, Haolin Gao, Mei Guo, Yu Liu 0081, Dan Zhang 0014, Yue Liu 0005 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Forward Long-distance 3D Reconstruction in Rail Transit Scenarios based on Occupancy NetworksabstractIn rail transit autonomous driving scenarios, real-time three-dimensional (3D) reconstruction is crucial for understanding scenes and ensuring the safety of the driving environment. Urban rail trains, with their substantial weight and long braking distances, necessitate an extended forward perception range in 3D space. To address this challenge, this paper proposes a method for forward long-distance 3D scene reconstruction tailored for rail transit scenarios based on occupancy networks. Firstly, a 3D feature representation method using three mutually perpendicular spatial planes is proposed to mitigate the high computational complexity of spatial voxel features. Secondly, considering the characteristics of forward binocular vision in rail transit scenarios, we employ self-attention and cross-attention mechanisms to fuse features between different images. Thirdly, due to the limitations in projection distance of the Light Detection and Ranging (LiDAR) point cloud used for supervision, we introduce a method to generate long-distance dense ground truth during the training stage. By pioneering the application of occupancy networks in rail transit scenarios, this approach significantly extends the forward perception range of autonomous driving trains, achieving an impressive 77.79% mean Intersection over Union (mIoU) accuracy. Songyue Yang, Guizhen Yu |
INDIN | 3 |
| 2024 | Real-Time 3D Object Detection Based on Dynamic Sparse Voxel Transformer in Mining AreaabstractTo improve the safety and efficiency of mining operation, the development of unmanned transportation technologies for mining trucks is extraordinarily rapid. The application of LiDAR in autonomous mining trucks has also become increasingly popular. However, due to the significant difference in the size of objects and the complex noise, the problem of point cloud object detection in mining scenarios remains a major challenge. Existing methods, such as those based on point cloud clustering algorithms, have proved to be difficult to meet practical application requirements. In this paper, a real-time point cloud object detection based on Dynamic Sparse Voxel Transformer (DSVT) is introduced. A series of experiments were conducted to examine the efficacy of the proposed methodology. The experimental results demonstrated that the proposed method attained a mean average precision of 75%. Additionally, the detection latency of the proposed method was observed to be 62 ms. Collectively, the detection speed and accuracy both align with the practical application requirements. Shengdi Sun, Runsen Liu, Songyue Yang, Liyun Wang |
INDIN | 4 |
| 2024 | A Vision-Based Bird's Eye View Representation Network for 3D Objects in Open-pit Mining AreaabstractAutonomous transportation systems, which have reconstructed open-pit mining operations, depend on accurate and real-time obstacle detection in challenging environments. Existing methods often exhibit limitations in accuracy due to their reliance on traditional image-based techniques, leading to errors and incomplete results. To address these issues, we developed a vision-based 3D object detection algorithm designed for open-pit mining environments. Our approach leverages the BEVDepth model, enabling accurate 3D object recognition using monocular camera input. Moreover, camera parameters are integrated to enhance resilience and flexibility across diverse mining environments. Last, our algorithm was implemented on a newly customized open-pit mining dataset. The experimental results verify the algorithm efficiency by attaining a high mean Average Precision score (m$A$P) of 67% while offering real-time performance with inference times as short as 25ms per frame. It enables accurate and efficient obstacle detection in autonomous mining vehicles and contributes to the development of safer and more productive mining operations. Mengen Tai, Bin Zhou 0007, Guizhen Yu, Songyue Yang |
INDIN | 6 |
| 2024 | Cable Segmentation Based on Mask2Former in Open-Pit Mining AreaabstractThe development of unmanned transportation technology has improved operational efficiency and safety in open-pit mining areas. However, there remain significant challenges to be addressed. One pressing issue is the need for mining trucks to pass through the area where the cable is laid on the ground. Accidentally crushing or damaging these cables would lead to significant risk to the mining area operation. However, due to the complexity of the mining environment and the characteristics of cables being thin and curved, the existing methods such as edge segmentation are difficult to meet the requirements for cable segmentation. This paper introduces a cable segmentation method based on Mask2Former and carries out comprehensive experiments to examine the effectiveness of the method. Experimental results show that the proposed method achieves IoU of 66.04% and PA of 80.08%, which can meet the requirements of practical applications. Liyun Wang, Bin Zhou 0007, Songyue Yang, Huazhi Li, Shengdi Sun |
INDIN | 3 |
| 2024 | Exploring Depth-based Perception Conflicts in Virtual Reality through Error-Related PotentialsabstractVirtual Reality (VR) offers a valuable platform for real-life skills training. However, previous research has indicated that human’s perception of depth in VR differs from that of the real world. Such perceptual conflicts can impact immersion and the learning of skills, thus attracting widespread attention. Various methods have been proposed to enhance users’ depth perception, yet the underlying mechanisms of depth perception conflicts still require further research. In this paper, we used Error-Related Potentials (ErrPs) from electroencephalography (EEG) data to investigate the differences in participants’ perceptions at varying depths within the near-field. We designed a within-subjects experiment to successfully introduce depth perception conflicts. From participants exposed to three distinct depths, we collected questionnaire results, performance data, and EEG data. Our findings showed that EEG can effectively detect depth perception conflicts and, following each conflict, participants’ behavioral patterns showed significant changes. In situations with shallower depths, participants exhibited stronger responses to the designed conflicts. This increased sensitivity correlates with their accuracy in depth estimation. This study represents a novel approach to depth perception in VR using ErrPs, setting the stage for further use of physiological signals to measure the granularity of depth perception in VR/AR environments. Haolin Gao, Kang Yue, Songyue Yang, Yu Liu 0081, Mei Guo, Yue Liu 0005 |
VR | 3 |
| 2024 | Key Point Estimate Network for Rail-Track DetectionabstractRail-track detection is a crucial function for an active obstacle avoidance system in trains. However, existing methods face challenges in effectively detecting rail-tracks, particularly in turnout scenarios. This study introduces a novel rail-track detection approach using a key-point estimate network. The network treats the rail-track as a pair and constructs a dedicated model for detection. Additionally, a pseudo-attention mechanism leverages the detection output from previous stages, enabling the network to focus on the rail-track region. Also, a dislocation assignment mechanism is proposed to address label assignment confusion at turnouts. Moreover, a rail-track generalized IoU is also introduced, treating the rail-track as a pair and adds a correction term to enhance detection performance. Experimental results demonstrate that the proposed method achieves a remarkable mF1 score of 69.42%, establishing it as the state-of-the-art (SOTA) in this field. Furthermore, the effectiveness of the proposed method has been validated and applied in real-world testing on the Hong Kong Metro Tsuen Wan Line. Songyue Yang, Guizhen Yu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | FarNet: An Attention-Aggregation Network for Long-Range Rail Track Point Cloud SegmentationabstractRail track segmentation is key to environmental perception of autonomous train. However, due to the complexity of railway track environment, critical issues such as the detection of rail tracks with different curvatures remain to be overcome. In this study, a novel architecture called FarNet is proposed for long-range railway track point cloud segmentation. The proposed FarNet is mainly divided into three parts, i.e., spherical projection, attention-aggregation network and results refinement. Specifically, spherical projection converts the LiDAR point cloud into a pseudo range image, and attention-aggregation network enables railway track detection using the pseudo range image. Furthermore, in the attention-aggregation network two components, i.e., spatial attention module and information aggregation module, are proposed to enhance the capability of rail track segmentation. Last, the results refinement helps further filter out the noise points after segmentation. Experimental results show that the proposed FarNet achieved 98.0% mean intersection-over-union (MIoU) and 98.9% mean pixel accuracy (MPA) for rail track segmentation. Guizhen Yu, Peng Chen 0021, Bin Zhou 0007, Songyue Yang |
IEEE Trans. Intell. Transp. Syst. | 5 |