Daosong Hu

dblp:315/2691 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-5118-7202ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Respiratory Motion Compensation Based on Mid-axis Plane for Dynamic Human Point Cloud Inpainting
Shengtao Li, Jiadun Wang, Daosong Hu, Kai Huang 0001
ICIC (17)3
2026 A frequency-guided denoising framework based on convolutional transformer for electrocardiogram signals
Mingyue Cui, Yewei Gan, Jiepeng Chen, Yanchong Xie, Daosong Hu, Yuning Cui 0001, Kai Huang 0001
Eng. Appl. Artif. Intell.6
2026 LIGA-Net: A lightweight in-vehicle gaze estimation framework based on anisotropic attention
Zexuan Yu, Daosong Hu
Pattern Recognit.3
2026 A Self-Attention-Based LiDAR Point Cloud Compression Framework in Autonomous Driving Environments
abstract
Light detection and ranging (LiDAR) sensors are crucial for autonomous vehicles to accurately perceive the surrounding environment. However, the sparsity and irregularity of large-scale LiDAR point clouds (LPCs) bring challenges for storage and transmission. Meanwhile, existing works usually adopt insufficient context and bring intolerable computation complexity, especially for high-precision LPC reconstruction. To address these problems, we propose a novel self-attention-based framework for LPC compression and reconstruction in autonomous driving environments. Specifically, our approach employs a robust backbone for octree-based feature extraction, which can be pretrained and easily extended to various tasks, thereby reducing the need for extensive task-specific architectural modifications. The backbone constructs node sequences of octree by nonoverlapping context windows and shares the result of a multihead self-attention (MSA) operation among them. Considering the similarity in features among sibling nodes, we design a locally enhanced module for exploiting sibling features and a positional encoding generator for enhancing the translation invariance of the octree node sequence. During postprocessing, we further propose an offset prediction model to reduce coordinate distortions caused by voxelization. Experimental results indicate that compared to the benchmark geometry-based point cloud compression (GPCC), our approach achieves gains of up to 54.4% for geometry and 6.8% for intensity, while compared to the attention-based baseline, we achieve up to 99% reduction in coding time. We believe that our approach effectively mines the spatial geometric features in LPCs and has low coupling for specific tasks, which will boost the related applications from algorithm optimization to industrial products.
Mingyue Cui, Junhua Long, Mingjian Feng, Juncheng Tao, Yuyang Zhong, Yehua Ling, Daosong Hu, Kai Huang 0001
IEEE Trans. Ind. Informatics7
2026 LNet: Lightweight Network for Driver Attention Estimation via Scene and Gaze Consistency
abstract
In resource-constrained vehicle systems, establishing consistency between multi-view scenes and driver gaze remains challenging. Prior methods mainly focus on cross-source data fusion, estimating gaze or attention maps through unidirectional implicit links between scene and facial features. Although bidirectional projection can correct misalignment between predictions and ground truth, the high resolution of scene images and complex semantic extraction incur heavy computational loads. To address these issues, we propose a lightweight driver-attention estimation framework that leverages geometric consistency between scene and gaze to guide feature extraction bidirectionally, thereby strengthening representation. Specifically, we first introduce a lightweight feature extraction module that captures global and local information in parallel through dual asymmetric branches to efficiently extract facial and scene features. An information cross fusion module is then designed to promote interaction between the scene and gaze streams. The multi-branch architecture extracts gaze and geometric cues at multiple scales, reducing the computational redundancy caused by mixed features when modeling geometric consistency across both views. Experiments on a large public dataset show that incorporating scene information introduces no significant computational overhead and yields a better trade-off between accuracy and efficiency. Moreover, leveraging bidirectional projection and the temporal continuity of gaze, we preliminarily explore the framework's potential for predicting attention trends.
Daosong Hu, Mingyue Cui, Kai Huang 0001
IEEE Trans. Image Process.1
2025 FIFA: Fine-grained Inter-frame Attention for Driver's Video Gaze Estimation
abstract
Gaze direction serves as a pivotal indicator for assessing the level of driver attention. While image-based gaze estimation has been extensively researched, there has been a recent shift towards capturing gaze direction from video sequences. This approach encounters notable challenges, including the comprehension of the dynamic pupil evolution across frames and the extraction of head pose information from a relatively static background. To surmount these challenges, we introduce a dual-stream deep learning framework that explicitly models the displacement changes of the pupil through a fine-grained inter-frame attention mechanism and generates weights to adjust gaze embeddings. This technique transforms the face into a set of distinct patches and employs cross-attention to ascertain the correlation between pixel displacements in various patches and adjacent frames, thereby tracking spatial dynamics within the sequence. Our method is validated using two publicly available driver gaze datasets, and the results indicate that it achieves state-of-the-art performance or is on par with the best outcomes while reducing the parameters.
Daosong Hu, Mingyue Cui, Kai Huang 0001
CVPR1
2025 Region Expansion: Optimization of Patch-Fetching Method for Point Cloud Denoising
Shengtao Li, Jiadun Wang, Daosong Hu, Kai Huang 0001
ICANN (2)4
2025 Planar KNN for Multi-camera Interference Mitigation of Point Cloud
Shengtao Li, Jiadun Wang, Daosong Hu, Kai Huang 0001
ICIC (1)3
2025 A Two-Stage Method for Specular Highlight Detection and Removal in Medical Images
Zefeng Li, Mingyue Cui, Daosong Hu, Jin Gong, Jingchong Weng, Lele Tian, Kai Huang 0001
MICCAI (10)3
2025 GHR-2D: Gaze and head redirection via disentanglement and diffusion for gaze estimation
Daosong Hu, Mingyue Cui, Kai Huang 0001
Eng. Appl. Artif. Intell.1
2025 UnMoDE: Uncertainty Modeling for Driver Gaze Estimation via Feature Disentanglement
abstract
Gaze estimation can be used for assessing the attention level of drivers. Current works predominantly focus on enhancing model accuracy, often overlooking the influence of input sample and label uncertainty. In this paper, we propose a framework for uncertainty modeling in driver gaze estimation via feature disentanglement, referred to as UnMoDE. Our approach begins by extracting facial information into distinct feature spaces using an asymmetric dual-branch encoder to obtain gaze features. Subsequently, a multi-layer perceptron (MLP) is employed to project gaze features and labels into an embedding space, representing them as Gaussian distributions. The uncertainty is described using a covariance matrix. Random sampling is applied to derive samples from the gaze embedding distribution to estimate the most probable embedding representation. This estimated representation is then used to regress the gaze direction and is projected back into the gaze feature space, along with identity information, to facilitate facial reconstruction. Extensive experimental evaluations demonstrate that UnMoDE significantly outperforms baseline and state-of-the-art methods on the latest benchmark datasets collected for drivers, particularly in reducing the number of samples with significant errors.
Daosong Hu, Mingyue Cui, Kai Huang 0001
IEEE Trans. Intell. Transp. Syst.1
2024 4D-CAT: Synthesis of 4D Coronary Artery Trees from Systole and Diastole
abstract
The three-dimensional vascular model reconstructed from CT images is widely used in medical diagnosis. At different phases, the beating of the heart can cause deformation of vessels, resulting in different vascular imaging states and false positive diagnostic results. The 4D model can simulate a complete cardiac cycle. Due to the dose limitation of contrast agent injection in patients, it is valuable to synthesize a 4D coronary artery trees through finite phases imaging. In this paper, we propose a method for generating a 4D coronary artery trees, which maps the systole to the diastole through deformation field prediction, interpolates on the timeline, and the motion trajectory of points are obtained. Specifically, the centerline is used to represent vessels and to infer deformation fields using cube-based sorting and neural networks. Adjacent vessel points are aggregated and interpolated based on the deformation field of the centerline point to obtain displacement vectors of different phases. Finally, the proposed method is validated through experiments to achieve the registration of non-rigid vascular points and the generation of 4D coronary trees.
Daosong Hu, Ruomeng Wang, Mingyue Cui, Kai Huang 0001
BIBM1
2024 A Du-Octree based Cross-Attention Model for LiDAR Geometry Compression
abstract
Point cloud compression is an essential technology for efficient storage and transmission of 3D data. Previous methods usually use hierarchical tree data structures for encoding the spatial sparseness of point clouds. However, the node context within the tree is not fully discovered since the feature space among nodes varies significantly. To address this problem, we innovatively represent the LiDAR points in a two-octree structure instead of using traditional single-octree coding, and then design the cross-attention model to capture the hierarchical features between different octrees, of which each octree incorporates a transformer-based deep entropy model and an arithmetic encoder. Besides, we introduce the untied cross-aware position encoding with principal component analysis and different projection matrices, which enhances the correlations over two octrees’ attention feature embeddings. Experimental results show that our method outperforms the previous state-of-the-art works, achieving up to 8.2% Bpp savings on point cloud benchmark datasets with different lasers.
Mingyue Cui, Mingjian Feng, Junhua Long, Daosong Hu, Shuai Zhao 0004, Kai Huang 0001
ICRA4
2024 Semi-Supervised Multitask Learning Using Gaze Focus for Gaze Estimation
abstract
Gaze estimation can be applied in various scenarios, seeking to comprehend human visual attention through camera images. Contemporary research predominantly employs deep learning to directly output gaze from facial or ocular images. however, most methods concentrate solely on estimating gaze direction, overlooking gaze point. We propose two multitask learning frameworks for estimating gaze point and gaze direction, with the objective of achieving unsupervised learning of gaze point and supervised gaze estimation via gaze intersection. Two attention layers are proposed to guide the generation of facial features, addressing the challenge posed by unlabeled gaze point. The focus attention layer employs the eyes to guide facial features, connecting both features and utilizing similarity to enhance eye information. Another approach utilizes only the full face image, employing self-attention to enhance pertinent information. Four loss functions are employed to constrain networks in 2D and 3D spaces. The combination of eye position constraints and attention layers ensures the accuracy of gaze point prediction. Gaze intersection can be used to obtain gaze depth, thereby solving the problem of depth-overlapping. The advantages of the proposed method in gaze tracking are verified through comprehensive experiments.
Daosong Hu, Kai Huang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 GFNet: Gaze Focus Network using Attention for Gaze Estimation
abstract
Gaze estimation can be applied to human visual attention understanding. The current methods mainly obtain gaze mapping from facial or eye images, and most of them only focus on gaze point or gaze direction estimation. In this paper, we propose a multitask gaze focus network for gaze point and gaze direction estimation. Focus attention layer is used to guide the generation of facial features. By connecting eye and face features, feature similarity is used to get attention weights, and make it tend to eyes position. We propose four loss functions to constrain the network in 2D and 3D spaces. The combination of eye position constraint and focus attention layer ensures the accuracy of gaze point estimation. Gaze focus is used to obtain gaze depth. Through comprehensive experiments, the advantages of proposed method in gaze tracking are verified. In addition, the application prospect of proposed method in depth-overlapping is proved.
Daosong Hu, Kai Huang 0001
ICME1
2022 Fast outdoor hazy image dehazing based on saturation and brightness
abstract
Abstract Haze usually limits the visibility and reduces the contrast of outdoor images. The removal of haze in a single image has always been a challenging problem. Recently, many methods have been proposed to effectively remove haze in harsh conditions. However, these methods fail to balance the dehazing effect and time consumption. This paper proposes a method based on HSV colour space to restore the visibility of uniform scattering medium. It can prevent the atmospheric light and transmission from being miscalculated. This method uses the brightness component of haze image to estimate the global atmospheric light, which reduces the influence of luminescent objects on atmospheric light estimation. Then, this paper deduces the estimation model of the saturation of scene radiance based on the atmospheric scattering model. In the meantime, according to the advantages and problems of the model, the fast estimation of the transmission is realized by using the stretching function. Finally, this paper solves the parameters in the model by iterative method. Since this method can estimate different media transmission for each pixel, better results can be obtained. The simulation results show that the algorithm is superior to the existing algorithms in terms of dehazing effect and time consumption.
Daosong Hu, Huiming Tang
IET Image Process.1