EDBT 2026 Demo / reviewers in the wild / expert
Hang Shao 0001
dblp:33/11442-1
· DBLP profile ↗
15ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TranSpike: Pixel-wise frequency reconstruction and spike interaction for remote photoplethysmography
Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Chuanfei Hu, Shuo Chen 0003, Jian Yang 0003 |
Pattern Recognit. | 1 |
| 2025 | Remote Photoplethysmography in Real-World and Extreme Lighting ScenariosabstractPhysiological activities can be manifested by the sensitive changes in facial imaging. While they are barely observable to our eyes, computer vision manners can, and the derived remote photoplethysmography (rPPG) has shown considerable promise. However, existing studies mainly rely on spatial skin recognition and temporal rhythmic interactions, so they focus on identifying explicit features under ideal light conditions, but perform poorly in-the-wild with intricate obstacles and extreme illumination exposure. In this paper, we propose an end-to-end video transformer model for rPPG. It strives to eliminate complex and unknown external time-varying interferences, whether they are sufficient to occupy subtle biosignal amplitudes or exist as periodic perturbations that hinder network training. In the specific implementation, we utilize global interference sharing, subject background reference, and self-supervised disentanglement to eliminate interference, and further guide learning based on spatiotemporal filtering, reconstruction guidance, and frequency domain and biological prior constraints to achieve effective rPPG. To the best of our knowledge, this is the first robust rPPG model for real outdoor scenarios based on natural face videos, and is lightweight to deploy. Extensive experiments show the competitiveness and performance of our model in rPPG prediction across datasets and scenes. Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Mengkai Yan, Shuo Chen 0003, Jian Yang 0003 |
CVPR | 1 |
| 2025 | Trusted Video-Based Sewer Inspection via Support Clip-Based Pareto-Optimal Evidential NetworkabstractAn automatic vision-based sewer inspection plays a vital role of sewage system in a modern city. Existing methods have utilized evidential deep learning to construct trusted models. Although the acceptable performance has been achieved in sewer defect classification, the fine-grained information of sewer defects in videos is ignored. Meanwhile, the trade-off between multi-label classification and uncertainty estimation remains challenging. In this paper, support clip-based pareto-optimal evidential network (POEN) is proposed for trusted video-based sewer inspection. Specifically, support clip module (SCM) is designed to capture the fine-grained visual representation of defects from local scale segments. Then, evidential deep learning is introduced to quantify the uncertainty for out-of-distribution detection. Furthermore, Pareto-optimal weighting scheme (PWS) is designed to solve the common trade-off dilemma in multi-task learning. Extensive experiments are conducted on VideoPipe, in which the superiority of POEN is demonstrated compared with the state-of-the-art methods. Chenyang Zhao 0009, Chuanfei Hu, Hang Shao 0001, Fir Dunkin, Yongxiong Wang |
IEEE Signal Process. Lett. | 3 |
| 2025 | Self-Supervised Temperature Representation Learning for Fever ScreeningabstractUtilizing thermal infrared facial imaging for fever screening in public spaces has become a common strategy to curb the spread of influenza viruses. However, it is difficult to capture larger number of faces with fever labels, which makes learning facial temperature representation extremely difficult. To overcome this limitation, we propose a self-supervised fever screening framework (SelfFS) to learn temperature representation from infrared face images. Specifically, SelfFS employs rate reduction theory to guide the network to focus on temperature features by expanding the coding rate of faces with different temperatures and compressing the coding rate of faces with the same temperature but different appearances. Furthermore, we impose sparsity constraints on the network parameters, which facilitates the extraction of simple temperature features with a limited number of neurons while filtering complex appearance features. Experiments demonstrate that our SelfFS framework outperforms existing fever screening techniques and achieves the comparable results with the supervised methods. Mengkai Yan, Jianjun Qian, Hang Shao 0001, Lei Luo 0001, Jian Yang 0003 |
IEEE Trans. Cybern. | 3 |
| 2025 | Video-Based Multiphysiological Disentanglement and Remote Robust Estimation for RespirationabstractRemote noncontact respiratory rate estimation by facial visual information has great research significance, providing valuable priors for health monitoring, clinical diagnosis, and anti-fraud. However, existing studies suffer from disturbances in epidermal specular reflections induced by head movements and facial expressions. Furthermore, diffuse reflections of light in the skin-colored subcutaneous tissue caused by multiple time-varying physiological signals independent of breathing are entangled with the intention of the respiratory process, leading to confusion in current research. To address these issues, this article proposes a novel network for natural light video-based remote respiration estimation. Specifically, our model consists of a two-stage architecture that progressively implements vital measurements. The first stage adopts an encoder-decoder structure to recharacterize the facial motion frame differences of the input video based on the gradient binary state of the respiratory signal during inspiration and expiration. Then, the obtained generative mapping, which is disentangled from various time-varying interferences and is only linearly related to the respiratory state, is combined with the facial appearance in the second stage. To further improve the robustness of our algorithm, we design a targeted long-term temporal attention module and embed it between the two stages to enhance the network's ability to model the breathing cycle that occupies ultra many frames and to mine hidden timing change clues. We train and validate the proposed network on a series of publicly available respiration estimation datasets, and the experimental results demonstrate its competitiveness against the state-of-the-art breathing and physiological prediction frameworks. Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Mengkai Yan, Shangbing Gao, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | ASD: Towards Attribute Spatial Decomposition for Prior-Free Facial Attribute RecognitionabstractRepresenting the spatial properties of facial attributes is a vital challenge for facial attribute recognition (FAR). Recent advances have achieved the reliable performances for FAR, benefiting from the description of spatial properties via extra prior information. However, the extra prior information might not be always available, resulting in the restricted application scenario of the prior-based methods. Meanwhile, the spatial ambiguity of facial attributes caused by inherent spatial diversities of facial parts is ignored. To address these issues, we propose a prior-free method for attribute spatial decomposition (ASD), mitigating the spatial ambiguity of facial attributes. The attribute components could be formally described in terms of the spatial locations without any extra prior information. Experimental results demonstrate the superiority of ASD compared with state-of-the-art prior-based methods on both CelebA and LFWA. Chuanfei Hu, Hang Shao 0001, Bo Dong 0001, Zhe Wang 0027, Yongxiong Wang |
ICME | 2 |
| 2024 | FeverNet: Enabling accurate and robust remote fever screening
Mengkai Yan, Jianjun Qian, Hang Shao 0001, Lei Luo 0001, Jian Yang 0003 |
Pattern Recognit. | 3 |
| 2024 | TMFF: Trustworthy Multi-Focus Fusion Framework for Multi-Label Sewer Defect Classification in Sewer Inspection VideosabstractAn automatic vision-based sewer inspection plays a vital role of sewage system in a modern city. Recent advances focus on modeling a deep learning-based method to realize the sewer inspection system, benefiting from the capability of data-driven feature extraction. Although the acceptable performances of sewer defect classification are achieved, there is still a gap between the emerged methods and actual application scenarios. The first issue is that the multi-focus complementarity is ignored to represent the sewer defect, resulting in capturing the multi-scale information of sewer defect inefficiently. Second, the inherent uncertainty of sewer defect is not considered, while the serious unknown sewer defect categories would be missed, resulting in the untrustworthy sewer inspection. In this paper, we focus on quick-view (QV)-based sewer inspection, while a trustworthy multi-focus fusion framework (TMFF) is proposed, jointly combining multi-label classification and uncertainty estimation. Specifically, focal segment module (FSM) is designed based on optical flow to split the QV sewer video into long-focus and short-focus segments, where the multi-focus segments can be modeled to represent the multi-scale information of sewer defect. Then, evidential deep learning (EDL) is introduced to quantify the uncertainty, while joint expert scheme (JES) is designed to aggregate the expert opinions of multi-focus segments. Moreover, evidential disambiguating strategy (EDS) is proposed to alleviate the ambiguity of uncertainty estimation. Extensive experiments are conducted on VideoPipe, in which the superiority of TMFF is demonstrated compared with the state-of-the-art methods. Furthermore, we validate the potential capability of TMFF against the unknown cases of sewer defects. Chuanfei Hu, Chenyang Zhao 0009, Hang Shao 0001, Jin Deng, Yongxiong Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | TranPhys: Spatiotemporal Masked Transformer Steered Remote Photoplethysmography EstimationabstractSubtle variations are invisible to the naked eyes in human physiological signals can reflect important biological and health indicators. Although numerous computer vision methods have been proposed to recover and magnify these changes, most of them either only focus on identifying and recognizing explicit features such as shapes and textures, or are weak in long-term temporal modeling and spatiotemporal interactive perception of implicit biometrics. Therefore, it is difficult for them to robustly overcome various disturbances that affect detection performance. To address these issues, this paper presents TranPhys, a novel remote photoplethysmography (rPPG) network for facial video-based heart rate estimation. Specifically, first, we argue that facial subregions vary over time due to their biological personalities. So we split the input face video into multiple spatiotemporal tubes, build the 3D vision transformer with encoders and decoders to adequately model the high-dimensional representations of the respective regulars in each subregion, and globally coordinate their feedback on the cardiac pulsing waveform. Second, we design the temporal pooling attention to more finely mine the subtle changes hidden in the skin color over time and their long-term contextual rhythm cues. Third, we leverage the self-supervised masked autoencoding paradigm to overcome redundancy to enhance the robustness of our model, and construct the targeted spatiotemporal sampling maps instead of raw input sequences as the pretrained constraint labels to fully inspire self-supervision. We train, validate, and practice our TranPhys on multiple public datasets to demonstrate that our method achieves the competitive performance in remote heart rate estimation. Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Shuo Chen 0003, Chuanfei Hu, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Towards Trustworthy Multi-Label Sewer Defect Classification via Evidential Deep LearningabstractAn automatic vision-based sewer inspection plays a key role of sewage system in a modern city. Recent advances focus on utilizing deep learning model to realize the sewer inspection system, benefiting from the capability of data-driven feature representation. However, the inherent uncertainty of sewer defects is ignored, resulting in the missed detection of serious unknown sewer defect categories. In this paper, we propose a trustworthy multi-label sewer defect classification (TMSDC) method, which can quantify the uncertainty of sewer defect prediction via evidential deep learning. Meanwhile, a novel expert base rate assignment (EBRA) is proposed to introduce the expert knowledge for describing reliable evidences in practical situations. Experimental results demonstrate the effectiveness of TMSDC and the superior capability of uncertainty estimation is achieved on the latest public benchmark. Chenyang Zhao 0009, Chuanfei Hu, Hang Shao 0001, Zhe Wang 0027, Yongxiong Wang |
ICASSP | 3 |
| 2023 | Hyperbolic embedding steered spatiotemporal graph convolutional network for video-based remote heart rate estimation
Hang Shao 0001, Lei Luo 0001, Shuo Chen 0003, Chuanfei Hu, Jian Yang 0003 |
Eng. Appl. Artif. Intell. | 1 |
| 2021 | A Semantic-Enhanced Method Based On Deep SVDD for Pixel-Wise Anomaly DetectionabstractDetecting the anomalous information in multimedia is valuable to many computer vision applications. Recently, many pixel-wise methods modeling by deep learning model have been presented, which can be divided in reconstruction-based and distance-based methods. However, reconstruction-based methods suffer from the low precision of pixel reconstructions. Distance-based methods extract the hierarchical features by a pre-trained model, in order to estimate the anomalies by distances between normal and anomalous features. Nevertheless, multi-level features are ignored in these methods, and semantic information is not considered which is important to enhance the description of anomalies. To over-come the problems, we propose a novel semantic-enhanced anomaly detection method based on deep Support Vector Data Description (SVDD). A new semantic correlation module (SCB) is introduced to enhance the semantic information of the feature representations by cosine similarity. Mean-while, the multi-level architecture is utilized to estimate the final pixel-wise anomaly score. Experimental results demonstrate the proposed method outperforms state-of-the-art methods on MVTec and STC dataset. Chuanfei Hu, Hang Shao 0001 |
ICME | 3 |
| 2021 | Generative image inpainting with salient prior and relative total variation
Hang Shao 0001, Yongxiong Wang |
J. Vis. Commun. Image Represent. | 1 |
| 2020 | Salient Object Detection with Boundary InformationabstractHow to distinguish the low-contrast area near boundaries is a basic challenge in salient object detection. Most of recent state-of-the-art methods can achieve a good performance but still can't work well near boundaries. In this paper, we propose a novel network based on multi-level feature fusion with boundary information to solve this problem. Our model includes two separate decoding sub-networks, one is object sub-network to detect salient objects and another is boundary sub-network which outputs error maps to get boundary information by boundary maps. Moreover, we design a connection and fusion module to exchange and fuse information of objects and boundaries. In addition, to balance the two subnetworks, the optimal weight of loss function is obtained by experiments. The experimental results show that our model can distinguish the low-contrast area near boundaries well by boundary information and achieves the state-of-the-art performance on five common datasets. Yongxiong Wang, Chuanfei Hu, Hang Shao 0001 |
ICME | 4 |
| 2020 | Generative image inpainting via edge structure and color aware fusion
Hang Shao 0001, Yongxiong Wang, Yinghua Fu |
Signal Process. Image Commun. | 1 |