EDBT 2026 Demo / reviewers in the wild / expert
Menghan Hu
dblp:203/4120
· DBLP profile ↗
48ranked-venue papers
0as first author
37since 2021 · last 2026
0000-0002-8557-8930ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 23 since 2021Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Live Demonstration: A Portable Pressure Imaging System for Convenient and Ubiquitous Lumbar Disc Herniation Screening
Jin Ai, Sixu Tao, Menghan Hu, Jian Zhang 0060 |
ISCAS | 3 |
| 2026 | Screening of Lumbar Disc Herniation Using Buttock Pressure Imaging System
Jin Ai, Sixu Tao, Menghan Hu, Jian Zhang 0060 |
ISCAS | 3 |
| 2026 | Live Demonstration: Low-Power Wearable System for Real-Time Myopia Prevention and Monitoring
Wanli Bai, Yunmian Li, Menghan Hu |
ISCAS | 6 |
| 2026 | Live Demonstration: Flexible Airbag-Based Intraocular Pressure Monitoring System
Chaoyi Liu, Jian Zhang 0060, Menghan Hu |
ISCAS | 7 |
| 2026 | Video-Based Gait Analysis for Lumbar Disc Herniation Screening
Sixu Tao, Jin Ai, Menghan Hu, Jian Zhang 0060 |
ISCAS | 3 |
| 2026 | Phase Transition Hypothesis of Perception and Cognition in the Visually ImpairedabstractPerception and cognition are core processes that transform external sensory signals into internal representations for knowledge construction and understanding, and in visually impaired individuals, this transformation is reorganized through auditory and tactile feedback. To explain how perceptual information evolves into stable cognitive representations under limited sensory bandwidth, this study proposes Phase Transition Hypothesis of Perception and Cognition. The proposed hypothesis models the perceptual–cognitive process as a dynamic phase transition, in which sensory information evolves from fragmented perception into organized cognition. To counteract perceptual bias induced by information collapse, the Perceptual Dynamic Optimization Mechanism adaptively regulates sensory deviations to stabilize the perceptual–cognitive transition, whereas the Cognitive Potential Model, derived from the Free-Energy Principle, elucidates how stable and self-organizing cognition emerges from this dynamic process. A cognitive simulation system and a blind writing navigation experiment are conducted to validate the hypothesis. Experiments demonstrate the proposed adaptive correction of perceptual bias and the phase transition mechanism from perception to cognition. Ji-Feng Luo, Zhengqiang Jiang, Jian Zhang 0060, Guangtao Zhai, Menghan Hu |
IEEE Signal Process. Lett. | 8 |
| 2026 | Video Respiratory Rate Measurement in Walking Scenarios Using Multi-Strategy Adaptive Denoising
Gan Pei, Junhao Ning, Chenrui Niu, Siqiong Yao, Menghan Hu, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Quality Assessment and Distortion-Aware Saliency Prediction for AI-Generated Omnidirectional ImagesabstractWith the rapid advancement of Artificial Intelligence Generated Content (AIGC) techniques, AI generated images (AIGIs) have attracted widespread attention, among which AI generated omnidirectional images (AIGODIs) hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications. AI generated omnidirectional images exhibit unique quality issues, however, research on the quality assessment and optimization of AI-generated omnidirectional images is still lacking. To this end, this work first studies the quality assessment and distortion-aware saliency prediction problems for AIGODIs, and further presents a corresponding optimization process. Specifically, we first establish a comprehensive database to reflecthumanfeedback for AI-generatedomnidirectionals, termed OHF2024, which includes both subjective quality ratings evaluated from three perspectives and distortion-aware salient regions. Based on the constructed OHF2024 database, we propose two models with shared encoders based on the BLIP-2 model to evaluate the human visual experience and predict distortion-aware saliency for AI-generated omnidirectional images, which are named as BLIP2OIQA and BLIP2OISal, respectively. Finally, based on the proposed models, we present an automatic optimization process that utilizes the predicted visual experience scores and distortion regions to further enhance the visual quality of an AI-generated omnidirectional image. Extensive experiments show that our BLIP2OIQA model and BLIP2OISal model achieve state-of-the-art (SOTA) results in the human visual experience evaluation task and the distortion-aware saliency prediction task for AI generated omnidirectional images, and can be effectively used in the optimization process. The database and codes will be released on https://github.com/IntMeGroup/AIGCOIQA to facilitate future research. Huiyu Duan, Jing Liu 0002, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Clinically Interpretable Geometric Constraints for Optic Cup and Disc SegmentationabstractWe present GCN-CIS, a segmentation framework that embeds clinically interpretable geometric constraints into deep networks for optic cup and disc delineation in fundus images. A Prior Boundary Attention Module (PBAM) sharpens boundary predictions, while an ISNT-guided geometric loss enforces anatomically plausible regional ordering. Evaluated on an internal glaucoma screening set and three public benchmarks (REFUGE, ORIGA, Drishti-GS), GCN-CIS yields consistent gains in cup/disc Dice and reduces vCDR error compared to strong baselines. The method improves clinical reliability of segmentation outputs and facilitates efficient vCDR estimation for screening workflows. Code will be released. Bailiang Zhao, Peng Gao 0007, Menghan Hu |
BIBM | 6 |
| 2025 | CoughSlowFast: Cough Recognition With Audio and Video Signal Fusion
Mingke Feng, Guangtao Zhai, Xiao-Ping Zhang 0002, Menghan Hu |
IEEE Signal Process. Lett. | 4 |
| 2025 | Real-Time Respiration Monitoring via Motion Artifact Suppression and Quality-Guided Peak DetectionabstractReal-time respiration monitoring faces several challenges including network latency in remote settings, limited computational resources, and increased motion artifacts. Although many existing non-contact respiration algorithms are designed for offline processing and thus overlook these limitations, real-time applications demand greater robustness, efficiency, and adaptability to dynamic conditions. In this study, a lightweight framework called Quality-Guided Respiration Monitoring (QGRM) is proposed. This framework integrates a two-stage motion artifact suppression module and a quality-guided peak detection (QGPD) module. The former enhances signal stability through FIR filtering and amplitude limiting, while the latter improves the estimation of the respiration rate by filtering false peaks based on amplitude and zero-crossing constraints. The experimental results obtained with both the public OVRM dataset and a self-constructed simulated dataset demonstrate that QGRM achieves superior accuracy and robustness compared to state-of-the-art methods. The dataset and code are available athttps://github.com/zxx5058/QGRM. Chenrui Niu, Zhanzhan Cheng, Nengfeng Qian, Changyin Wu, Guangtao Zhai, Menghan Hu |
IEEE Signal Process. Lett. | 7 |
| 2025 | Human-Centered Financial Signal Analysis Based on Visual Patterns in Stock ChartsabstractThe study adopted a human-centered perspective to research the financial markets, focusing on identifying variations in eye movement patterns between professional and non-professional traders as they analyze a series of stock charts. Eye movement data was selected as the analysis target based on the hypothesis that it represents a behavioral phenotype indicative of stock analysts' cognitive processes during market analysis. Disparities were identified by conducting variance analysis and the Wilcoxon signed-rank test on statistical metrics derived from eye fixations and saccades. Psychological and behavioral economic interpretations were provided to understand the underlying reasons for these observed patterns. To showcase the practical application potential of the human-centered perspective, eye movement data and human visual characteristics were used to construct visual saliency prediction models of professional stock analysts. Leveraging this human-centered model, we developed two practical application demonstrations specifically designed to support and instruct novice traders. Based on the above demonstrations, a training program was designed that demonstrates how, with ongoing training, the non-professional traders' ability to observe stock charts improves progressively. Ji-Feng Luo, Kaixun Zhang, Xudong An, Menghan Hu, Guangtao Zhai, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 5 |
| 2025 | Optimizing Video-Based Respiration Monitoring: Motion Artifact Reduction and Adaptive ROI SelectionabstractIn non-contact respiratory monitoring, reducing motion artifact and selecting the appropriate Region of Interest (ROI) pose significant challenges. Most motion artifact removal methods rely on signal periodicity assumptions, while respiratory signals usually are non-periodic in real-world scenarios. Existing automated ROI selection approaches are mostly primarily impacted by the texture of clothing, absence of chest landmarks, and obstruction of face. To improve the quality of respiratory signals, in this study, we propose a framework for automatic respiratory ROI selection based on video, namely, Optimizing Video-based Respiration Monitoring (OVRM), which consists of peak-trough adaptive motion artifact removal and characteristic-driven adaptive ROI selection. This motion artifact removal strategy removes motion artifacts by using a dynamic ratio-based judgment mechanism, and reconstructs signals using sinusoidal interpolation. The adaptive ROI method scores signals based on periodicity, similarity, smoothness, and energy, selecting the highest-scoring blocks as the ROIs to match respiratory signals efficiently. Experimental results, validated across four datasets, demonstrate that OVRM effectively reduces signal noise caused by subject movement and outperforms state-of-the-art non-contact respiratory monitoring algorithms. The dataset and code are publicly available at:https://github.com/zxx5058/OVRM. Xudong Tan, Mei Zhou, Menghan Hu, Zhanzhan Cheng, Nengfeng Qian, Changyin Wu, Guangtao Zhai, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 5 |
| 2024 | A Semi-Supervised Approach with Error Reflection for Echocardiography SegmentationabstractSegmenting internal structure from echocardiography is essential for the diagnosis and treatment of various heart diseases. Semi-supervised learning shows its ability in alleviating annotations scarcity. While existing semi-supervised methods have been successful in image segmentation across various medical imaging modalities, few have attempted to design methods specifically addressing the challenges posed by the poor contrast, blurred edge details and noise of echocardiography. These characteristics pose challenges to the generation of high-quality pseudo-labels in semi-supervised segmentation based on Mean Teacher. Inspired by human reflection on erroneous practices, we devise an error reflection strategy for echocardiography semi-supervised segmentation architecture. The process triggers the model to reflect on inaccuracies in unlabeled image segmentation, thereby enhancing the robustness of pseudo-label generation. Specifically, the strategy is divided into two steps. The first step is called reconstruction reflection. The network is tasked with reconstructing authentic proxy images from the semantic masks of unlabeled images and their auxiliary sketches, while maximizing the structural similarity between the original inputs and the proxies. The second step is called guidance correction. Reconstruction error maps decouple unreliable segmentation regions. Then, reliable data that are more likely to occur near high-density areas are leveraged to guide the optimization of unreliable data potentially located around decision boundaries. Additionally, we introduce an effective data augmentation strategy, termed as multi-scale mixing up strategy, to minimize the empirical distribution gap between labeled and unlabeled images and perceive diverse scales of cardiac anatomical structures. Extensive experiments on a public echocardiography dataset CAMUS, and a private clinical echocardiography dataset demonstrate the competitiveness of the proposed method. Xiaoxiang Han 0001, Yiman Liu, Jiang Shang, Qingli Li, Menghan Hu, Qi Zhang 0003, Yan Wang 0033 |
BIBM | 6 |
| 2024 | AIGCOIQA2024: Perceptual Quality Assessment of AI Generated Omnidirectional Imagesabstract[?]In recent years, the rapid advancement of Artificial Intelligence Generated Content (AIGC) has attracted widespread attention. Among the AIGC, AI generated omnidirectional images hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications, hence omnidirectional AIGC techniques have also been widely studied. AI-generated omnidirectional images exhibit unique distortions compared to natural omnidirectional images, however, there is no dedicated Image Quality Assessment (IQA) criteria for assessing them. This study addresses this gap by establishing a large-scale AI generated omnidirectional image IQA database named AIGCOIQA2024 and constructing a comprehensive benchmark. We first generate 300 omnidirectional images based on 5 AIGC models utilizing 25 text prompts. A subjective IQA experiment is conducted subsequently to assess human visual preferences from three perspectives including quality, comfortability, and correspondence. Finally, we conduct a benchmark experiment to evaluate the performance of state-of-the-art IQA models on our database. The AIGCOIQA2024 database is released to facilitate future research on https://github.com/IntMeGroup/AIGCOIQA. Huiyu Duan, Yucheng Zhu, Xiaohong Liu 0001, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
ICIP | 6 |
| 2024 | Lightweight Video-Based Respiration Rate Detection Algorithm: An Application Case on Intensive CareabstractThe video-based non-contact respiration detection technology can be used in many application scenarios to unobtrusively and ubiquitously monitor the physical state of living beings, and various researchers are currently working on this technology. The optical flow method in tandem with crossover point method is rather effective for respiration rate extraction. However, each method has one disadvantage: 1) the redundant feature points in the traditional optical flow method increase the computational effort and reduce the estimation accuracy; and 2) the traditional crossover point method suffers from crossover points unrelated to breathing movements. For these two challenges, two optimization points are proposed in this work: 1) optimize feature point space by combining spatio-temporal information; and 2) use negative feedback design to adaptively remove crossovers that are not related to respiratory movements. The performance of the proposed algorithm is validated by the Large-scale Bedside Respiration Dataset for Intensive Care (LBRD-IC), which is established using the actual surveillance videos acquired from ICU wards. The validity of the above two optimization points is verified by the ablation experiments. The influential analysis of computation time and video resolution on the performance of the proposed algorithm demonstrates that the proposed algorithm can be deployed to various application terminals to monitor the respiration rate of living organisms in real-time and with high accuracy. In addition, field measurements in the ICU ward have shown that our algorithm can measure respiratory signals of the single patient and multiple patients when only one surveillance camera is present. Xudong Tan, Menghan Hu, Guangtao Zhai, Wenfang Li, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2024 | Intelligent diagnosis of atrial septal defect in children using echocardiography with deep learningabstractAtrial septal defect (ASD) is one of the most common congenital heart diseases. The diagnosis of ASD via transthoracic echocardiography is subjective and time-consuming. The objective of this study was to evaluate the feasibility and accuracy of automatic detection of ASD in children based on color Doppler echocardiographic static images using end-to-end convolutional neural networks. The proposed depthwise separable convolution model identifies ASDs with static color Doppler images in a standard view. Among the standard views, we selected two echocardiographic views, i.e., the subcostal sagittal view of the atrium septum and the low parasternal four-chamber view. The developed ASD detection system was validated using a training set consisting of 396 echocardiographic images corresponding to 198 cases. Additionally, an independent test dataset of 112 images corresponding to 56 cases was used, including 101 cases with ASDs and 153 cases with normal hearts. The average area under the receiver operating characteristic curve, recall, precision, specificity, F1-score, and accuracy of the proposed ASD detection model were 91.99, 80.00, 82.22, 87.50, 79.57, and 83.04, respectively. The proposed model can accurately and automatically identify ASD, providing a strong foundation for the intelligent diagnosis of congenital heart diseases. Yiman Liu, Size Hou, Xiaoxiang Han 0001, Tongtong Liang, Menghan Hu, Qingli Li |
Virtual Real. Intell. Hardw. | 5 |
| 2024 | A review of medical ocular image segmentationabstractDeep learning has been extensively applied to medical image segmentation, resulting in significant advancements in the field of deep neural networks for medical image segmentation since the notable success of U-Net in 2015. However, the application of deep learning models to ocular medical image segmentation poses unique challenges, especially compared to other body parts, due to the complexity, small size, and blurriness of such images, coupled with the scarcity of data. This article aims to provide a comprehensive review of medical image segmentation from two perspectives: the development of deep network structures and the application of segmentation in ocular imaging. Initially, the article introduces an overview of medical imaging, data processing, and performance evaluation metrics. Subsequently, it analyzes recent developments in U-Net-based network structures. Finally, for the segmentation of ocular medical images, the application of deep learning is reviewed and categorized by the type of ocular tissue. Menghan Hu |
Virtual Real. Intell. Hardw. | 2 |
| 2024 | ARGA-Unet: Advanced U-net segmentation model using residual grouped convolution and attention mechanism for brain tumor MRI image segmentationabstractMagnetic resonance imaging (MRI) has played an important role in the rapid growth of medical imaging diagnostic technology, especially in the diagnosis and treatment of brain tumors owing to its non-invasive characteristics and superior soft tissue contrast. However, brain tumors are characterized by high non-uniformity and non-obvious boundaries in MRI images because of their invasive and highly heterogeneous nature. In addition, the labeling of tumor areas is time-consuming and laborious. To address these issues, this study uses a residual grouped convolution module, convolutional block attention module, and bilinear interpolation upsampling method to improve the classical segmentation network U-net. The influence of network normalization, loss function, and network depth on segmentation performance is further considered. In the experiments, the Dice score of the proposed segmentation model reached 97.581%, which is 12.438% higher than that of traditional U-net, demonstrating the effective segmentation of MRI brain tumor images. In conclusion, we use the improved U-net network to achieve a good segmentation effect of brain tumor MRI images. Siyi Xun, Sixu Duan, Tong Tong 0001, Qinquan Gao, Chan-Tong Lam, Menghan Hu, Tao Tan 0002 |
Virtual Real. Intell. Hardw. | 9 |
| 2023 | Energy Efficiency Optimization of Intelligent Reflective Surface-assisted Terahertz-RSMA SystemabstractThis paper examines the energy efficiency optimization problem of Intelligent Reflective Surface (IRS)assisted multi-user Rate-Splitting Multiple Access (RSMA) under terahertz propagation. Comparing Salp Swarm Algorithm (SSA) and Successive Convex Approximation (SCA), it is found that SCA requires multiple iterations to solve non-convex resource allocation problems. At the same time, SSA can consume less time to improve energy efficiency. Menghan Hu, Zihuai Lin |
APCC | 3 |
| 2023 | Cup-Disk Ratio Segmentation Joint with Key Retinal Vascular Information Under Diagnostic and Screening Scenarios
Yiqiao Shi, Wanli Bai, Menghan Hu |
CGI (4) | 8 |
| 2023 | CASCO: A Contactless Cough Screening System Based on Audio Signal Processing
Xinru Chen, Wenfang Li, Menghan Hu, Jian Zhang 0060 |
CGI (4) | 7 |
| 2023 | Unobtrusive Respiratory Monitoring System for Intensive CareabstractThe video-based non-contact respiration detection technology can be used in many application scenarios to unobtrusively and ubiquitously monitor the physical state of living beings, and various researchers are currently working on this technology. The optical flow method in tandem with crossover point method is rather effective for respiration rate extraction. However, each method has one disadvantage: 1) the redundant feature points in the traditional optical flow method increase the computational effort and reduce the estimation accuracy; and 2) the traditional crossover point method suffers from crossover points unrelated to breathing movements. For these two challenges, two optimization points are proposed 1) optimize feature point space by combining spatio-temporal information; and 2) use negative feedback design to adaptively remove crossovers unrelated to respiratory movements. The performance of the proposed algorithm is validated by the Large-scale Bedside Respiration Dataset for Intensive Care (LBRD-IC), which is established using the actual surveillance videos acquired from ICU wards. In addition, field measurements in the ICU ward have shown that our algorithm can measure respiratory signals of the single patient and multiple patients when only one surveillance camera is present. Xudong Tan, Menghan Hu, Guangtao Zhai, Wenfang Li, Xiao-Ping Zhang 0002 |
ICASSP | 2 |
| 2023 | Blind Image Quality Assessment for Pathological Microscopic Image Under Screen and Immersion ScenariosabstractThe high-quality pathological microscopic images are essential for physicians or pathologists to make a correct diagnosis. Image quality assessment (IQA) can quantify the visual distortion degree of images and guide the imaging system to improve image quality, thus raising the quality of pathological microscopic images. Current IQA methods are not ideal for pathological microscopy images due to their specificity. In this paper, we present deep learning-based blind image quality assessment model with saliency block and patch block for pathological microscopic images. The saliency block and patch block can handle the local and global distortions, respectively. To better capture the area of interest of pathologists when viewing pathological images, the saliency block is fine-tuned by eye movement data of pathologists. The patch block can capture lots of global information strongly related to image quality via the interaction between different image patches from different positions. The performance of the developed model is validated by the home-made Pathological Microscopic Image Quality Database under Screen and Immersion Scenarios (PMIQD-SIS) and cross-validated by the five public datasets. The results of ablation experiments demonstrate the contribution of the added blocks. The dataset and the corresponding code are publicly available at: https://github.com/mikugyf/PMIQD-SIS. Yifei Guo, Menghan Hu, Xiongkuo Min, Yan Wang 0036, Guangtao Zhai, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2023 | RIVIE: Robust Inherent Video Information EmbeddingabstractImagine an interesting situation when watching a movie, we can scan the screen using our smartphones to get some extra information about this movie such as the cast, the release date, the movie's homepage, etc. Our prospect is a world where each video contains invisible information that can be delivered to us through mobile devices with cameras. This paper proposes the first deep learning-based information hiding method for videos to achieve information transmission from screens to cameras. Compared with hiding information in single images, the methods for videos need to maintain visual quality in both spatial and temporal domains. Furthermore, the training of video models builds on a large video dataset, which needs much more computational resources than training models for images. To reduce the computational complexity, we propose to simulate data on-the-fly to generate simulated sequences from single images. Then, we use the simulated data to train a spatio-temporal generator that hides information in videos while maintaining visual quality. During training, a temporal loss function based on the simulated data is exploited to ensure the temporal consistency of generated videos. After embedding, we use a decoder to recover the hidden information. To simulate the imaging pipeline from screens to cameras in the real world, we insert a distortion network between the generator and decoder. The distortion network is based on differentiable 3D rendering to cover possible distortions introduced in the procedure of camera imaging. Experimental results show that the hidden information in videos can be extracted by cameras without impacting the visual quality. Our work can be applied to many fields, such as advertisement, entertainment, and education. Jun Jia, Zhongpai Gao, Dandan Zhu 0001, Xiongkuo Min, Menghan Hu, Guangtao Zhai |
IEEE Trans. Multim. | 5 |
| 2023 | Angel's Girl for Blind Painters: An Efficient Painting Navigation System Validated by Multimodal Evaluation ApproachabstractFor people who ardently love painting but unfortunately have visual impairments, holding a paintbrush to create a work is a very difficult task. People in this special group are eager to pick up the paintbrush, like Leonardo da Vinci, to create and make full use of their own talents. Therefore, to maximally bridge this gap, we propose a painting navigation system called “Angle’s Eyes” to assist blind people in artistic creation. The proposed system is composed of cognitive system and guidance system. The system adopts drawing board positioning based on QR code, brush navigation based on target detection and bush real-time positioning. Meanwhile, we design a simple yet efficient position information coding rule to remind the user of the current brush tip position. In addition, we design a criterion to efficiently judge whether the brush reaches the target or not. The numerous experiments are conducted to optimize and test the performance of the system. The results of real-world scenario experiments demonstrate that the developed system has great potential to help blind people with painting. This work also demonstrates that it is practicable for the blind people to feel the world through the brush in their hands. In the future, we plan to deploy “Angle’s Eyes” on the phone to make it more portable. The demo video of the proposed painting navigation system is available athttps://doi.org/10.6084/m9.figshare.9760004.v1. Menghan Hu, Qingli Li, Guangtao Zhai, Simon X. Yang, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | TransMRSR: transformer-based self-distilled generative prior for brain MRI super-resolution
Xiaohong Liu 0001, Tao Tan 0002, Menghan Hu, Xiaoer Wei, Tingli Chen, Bin Sheng 0001 |
Vis. Comput. | 4 |
| 2022 | Intelligent Reflection Elimination Imaging Device based on PolarizerabstractGlass reflection is a problem when taking photos through glass windows or showcases. As the visual quality of captured image can be enhanced by removing reflection, we develop an intelligent reflection elimination imaging device based on polarizer to minimize reflection effect on the images. The system mainly consists of a polarizing module, an image analysis module and a reflection removal module. The users can hold the device and capture images with minimum reflection whether in the day or night. The demo video is available at: https://doi.org/10.6084/m9.figshare.19687830.v1. Xinru Chen, Menghan Hu, Lejing Zhang, Yunmian Li |
VCIP | 3 |
| 2022 | Portable Eye Movement Feature Collection Device for Children with AutismabstractEye movement data has become an important char-acterization in the analysis of children with autism spectrum disorder (ASD). Current eye movement measurement meth-ods require specialized expensive equipment, calibration, and trained personnel, limiting their use in general ASD screening, especially in resource-scarce environments. Therefore, collecting eye movement features based on the standard RGB camera of a mobile phone or tablet has many advantages over professional equipment. The system design is based on the Android tablet design, and the screen is divided into two parts to display the normal children and the ASD children paintings. The eye movement data of children is obtained through the front camera, so as to provide data support for future data analysis. Taking the different cooperation degrees of children into account, two collection modes are designed: 1) directly displaying the stimuli in a loop (image mode); and 2) providing the background video interspersed with the stimulus display (video mode). The demo video of the proposed system is available at: https://doi.org/10.6084/m9.figshare.21346806.v1. Xinding Xia, Menghan Hu, Xiaojuan Xue, Qiaoyun Liu, Jian Zhang 0060, Guangtao Zhai |
VCIP | 2 |
| 2022 | Graph-Based Denoising for Respiration and Heart Rate Estimation During Sleep in Thermal VideoabstractQuality sleep is a basic human need for well-being, yet sleep deprivation has been a long-term global problem. A common type of sleep deprivation is obstrucive sleep apnea, where people repeatedly stop breathing during sleep with subsequent abnormal vital signs, namely, respiration rate and heart rate. While tremendous effort has been made for vital signs monitoring systems during sleep, existing works still lack portability for bulky and intrusive systems and reliability for consumer-level, nonintrusive systems. To bridge the gap between practicability and accuracy and facilitate Internet of Things for smart healthcare, in this article, we propose a vital signs estimation system during sleep via a thermal camera. The system first captures thermal image sequences of a sleeping subject and then processes the facial regions within the thermal images for vital signs signal extraction. Specifically, leveraging on the inherent graph structure among subregions of the facial area, we propose a graph-based, spatial–temporal signal denoising scheme. Experimental results show that the graph-based denoising scheme in our system effectively reduces the noise level introduced by cameras and subjects, and our proposed system outperforms state-of-the-art nonintrusive vital signs monitoring systems. Since the algorithm components in our system have relatively low time complexity and no model training is required, our system can be deployed efficiently at the edge devices in a smart home setting. The extracted vital signs can then be used for sleep abnormality detection and disease screening. Cheng Yang 0003, Menghan Hu, Guangtao Zhai, Xiao-Ping Zhang 0002 |
IEEE Internet Things J. | 2 |
| 2022 | RIHOOP: Robust Invisible Hyperlinks in Offline and Online PhotographsabstractIn the era of multimedia and Internet, the quick response (QR) code helps people obtain information from offline to online quickly. However, the QR code is often limited in many scenarios because of its random and dull appearance. Therefore, this article proposes a novel approach to embed hyperlinks into common images, making the hyperlinks invisible for human eyes but detectable for mobile devices equipped with a camera. Our approach is an end-to-end neural network with an encoder to hide messages and a decoder to extract messages. To maintain the hidden message resilient to cameras, we build a distortion network between the encoder and the decoder to augment the encoded images. The distortion network uses differentiable 3-D rendering operations, which can simulate the distortion introduced by camera imaging in both printing and display scenarios. To maintain the visual attraction of the image with hyperlinks, a loss function conforming to the human visual system (HVS) is used to supervise the training of the encoder. Experimental results show that the proposed approach outperforms the previous work on both robustness and quality. Based on the proposed approach, many applications become possible, for example, "image hyperlinks" for advertisement on TV, website, or poster, and "invisible watermark" for copyright protection on digital resources or product packagings. Jun Jia, Zhongpai Gao, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Cybern. | 4 |
| 2021 | Identification of Deep Breath While Moving Forward Based on Multiple Body Regions and Graph Signal AnalysisabstractThis paper presents an unobtrusive solution that can automatically identify deep breath when a person is walking past the global depth camera. Existing non-contact breath assessments achieve satisfactory results under restricted conditions when human body stays relatively still. When someone moves forward, the breath signals detected by depth camera are hidden within signals of trunk displacement and deformation, and the signal length is short due to the short stay time, posing great challenges for us to establish models. To over-come these challenges, multiple region of interests (ROIs) based signal extraction and selection method is proposed to automatically obtain the signal informative to breath from depth video. Subsequently, graph signal analysis (GSA) is adopted as a spatial-temporal filter to wipe the components unrelated to breath. Finally, a classifier for identifying deep breath is established based on the selected breath-informative signal. In validation experiments, the proposed approach outperforms the comparative methods with the accuracy, precision, recall and F1 of 75.5%, 76.2%, 75.0% and 75.2%, respectively. This system can be extended to public places to provide timely and ubiquitous help for those who may have or are going through physical or mental trouble. Yunlu Wang, Cheng Yang 0003, Menghan Hu, Jian Zhang 0060, Qingli Li, Guangtao Zhai, Xiao-Ping Zhang 0002 |
ICASSP | 3 |
| 2021 | Low-Cost and Unobtrusive Respiratory Condition Monitoring Based on Raspberry Pi and Recurrent Neural NetworkabstractThis paper presents a low-cost and unobtrusive intelligent respiratory monitoring system. To achieve low-cost and remote measurement of respiratory signal, an RGB camera collaborated with marker tracking is used as data acquisition sensor, and a Raspberry Pi is used as data processing platform. To overcome challenges in actual applications, the signal processing algorithms are designed for removing sudden body movements and smoothing the raw signal. To discover more specific information in the respiratory signal, respiratory rate is estimated by a translational cross point algorithm, and respiratory pattern is identified by recurrent neural network. Finally, the obtained decision-making information and some original information are sent to user's smartphone via a cloud service platform. For estimating respiratory rate, the Bland-Altman plot demonstrates the satisfactory results with agreement ranges of -0.13 ± 5.85 bpm. With respect to the classification of breathing patterns, the results validate that the system has the good performance with the accuracy, precision, recall, and F1 of 92.5%, 92.5%, 93.3%, and 92.9%, respectively. This work may contribute to the development of low-cost and non-contact respiratory monitoring products specific to home or work health care. Yunlu Wang, Menghan Hu, Jian Zhang 0060, Qingli Li, Guangtao Zhai, Simon X. Yang |
ISCAS | 2 |
| 2021 | Portable Congenital Glaucoma Detection SystemabstractCongenital glaucoma is an eye disease caused by embryonic developmental disorders, which damages the optic nerve. In this demo paper, we proposed a portable non-contact congenital glaucoma detection system, which can evaluate the condition of children's eyes by measuring the cornea size using the developed mobile application. The system consists of two modules viz. cornea identification module and diagnosis module. This system can be utilized by everyone with a smartphone, which is of wider application. It can be used as a convenient home self-examination tool for children in the large-scale screening of congenital glaucoma. The demo video of the proposed detection system is available at: https://doi.org/10.6084/m9.figshare.14728854.v1. Chunjun Hua, Menghan Hu |
VCIP | 2 |
| 2021 | Respiratory Consultant by Your Side: Affordable and Remote Intelligent Respiratory Rate and Respiratory Pattern Monitoring SystemabstractThe aim of this study is to develop an affordable and remote intelligent respiratory monitoring system. To achieve low-cost and remote measurement of respiratory signal, an RGB camera collaborated with marker tracking is used as a data acquisition sensor, and a Raspberry Pi is used as a data processing platform. To overcome challenges in actual applications, the signal processing algorithms are designed for removing sudden body movements and smoothing the raw signal. Subsequently, respiratory rate (RR) is estimated by a translational cross-point algorithm, and the respiratory pattern is identified by the recurrent neural network. For estimating RR, the translational cross-point algorithm performs better than other methods with root-mean-square error (RMSE) of 3.29 bpm. With respect to the classification of breathing patterns, the established neural network performs better than support vector machine-based classifiers with the accuracy, precision, recall, and F1 of 89.0%, 89.0%, 90.5%, and 89.0%, respectively. The obtained decision-making information and some original information are sent to the user’s smartphone via a cloud service platform. In a way, due to its low-price, noncontact, and portable merits, the established system can be seen as a “respiratory consultant” by your side. Yunlu Wang, Menghan Hu, Jian Zhang 0060, Qingli Li, Guangtao Zhai, Simon X. Yang, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Internet Things J. | 2 |
| 2021 | A new bio-inspired metric based on eye movement data for classifying ASD and typically developing children
Shuning Xu, Menghan Hu |
Signal Process. Image Commun. | 3 |
| 2021 | Identification of Melanoma From Hyperspectral Pathology Image Using 3D Convolutional NetworksabstractSkin biopsy histopathological analysis is one of the primary methods used for pathologists to assess the presence and deterioration of melanoma in clinical. A comprehensive and reliable pathological analysis is the result of correctly segmented melanoma and its interaction with benign tissues, and therefore providing accurate therapy. In this study, we applied the deep convolution network on the hyperspectral pathology images to perform the segmentation of melanoma. To make the best use of spectral properties of three dimensional hyperspectral data, we proposed a 3D fully convolutional network named Hyper-net to segment melanoma from hyperspectral pathology images. In order to enhance the sensitivity of the model, we made a specific modification to the loss function with caution of false negative in diagnosis. The performance of Hyper-net surpassed the 2D model with the accuracy over 92%. The false negative rate decreased by nearly 66% using Hyper-net with the modified loss function. These findings demonstrated the ability of the Hyper-net for assisting pathologists in diagnosis of melanoma based on hyperspectral pathology images. Qian Wang 0046, Li Sun 0012, Yan Wang 0033, Mei Zhou, Menghan Hu, Ying Wen 0003, Qingli Li |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Wearable Visually Assistive Device for Blind People to Appreciate Real-world Scene and Screen ImageabstractDue to the loss of vision, the appreciation of the realworld scene and the images displayed on the screen becomes almost impossible for blind people. In an effort to meet the needs of the blind community, we develop a wearable visually assistive device to help them perceive images. With the help of various multimedia information processing technologies, the proposed device can first acquire image information through a depth camera, then implement an image-to-text transformation using image caption technology, and finally the obtained text sequence is fed back to the user via voice. In this way, blind people are able to perceive the outside world, thus creating an unprecedented experience for them. The main technical specifications of the system are: distance perception range is 0.1m to 10m; RGB field of view is 69.4°×42.5°×77°; depth field of view is 91.2°×65.5°×100.6°; maximum weight is 3.05kg. Two demo videos of the proposed navigation system which are respectively recorded for real-world scene and screen image are available at: https://doi.org/10.6084/m9.figshare.12520499.v1. Jin Ai, Menghan Hu, Guangtao Zhai, Jian Zhang 0060, Qingli Li, Wendell Q. Sun |
VCIP | 2 |
| 2020 | Special Cane with Visual Odometry for Real-time Indoor Navigation of Blind PeopleabstractIndoor navigation is urgently needed by blind people in their everyday lives. In this paper, we design an assistive cane with visual odometry based on actual requirements of the blind to aid them in attaining safe indoor navigation. Compared to the state-of-the-art indoor navigation systems, the proposed device is portable, compact, and adaptable. The main specifications of the system are: the perception range is respectively from 0.10m to 2.10m, and 0.08m to 1.60m for width and length dimensions; the maximum weight is 2.1kg; the detection range is from 0.15m and 3.00m; the cruising ability is about 8h; and the objects whose heights are below 80cm can be detected. The demo video of the proposed navigation system is available at: https://doi.org/10.6084/m9.figshare.12399572.v1. Menghan Hu, Qingli Li, Jian Zhang 0060, Xiaofeng Zhou 0002, Guangtao Zhai |
VCIP | 2 |
| 2020 | Unobtrusive and Automatic Classification of Multiple People's Abnormal Respiratory Patterns in Real Time Using Deep Neural Network and Depth CameraabstractRespiratory pattern is a representation of human breathing activity, which can reflect people's physical and psychological condition. Capturing the unexpected abnormal respiratory pattern unobtrusively of the patient or the potential patient has great significance. In the current work, we attempt to capitalize on depth camera and deep learning architecture to achieve the accurate and unobtrusive measurement of abnormal respiratory patterns, and the whole system can classify multiple people's respiratory patterns in a real-time manner. The challenges in this task are threefold: 1) the real-time online system means that the Region of Interest (ROI) needs to be located and tracked automatically; 2) the amount of real-world data is not enough for training to obtain the robust deep neural network; and 3) the intraclass variation is large and the outer class variation is small. Consequently, human joints tracking is applied to determine the location of subjects shoulder and chest. Based on the characteristics of actual respiratory signals, a novel and efficient respiratory simulation model (RSM) is proposed to generate abundant and high-quality training data. Finally, we apply a gated recurrent unit (GRU) neural network with bidirectional and attentional mechanisms (BI-AT-GRU) to classify six clinically significant respiratory patterns (Eupnea, Tachypnea, Bradypnea, Biots, Cheyne-Stokes, and Central-Apnea). The performance of the obtained BI-AT-GRU is tested by the data that is actually measured by the depth camera. The experimental results demonstrate that the proposed model can classify six different respiratory patterns with the accuracy, precision, recall, and F1 of 94.5%, 94.4%, 95.1%, and 94.8%, respectively. In comparative experiments, the obtained BI-AT-GRU specific to respiratory pattern classification outperforms the existing state-of-the-art, viz., BI-AT-LSTM, GRU, long short-term memory (LSTM), and BI-AT-GRU. Moreover, other experimental results indicate that the proposed online measuring system, deep neural network, and the modeling ideas have the potential to be extended to the large-scale applications, such as public places, sleep scenario, and office environment. The demo videos of the proposed system are available at: https://doi.org/10.6084/m9.figshare.11493666.v1. Yunlu Wang, Menghan Hu, Yuwen Zhou, Qingli Li, Nan Yao, Guangtao Zhai, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Internet Things J. | 2 |
| 2019 | RGB-thermal Imaging System Collaborated with Marker Tracking for Remote Breathing Rate MeasurementabstractThe pixel variation signal extracted from the nasal region of RGB-thermal images can be used to achieve breathing rate (BR) measurement. However, this method fails when the nasal region is not detected in complicated motion scenarios. In this paper, we develop an RGB-thermal imaging system collaborated with marker sticker to achieve unobtrusive and accurate BR measurement. Pixel variation signal of Regions of interest (ROI) is extracted from the thermal video and chest movement signal is extracted from the RGB video with the assistance of marker stickers. Subsequently, a custom-made time-domain signal processing approach is developed for determining BR. We further propose a method of splicing computation to measure the BR after separate processing of signal segments. We construct an RGB-thermal video dataset with different head and body movements to evaluate the effectiveness of the proposed algorithm. After linear regression analysis, the determination coefficient (R2) of 0.905 has been observed for the estimated and reference BRs, indicating the feasibility of our proposed method in complex motion scenarios. Lushuang Chen, Menghan Hu, Guangtao Zhai |
VCIP | 3 |
| 2019 | Angel Girl of Visually Impaired Artists: Painting Navigation System for Blind or Visually Impaired PaintersabstractFor those who love painting but unfortunately have visual impairments, holding a paintbrush to create a work is really a difficult task. For the purpose of solving this problem, a painting navigation system for visually impaired painters is introduced through the live demonstration. When painting, the developed system can endow visually impaired persons with the ability to perceive the surrounding environment, thus helping them realize their dream of painting. To achieve this goal, we designed four main modules viz., QR code based drawing board positioning module, brush real-time positioning module, color recognition module and human-computer interaction module, and integrated them into the system. In the validation experiments, the blindfolded users can successfully create a painting with the help of the developed navigation system. Moreover, the users told us that this system provided them with good experience. In a way, this painting navigation system can be seen as "angel's eyes" of visually impaired painters. The demo video of the proposed painting navigation system is available at: https://doi.org/10.6084/m9.figshare.9760004.v1. Menghan Hu, Guangtao Zhai, Huijing Huang, Wa Zhang, Qingli Li, Yinghong Tian, Yanling Shi |
VCIP | 2 |
| 2019 | Physical Password Breaking via Thermal Sequence AnalysisabstractThe thermal camera can capture keyboard surface temperature change after a human's touch. This phenomenon may be used to steal users' passwords physically. In this paper, based on the study of thermal dynamics of keyboards, we design a password break system using an infrared thermal camera. First, we build a signal model to describe the dynamic process of temperature change on the keyboard using Newton's law of cooling. Next, we develop a maximum likelihood parameter estimation algorithm to estimate the keystroke time instants. Then, by maximizing the probability of key order arrangement, a novel password breaking algorithm is developed. Our algorithm is tested using simulated data as well as real-world data. Experiment results show that our algorithm is effective for physical password breaking using thermal characteristics. Based on our results, we discuss strategies for password protection at the end. Xiao-Ping Zhang 0002, Menghan Hu, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Multi-Channel Decomposition in Tandem With Free-Energy Principle for Reduced-Reference Image Quality AssessmentabstractThe visual quality of perceptions is highly correlated with the mechanisms of the human brain and visual system. Recently, the free-energy principle, which has been widely researched in brain theory and neuroscience, is introduced to quantize the perception, action, and learning in human brain. In the field of image quality assessment (IQA), on one hand, the free-energy principle can resort to the internal generative model to simulate the visual stimulus of the human beings. On the other hand, abundant psychological and neurobiological studies reveal that different frequency and orientation components of one visual stimulus arouse different neurons in the striate cortex, and the striate cortex processes visual information in the cerebral cortex. Motivated by these two aspects, a novel reduce-reference IQA metric called the multi-channel free-energy based reduced-reference quality metric is proposed in this paper. First, a two-level discrete Haar wavelet transform is used to decompose the input reference and distorted images. Next, to simulate the generative model in the human brain, the sparse representation is leveraged to extract the free-energy-based features in subband images. Finally, the overall quality metric is obtained through the support vector regressor. Extensive experimental comparisons on four benchmark image quality databases (LIVE, CSIQ, TID2008, and TID2013) demonstrate that the proposed method is highly competitive with the representative reduced-reference and classical full-reference models. Wenhan Zhu, Guangtao Zhai, Xiongkuo Min, Menghan Hu, Jing Liu 0002, Guodong Guo, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Saliency-induced reduced-reference quality index for natural scene and screen content images
Xiongkuo Min, Ke Gu 0001, Guangtao Zhai, Menghan Hu, Xiaokang Yang 0001 |
Signal Process. | 4 |
| 2018 | Arrow's Impossibility Theorem inspired subjective image quality assessment approach
Wenhan Zhu, Guangtao Zhai, Menghan Hu, Jing Liu 0002, Xiaokang Yang 0001 |
Signal Process. | 3 |
| 2017 | IPAD: Intensity potential for adaptive de-quantizationabstractDisplay devices at bit-depth of 10 or higher have been mature but the mainstream media source is still at bit-depth as low as 8. To accommodate the gap, the most economic solution is to render source at low bit-depth for high bit-depth display, which is essentially the procedure of de-quantization. Traditional methods, like zero-padding or bit replication, introduce annoying false contour artifacts. To better estimate the least-significant bits, later works use filtering or interpolation approaches, which exploit only limited neighbor information, can not thoroughly remove the false contours. In this paper, we propose a novel intensity potential field to model the complicated relationships among pixels. Then, an adaptive de-quantization algorithm is proposed to convert low bit-depth images to high bit-depth ones. To the best of our knowledge, this is the first attempt to apply potential field for natural images. The proposed potential field preserves local consistency and models the complicated contexts very well. Extensive experiments on natural image datasets validate the efficiency of the proposed intensity potential field. Significant improvements have been achieved over the state-of-the-art methods on both PSNR and SSIM. Jing Liu 0002, Guangtao Zhai, Xiaokang Yang 0001, Menghan Hu, Chang Wen Chen |
ICME | 4 |
| 2017 | Perceptual information hiding based on multi-channel visual masking
Guangtao Zhai, Xiaokang Yang 0001, Menghan Hu, Jing Liu 0002 |
Neurocomputing | 4 |