Litong Feng

dblp:133/4032 · DBLP profile ↗
← Back
35ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-6716-3520ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 15 · 10 since 2021Systems, architecture and hardware · 4 · 1 first-author
YearPublicationVenuePosition
2025 VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis
abstract
This paper develops a Versatile and Honest vision language Model (VHM) for remote sensing image analysis. VHM is built on a large-scale remote sensing image-text dataset with rich-content captions (VersaD), and an honest instruction dataset comprising both factual and deceptive questions (HnstD). Unlike prevailing remote sensing image-text datasets, in which image captions focus on a few prominent objects and their relationships, VersaD captions provide detailed information about image properties, object attributes, and the overall scene. This comprehensive captioning enables VHM to thoroughly understand remote sensing images and perform diverse remote sensing tasks. Moreover, different from existing remote sensing instruction datasets that only include factual questions, HnstD contains additional deceptive questions stemming from the non-existence of objects. This feature prevents VHM from producing affirmative answers to nonsense queries, thereby ensuring its honesty. In our experiments, VHM significantly outperforms various vision language models on common tasks of scene classification, visual question answering, and visual grounding. Additionally, VHM achieves competent performance on several unexplored tasks, such as building vectorizing, multi-label classification and honest question answering.
Chao Pang 0001, Xingxing Weng, Jiang Wu 0003, Yi Liu 0028, Jiaxing Sun 0001, Litong Feng, Gui-Song Xia, Conghui He
AAAI9
2025 Text4Seg: Reimagining Image Segmentation as Text Generation
abstract
Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks; however, effectively integrating image segmentation into these models remains a significant challenge. In this paper, we introduce Text4Seg, a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. This unified representation allows seamless integration into the auto-regressive training pipeline of MLLMs for easier optimization. We demonstrate that representing an image with $16\times16$ semantic descriptors yields competitive segmentation performance. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptors by 74\% and accelerating inference by $3\times$, without compromising performance. Extensive experiments across various vision tasks, such as referring expression segmentation and comprehension, show that Text4Seg achieves state-of-the-art performance on multiple datasets by fine-tuning different MLLM backbones. Our approach provides an efficient, scalable solution for vision-centric tasks within the MLLM framework.
Mengcheng Lan, Chaofeng Chen, Yue Zhou 0005, Jiaxing Xu, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang 0001
ICLR7
2024 ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang 0001
ECCV (47)5
2024 ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang, Litong Feng, Wayne Zhang 0001
ECCV (68)5
2023 Consistent-Teacher: Towards Reducing Inconsistent Pseudo-Targets in Semi-Supervised Object Detection
abstract
In this study, we dive deep into the inconsistency of pseudo targets in semi-supervised object detection (SSOD). Our core observation is that the oscillating pseudo-targets undermine the training of an accurate detector. It injects noise into the student's training, leading to severe overfitting problems. Therefore, we propose a systematic solution, termed Consistent-Teacher, to reduce the inconsistency. First, adaptive anchor assignment (ASA) substitutes the static IoU-based strategy, which enables the student network to be resistant to noisy pseudo-bounding boxes. Then we calibrate the subtask predictions by designing a 3D feature alignment module (FAM-3D). It allows each classification feature to adaptively query the optimal feature vector for the regression task at arbitrary scales and locations. Lastly, a Gaussian Mixture Model (GMM) dynamically revises the score threshold of pseudo-bboxes, which stabilizes the number of ground truths at an early stage and remedies the unreliable supervision signal during training. Consistent-Teacher provides strong results on a large range of SSOD evaluations. It achieves 40.0 mAP with ResNet-50 backbone given only 10% of annotated MS-COCO data, which surpasses previous base-lines using pseudo labels by around 3 mAP. When trained on fully annotated MS-COCO with additional unlabeled data, the performance further increases to 47.7 mAP. Our code is available at https://github.com/Adamdad/ConsistentTeacher.
Xinjiang Wang, Xingyi Yang, Yijiang Li, Litong Feng, Shijie Fang, Chengqi Lyu, Kai Chen 0002, Wayne Zhang 0001
CVPR5
2023 Revisiting Weak-to-Strong Consistency in Semi-Supervised Semantic Segmentation
abstract
In this work, we revisit the weak-to-strong consistency framework, popularized by FixMatch from semi-supervised classification, where the prediction of a weakly perturbed image serves as supervision for its strongly perturbed version. Intriguingly, we observe that such a simple pipeline already achieves competitive results against recent advanced works, when transferred to our segmentation scenario. Its success heavily relies on the manual design of strong data augmentations, however, which may be limited and inadequate to explore a broader perturbation space. Motivated by this, we propose an auxiliary feature perturbation stream as a supplement, leading to an expanded perturbation space. On the other, to sufficiently probe original image-level augmentations, we present a dual-stream perturbation technique, enabling two strong views to be simultaneously guided by a common weak view. Consequently, our overall Unified Dual-Stream Perturbations approach (UniMatch) surpasses all existing methods significantly across all evaluation protocols on the Pascal, Cityscapes, and COCO benchmarks. Its superiority is also demonstrated in remote sensing interpretation and medical image analysis. We hope our reproduced FixMatch and our results can inspire more future works.
Lihe Yang, Lei Qi 0001, Litong Feng, Wayne Zhang 0001, Yinghuan Shi
CVPR3
2023 Diverse Cotraining Makes Strong Semi-Supervised Segmentor
abstract
Deep co-training has been introduced to semi-supervised segmentation and achieves impressive results, yet few studies have explored the working mechanism behind it. In this work, we revisit the core assumption that supports co-training: multiple compatible and conditionally independent views. By theoretically deriving the generalization upper bound, we prove the prediction similarity between two models negatively impacts the model’s generalization ability. However, most current co-training models are tightly coupled together and violate this assumption. Such coupling leads to the homogenization of networks and confirmation bias which consequently limits the performance. To this end, we explore different dimensions of co-training and systematically increase the diversity from the aspects of input domains, different augmentations and model architectures to counteract homogenization. Our Diverse Co-training outperforms the state-of-the-art (SOTA) methods by a large margin across different evaluation protocols on the Pascal and Cityscapes. For example, we achieve the best mIoU of 76.2%, 77.7% and 80.2% on Pascal with only 92, 183 and 366 labeled images, surpassing the previous best results by more than 5%.
Yijiang Li, Xinjiang Wang, Lihe Yang, Litong Feng, Wayne Zhang 0001, Ying Gao 0004
ICCV4
2023 SmooSeg: Smoothness Prior for Unsupervised Semantic Segmentation
abstract
Unsupervised semantic segmentation is a challenging task that segments images into semantic groups without manual annotation. Prior works have primarily focused on leveraging prior knowledge of semantic consistency or priori concepts from self-supervised learning methods, which often overlook the coherence property of image segments. In this paper, we demonstrate that the smoothness prior, asserting that close features in a metric space share the same semantics, can significantly simplify segmentation by casting unsupervised semantic segmentation as an energy minimization problem. Under this paradigm, we propose a novel approach called SmooSeg that harnesses self-supervised learning methods to model the closeness relationships among observations as smoothness signals. To effectively discover coherent semantic segments, we introduce a novel smoothness loss that promotes piecewise smoothness within segments while preserving discontinuities across different segments. Additionally, to further enhance segmentation quality, we design an asymmetric teacher-student style predictor that generates smoothly updated pseudo labels, facilitating an optimal fit between observations and labeling outputs. Thanks to the rich supervision cues of the smoothness prior, our SmooSeg significantly outperforms STEGO in terms of pixel accuracy on three datasets: COCOStuff (+14.9\%), Cityscapes (+13.0\%), and Potsdam-3 (+5.7\%).
Mengcheng Lan, Xinjiang Wang, Yiping Ke, Jiaxing Xu, Litong Feng, Wayne Zhang 0001
NeurIPS5
2022 ViM: Out-Of-Distribution with Virtual-logit Matching
abstract
Most of the existing Out-Of-Distribution (OOD) detection algorithms depend on single input source: the feature, the logit, or the softmax probability. However, the immense diversity of the OOD examples makes such methods fragile. There are OOD samples that are easy to identify in the feature space while hard to distinguish in the logit space and vice versa. Motivated by this observation, we propose a novel OOD scoring method named Virtual-logit Matching (ViM), which combines the class-agnostic score from feature space and the In-Distribution (ID) class-dependent logits. Specifically, an additional logit representing the virtual OOD class is generated from the residual of the feature against the principal space, and then matched with the original logits by a constant scaling. The probability of this virtual logit after softmax is the indicator of OOD-ness. To facilitate the evaluation of large-scale OOD detection in academia, we create a new OOD dataset for ImageNet1K, which is human-annotated and is 8.8× the size of existing datasets. We conducted extensive experiments, including CNNs and vision transformers, to demonstrate the effectiveness of the proposed ViM score. In particular, using the BiT-S model, our method gets an average AUROC 90.91% on four difficult OOD benchmarks, which is 4% ahead of the best baseline. Code and dataset are available at https://github.com/haoqiwang/vim.
Zhizhong Li 0002, Litong Feng, Wayne Zhang 0001
CVPR3
2021 Semantically Coherent Out-of-Distribution Detection
abstract
Current out-of-distribution (OOD) detection benchmarks are commonly built by defining one dataset as in-distribution (ID) and all others as OOD. However, these benchmarks unfortunately introduce some unwanted and impractical goals, e.g., to perfectly distinguish CIFAR dogs from ImageNet dogs, even though they have the same semantics and negligible covariate shifts. These unrealistic goals will result in an extremely narrow range of model capabilities, greatly limiting their use in real applications. To overcome these drawbacks, we re-design the benchmarks and propose the semantically coherent out-of-distribution detection (SC-OOD). On the SC-OOD benchmarks, existing methods suffer from large performance degradation, suggesting that they are extremely sensitive to low-level discrepancy between data sources while ignoring their inherent semantics. To develop an effective SC-OOD detection approach, we leverage an external unlabeled set and design a concise framework featured by unsupervised dual grouping (UDG) for the joint modeling of ID and OOD data. The proposed UDG can not only enrich the semantic knowledge of the model by exploiting unlabeled data in an unsupervised manner, but also distinguish ID/OOD samples to enhance ID classification and OOD detection tasks simultaneously. Extensive experiments demonstrate that our approach achieves the state-of-the-art performance on SC-OOD benchmarks. Code and benchmarks are provided on our project page: https://jingkang50.github.io/projects/scood.
Litong Feng, Xiaopeng Yan, Huabin Zheng, Wayne Zhang 0001, Ziwei Liu 0002
ICCV3
2020 Scale-Equalizing Pyramid Convolution for Object Detection
abstract
Feature pyramid has been an efficient method to extract features at different scales. Development over this method mainly focuses on aggregating contextual information at different levels while seldom touching the inter-level correlation in the feature pyramid. Early computer vision methods extracted scale-invariant features by locating the feature extrema in both spatial and scale dimension. Inspired by this, a convolution across the pyramid level is proposed in this study, which is termed pyramid convolution and is a modified 3-D convolution. Stacked pyramid convolutions directly extract 3-D (scale and spatial) features and outperforms other meticulously designed feature fusion modules. Based on the viewpoint of 3-D convolution, an integrated batch normalization that collects statistics from the whole feature pyramid is naturally inserted after the pyramid convolution. Furthermore, we also show that the naive pyramid convolution, together with the design of RetinaNet head, actually best applies for extracting features from a Gaussian pyramid, whose properties can hardly be satisfied by a feature pyramid. In order to alleviate this discrepancy, we build a scale-equalizing pyramid convolution (SEPC) that aligns the shared pyramid convolution kernel only at high-level feature maps. Being computationally efficient and compatible with the head design of most single-stage object detectors, the SEPC module brings significant performance improvement (>4AP increase on MS-COCO2017 dataset) in state-of-the-art one-stage object detectors, and a light version of SEPC also has ~3.5AP gain with only around 7% inference time increase. The pyramid convolution also functions well as a stand-alone module in two-stage object detectors and is able to improve the performance by ~2AP. The source code can be found at https://github.com/jshilong/SEPC.
Xinjiang Wang, Zhuoran Yu, Litong Feng, Wayne Zhang 0001
CVPR4
2020 Webly Supervised Image Classification with Self-contained Confidence
Litong Feng, Xiaopeng Yan, Huabin Zheng, Ping Luo 0002, Wayne Zhang 0001
ECCV (8)2
2020 Webly Supervised Image Classification with Metadata: Automatic Noisy Label Correction via Visual-Semantic Graph
abstract
Webly supervised learning becomes attractive recently for its efficiency in data expansion without expensive human labeling. However, adopting search queries or hashtags as web labels of images for training brings massive noise that degrades the performance of DNNs. Especially, due to the semantic confusion of query words, the images retrieved by one query may contain tremendous images belonging to other concepts. For example, searching 'tiger cat' on Flickr will return a dominating number of tiger images rather than the cat images. These realistic noisy samples usually have clear visual semantic clusters in the visual space that mislead DNNs from learning accurate semantic labels. To correct real-world noisy labels, expensive human annotations seem indispensable. Fortunately, we find that metadata can provide extra knowledge to discover clean web labels in a labor-free fashion, making it feasible to automatically provide correct semantic guidance among the massive label-noisy web data. In this paper, we propose an automatic label corrector VSGraph-LC based on the visual-semantic graph. VSGraph-LC starts from anchor selection referring to the semantic similarity between metadata and correct label concepts, and then propagates correct labels from anchors on a visual graph using graph neural network (GNN). Experiments on realistic webly supervised learning datasets Webvision-1000 and NUS-81-Web show the effectiveness and robustness of VSGraph-LC. Moreover, VSGraph-LC reveals its advantage on the open-set validation set.
Litong Feng, Xiaopeng Yan, Huabin Zheng, Wayne Zhang 0001
ACM Multimedia3
2019 Learning Efficient Detector with Semi-supervised Adaptive Distillation
Shitao Tang, Litong Feng, Wenqi Shao, Zhanghui Kuang, Wayne Zhang 0001
BMVC2
2019 Gradual Network for Single Image De-raining
abstract
Most advances in single image de-raining meet a key challenge, which is removing rain streaks with different scales and shapes while preserving image details. Existing single image de-raining approaches treat rain-streak removal as a process of pixel-wise regression directly. However, they are lacking in mining the balance between over-de-raining (e.g. removing texture details in rain-free regions) and under-de-raining (e.g. leaving rain streaks). In this paper, we firstly propose a coarse-to-fine network called Gradual Network (GraNet) consisting of coarse stage and fine stage for delving into single image de-raining with different granularities. Specifically, to reveal coarse-grained rain-streak characteristics (e.g. long and thick rain streaks/raindrops), we propose a coarse stage by utilizing local-global spatial dependencies via a local-global sub-network composed of region-aware blocks. Taking the residual result (the coarse de-rained result) between the rainy image sample (i.e. the input data) and the output of coarse stage (i.e. the learnt rain mask) as input, the fine stage continues to de-rain by removing the fine-grained rain streaks (e.g. light rain streaks and water mist) to get a rain-free and well-reconstructed output image via a unified contextual merging sub-network with dense blocks and a merging block. Solid and comprehensive experiments on synthetic and real data demonstrate that our GraNet can significantly outperform the state-of-the-art methods by removing rain streaks with various densities, scales and shapes while keeping the image details of rain-free regions well-preserved.
Weijiang Yu, Zhe Huang 0006, Wayne Zhang 0001, Litong Feng
ACM Multimedia4
2019 Video copy detection by conducting fast searching of inverted files
Mengyang Liu, Lai-Man Po, Yasar Abbas Ur Rehman, Xuyuan Xu, Litong Feng
Multim. Tools Appl.6
2018 Fast Video Shot Transition Localization with Deep Structured Models
Shitao Tang, Litong Feng, Zhanghui Kuang, Wayne Zhang 0001
ACCV (1)2
2018 Extractive Video Summarizer with Memory Augmented Neural Networks
abstract
Online videos have been growing explosively in recent years. How to help human users efficiently browse videos becomes more and more important. Video summarization can automatically shorten a video through extracting key-shots from the raw video, which is helpful for digesting video data. State-of-the-art supervised video summarization algorithms directly learn from manually-created summaries to mimic the key-frame/key-shot selection criterion of humans. Humans usually create a summary after viewing and understanding the whole video, and the global attention mechanism capturing information from all video frames plays a key role in the summarization process. However, previous supervised approaches ignored the temporal relations or simply modeled local inter-dependency across frames. Motivated by this observation, we proposed a memory augmented extractive video summarizer, which utilizes an external memory to record visual information of the whole video with high capacity. With the external memory, the video summarizer simply predicts the importance score of a video shot based on the global understanding of the video frames. The proposed method outperforms previous state-of-the-art algorithms on the public SumMe and TVSum datasets. More importantly, we demonstrate that the global attention modeling has two advantages: good transferring ability across datasets and high robustness to noisy videos.
Litong Feng, Ziyin Li, Zhanghui Kuang, Wayne Zhang 0001
ACM Multimedia1
2018 Temporal Sequence Distillation: Towards Few-Frame Action Recognition in Videos
abstract
Video Analytics Software as a Service (VA SaaS) has been rapidly growing in recent years. VA SaaS is typically accessed by users using a lightweight client. Because the transmission bandwidth between the client and cloud is usually limited and expensive, it brings great benefits to design cloud video analysis algorithms with a limited data transmission requirement. Although considerable research has been devoted to video analysis, to our best knowledge, little of them has paid attention to the transmission bandwidth limitation in SaaS. As the first attempt in this direction, this work introduces a problem of few-frame action recognition, which aims at maintaining high recognition accuracy, when accessing only a few frames during both training and test. Unlike previous work that processed dense frames, we present Temporal Sequence Distillation (TSD), which distills a long video sequence into a very short one for transmission. By end-to-end training with 3D CNNs for video action recognition, TSD learns a compact and discriminative temporal and spatial representation of video frames. On Kinetics dataset, TSD+I3D typically requires only 50% of the number of frames compared to I3D, a state-of-the-art video action recognition algorithm, to achieve almost the same accuracies. The proposed TSD has three appealing advantages. Firstly, TSD has a lightweight architecture and can be deployed in the client, eg., mobile devices, to produce compressed representative frames to save transmission bandwidth. Secondly, TSD significantly reduces the computations to run video action recognition with compressed frames on the cloud, while maintaining high recognition accuracies. Thirdly, TSD can be plugged in as a preprocessing module of any existing 3D CNNs. Extensive experiments show the effectiveness and characteristics of TSD.
Zhaoyang Zhang 0004, Zhanghui Kuang, Ping Luo 0002, Litong Feng, Wayne Zhang 0001
ACM Multimedia4
2018 Block-based adaptive ROI for remote photoplethysmography
Lai-Man Po, Litong Feng, Xuyuan Xu, Chun-Ho Cheung, Kwok-Wai Cheung 0002
Multim. Tools Appl.2
2016 Face liveness detection and recognition using shearlet based feature descriptors
abstract
Face recognition is a widely used biometric technology due to its convenience but it is vulnerable to spoofing attacks made by non-real faces such as a photograph or video of valid user. Face liveness detection is a core technology to make sure that the input face is a live person. However, this is still very challenging using conventional liveness detection approaches of texture analysis and motion detection. The aim of this paper is to develop a multifunctional feature descriptor and an efficient framework which can be used to deal with both face liveness detection and recognition. In this framework, new feature descriptors are defined using a multiscale directional transform (shearlet transform). Then, stacked autoencoders and softmax classifier are concatenated to detect face liveness and identify person. We evaluated this approach using CASIA Face Anti-Spoofing Database and the results show that our approach performs better than state-of-the-art techniques following the provided evaluation protocols of this database, and is possible to significantly enhance the security of face recognition biometric system.
Lai-Man Po, Xuyuan Xu, Litong Feng
ICASSP4
2016 Integration of image quality and motion cues for face anti-spoofing: A neural network approach
Litong Feng, Lai-Man Po, Xuyuan Xu, Chun-Ho Cheung, Kwok-Wai Cheung 0002
J. Vis. Commun. Image Represent.1
2016 No-Reference Video Quality Assessment With 3D Shearlet Transform and Convolutional Neural Networks
abstract
In this paper, we propose an efficient general-purpose no-reference (NR) video quality assessment (VQA) framework that is based on 3D shearlet transform and convolutional neural network (CNN). Taking video blocks as input, simple and efficient primary spatiotemporal features are extracted by 3D shearlet transform, which are capable of capturing natural scene statistics properties. Then, CNN and logistic regression are concatenated to exaggerate the discriminative parts of the primary features and predict a perceptual quality score. The resulting algorithm, which we name shearlet- and CNN-based NR VQA (SACONVA), is tested on well-known VQA databases of Laboratory for Image & Video Engineering, Image & Video Processing Laboratory, and CSIQ. The testing results have demonstrated that SACONVA performs well in predicting video quality and is competitive with current state-of-the-art full-reference VQA methods and general-purpose NR-VQA algorithms. Besides, SACONVA is extended to classify different video distortion types in these three databases and achieves excellent classification accuracy. In addition, we also demonstrate that SACONVA can be directly applied in real applications such as blind video denoising.
Lai-Man Po, Chun-Ho Cheung, Xuyuan Xu, Litong Feng, Kwok-Wai Cheung 0002
IEEE Trans. Circuits Syst. Video Technol.5
2015 Dynamic ROI based on K-means for remote photoplethysmography
abstract
Remote imaging photoplethysmography (RIPPG) can achieve contactless human vital signs monitoring. Though the remote operation mode brings a great convenience for RIPPG applications, the RIPPG signal quality is limited by the remote nature. Improving the RIPPG signal quality becomes an essential task in the clinical application of RIPPG. Since the region of interest (ROI) of the RIPPG transforms from a point to an area, there is a new approach to improving the RIPPG signal quality through refining the ROI. In this paper, we propose a dynamic ROI for RIPPG, which can automatically select the skin regions corresponding to good quality RIPPG signals. First, a fixed ROI is divided into non-overlapped blocks. Then two features are proposed to perform no-reference quality assessment for RIPPG signals from different blocks. After that, K-means clustering operates in a two dimensional feature space. A dynamic ROI can be selected for a video segment based on the clustering result, updated every two seconds. Nineteen healthy subjects were enrolled to test the proposed ROI selection method on both the facial region and the palmar region. Experimental results of heart rate measurement show that the proposed dynamic ROI method for RIPPG can effectively improve the RIPPG signal quality, compared with the state-of-the-art ROI methods for RIPPG.
Litong Feng, Lai-Man Po, Xuyuan Xu, Chun-Ho Cheung, Kwok-Wai Cheung 0002
ICASSP1
2015 No-reference image quality assessment using shearlet transform and stacked autoencoders
abstract
In this work, we describe an efficient generalpurpose no-reference (NR) image quality assessment (IQA) algorithm that is based on a new multiscale directional transform (shearlet transform) with a strong ability to localize distributed discontinuities. The algorithm relies on utilizing the sum of subband coefficient amplitudes (SSCA) as primary features to describe the behavior of natural images and distorted images. Then, stacked autoencoders are applied to exaggerate the discriminative parts of the primary features. Finally, by translating the NR-IQA problem into classification problem, the differences of evolved features are identified by softmax classifier. The resulting algorithm, which we name SESANIA (ShEarlet and Stacked Autoencoders based No-reference Image quality Assessment), is tested on several databases (LIVE, Multiply Distorted LIVE and TID2008) and shown to be suitable to many common distortions, consistent with subjective assessment and comparable to full-reference IQA methods and state-of-the-art general purpose NR-IQA algorithms.
Lai-Man Po, Xuyuan Xu, Litong Feng, Chun-Ho Cheung, Kwok-Wai Cheung 0002
ISCAS4
2015 Frame adaptive ROI for photoplethysmography signal extraction from fingertip video captured by smartphone
abstract
Photoplethysmography (PPG) has been widely used in clinical applications for monitoring vital signs especially heart rate by pulse oximeter. Recent researches have demonstrated the possibility of using fingertip video based PPG approach to estimate heart rate by smartphones. However, due to the variation of camera sensor characteristics in difference smartphones, the conventional fixed region-of-interest (ROI) for PPG signal extraction technique is not reliable. In this paper, a novel frame adaptive ROI method is proposed to detour the color saturation or cut-off distortion in the fingertip video capturing process for improving the reliability due to variation and limited dynamic range of the camera sensors in different smartphone models. Experimental results demonstrate that the proposed method can produce good pulsatile waveform and achieve high heart rate estimation accuracy using different smartphone models as compared with a FDA (U.S. Food and Drug Administration) approved commercial pulse oximeter.
Lai-Man Po, Xuyuan Xu, Litong Feng, Kwok-Wai Cheung 0002, Chun-Ho Cheung
ISCAS3
2015 No-reference image quality assessment with shearlet transform and deep neural networks
Lai-Man Po, Xuyuan Xu, Litong Feng, Chun-Ho Cheung, Kwok-Wai Cheung 0002
Neurocomputing4
2015 Motion-Resistant Remote Imaging Photoplethysmography Based on the Optical Properties of Skin
abstract
Remote imaging photoplethysmography (RIPPG) can achieve contactless monitoring of human vital signs. However, the robustness to a subject's motion is a challenging problem for RIPPG, especially in facial video-based RIPPG. The RIPPG signal originates from the radiant intensity variation of human skin with pulses of blood and motions can modulate the radiant intensity of the skin. Based on the optical properties of human skin, we build an optical RIPPG signal model in which the origins of the RIPPG signal and motion artifacts can be clearly described. The region of interest (ROI) of the skin is regarded as a Lambertian radiator and the effect of ROI tracking is analyzed from the perspective of radiometry. By considering a digital color camera as a simple spectrometer, we propose an adaptive color difference operation between the green and red channels to reduce motion artifacts. Based on the spectral characteristics of photoplethysmography signals, we propose an adaptive bandpass filter to remove residual motion artifacts of RIPPG. We also combine ROI selection on the subject's cheeks with speeded-up robust features points tracking to improve the RIPPG signal quality. Experimental results show that the proposed RIPPG can obtain greatly improved performance in accessing heart rates in moving subjects, compared with the state-of-the-art facial video-based RIPPG methods.
Litong Feng, Lai-Man Po, Xuyuan Xu, Ruiyi Ma
IEEE Trans. Circuits Syst. Video Technol.1
2014 Adaptive block truncation filter for MVC depth image enhancement
abstract
In Multiview Video plus Depth (MVD) format, virtual views are generated from decoded texture videos with decoded depth images through Depth Image based Rendering (DIBR). 3DV-ATM is a reference model for H.264/AVC based Multiview Video Coding (MVC) and aims at achieving high coding efficiency for 3D video in MVD format. Depth images are first downsampled then coded by 3DV-ATM. However, sharp object boundary characteristic of depth images does not well match with the transform coding of 3DV-ATM. Depth boundaries are often blurred with ringing artifacts in the decoded depth images that result in noticeable artifacts in synthesized views. This paper presents a low complexity adaptive block truncation filter to recover the sharp object boundaries of depth images using adaptive block repositioning and expansion for increasing the depth values refinement accuracy. This new approach is very efficient and can avoid false depth boundary refinement when block boundaries lie around the depth edge regions and ensure sufficient information within the processing block for depth layers classification. Experimental results show that sharp depth edges can be recovered using the proposed filter and boundary artifacts in the synthesized views can be removed. The proposed method can provide improvement up to 3.25dB in the depth map enhancement and bitrate reduction of 3.06% in the synthesized views.
Xuyuan Xu, Lai-Man Po, Chun-Ho Cheung, Litong Feng, Kwok-Wai Cheung 0002, Chi-Wang Ting, Ka-Ho Ng
ICASSP4
2014 No-reference image quality assessment using statistical characterization in the shearlet domain
Lai-Man Po, Xuyuan Xu, Litong Feng
Signal Process. Image Commun.4
2014 Adaptive depth truncation filter for MVC based compressed depth image
Xuyuan Xu, Lai-Man Po, Chun-Ho Cheung, Kwok-Wai Cheung 0002, Litong Feng, Chi-Wang Ting, Ka-Ho Ng
Signal Process. Image Commun.5
2013 Watershed based depth map misalignment correction and foreground biased dilation for DIBR view synthesis
abstract
The quality of the synthesized views by Depth Image Based Rendering (DIBR) highly depends on the accuracy of the depth map, especially the alignment of object boundaries of texture image. In practice, the misalignment of sharp depth map edges is the major cause of the annoying artifacts at the disoccluded regions of the synthesized views. In this paper, a new depth map preprocessing method using Watershed misalignment correction and dilation filter is proposed to align the foreground depth edges to cover the whole transitional color edge regions. This approach can handle the sharp depth map edges lying inside or outside the object boundaries in 2D sense. The quality of the disoccluded regions of the synthesized views can be significantly improved. Experimental results show that the proposed method achieves superior performance for view synthesis by DIBR especially for generating large baseline virtual views.
Xuyuan Xu, Lai-Man Po, Kwok-Wai Cheung 0002, Litong Feng, Chun-Ho Cheung
ICIP4
2013 An adaptive background biased depth map hole-filling method for Kinect
abstract
The launch of Kinect provides a convenient way to access the depth information in real time. However the depth map quality still needs to be enhanced for 3D visual applications. In this paper, an adaptive background biased depth map hole-filling method is proposed. First, depth holes caused by abnormal reflection are filled by color similarity in-painting, and a soft decision for color similarity checking is performed by the use of probabilities in random walks color segmentation. Afterwards it is assumed that the lost information in the rest of depth holes belongs to the background. The background depth information is extracted by automatic thresholding in the neighborhood of each hole. Depth holes are in-painted with the background information in their local neighborhood. Combination of color similarity in-painting and background biased in-painting is able to perform depth map hole-filling adaptively for different kinds of depth holes for Kinect. The hole-filling results and virtual view synthesis results show that the Kinect depth map quality can be improved significantly by the proposed method.
Litong Feng, Lai-Man Po, Xuyuan Xu, Ka-Ho Ng, Chun-Ho Cheung, Kwok-Wai Cheung 0002
IECON1
2013 Depth-aided exemplar-based hole filling for DIBR view synthesis
abstract
Quality of synthesized view by Depth-Image-Based Rendering (DIBR) highly depends on hole filling, especially for synthesized view with large disocclusion. Many hole filling methods are proposed to improve the synthesized view quality and inpainting is the most popular approach to recover the disocclusions. However, the conventional inpainting either makes the hole regions blurred via diffusion or propagates the foreground information to the disoclusion regions. Annoying artifacts are created in the synthesized virtual views. This paper proposes a depth-aided exemplar-based inpainting method for recovering large disoclusion. It consists of two processes, warped depth map filling and warped color image filling. Since depth map can be considered as a grey-scale image without texture, it is much easier to be filled. Disoccluded regions of color image are predicted based on its associated filled depth map information. Regions with texture lying around the background have higher priority to be filled than other regions and disoccluded regions are filled by propagating the background texture through the exemplar-based inpainting. Thus artifacts created by diffusion or using foreground information for prediction can be eliminated. Experimental results show texture can be recovered in large disocclusions and the proposed method has better visual quality compared to existing methods.
Xuyuan Xu, Lai-Man Po, Chun-Ho Cheung, Litong Feng, Ka-Ho Ng, Kwok-Wai Cheung 0002
ISCAS4
2013 Depth map misalignment correction and dilation for DIBR view synthesis
Xuyuan Xu, Lai-Man Po, Ka-Ho Ng, Litong Feng, Kwok-Wai Cheung 0002, Chun-Ho Cheung, Chi-Wang Ting
Signal Process. Image Commun.4