EDBT 2026 Demo / reviewers in the wild / expert
Jingang Shi
dblp:84/1047
· DBLP profile ↗
49ranked-venue papers
10as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 9 first-author · 17 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatial-frequency hybrid Mamba with structure-aware scanning for face super-resolution
Jianan Cao, Shuyang Chu, Jingang Shi |
Neurocomputing | 5 |
| 2026 | SaTPhys: Sandglass Transformer for Efficient Video-Based Remote Physiological MeasurementabstractThe effectiveness of Transformers has been proven in video-based remote photoplethysmography (rPPG) measurement. However, the inherently high computational cost of the Transformer poses limitations of these methods on resource-constrained devices. This paper presents a novelaggregating-and-distributingsandglass-like framework, called SaTPhys, for efficient Transformer-based rPPG measurement. Our SaTPhys initiates by clustering physiological tokens that possess redundant spatio-temporal information and concludes with the recovery of full-length tokens. This process leads to fewer cluster centers passing through the intermediate Transformer block, consequently enhancing the model’s efficiency. To accomplish this effectively, we manually design a physiological context aggregation (PCA) module to generate representative cluster centers, thereby eliminating spatio-temporal redundancy. Subsequently, we employ an inter-cluster Transformer (ICT) to efficiently interact with these cluster centers on a global scale. Finally, we introduce a physiological context distributing (PCD) module to restore full-length tokens and distribute the aggregated global information. Furthermore, we develop a frequency modulator (FM) block to enhance the frequency information, thereby improving the periodic fidelity of the estimated rPPG signal. Comprehensive experiments across multiple benchmark datasets have shown that the proposed method achieves superior performance with minimal computational cost. For example, compared to the SOTA rPPG method, our method achieves a lower MAE on the VIPL-HR dataset (3.96 bpm vs. 4.32 bpm) with a significantly lower computational cost (3.94 GMACs vs. 12.9 GMACs). The code is available at https://github.com/xjtucsy/SaTPhys. Shuyang Chu, Jingang Shi, Mengyao Yuan, Xuqi Li, Zhengdong Jiang, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | HemNet: Hemoglobin-Assistant Network for Video-Based Remote Photoplethysmography MeasurementabstractTraditional skin-contact physical sensors typically detect changes of blood volume to predict the periodicity of heartbeat by analyzing the absorption spectra of hemoglobin. However, the contact on human skin may cause uncomfortable feeling and induce difficulty for long-term monitoring. Recently, video-based remote photoplethysmography (rPPG) estimation approaches analyze the periodic facial color changes for matching cardiac cycle in a contactless manner. Nevertheless, the inherent relationship between the changes of facial color and blood volume is not fully exploited. Besides the influence of blood volume (i.e., hemoglobin), there are also other factors such as lighting and reflection that cause the change on facial color. We exploit the physical principles that cause skin color variations to separate the hemoglobin factor driven by blood volume. Based on the physical prior of the reflection of human skin, we introduce an rPPG estimation network assisted by decoupled hemoglobin sequence, named HemNet, which first explicitly leverages hemoglobin to assist rPPG signal estimation. To obtain meaningful hemoglobin from facial video, we design a human skin color disentangler that decouples the facial color variations into four significant features, i.e., hemoglobin, melanin, shading, and specular. We then present a multi-modality rPPG estimator that utilizes cross-covariance attention to extract fused feature from hemoglobin and RGB video inputs. Finally, an adaptive negative Pearson loss is proposed to effectively address phase misalignment between the blood volume in the finger and facial region during the training phase. We evaluate our HemNet on four widely used public benchmark datasets. The superiority of our method is demonstrated in both intra-dataset and cross-dataset test settings. The code is available at https://github.com/jingang-cv/hemnet. Ruize Wu, Jingang Shi, Xin Liu 0012, LinLin Shen, Yihong Gong, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | ResoPhys: Unsupervised Plug-and-Play Remote Physiological Measurement via Facial Videos of Arbitrary ResolutionabstractRemote photoplethysmography (rPPG) is a non-contact method that detects blood volume changes in facial tissues from video. The non-invasiveness of rPPG makes it promising for applications in remote health monitoring and telemedicine. However, its real-world application is hindered by a fundamental challenge. Existing models are typically designed for high-resolution, fixed-size inputs, making them ill-suited for the arbitrary-resolution videos commonly encountered in practical scenarios due to dynamic camera-to-subject distances. To address this challenge, we propose ResoPhys, an unsupervised plug-and-play rPPG measurement method designed for facial videos of arbitrary resolution. This method first generates video pairs via random scaling and then employs specialized modules for arbitrary-resolution feature extraction and upsampling to analyze the resulting multi-scale features. The framework is optimized via an unsupervised contrastive learning approach using our proposed multi-resolution contrastive loss. To validate its performance across a spectrum of resolutions, we evaluated ResoPhys on several public datasets. The results demonstrate the superiority of our method over previous unsupervised approaches, exhibiting particular strength in challenging low-resolution scenarios, which underscores its robustness to resolution changes. Crucially, ResoPhys acts as a universal front-end that decouples resolution handling from signal extraction, empowering existing rPPG networks for effective deployment in arbitrary-resolution conditions. Zhongtian He, Shuyang Chu, Xuqi Li, Zhengdong Jiang, Guoying Zhao 0001, Jingang Shi |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | To Remember, To Adapt, To Preempt: A Stable Continual Test-Time Adaptation Framework for Remote Physiological Measurement in Dynamic Domain ShiftsabstractRemote photoplethysmography (rPPG) aims to extract non-contact physiological signals from facial videos and has shown great potential. However, existing rPPG approaches struggle to bridge the gap between source and target domains. Recent test-time adaptation (TTA) solutions typically optimize rPPG model for the incoming test videos using self-training loss under an unrealistic assumption that the target domain remains stationary. However, time-varying factors like weather and lighting in dynamic environments often cause continual domain shifts. The erroneous gradients accumulation from these shifts may corrupt the model's key parameters for physiological information, leading to catastrophic forgetting. Therefore, We propose a physiology-related parameters freezing strategy to retain such knowledge. It isolates physiology-related and domain-related parameters by assessing the model's uncertainty to current domain and freezes the physiology-related parameters during adaptation to prevent catastrophic forgetting. Moreover, the dynamic domain shifts with various non-physiological characteristics may lead to conflicting optimization objectives during TTA, which is manifested as the over-adapted model losing its adaptability to future domains. To fix over-adaptation, we propose a preemptive gradient modification strategy. It preemptively adapts to future domains and uses the acquired gradients to modify current adaptation, thereby preserving the model's adaptability. In summary, we propose a stable continual test-time adaptation (CTTA) framework for rPPG measurement, called PhysRAP, which Remembers the past, Adapts to the present, and Preempts the future. Extensive experiments show its state-of-the-art performance, especially in domain shifts. The code is available at https://github.com/xjtucsy/PhysRAP. Shuyang Chu, Jingang Shi, Xu Cheng 0003, Haoyu Chen 0001, Xin Liu 0012, Guoying Zhao 0001 |
ACM Multimedia | 2 |
| 2025 | FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
Zhuozhao Hu, Kaishen Yuan, Xin Liu 0012, Zitong Yu, Yuan Zong, Jingang Shi, Huanjing Yue, Jing-Yu Yang 0002 |
ACM Multimedia | 6 |
| 2025 | Multi-Scale Promoted Self-Adjusting Correlation Learning for Facial Action Unit DetectionabstractFacial Action Unit (AU) detection is a crucial task in affective computing and social robotics as it helps to identify emotions expressed through facial expressions. Anatomically, there are innumerable correlations between AUs, which contain rich information and are vital for AU detection. Previous methods used fixed AU correlations based on expert experience or statistical rules on specific benchmarks, but it is challenging to comprehensively reflect complex correlations between AUs via hand-crafted settings. There are alternative methods that employ a fully connected graph to learn these dependencies exhaustively. However, these approaches can result in a computational explosion and high dependency with a large dataset. To address these challenges, this paper proposes a novel self-adjusting AU-correlation learning (SACL) method with less computation for AU detection. This method adaptively learns and updates AU correlation graphs by efficiently leveraging the characteristics of different levels of AU motion and emotion representation information extracted in different stages of the network. Moreover, this paper explores the role of multi-scale learning in correlation information extraction, and design a simple yet effective multi-scale feature learning (MSFL) method to promote better performance in AU detection. By integrating AU correlation information with multi-scale features, the proposed method obtains a more robust feature representation for the final AU detection. Extensive experiments show that the proposed method outperforms the state-of-the-art methods on widely used AU detection benchmark datasets, with only 28.7% and 12.0% of the parameters and FLOPs of the best method, respectively. Xin Liu 0012, Kaishen Yuan, Xuesong Niu, Jingang Shi, Zitong Yu, Huanjing Yue, Jing-Yu Yang 0002 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Learning to Rank Onset-Occurring-Offset Representations for Micro-Expression RecognitionabstractThis paper focuses on the research of micro-expression recognition (MER) and proposes a flexible and reliable deep learning method called learning to rank onset-occurring-offset representations (LTR3O). The LTR3O method introduces a dynamic and reduced-size sequence structure known as 3O, which consists of onset, occurring, and offset frames, for representing micro-expressions (MEs). This structure facilitates the subsequent learning of ME-discriminative features. A noteworthy advantage of the 3O structure is its flexibility, as the occurring frame is randomly extracted from the original ME sequence without the need for accurate frame spotting methods. Based on the 3O structures, LTR3O generates multiple 3O representation candidates for each ME sample and incorporates well-designed modules based on learning to rank (LTR) to measure and calibrate their emotional expressiveness. This calibration process implicitly enhances the visibility of MEs by amplifying the originally narrow emotional expressiveness gap among ME frames caused by their low-intensity characteristics, thereby facilitating the reliable learning of more discriminative features for MER. Extensive experiments were conducted to evaluate the performance of LTR3O using four widely-used ME databases: CASME II, SMIC, SAMM, and MEVIEW. The experimental results demonstrate the effectiveness and superior performance of LTR3O, particularly in terms of its flexibility and reliability, when compared to recent state-of-the-art MER methods. Yuan Zong, Jingang Shi, Cheng Lu 0005, Hongli Chang, Wenming Zheng |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Towards Domain-Specific Cross-Corpus Speech Emotion Recognition ApproachabstractCross-corpus speech emotion recognition (SER) poses a challenge due to feature distribution mismatch between the training and testing speech samples, potentially degrading the performance of established SER methods. In this article, we tackle this challenge by proposing a novel transfer subspace learning method called acoustic knowledge-guided transfer linear regression (AKTLR). Unlike existing approaches, which often overlook domain-specific knowledge related to SER and simply treat cross-corpus SER as a generic transfer learning task, our AKTLR method is built upon a well-designed acoustic knowledge-guided dual sparsity constraint mechanism. This mechanism emphasizes the potential of minimalistic acoustic parameter feature sets to alleviate classifier over-adaptation, which is empirically validated acoustic knowledge in SER, enabling superior generalization in cross-corpus SER tasks compared to using large feature sets. Through this mechanism, we extend a simple transfer linear regression model to AKTLR. This extension harnesses its full capability to seek emotion-discriminative and corpus-invariant features from established acoustic parameter feature sets used for describing speech signals across two scales: contributive acoustic parameter groups and constituent elements within each contributive group. We evaluate our method through extensive cross-corpus SER experiments on three widely used speech emotion corpora: EmoDB, eNTERFACE, and CASIA. The proposed AKTLR achieves an average UAR of 42.12% across six tasks using the eGeMAPS feature set, outperforming many recent state-of-the-art transfer subspace learning and deep transfer learning methods. This demonstrates the effectiveness and superior performance of our approach. Furthermore, our work provides experimental evidence supporting the feasibility and superiority of incorporating domain-specific knowledge into the transfer learning model to address cross-corpus SER tasks. Yan Zhao 0037, Yuan Zong, Hailun Lian, Cheng Lu 0005, Jingang Shi, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Understanding the Dimensional Need of Noncontrastive LearningabstractNoncontrastive self-supervised learning methods offer an effective alternative to contrastive approaches by avoiding the need for negative samples to avoid representation collapse. Noncontrastive learning methods explicitly or implicitly optimize the representation space, yet they often require large representation dimensions, leading to dimensional inefficiency. To provide negative samples, contrastive learning methods often require large batch sizes, thus regarded as sample inefficient, while noncontrastive learning methods require large representation dimensions, thus regarded as dimension inefficient. Although we have some understanding of the noncontrastive learning method, theoretical analysis of such phenomenon still remains largely unexplored. We present a theoretical analysis of the dimensional need for noncontrastive learning. We investigate the transfer between upstream representation learning and downstream tasks' performance, demonstrating how noncontrastive methods implicitly increase interclass distances within the representation space and how the distance affects the model performance of evaluation performance. We prove that the performance of noncontrastive methods is affected by the output dimension and the number of latent classes, and illustrate why performance degrades significantly when the output dimension is substantially smaller than the number of latent classes. We demonstrate our findings through experiments on image classification experiments, and enrich the verification in audio, graph and text modalities. We also perform empirical evaluation for image models on extensive detection and segmentation tasks beyond classification that show satisfactory correspondence to our theorem. Zhexiao Cao, Lei Huang 0015, Tian Wang 0002, Yinquan Wang, Jingang Shi, Aichun Zhu, Tianyun Shi, Hichem Snoussi |
IEEE Trans. Cybern. | 5 |
| 2025 | CodePhys: Robust Video-Based Remote Physiological Measurement Through Latent Codebook QueryingabstractRemote photoplethysmography (rPPG) aims to measure non-contact physiological signals from facial videos, which has shown great potential in many applications. Most existing methods directly extract video-based rPPG features by designing neural networks for heart rate estimation. Although they can achieve acceptable results, the recovery of rPPG signal faces intractable challenges when interference from real-world scenarios takes place on facial video. Specifically, facial videos are inevitably affected by non-physiological factors (e.g., camera device noise, defocus, and motion blur), leading to the distortion of extracted rPPG signals. Recent rPPG extraction methods are easily affected by interference and degradation, resulting in noisy rPPG signals. In this paper, we propose a novel method named CodePhys, which innovatively treats rPPG measurement as a code query task in a noise-free proxy space (i.e., codebook) constructed by ground-truth PPG signals. We consider noisy rPPG features as queries and generate high-fidelity rPPG features by matching them with noise-free PPG features from the codebook. Our approach also incorporates a spatial-aware encoder network with a spatial attention mechanism to highlight physiologically active areas and uses a distillation loss to reduce the influence of non-periodic visual interference. Experimental results on four benchmark datasets demonstrate that CodePhys outperforms state-of-the-art methods in both intra-dataset and cross-dataset settings. Shuyang Chu, Menghan Xia, Mengyao Yuan, Xin Liu 0012, Tapio Seppänen, Guoying Zhao 0001, Jingang Shi |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | You Can Wash Hands Better: Accurate Daily Handwashing Assessment With a SmartwatchabstractHand hygiene is among the most effective daily practices for preventing infectious diseases such as influenza, malaria, and skin infections. While professional guidelines emphasize proper handwashing to reduce the risk of viral infections, surveys reveal that adherence to these recommendations remains low. To address this gap, we propose UWash, a wearable solution leveraging smartwatches to evaluate handwashing procedures, aiming to raise awareness and cultivate high-quality handwashing habits. We frame the task of handwashing assessment as an action segmentation problem, similar to those in computer vision, and introduce a simple yet efficient two-stream UNet-like network to achieve this goal. Experiments involving 51 subjects demonstrate that UWash achieves 92.27% accuracy in handwashing gesture recognition, an error of$\lt $0.5 seconds in onset/offset detection, and an error of$\lt $5 points in gesture scoring under user-dependent settings. The system also performs robustly in user-independent and user-independent-location-independent evaluations. Remarkably, UWash maintains high performance in real-world tests, including evaluations with 10 random passersby at a hospital 9 months later and 10 passersby in an in-the-wild test conducted 2 years later. UWash is the first system to score handwashing quality based on gesture sequences, offering actionable guidance for improving daily hand hygiene. The code and dataset are publicly available athttps://github.com/aiotgroup/UWash. Fei Wang 0037, Xilei Wu, Xin Wang 0195, Han Ding 0002, Jingang Shi, Jinsong Han |
IEEE Trans. Mob. Comput. | 7 |
| 2025 | Dual-Path Imbalanced Feature Compensation Network for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) presents significant challenges on account of the substantial cross-modality gap and intra-class variations. Most existing methods primarily concentrate on aligning cross-modality at the feature or image levels and training with an equal number of samples from different modalities. However, in the real world, there exists an issue of modality imbalance between visible and infrared data. Besides, imbalanced samples between train and test impact the robustness and generalization of the VI-ReID. To alleviate this problem, we propose a dual-path imbalanced feature compensation network (DICNet) for VI-ReID, which provides equal opportunities for each modality to learn inconsistent information from different identities of others, enhancing identity discrimination performance and generalization. First, a modality consistency perception (MCP) module is designed to assist the backbone focus on spatial and channel information, extracting diverse and salient features to enhance feature representation. Second, we propose a cross-modality features re-assignment strategy to simulate modality imbalance by grouping and re-organizing the cross-modality features. Third, we perform bidirectional heterogeneous cooperative compensation with cross-modality imbalanced feature interaction modules (CIFIMs), allowing our network to explore the identity-aware patterns from imbalanced features of multiple groups for cross-modality interaction and fusion. Further, we design a feature re-construction difference loss to reduce cross-modality discrepancy and enrich feature diversity within each modality. Extensive experiments on three mainstream datasets show the superiority of the DICNet. Additionally, competitive results in corrupted scenarios verify its generalization and robustness. Xu Cheng 0003, Hao Yu 0015, Jingang Shi, Zitong Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | DiffFAS: Face Anti-spoofing via Generative Diffusion Models
Xinxu Ge, Xin Liu 0012, Zitong Yu, Jingang Shi, Chun Qi, Heikki Kälviäinen |
ECCV (54) | 4 |
| 2024 | Target-Specific Domain Adaptation via Geometry-Correlation Prediction for Point Cloud
Junqiao Li, Leyan Zhu, Tian Wang 0002, Jingang Shi, Hichem Snoussi |
PRCV (4) | 5 |
| 2024 | U-Shape Networks Are Unified Backbones for Human Action Understanding From Wi-Fi SignalsabstractWi-Fi is well-known in communication and deployment for indoor localization. Recently, Wi-Fi has been exploited for human action understanding, e.g., elder fall detection, smoking detection, and hand-gesture recognition. Many researchers have made great efforts to associate their own expertise on Wi-Fi signals with deep networks like convolutional neural networks, recurrent neural networks, and Transformers for representation learning, and demonstrate that expertise can promote action understanding accuracy. However, expert knowledge always relies on the existing personal understanding and assumptions of Wi-Fi signals, which limits the scalability of associated models and may introduce subjective bias to the models. Besides, requiring expertise raises an inescapable barrier and cost to the model design. Recent years have witnessed the great value of backbone networks, such as ResNet, in advancing the research progress in computer vision. We believe a backbone network for Wi-Fi signals will also play a crucial role. In this article, instead of proposing novel algorithms with expertise, we present that simple U-shape deep networks, such as FCN, U-Net, and U-Net++, are efficient and unified backbones for Wi-Fi-based human action understanding tasks, i.e., action recognition, action detection, and action segmentation. Results on three public data sets, Wi-Fi activity recognition, action recognition and indoor localization, and human-to-human interaction, show that all these U-shape deep networks have superior or competitive performance compared with original papers as well as state-of-the-art approaches, e.g., >97% recognition accuracy,90% segmentation accuracy. We envision this work breaks the barrier of network design and facilitates human action understanding from Wi-Fi signals. Fei Wang 0037, Yiao Gao, Han Ding 0002, Jingang Shi, Jinsong Han |
IEEE Internet Things J. | 5 |
| 2024 | TVRPCA+: Low-rank and sparse decomposition based on spectral norm and structural sparsity-inducing norm
Ruibo Fan, Mingli Jing, Jingang Shi |
Signal Process. | 3 |
| 2024 | MaskFusionNet: A Dual-Stream Fusion Model With Masked Pre-Training Mechanism for rPPG MeasurementabstractRemote photoplethysmography (rPPG) has considerable significance in areas such as disease diagnosis and emotion analysis. Recent rPPG models have demonstrated excellent performance due to their powerful heart rate information extraction capabilities. However, these models often focus on limited regions of interest (ROI) on facial image, which makes them sensitive to interference. If the ROI is affected by muscle movement, lighting variation and noise, the model’s performance would degrade significantly. To address this limitation, we propose a two-stage model called MaskFusionNet. The model includes two stages: 1) During the pre-training stage, the mask-reconstruction mechanism drives MaskFusionNet to learn rPPG information from various facial regions by applying a tube masking strategy. This enhances the model’s ability to resist interference. Based on the periodicity and continuity of the heart rate signal, we also design a novel spatio-temporal reconstruction loss function that focuses on the data’s spatial features and temporal continuity. 2) In the fine-tuning stage, we propose the Multi-Scale Fusion Block (MFB) to combine multi-scale features from the dual-stream network. It allows the model to detect subtle heart rate variations in adjacent frames while minimizing the impact of interference by extracting features within longer segments. The transformer-based MaskFusionNet can extract multi-scale fused heart rate features from a wide range of skin regions while preserving the modeling capability of long-range sequence information. To validate its effectiveness, we extensively evaluate our model on three benchmark datasets (VIPL-HR, COHFACE, and PURE), demonstrating its superior performance in both intra-dataset and cross-dataset testing scenarios. Yizhu Zhang, Jingang Shi, Yuan Zong, Wenming Zheng, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Benchmarking Joint Face Spoofing and Forgery Detection With Visual and Physiological CuesabstractFace anti-spoofing (FAS) and face forgery detection play vital roles in securing face biometric systems from presentation attacks (PAs) and vicious digital manipulation (e.g., deepfakes). Despite satisfactory performance upon large-scale data and powerful deep models, recent advances in face spoofing and forgery detection approaches usually focus on 1) unimodal visual appearance or physiological (i.e., remote photoplethysmography (rPPG)) cues; and 2) separated feature representation for FAS or face forgery detection. On one side, unimodal appearance and rPPG features are respectively vulnerable to high-fidelity face 3D mask and video replay attacks, inspiring us to design reliable multi-modal fusion mechanisms for generalized FAS. On the other side, there are rich common features across FAS and face forgery detection tasks (e.g., periodic rPPG rhythms and vanilla appearance for bonafides), providing solid evidence to design a joint FAS and face forgery detection system in a multi-task learning fashion. In this paper, we establish the first joint face spoofing and forgery detection benchmark using both visual appearance and physiological rPPG cues. To enhance the rPPG periodicity discrimination, we design a two-branch physiological network using both facial spatio-temporal rPPG signal map and its continuous wavelet transformed counterpart as inputs. To mitigate the modality bias and improve the fusion efficacy, we conduct a weighted batch and layer normalization for both appearance and rPPG features before multi-modal fusion. We also investigate prevalent deep models, feature fusion strategies and multi-task learning configurations for joint face spoofing and forgery detection. We find that the generalization capacities of both unimodal (appearance or rPPG) and multi-modal (appearance+rPPG) models can be obviously improved via joint training on these two tasks. We hope this new benchmark will facilitate the future research of both FAS and deepfake detection communities. The codes will be released athttps://github.com/ZitongYu/Benchmarking. Zitong Yu, Rizhao Cai, Zhi Li 0054, Wenhan Yang, Jingang Shi, Alex Chichung Kot |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | ST-Phys: Unsupervised Spatio-Temporal Contrastive Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) is a non-contact method that employs facial videos for measuring physiological parameters. Existing rPPG methods have achieved remarkable performance. However, the success mainly profits from supervised learning over massive labeled data. On the other hand, existing unsupervised rPPG methods fail to fully utilize spatio-temporal features and encounter challenges in low-light or noise environments. To address these problems, we propose an unsupervised contrast learning approach, ST-Phys. We incorporate a low-light enhancement module, a temporal dilated module, and a spatial enhanced module to better deal with long-term dependencies under the random low-light conditions. In addition, we design a circular margin loss, wherein rPPG signals originating from identical videos are attracted, while those from distinct videos are repelled. Our method is assessed on six openly accessible datasets, including RGB and NIR videos. Extensive experiments reveal the superior performance of our proposed ST-Phys over state-of-the-art unsupervised rPPG methods. Moreover, it offers advantages in parameter reduction and noise robustness. Mingyue Cao, Xu Cheng 0003, Hao Yu 0015, Jingang Shi |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | Exploiting Multi-Scale Parallel Self-Attention and Local Variation via Dual-Branch Transformer-CNN Structure for Face Super-ResolutionabstractRecently, deep learning technique has been widely employed to deal with face super-resolution (FSR) problem. It aims to predict the nonlinear relationship between the low-resolution (LR) face images and corresponding high-resolution (HR) ones, which could recover the high-frequency details from the LR degraded textures. However, either CNN-based or Transformer-based approaches mostly enhance the details by exploiting the relationship of local pixels or patches on LR features, the nonlocal features are not fully taken into account for producing high-frequency textures. To improve the above problem, we design a novel dual-branch module which consists of Transformer and CNN respectively. The Transformer branch extracts multiple scale feature embeddings and explores local and nonlocal self-attention simultaneously. Thus, the parallel self-attention mechanism has superior capabilities to capture the local and nonlocal dependencies on face image in the face reconstruction. Furthermore, the traditional CNNs usually extract features by combining pixels in a local convolutional kernel, it may be not effective to recover lost high-frequency details since the variations of local pixels are not well measured, which is important in recovering vivid edges and contours. To this end, we propose the local variation based attention block on the CNN branch, which could enhance the capabilities by directly extracting features from the variation of neighboring pixels. Finally, the Transformer-branch and CNN-branch are combined together by the modulation block to fuse both nonlocal and local advantages from two branches. Experimental results demonstrate the effectiveness of the proposed method when compared with state-of-the-art approaches. Jingang Shi, Yusi Wang, Zitong Yu, Guanxin Li, Xiaopeng Hong, Fei Wang 0037, Yihong Gong |
IEEE Trans. Multim. | 1 |
| 2024 | Brain Cognition-Inspired Dual-Pathway CNN Architecture for Image ClassificationabstractInspired by the global-local information processing mechanism in the human visual system, we propose a novel convolutional neural network (CNN) architecture named cognition-inspired network (CogNet) that consists of a global pathway, a local pathway, and a top-down modulator. We first use a common CNN block to form the local pathway that aims to extract fine local features of the input image. Then, we use a transformer encoder to form the global pathway to capture global structural and contextual information among local parts in the input image. Finally, we construct the learnable top-down modulator where fine local features of the local pathway are modulated by global representations of the global pathway. For ease of use, we encapsulate the dual-pathway computation and modulation process into a building block, called the global-local block (GL block), and a CogNet of any depth can be constructed by stacking a necessary number of GL blocks one after another. Extensive experimental evaluations have revealed that the proposed CogNets have achieved the state-of-the-art performance accuracies on all the six benchmark datasets and are very effective for overcoming the "texture bias" and the "semantic confusion" problems faced by many CNN models. Songlin Dong, Yihong Gong, Jingang Shi, Miao Shang, Xing Wei 0001, Xiaopeng Hong, Tiangang Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Learning Motion-Robust Remote Photoplethysmography through Arbitrary Resolution VideosabstractRemote photoplethysmography (rPPG) enables non-contact heart rate (HR) estimation from facial videos which gives significant convenience compared with traditional contact-based measurements. In the real-world long-term health monitoring scenario, the distance of the participants and their head movements usually vary by time, resulting in the inaccurate rPPG measurement due to the varying face resolution and complex motion artifacts. Different from the previous rPPG models designed for a constant distance between camera and participants, in this paper, we propose two plug-and-play blocks (i.e., physiological signal feature extraction block (PFE) and temporal face alignment block (TFA)) to alleviate the degradation of changing distance and head motion. On one side, guided with representative-area information, PFE adaptively encodes the arbitrary resolution facial frames to the fixed-resolution facial structure features. On the other side, leveraging the estimated optical flow, TFA is able to counteract the rPPG signal confusion caused by the head movement thus benefit the motion-robust rPPG signal recovery. Besides, we also train the model with a cross-resolution constraint using a two-stream dual-resolution framework, which further helps PFE learn resolution-robust facial rPPG features. Extensive experiments on three benchmark datasets (UBFC-rPPG, COHFACE and PURE) demonstrate the superior performance of the proposed method. One highlight is that with PFE and TFA, the off-the-shelf spatio-temporal rPPG models can predict more robust rPPG signals under both varying face resolution and severe head movement scenarios. The codes are available at https://github.com/LJWGIT/Arbitrary_Resolution_rPPG. Zitong Yu, Jingang Shi |
AAAI | 3 |
| 2023 | Learning Attention from Attention: Efficient Self-Refinement Transformer for Face Super-ResolutionabstractRecently, Transformer-based architecture has been introduced into face super-resolution task due to its advantage in capturing long-range dependencies. However, these approaches tend to integrate global information in a large searching region, which neglect to focus on the most relevant information and induce blurry effect by the irrelevant textures. Some improved methods simply constrain self-attention in a local window to suppress the useless information. But it also limits the capability of recovering high-frequency details when flat areas dominate the local searching window. To improve the above issues, we propose a novel self-refinement mechanism which could adaptively achieve texture-aware reconstruction in a coarse-to-fine procedure. Generally, the primary self-attention is first conducted to reconstruct the coarse-grained textures and detect the fine-grained regions required further compensation. Then, region selection attention is performed to refine the textures on these key regions. Since self-attention considers the channel information on tokens equally, we employ a dual-branch feature integration module to privilege the important channels in feature extraction. Furthermore, we design the wavelet fusion module which integrate shallow-layer structure and deep-layer detailed feature to recover realistic face images in frequency domain. Extensive experiments demonstrate the effectiveness on a variety of datasets. Guanxin Li, Jingang Shi, Yuan Zong, Fei Wang 0037, Tian Wang 0002, Yihong Gong |
IJCAI | 2 |
| 2023 | PhysFormer++: Facial Video-Based Physiological Measurement with SlowFast Temporal Difference TransformerabstractAbstract Remote photoplethysmography (rPPG), which aims at measuring heart activities and physiological signals from facial video without any contact, has great potential in many applications (e.g., remote healthcare and affective computing). Recent deep learning approaches focus on mining subtle rPPG clues using convolutional neural networks with limited spatio-temporal receptive fields, which neglect the long-range spatio-temporal perception and interaction for rPPG modeling. In this paper, we propose two end-to-end video transformer based architectures, namely PhysFormer and PhysFormer++, to adaptively aggregate both local and global spatio-temporal features for rPPG representation enhancement. As key modules in PhysFormer, the temporal difference transformers first enhance the quasi-periodic rPPG features with temporal difference guided global attention, and then refine the local spatio-temporal representation against interference. To better exploit the temporal contextual and periodic rPPG clues, we also extend the PhysFormer to the two-pathway SlowFast based PhysFormer++ with temporal difference periodic and cross-attention transformers. Furthermore, we propose the label distribution learning and a curriculum learning inspired dynamic constraint in frequency domain, which provide elaborate supervisions for PhysFormer and PhysFormer++ and alleviate overfitting. Comprehensive experiments are performed on four benchmark datasets to show our superior performance on both intra- and cross-dataset testings. Unlike most transformer networks needed pretraining from large-scale datasets, the proposed PhysFormer family can be easily trained from scratch on rPPG datasets, which makes it promising as a novel transformer baseline for the rPPG community. Zitong Yu, Yuming Shen, Jingang Shi, Hengshuang Zhao, Yawen Cui, Philip Torr 0001, Guoying Zhao 0001 |
Int. J. Comput. Vis. | 3 |
| 2023 | Model Behavior Preserving for Class-Incremental LearningabstractDeep models have shown to be vulnerable to catastrophic forgetting, a phenomenon that the recognition performance on old data degrades when a pre-trained model is fine-tuned on new data. Knowledge distillation (KD) is a popular incremental approach to alleviate catastrophic forgetting. However, it usually fixes the absolute values of neural responses for isolated historical instances, without considering the intrinsic structure of the responses by a convolutional neural network (CNN) model. To overcome this limitation, we recognize the importance of the global property of the whole instance set and treat it as a behavior characteristic of a CNN model relevant to model incremental learning. On this basis: 1) we design an instance neighborhood-preserving (INP) loss to maintain the order of pair-wise instance similarities of the old model in the feature space; 2) we devise a label priority-preserving (LPP) loss to preserve the label ranking lists within instance-wise label probability vectors in the output space; and 3) we introduce an efficient derivable ranking algorithm for calculating the two loss functions. Extensive experiments conducted on CIFAR100 and ImageNet show that our approach achieves the state-of-the-art performance. Xiaopeng Hong, Songlin Dong, Jingang Shi, Yihong Gong |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | PhysFormer: Facial Video-based Physiological Measurement with Temporal Difference TransformerabstractRemote photoplethysmography (rPPG), which aims at measuring heart activities and physiological signals from facial video without any contact, has great potential in many applications. Recent deep learning approaches focus on mining subtle rPPG clues using convolutional neural networks with limited spatio-temporal receptive fields, which neglect the long-range spatio-temporal perception and interaction for rPPG modeling. In this paper, we propose the PhysFormer, an end-to-end video transformer based architecture, to adaptively aggregate both local and global spatio-temporal features for rPPG representation enhancement. As key modules in PhysFormer, the temporal difference transformers first enhance the quasi-periodic rPPG features with temporal difference guided global attention, and then refine the local spatio-temporal representation against interference. Furthermore, we also propose the label distribution learning and a curriculum learning inspired dynamic constraint in frequency domain, which provide elaborate supervisions for PhysFormer and alleviate overfitting. Comprehensive experiments are performed on four benchmark datasets to show our superior performance on both intra- and cross-dataset testings. One highlight is that, unlike most transformer networks needed pretraining from large-scale datasets, the proposed PhysFormer can be easily trained from scratch on rPPG datasets, which makes it promising as a novel transformer baseline for the rPPG community. The codes are available at https://github.com/ZitongYu/PhysFormer. Zitong Yu, Yuming Shen, Jingang Shi, Hengshuang Zhao, Philip Torr 0001, Guoying Zhao 0001 |
CVPR | 3 |
| 2022 | Information-Growth Swin Transformer Network for Image Super-ResolutionabstractSuper-resolution (SR) reconstruction is a typical ill-posed problem and therefore can be considered as an information-growth process. The regions with dramatic information increase in the stage of extracting depth features often contain more high-frequency details. So giving more attention to these regions will improve the performance of super-resolution reconstruction. Recently, Transformer-based models have shown remarkable performance in SR. However, current Transformer-based models focus on processing for the features of the current layer input and cannot capture the degree of informational growth crossing successive layers. For this reason, we propose an information-growth Swin Transformer network (IGSTN) for single image super-resolution. The IGSTN can adaptively extract information-growth global dependencies to generate spatial attention, and then this spatial attention will be fused with the feature self-attention in the Transformer to produce the final attention, which allows the model to focus more on high-frequency regions and learn more high-frequency details from them. Extensive experimental results on publicly benchmark datasets show the effectiveness of our IGSTN. Yantao Ji, Peilin Jiang, Jingang Shi, Ruiteng Zhang |
ICIP | 3 |
| 2022 | IDPT: Interconnected Dual Pyramid Transformer for Face Super-ResolutionabstractFace Super-resolution (FSR) task works for generating high-resolution (HR) face images from the corresponding low-resolution (LR) inputs, which has received a lot of attentions because of the wide application prospects. However, due to the diversity of facial texture and the difficulty of reconstructing detailed content from degraded images, FSR technology is still far away from being solved. In this paper, we propose a novel and effective face super-resolution framework based on Transformer, namely Interconnected Dual Pyramid Transformer (IDPT). Instead of straightly stacking cascaded feature reconstruction blocks, the proposed IDPT designs the pyramid encoder/decoder Transformer architecture to extract coarse and detailed facial textures respectively, while the relationship between the dual pyramid Transformers is further explored by a bottom pyramid feature extractor. The pyramid encoder/decoder structure is devised to adapt various characteristics of textures in different spatial spaces hierarchically. A novel fusing modulation module is inserted in each spatial layer to guide the refinement of detailed texture by the corresponding coarse texture, while fusing the shallow-layer coarse feature and corresponding deep-layer detailed feature simultaneously. Extensive experiments and visualizations on various datasets demonstrate the superiority of the proposed method for face super-resolution tasks. Jingang Shi, Yusi Wang, Songlin Dong, Xiaopeng Hong, Zitong Yu, Fei Wang 0037, Changxin Wang, Yihong Gong |
IJCAI | 1 |
| 2021 | Structural Knowledge Organization and Transfer for Class-Incremental LearningabstractDeep models are vulnerable to catastrophic forgetting when fine-tuned on new data. Popular distillation-based methods usually neglect the relations between data samples and may eventually forget essential structural knowledge. To solve these shortcomings, we propose a structural graph knowledge distillation based incremental learning framework to preserve both the positions of samples and their relations. Firstly, a memory knowledge graph (MKG) is generated to fully characterize the structural knowledge of historical tasks. Secondly, we develop a graph interpolation mechanism to enrich the domain of knowledge and alleviate the inter-class sample imbalance issue. Thirdly, we introduce structural graph knowledge distillation to transfer the knowledge of historical tasks. Comprehensive experiments on three datasets validate the proposed method. Xiaopeng Hong, Songlin Dong, Jingang Shi, Yihong Gong |
MMAsia | 5 |
| 2021 | Rethinking the ST-GCNs for 3D skeleton-based human action recognitionabstractThe skeletal data has been an alternative for the human action recognition task as it provides more compact and distinct information compared to the traditional RGB input. However, unlike the RGB input, the skeleton data lies in a non-Euclidean space that traditional deep learning methods are not able to use their fullest potential. Fortunately, with the emerging trend of Geometric deep learning, the spatial-temporal graph convolutional network (ST-GCN) has been proposed to deal with the action recognition problem from skeleton data. ST-GCN and its variants fit well with skeleton-based action recognition and are becoming the mainstream frameworks for this task. However, the efficiency and the performance of the task are hindered by either fixing the skeleton joint correlations or providing a computational expensive strategy to construct a dynamic topology for the skeleton. We argue that many of these operations are either unnecessary or even harmful for the task. By theoretically and experimentally analysing the state-of-the-art ST-GCNs, we provide a simple but efficient strategy to capture the global graph correlations and thus efficiently model the representation of the input graph sequences. Moreover, the global graph strategy also reduces the graph sequence into the Euclidean space, thus a multi-scale temporal filter is introduced to efficiently capture the dynamic information. With the method, we are not only able to better extract the graph correlations with much fewer parameters (only 12.6% of the current best), but we also achieve a superior performance. Extensive experiments on current largest 3D datasets, NTU-RGB+D and NTU-RGB+D 120, demonstrate the ability of our network to perform efficient and lightweight priority on this task. Wei Peng 0009, Jingang Shi, Tuomas Varanka, Guoying Zhao 0001 |
Neurocomputing | 2 |
| 2021 | Spatial Temporal Graph Deconvolutional Network for Skeleton-Based Human Action RecognitionabstractBenefited from the powerful ability of spatial temporal Graph Convolutional Networks (ST-GCNs), skeleton-based human action recognition has gained promising success. However, the node interaction through message propagation does not always provide complementary information. Instead, it May even produce destructive noise and thus make learned representations indistinguishable. Inevitably, the graph representation would also become over-smoothing especially when multiple GCN layers are stacked. This paper proposes spatial-temporal graph deconvolutional networks (ST-GDNs), a novel and flexible graph deconvolution technique, to alleviate this issue. At its core, this method provides a better message aggregation by removing the embedding redundancy of the input graphs from either node-wise, frame-wise or element-wise at different network layers. Extensive experiments on three current most challenging benchmarks verify that ST-GDN consistently improves the performance and largely reduce the model size on these datasets. Wei Peng 0009, Jingang Shi, Guoying Zhao 0001 |
IEEE Signal Process. Lett. | 2 |
| 2020 | Face Anti-Spoofing with Human Material Perception
Zitong Yu, Xuesong Niu, Jingang Shi, Guoying Zhao 0001 |
ECCV (7) | 4 |
| 2020 | Mix Dimension in Poincaré Geometry for 3D Skeleton-based Action RecognitionabstractGraph Convolutional Networks (GCNs) have already demonstrated their powerful ability to model the irregular data, e.g., skeletal data in human action recognition, providing an exciting new way to fuse rich structural information for nodes residing in different parts of a graph. In human action recognition, current works introduce a dynamic graph generation mechanism to better capture the underlying semantic skeleton connections and thus improves the performance. In this paper, we provide an orthogonal way to explore the underlying connections. Instead of introducing an expensive dynamic graph generation paradigm, we build a more efficient GCN on a Riemann manifold, which we think is a more suitable space to model the graph data, to make the extracted representations fit the embedding matrix. Specifically, we present a novel spatial-temporal GCN (ST-GCN) architecture which is defined via the Poincaré geometry such that it is able to better model the latent anatomy of the structure data. To further explore the optimal projection dimension in the Riemann space, we mix different dimensions on the manifold and provide an efficient way to explore the dimension for each ST-GCN layer. With the final resulted architecture, we evaluate our method on two current largest scale 3D datasets, i.e., NTU RGB+D and NTU RGB+D 120. The comparison results show that the model could achieve a superior performance under any given evaluation metrics with only 40% model size when compared with the previous best GCN method, which proves the effectiveness of our model. Wei Peng 0009, Jingang Shi, Zhaoqiang Xia, Guoying Zhao 0001 |
ACM Multimedia | 2 |
| 2020 | AutoHR: A Strong End-to-End Baseline for Remote Heart Rate Measurement With Neural SearchingabstractRemote photoplethysmography (rPPG), which aims at measuring heart activities without any contact, has great potential in many applications (e.g., remote healthcare). Existing end-to-end rPPG and heart rate (HR) measurement methods from facial videos are vulnerable to the less-constrained scenarios (e.g., with head movement and bad illumination). In this letter, we explore the reason why existing end-to-end networks perform poorly in challenging conditions and establish a strong end-to-end baseline (AutoHR) for remote HR measurement with neural architecture search (NAS). The proposed method includes three parts: 1) a powerful searched backbone with novel Temporal Difference Convolution (TDC), intending to capture intrinsic rPPG-aware clues between frames; 2) a hybrid loss function considering constraints from both time and frequency domains; and 3) spatio-temporal data augmentation strategies for better representation learning. Comprehensive experiments are performed on three benchmark datasets, and we achieved superior performance on both intra- and cross-dataset testings. Zitong Yu, Xuesong Niu, Jingang Shi, Guoying Zhao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2020 | Atrial Fibrillation Detection From Face Videos by Fusing Subtle VariationsabstractAtrial fibrillation (AF) is one of the most common cardiac arrhythmias, which particularly occurs in the elderly individuals with heart disease. Though AF is often asymptomatic during normal activities, it has huge potential risks for stroke and other severe diseases. Thus, early detection of AF has great importance in the field of public health. Currently, electrocardiography (ECG) is the commonly used measure for the diagnosis of AF, which presents the irregular rhythm of waveform for AF patients. However, the measurement of the ECG signal requires special medical acquisition devices, which are not comfortable for practical monitoring in daily life. In this paper, we explore a very promising algorithm to detect AF from remote face videos by analyzing the color variations of face skin. The main challenge is that the current remote photoplethysmography (rPPG) technique is rather immature, which causes difficulty in extracting accurate pulse signals for describing the cardiac rhythm. To solve this problem, we first utilize various rPPG algorithms to capture pulse rhythms from different regions on the face video. We then investigate biomedical statistical methods to extract suitable features from each pulse signal. Due to the imprecision of video-extracted pulse signals, some traditional physiological features may lose their utility since they were originally proposed for ECG signals. Furthermore, some of them are very susceptible to the influence of noise. Thus, we propose a feature fusion algorithm to select and combine reasonable information from multiple physiological features, which aims to preserve the discriminability of detecting AF in the presence of the noise and outlier disturbances. The experimental results on a real-world database demonstrate the effectiveness of the proposed method in providing useful information for AF detection. Jingang Shi, Iman Alikhani, Zitong Yu, Tapio Seppänen, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Temporal Hierarchical Dictionary Guided Decoding for Online Gesture Segmentation and RecognitionabstractOnline segmentation and recognition of skeleton- based gestures are challenging. Compared with offline cases, the inference of online settings can only rely on the current few frames and always completes before whole temporal movements are performed. However, incompletely performed gestures are ambiguous and their early recognition is easy to fall into local optimum. In this work, we address the problem with a temporal hierarchical dictionary to guide the hidden Markov model (HMM) decoding procedure. The intuition is that, gestures are ambiguous with high uncertainty at early performing phases, and only become discriminate after certain phases. This uncertainty naturally can be measured by entropy. Thus, we propose a measurement called "relative entropy map" (REM) to encode this temporal context to guide HMM decoding. Furthermore, we introduce a progressive learning strategy with which neural networks could learn a robust recognition of HMM states in an iterative manner. The performance of our method is intensively evaluated on three challenging databases and achieves state-of-the-art results. Our method shows the abilities of both extracting the discriminate connotations and reducing large redundancy in the HMM transition process. It is verified that our framework can achieve online recognition of continuous gesture streams even when they are halfway performed. Haoyu Chen 0001, Xin Liu 0012, Jingang Shi, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Face Hallucination via Coarse-to-Fine Recursive Kernel Regression StructureabstractIn recent years, patch-based face hallucination algorithms have attracted considerable interest due to their effectiveness. These approaches produce a high-resolution (HR) face image according to the corresponding low-resolution (LR) input by learning a reconstruction model from the given training image set. The critical problem in these algorithms is establishing the underlying relationship between LR and HR patch pairs. Most previous methods aim to denote each input LR patch by the linear combination of the training set in the LR space while utilizing the combination weights to reconstruct the target HR patch. However, this assumes that the same combination weights should be shared between various resolution spaces, which is truly difficult to satisfy because of the one-to-many mapping relation between LR and HR patches. In this paper, we directly train a series of adaptive kernel regression mappings for predicting the lost high-frequency information from the LR patch, which avoids dealing with the above difficult problem. During the training process, we first establish a local optimization function on each LR/HR training pair according to the geometric structure of neighboring patches. The objective of local optimization can be presented in two aspects: 1) ensure the reconstruction consistency between each LR patch and the corresponding HR patch and 2) preserve the intrinsic geometry between each HR training patch and its original neighbors after the reconstruction process. The local optimizations are finally incorporated as the global optimization for calculating the optimal kernel regression function. To better approximate the target HR patch, we further propose a recursive structure to compensate for the residual reconstruction error of high-frequency details by a series of regression mappings. The proposed method is rather fast yet very effective in producing HR face images. Experimental results show that the proposed approach achieves superior performance with reasonable computational time compared with the state-of-the-art methods. Jingang Shi, Guoying Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Exploiting Nonslip Wall Contacts to Position Two Particles Using the Same Control InputabstractSteered particles offer a method for targeted therapy, interventions, and drug delivery in regions inaccessible by large robots. For example, magnetic actuation of particles has the benefits of requiring no tethers, being able to operate from a distance, and in some cases allows imaging for feedback (e.g., MRI). This paper investigates position control of particles using uniform forces (the same force is applied everywhere in the workspace). Given a controllable field that can generate bidirectional forces in three orthogonal directions, steering one particle in three-dimensional (3-D) is trivial. Adding additional particles to steer makes the system underactuated because there are more states than control inputs. However, the walls of in vivo and artificial environments often have surface roughness such that the particles do not move unless actuation pulls them away from the wall. In the previous work, we showed that the individual two-dimensional (2-D) position of two particles is controllable using global inputs in a square workspace with nonslip wall contact [1]. Because in vivo environments are usually not square, this paper extends the previous work to all convex workspaces, and shows how this could be extended to 3-D positioning of neutrally buoyant particles. We investigate analytically an idealized variant of this problem with nonslip boundaries and control inputs that are applied uniformly to all particles in the workspace. This paper also implements the algorithms in 2-D using a hardware setup inspired by the gastrointestinal tract. Shiva Shahrokhi, Jingang Shi, Benedict Isichei, Aaron T. Becker |
IEEE Trans. Robotics | 2 |
| 2018 | The OBF Database: A Large Face Video Database for Remote Physiological Signal Measurement and Atrial Fibrillation DetectionabstractPhysiological signals, including heart rate (HR), heart rate variability (HRV), and respiratory frequency (RF) are important indicators of our health, which are usually measured in clinical examinations. Traditional physiological signal measurement often involves contact sensors, which may be inconvenient or cause discomfort in long-term monitoring sessions. Recently, there were studies exploring remote HR measurement from facial videos, and several methods have been proposed. However, previous methods cannot be fairly compared, since they mostly used private, self-collected small datasets as there has been no public benchmark database for the evaluation. Besides, we haven't found any study that validates such methods for clinical applications yet, e.g., diagnosing cardiac arrhythmias/disease, which could be one major goal of this technology. In this paper, we introduce the Oulu Bio-Face (OBF) database as a benchmark set to fill in the blank. The OBF database includes large number of facial videos with simultaneously recorded reference physiological signals. The data were recorded both from healthy subjects and from patients with atrial fibrillation (AF), which is the most common sustained and widespread cardiac arrhythmia encountered in clinical practice. Accuracy of HR, HRV and RF measured from OBF videos are provided as the baseline results for future evaluation. We also demonstrated that the video-extracted HRV features can achieve promising performance for AF detection, which has never been studied before. From a wider outlook, the remote technology may lead to convenient self-examination in mobile condition for earlier diagnosis of the arrhythmia. Iman Alikhani, Jingang Shi, Tapio Seppänen, Juhani Junttila, Kirsi Majamaa-Voltti, Mikko Tulppo, Guoying Zhao 0001 |
FG | 3 |
| 2018 | Hallucinating Face Image by Regularization Models in High-Resolution Feature SpaceabstractIn this paper, we propose two novel regularization models in patch-wise and pixel-wise respectively, which are efficient to reconstruct high-resolution (HR) face image from low-resolution (LR) input. Unlike the conventional patch-based models which depend on the assumption of local geometry consistency in LR and HR spaces, the proposed method directly regularizes the relationship between the target patch and corresponding training set in the HR space. It avoids to deal with the tough problem of preserving local geometry in various resolutions. Taking advantage of kernel function in efficiently describing intrinsic features, we further conduct the patch-based reconstruction model in the high-dimensional kernel space for capturing nonlinear characteristics. Meanwhile, a pixel-based model is proposed to regularize the relationship of pixels in the local neighborhood, which can be employed to enhance the fuzzy details in the target HR face image. It privileges the reconstruction of pixels along the dominant orientation of structure, which is useful for preserving high-frequency information on complex edges. Finally, we combine the two reconstruction models into a unified framework. The output HR face image can be finally optimized by performing an iterative procedure. Experimental results demonstrate that the proposed face hallucination method produces superior performance than the state-of-the-art methods. Jingang Shi, Xin Liu 0012, Yuan Zong, Chun Qi, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Domain Regeneration for Cross-Database Micro-Expression RecognitionabstractRecently, micro-expression recognition has attracted lots of researchers' attention due to its potential value in many practical applications, e.g., lie detection. In this paper, we investigate an interesting and challenging problem in micro-expression recognition, i.e., cross-database micro-expression recognition, in which the training and testing samples come from different micro-expression databases. Under this problem setting, the consistent feature distribution between the training and testing samples originally existing in conventional micro-expression recognition would be seriously broken and hence the performance of most current well-performing micro-expression recognition methods may sharply drop. In order to overcome it, we propose a simple yet effective framework called Domain Regeneration (DR) in this paper. DR framework aims at learning a domain regenerator to regenerate the micro-expression samples from source and target databases respectively such that they can abide by the same or similar feature distributions. Thus, we are able to use the classifier learned based on the labeled source micro-expression samples to predict the label information of the unlabeled target micro-expression samples. To evaluate the proposed DR framework, we conduct extensive cross-database micro-expression recognition experiments designed based on SMIC and CASME II databases. Experimental results show that compared with recent state-of-the-art cross-database emotion recognition methods, the proposed DR framework has more promising performance. Yuan Zong, Wenming Zheng, Xiaohua Huang 0003, Jingang Shi, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Image denoising via group sparsity residual constraintabstractGroup sparsity has shown great potential in various low-level vision tasks (e.g, image denoising, deblurring and inpainting). In this paper, we propose a new prior model for image denoising via group sparsity residual constraint (GSRC). To enhance the performance of group sparse-based image denoising, the concept of group sparsity residual is proposed, and thus, the problem of image denoising is translated into one that reduces the group sparsity residual. To reduce the residual, we first obtain some good estimation of the group sparse coefficients of the original image by the first-pass estimation of noisy image, and then centralize the group sparse coefficients of noisy image to the estimation. Experimental results have demonstrated that the proposed method not only outperforms many state-of-the-art denoising methods such as BM3D and WNNM, but results in a faster speed. Zhiyuan Zha, Xin Liu 0012, Ziheng Zhou 0003, Xiaohua Huang 0003, Jingang Shi, Zhenhong Shang, Lan Tang, Yechao Bai, Qiong Wang 0002, Xinggan Zhang |
ICASSP | 5 |
| 2015 | From Local Geometry to Global Structure: Learning Latent Subspace for Low-resolution Face Image RecognitionabstractIn this letter, we propose a novel approach for learning coupled mappings to improve the performance of low-resolution (LR) face image recognition. The coupled mappings aim to project the LR probe images and high-resolution (HR) gallery images into a unified latent subspace, which is efficient to measure the similarity of face images with different resolutions. In the training phase, we first construct local optimization for each training sample according to the relationship of neighboring data points. The local optimization aims to: (1) ensure the consistency for each LR face image and corresponding HR one; (2) model the intrinsic geometric structure between each given sample and its neighbors; and (3) preserve the discriminative information across different subjects. We finally incorporate the local optimizations together for building the global structure. The coupled mappings can be learned by solving a standard eigen-decomposition problem, which avoids the small-sample-size problem. Experimental results demonstrate the effectiveness of the proposed method on public face databases. Jingang Shi, Chun Qi |
IEEE Signal Process. Lett. | 1 |
| 2015 | Kernel-Based Face Hallucination via Dual Regularization PriorsabstractRecently, patch-based face hallucination methods have shown the ability for achieving high-quality face images. The high-resolution (HR) patches can be reconstructed by a linear combination of training patches, while the combination coefficients are learned according to the corresponding low-resolution (LR) patches. In order to reflect the local features, face images are usually divided into very small patches, e.g.$3 \times 3$for LR case. Though we assume the linear relationship between training patches, it may fail to obtain suitable combination coefficients due to the low dimension of LR patches. In this letter, the kernel function is utilized for mapping the LR patches into a high-dimensional feature space. By taking into account the nonlinear structures, it is more effective to estimate the combination coefficients in the kernel feature space. Furthermore, a pixel-based model is also employed according to the characteristics of face images, which is useful to compensate the local textures. The final HR face images are obtained from a global optimization function by an iterative process. Experimental results show the advantage of the proposed approach in both reconstruction error and visual quality. Jingang Shi, Chun Qi |
IEEE Signal Process. Lett. | 1 |
| 2014 | Global consistency, local sparsity and pixel correlation: A unified framework for face hallucination
Jingang Shi, Xin Liu 0012, Chun Qi |
Pattern Recognit. | 1 |
| 2013 | Face hallucination based on PCA dictionary pairsabstractThis paper presents a new position-based face hallucination algorithm based on PCA dictionary pairs. The high-resolution (HR) face image is generated in patch-wise, while each patch is hallucinated from a low-resolution (LR) observation with the training patches on the same position of face images. Different from the previous literatures which reconstruct the HR patch with raw position-patches, a set of dictionary pairs are adaptively learned according to the patch location in the proposed algorithm. We joint the LR-HR position-patches together and project the dataset into principal directions by principal component analysis (PCA). The principal components are applied to generate the coupled LR-HR dictionaries. Moreover, the corresponding eigenvalues are also served as a constraint in the reconstruction. Experimental results demonstrate that the proposed approach achieves superior performance when compared with the state-of-the-art algorithms. Jingang Shi, Chun Qi |
ICIP | 1 |
| 2013 | Sparse modeling based image inpainting with local similarity constraintabstractIn this paper, we propose an efficient exemplar-based inpainting algorithm via sparse modeling and local similarity constraint. The inpainting procedure contains two steps: calculating the filling order and reconstructing the target patch. The filling order is decided by patch priority, which privileges the patch located at edge or corner. The target patch is then estimated by a combination of candidate patches. In the proposed method, three regularization terms are introduced to improve the patch reconstruction step. The first term ensures the compatibility between the target patch and the estimated one. The second term assigns larger combination coefficients for the candidate patches which are most similar with the target patch. The third term penalties the combination coefficients for the outliers in the candidate patches. Finally, the three regularization terms are incorporated into a unified sparse representation framework for reconstructing the target patch. Experiments show that the proposed algorithm can effectively fill in missing pixels in a visually plausible way. Jingang Shi, Chun Qi |
ICIP | 1 |
| 2011 | Requirement-Based Query and Update Scheduling in Real-Time Data Warehouses
Fangling Leng, Yubin Bao, Ge Yu 0001, Jingang Shi, Xiaoyan Cai |
WAIM | 4 |