Bochao Zou

dblp:197/9774 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0002-2126-8159ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Implicit alignment and query refinement for RGB-T semantic segmentation
Chang Liu 0136, Haizhuang Liu, Junbao Zhuo, Bochao Zou, Jiansheng Chen 0001, Qianchuan Zhao, Huimin Ma 0001
Pattern Recognit.4
2026 ME-TST+: Micro-Expression Analysis via Temporal State Transition With ROI Relationship Awareness
abstract
Micro-expressions (MEs) are regarded as important indicators of an individual’s intrinsic emotions, preferences, and tendencies. ME analysis requires spotting of ME intervals within long video sequences and recognition of their corresponding emotional categories. Previous deep learning approaches commonly employ sliding-window classification networks. However, the use of fixed window lengths and hard classification presents notable limitations in practice. Furthermore, these methods typically treat ME spotting and recognition as two separate tasks, overlooking the essential relationship between them. To address these challenges, this paper proposes two state space model-based architectures, namely ME-TST and ME-TST+, which utilize temporal state transition mechanisms to replace conventional window-level classification with video-level regression. This enables a more precise characterization of the temporal dynamics of MEs and supports the modeling of MEs with varying durations. In ME-TST+, we further introduce multi-granularity ROI modeling and the SlowFast Mamba framework to alleviate information loss associated with treating ME analysis as a time series task. Additionally, we propose a synergy strategy for spotting and recognition at both the feature and result levels, leveraging their intrinsic relationship to enhance overall analysis performance. Extensive experiments demonstrate that the proposed methods achieve state-of-the-art performance. The code is available at https://github.com/zizheng-guo/ME-TST.
Zizheng Guo 0002, Bochao Zou, Junbao Zhuo, Huimin Ma 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 MLSTP: A Mamba-Based Long-Short Term Fusion Network for Improved Trajectory Prediction in Autonomous Driving
abstract
Traffic trajectory prediction plays a crucial role in ensuring the safety and efficiency of autonomous driving systems. With increasing traffic complexity, accurate and real-time prediction of surrounding vehicle trajectories has become a major challenge. Existing methods often define trajectory prediction as a static task, predicting future trajectories at once based solely on historical states. However, this approach leads to significant long-term prediction errors and struggles to capture potential hazards in complex, dynamic traffic conditions. In this paper, we propose a lightweight end-to-end trajectory prediction model integrating a long-term guidance module and a short-term feedback mechanism. The long-term guidance module predicts the final driving goal that remains unchanged over time, guiding the vehicle’s eventual position. The short-term feedback mechanism captures the dynamic driving environment, continuously updating the model with immediate changes in vehicle interactions and road conditions to avoid unexpected dangers. At each time step, our approach leverages accumulated historical information and incorporates both short-term and long-term future information to iteratively predict multi-modal trajectories, effectively reducing long-term prediction errors. Additionally, our method introduces the Mamba module, which enhances real-time prediction performance and reduces historical information loss during long-term prediction. Experiments on the Argoverse 1 and Argoverse 2 datasets show that our method, with fewer parameters and higher inference speed, achieves results comparable to or surpassing previous state-of-the-art methods.
Jiaxing Ren, Bochao Zou, Juntao Lyu, Huimin Ma 0001
IEEE Trans. Intell. Transp. Syst.2
2025 A²RNet: Adversarial Attack Resilient Network for Robust Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion (IVIF) is a crucial technique for enhancing visual performance by integrating unique information from different modalities into one fused image. Exiting methods pay more attention to conducting fusion with undisturbed data, while overlooking the impact of deliberate interference on the effectiveness of fusion results. To investigate the robustness of fusion models, in this paper, we propose a novel adversarial attack resilient network, called A2RNet. Specifically, we develop an adversarial paradigm with an anti-attack loss function to implement adversarial attacks and training. It is constructed based on the intrinsic nature of IVIF and provide a robust foundation for future research advancements. We adopt a Unet as the pipeline with a transformer-based defensive refinement module (DRM) under this paradigm, which guarantees fused image quality in a robust coarse-to-fine manner. Compared to previous works, our method mitigates the adverse effects of adversarial perturbations, consistently maintaining high-fidelity fusion results. Furthermore, the performance of downstream tasks can also be well maintained under adversarial attacks.
Jiawei Li 0016, Jiansheng Chen 0002, Xinlong Ding, Jinyuan Liu 0001, Bochao Zou, Huimin Ma 0001
AAAI7
2025 ProtoCar: Learning 3D Vehicle Prototypes from Single-View and Unconstrained Driving Scene Images
abstract
Reconstructing 3D models from sensor data is a valuable and promising direction for developing testing and validation environments in applications like autonomous driving. However, existing methods for 3D modeling often rely on extensive multi-view data or controlled conditions, making them difficult and expensive to scale. Furthermore, these methods, particularly those based on neural radiance fields, typically produce implicit models that can be challenging to manipulate and suffer from slow rendering speeds. In this paper, we introduce ProtoCar, a novel approach that overcomes these limitations by learning 3D vehicle prototypes from single-view images with diverse and unconstrained visual conditions. ProtoCar uses real-world driving data from LiDAR and image sensors, and employs 3D Gaussian splatting techniques to represent explicit geometric and texture. Extensive experiments demonstrate that ProtoCar generates high-quality 3D models and adapts well to various vehicle types and challenging visual scenarios, offering a scalable and effective solution for 3D modeling in environments with limited and variable visual information.
Hongyuan Liu 0007, Haochen Yu, Bochao Zou, Juntao Lyu, Qi Mei, Huimin Ma 0001
AAAI3
2025 RhythmMamba: Fast, Lightweight, and Accurate Remote Physiological Measurement
abstract
Remote photoplethysmography (rPPG) is a method for non-contact measurement of physiological signals from facial videos, holding great potential in various applications such as healthcare, affective computing, and anti-spoofing. Existing deep learning methods struggle to address two core issues of rPPG simultaneously: understanding the periodic pattern of rPPG among long contexts and addressing large spatiotemporal redundancy in video segments. These represent a trade-off between computational complexity and the ability to capture long-range dependencies. In this paper, we introduce RhythmMamba, a state space model-based method that captures long-range dependencies while maintaining linear complexity. By viewing rPPG as a time series task through the proposed frame stem, the periodic variations in pulse waves are modeled as state transitions. Additionally, we design multi-temporal constraint and frequency domain feed-forward, both aligned with the characteristics of rPPG time series, to improve the learning capacity of Mamba for rPPG signals. Extensive experiments show that RhythmMamba achieves state-of-the-art performance with 319% throughput and 23% peak GPU memory.
Bochao Zou, Zizheng Guo 0002, Xiaocheng Hu, Huimin Ma 0001
AAAI1
2025 Synergistic Spotting and Recognition of Micro-Expression via Temporal State Transition
abstract
Micro-expressions are involuntary facial movements that cannot be consciously controlled, conveying subtle cues with substantial real-world applications. The analysis of micro-expressions generally involves two main tasks: spotting micro-expression intervals in long videos and recognizing the emotions associated with these intervals. Previous deep-learning methods have primarily relied on classification networks utilizing sliding windows. However, fixed window sizes and window-level hard classification introduce numerous constraints. Additionally, these methods have not fully exploited the potential of complementary pathways for spotting and recognition. In this paper, we present a novel temporal state transition architecture grounded in the state space model, which replaces conventional window-level classification with video-level regression. Furthermore, by leveraging the inherent connections between spotting and recognition tasks, we propose a synergistic strategy that enhances overall analysis performance. Extensive experiments demonstrate that our method achieves state-of-the-art performance. The codes are available at https://github.com/zizheng-guo/ME-TST.
Bochao Zou, Zizheng Guo 0002, Wenfeng Qin, Xin Li 0034, Kangsheng Wang, Huimin Ma 0001
ICASSP1
2025 Kaleidoscopic Background Attack: Disrupting Pose Estimation With Multi-Fold Radial Symmetry Textures
abstract
Camera pose estimation is a fundamental computer vision task that is essential for applications like visual localization and multi-view stereo reconstruction. In the object-centric scenarios with sparse inputs, the accuracy of pose estimation can be significantly influenced by background textures that occupy major portions of the images across different viewpoints. In light of this, we introduce the Kaleidoscopic Background Attack (KBA), which uses identical segments to form discs with multi-fold radial symmetry. These discs maintain high similarity across different viewpoints, enabling effective attacks on pose estimation models even with natural texture segments. Additionally, a projected orientation consistency loss is proposed to optimize the kaleidoscopic segments, leading to significant enhancement in the attack effectiveness. Experimental results show that optimized adversarial kaleidoscopic backgrounds can effectively attack various camera pose estimation models.
Xinlong Ding, Jiawei Li 0016, Bochao Zou, Huimin Ma 0001
ICCV6
2025 Puzzle-MAE: A Puzzle-Inspired Mask Autoencoder for Multi-Modal Fusion
abstract
Most unsupervised methods in the video domain rely on simple encoder-decoder structures, often resulting in discrepancies between the features extracted from unmasked patches and those from the original patches. To address this issue, we propose a novel self-supervised learning framework, PuzzleMAE, which extracts features from both masked and unmasked patches and aligns them with original image representations to improve feature consistency. Inspired by the human ability to solve puzzles through holistic image recognition and the exploitation of spatial adjacency, we propose the Global-Local Attention Module, which effectively integrates global contextual information with local feature representations. Furthermore, we introduce 3D Relative Position Embedding and Structural Position Embedding to emulate human-like spatial and structural awareness of positional relationships during the puzzle-solving process. The effectiveness of our method is validated on two downstream tasks: the First Impression V2 and DFEW datasets.
Xin Li 0034, Bochao Zou, Rongquan Wang, Huimin Ma 0001
ICME2
2025 From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models
abstract
As large language models evolve, there is growing anticipation that they will emulate human-like Theory of Mind (ToM) to assist with routine tasks. However, existing methods for evaluating machine ToM focus primarily on unimodal models and largely treat these models as black boxes, lacking an interpretative exploration of their internal mechanisms. In response, this study adopts an approach based on internal mechanisms to provide an interpretability-driven assessment of ToM in multimodal large language models (MLLMs). Specifically, we first construct a multimodal ToM test dataset, GridToM, which incorporates diverse belief testing tasks and perceptual information from multiple perspectives. Next, our analysis shows that attention heads in multimodal large models can distinguish cognitive information across perspectives, providing evidence of ToM capabilities. Furthermore, we present a lightweight, training-free approach that significantly enhances the model’s exhibited ToM by adjusting in the direction of the attention head.
Siqi Liu 0010, Bochao Zou, Jiansheng Chen 0001, Huimin Ma 0001
ICML3
2025 Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression
abstract
Micro-expressions (MEs) are involuntary, low-intensity, and short-duration facial expressions that often reveal an individual's genuine thoughts and emotions. Most existing ME analysis methods rely on window-level classification with fixed window sizes and hard decisions, which limits their ability to capture the complex temporal dynamics of MEs. Although recent approaches have adopted video-level regression frameworks to address some of these challenges, interval decoding still depends on manually predefined, window-based methods, leaving the issue only partially mitigated. In this paper, we propose a prior-guided video-level regression method for ME analysis. We introduce a scalable interval selection strategy that comprehensively considers the temporal evolution, duration, and class distribution characteristics of MEs, enabling precise spotting of the onset, apex, and offset phases. In addition, we introduce a synergistic optimization framework, in which the spotting and recognition tasks share parameters except for the classification heads. This fully exploits complementary information, makes more efficient use of limited data, and enhances the model's capability. Extensive experiments on multiple benchmark datasets demonstrate the state-of-the-art performance of our method, with an STRS of 0.0562 on CAS(ME)3 and 0.2000 on SAMMLV. The code is available at https://github.com/zizheng-guo/BoostingVRME.
Zizheng Guo 0002, Bochao Zou, Yinuo Jia, Huimin Ma 0001
ACM Multimedia2
2025 AGC-Drive: A Large-Scale Dataset for Real-World Aerial-Ground Collaboration in Driving Scenarios
abstract
By sharing information across multiple agents, collaborative perception helps autonomous vehicles mitigate occlusions and improve overall perception accuracy. While most previous work focus on vehicle-to-vehicle and vehicle-to-infrastructure collaboration, with limited attention to aerial perspectives provided by UAVs, which uniquely offer dynamic, top-down views to alleviate occlusions and monitor large-scale interactive environments. A major reason for this is the lack of high-quality datasets for aerial-ground collaborative scenarios. To bridge this gap, we present AGC-Drive, the first large-scale real-world dataset for Aerial-Ground Cooperative 3D perception. The data collection platform consists of two vehicles, each equipped with five cameras and one LiDAR sensor, and one UAV carrying a forward-facing camera and a LiDAR sensor, enabling comprehensive multi-view and multi-agent perception. Consisting of approximately 80K LiDAR frames and 360K images, the dataset covers 14 diverse real-world driving scenarios, including urban roundabouts, highway tunnels, and on/off ramps. Notably, 17\% of the data comprises dynamic interaction events, including vehicle cut-ins, cut-outs, and frequent lane changes. AGC-Drive contains 350 scenes, each with approximately 100 frames and fully annotated 3D bounding boxes covering 13 object categories. We provide benchmarks for two 3D perception tasks: vehicle-to-vehicle collaborative perception and vehicle-to-UAV collaborative perception. Additionally, we release an open-source toolkit, including spatiotemporal alignment verification tools, multi-agent visualization systems, and collaborative annotation utilities. The dataset and code are available at https://github.com/PercepX/AGC-Drive.
Yunhao Hou, Bochao Zou, Shangdong Yang, Junbao Zhuo, Siheng Chen, Jiansheng Chen 0001, Huimin Ma 0001
NeurIPS2
2025 SparseComm: An Efficient Sparse Communication Framework for Vehicle-Infrastructure Cooperative 3D Detection
Haizhuang Liu, Huazhen Chu, Junbao Zhuo, Bochao Zou, Jiansheng Chen 0001, Huimin Ma 0001
Pattern Recognit.4
2025 RhythmFormer: Extracting patterned rPPG signals based on periodic sparse attention
Bochao Zou, Zizheng Guo 0002, Jiansheng Chen 0001, Junbao Zhuo, Weiran Huang 0001, Huimin Ma 0001
Pattern Recognit.1
2024 EMo Transformer: Transformer-Based Depression Detection via Eye Movements
abstract
Depressive disorder has become a prevalent psychological illness that significantly impacts individuals’ daily lives. Traditional questionnaire assessment and clinical interviews suffer from issues such as subjectivity and a high consumption of medical resources. With the advancement of artificial intelligence, there is a growing number of depression detection methods based on statistical features. However, these methods have problems of insufficient stimulus extraction and neglecting temporal information. In order to solve these problems, we propose a transformer-based model named EMo Transformer, designed for detecting depression by effectively extracting features from stimuli and combining them with eye movements. Additionally, due to challenge in collecting data from depression patients, we design a simple and effective data augmentation method to solve this challenge. Subsequently, we design an ensemble model using the models with and without data augmentation. The experimental results of accuracy 91.95% demonstrate that our method is effective.
Xin Li 0034, Haizhuang Liu, Rongquan Wang, Bochao Zou, Huimin Ma 0001
ICME4
2024 Unveiling the Dynamics of Information Interplay in Supervised Learning
abstract
In this paper, we use matrix information theory as an analytical tool to analyze the dynamics of the information interplay between data representations and classification head vectors in the supervised learning process. Specifically, inspired by the theory of Neural Collapse, we introduce matrix mutual information ratio (MIR) and matrix entropy difference ratio (HDR) to assess the interactions of data representation and class classification heads in supervised learning, and we determine the theoretical optimal values for MIR and HDR when Neural Collapse happens. Our experiments show that MIR and HDR can effectively explain many phenomena occurring in neural networks, for example, the standard supervised training dynamics, linear mode connectivity, and the performance of label smoothing and pruning. Additionally, we use MIR and HDR to gain insights into the dynamics of grokking, which is an intriguing phenomenon observed in supervised training, where the model demonstrates generalization capabilities long after it has learned to fit the training data. Furthermore, we introduce MIR and HDR as loss terms in supervised and semi-supervised learning to optimize the information interactions among samples and classification heads. The empirical results provide evidence of the method’s effectiveness, demonstrating that the utilization of MIR and HDR not only aids in comprehending the dynamics throughout the training process but can also enhances the training procedure itself.
Kun Song 0004, Zhiquan Tan, Bochao Zou, Huimin Ma 0001, Weiran Huang 0001
ICML3
2024 Enhancing pseudo label quality for pedestrian and cyclist in weakly supervised 3D object detection
Haizhuang Liu, Huazhen Chu, Bochao Zou, Huimin Ma 0001
Neurocomputing4
2023 Micro-Expression Spotting with Face Alignment and Optical Flow
abstract
Facial expression spotting holds significant importance as it can signify emotional changes. Particularly, micro-expressions possess the potential to reveal genuine emotions, making them even more valuable in practical domains such as public safety and finance. However, spotting micro-expressions proves challenging due to their subtle movements and brief duration. This paper proposes an expression spotting method based on face alignment and optical flow. We first use a finer crop-align technique to preprocess the facial videos by aligning the face and the nose tip. Then, regions of interest (ROIs) are defined by analyzing the statistics of action units. The optical flow features are then extracted and subjected to low-pass filtering to eliminate high-frequency noise. Furthermore, candidate expression segments are identified based on the magnitude of the processed optical flows. Finally, non-maximum suppression is utilized to remove overlapping segments. The effectiveness of the proposed method is evaluated on the challenge test set, resulting in an overall F1-score of 0.19. Additional results obtained from CAS(ME)2 and SAMM Long videos provide further verification of the method's efficacy. The code is available online.
Wenfeng Qin, Bochao Zou, Xin Li 0034, Weiping Wang 0007, Huimin Ma 0001
ACM Multimedia2
2023 FD-Align: Feature Discrimination Alignment for Fine-tuning Pre-Trained Models in Few-Shot Learning
abstract
Due to the limited availability of data, existing few-shot learning methods trained from scratch fail to achieve satisfactory performance. In contrast, large-scale pre-trained models such as CLIP demonstrate remarkable few-shot and zero-shot capabilities. To enhance the performance of pre-trained models for downstream tasks, fine-tuning the model on downstream data is frequently necessary. However, fine-tuning the pre-trained model leads to a decrease in its generalizability in the presence of distribution shift, while the limited number of samples in few-shot learning makes the model highly susceptible to overfitting. Consequently, existing methods for fine-tuning few-shot learning primarily focus on fine-tuning the model's classification head or introducing additional structure. In this paper, we introduce a fine-tuning approach termed Feature Discrimination Alignment (FD-Align). Our method aims to bolster the model's generalizability by preserving the consistency of spurious features across the fine-tuning process. Extensive experimental results validate the efficacy of our approach for both ID and OOD tasks. Once fine-tuned, the model can seamlessly integrate with existing methods, leading to performance improvements. Our code can be found in https://github.com/skingorz/FD-Align.
Kun Song 0004, Huimin Ma 0001, Bochao Zou, Huishuai Zhang, Weiran Huang 0001
NeurIPS3
2023 Semi-Structural Interview-Based Chinese Multimodal Depression Corpus Towards Automatic Preliminary Screening of Depressive Disorders
abstract
Depression is a common psychiatric disorder worldwide. However, in China, a considerable number of patients with depression are not diagnosed, and most of them are not aware of their depression. Despite increasing efforts, the goal of automatic depression screening from behavioral indicators has not been achieved. A major limitation is the lack of available multimodal depression corpus in Chinese since linguistic knowledge is crucial in clinical practice. Therefore, we first carried out a comprehensive survey with psychiatrists from a renowned psychiatric hospital to identify key interview topics which are highly related to the diagnosis of depression. Then, a semi-structural interview study was conducted over a year with subjects who have undergone clinical diagnosis and professional assessment. After that, Visual, acoustic, and textual features were extracted and analyzed between the two groups, statistically significant differences were observed in all three modalities. Benchmark evaluations of both single modal and multimodal fusion methods of depression assessment were also performed. A multimodal transformer-based fusion approach achieved the best performance. Finally, the proposed Chinese Multimodal Depression Corpus (CMDC) was made publicly available after de-identification and annotation. Hopefully, the release of this corpus would promote the research progress and practical applications of automatic depression screening.
Bochao Zou, Jiali Han, Xiangwen Lyu, Huimin Ma 0001
IEEE Trans. Affect. Comput.1
2022 Feature Signal Resampling for rPPG-based Remote Cardiac Pulse Measurement with Streaming Video
abstract
Remote photoplethysmography (rPPG) which can realize contactless measurements of cardiac activity has applications in various fields. However, during practical usage with network cameras, non-uniform sampling caused by the unstable frame rate of the imaging equipment and transmission delay variations of the streaming signal may significantly reduce the performance of existing rPPG algorithms. Therefore, this paper proposes a feature signal resampling method with cubic spline interpolation, combined with an adaptive weight fusion to achieve a robust extraction of pulse waves. The evaluation experiment shows that the proposed method improves the results with streaming videos.
Bochao Zou, Dongyue Lv, Xiangwen Lyu
BIBM1
2022 Eliminating Spatial Ambiguity for Weakly Supervised 3D Object Detection without Spatial Labels
abstract
Previous weakly-supervised methods of 3D object detection in driving scenes mainly rely on spatial labels, which provide the location, dimension, or orientation information. The annotation of 3D spatial labels is time-consuming. There also exist methods that do not require spatial labels, but their detections may fall on object parts rather than entire objects or backgrounds. In this paper, a novel cross-modal weakly-supervised 3D progressive refinement framework (WS3DPR) for 3D object detection that only needs image-level class annotations is introduced. The proposed framework consists of two stages: 1) classification refinement for potential objects localization and 2) regression refinement for spatial pseudo labels reasoning. In the first stage, a region proposal network is trained by cross-modal class knowledge transferred from 2D image to 3D point cloud and class information propagation. In the second stage, the locations, dimensions, and orientations of 3D bounding boxes are further refined with geometric reasoning based on 2D frustum and 3D region. When only image-level class labels are available, proposals with different 3D locations become overlapped in 2D, leading to the misclassification of foreground objects. Therefore, a 2D-3D semantic consistency block is proposed to disentangle different 3D proposals after projection. The overall framework progressively learns features in a coarse to fine manner. Comprehensive experiments on the KITTI3D dataset demonstrate that our method achieves competitive performance compared with previous methods with a lightweight labeling process.
Haizhuang Liu, Huimin Ma 0001, Bochao Zou, Rongquan Wang, Jiansheng Chen 0001
ACM Multimedia4
2022 Concordance between facial micro-expressions and physiological signals under emotion elicitation
Bochao Zou, Xiangwen Lyu, Huimin Ma 0001
Pattern Recognit. Lett.1
2021 Effect of Depression Severity on Emotion Context Insensitivity Revealed by Facial Activities Analysis
abstract
Major Depression Disorder (MDD) is defined as a mood condition. The Emotion Context Insensitivity (ECI) hypothesis of depression argues that the reaction sensitivity of depressed people to the changes of emotional context is attenuated. This paper designs an experiment under this hypothesis to more effectively distinguish facial cues for different levels of depression severity. By adopting publicly validated video stimuli for emotion elicitation, statistical analysis was performed to reveal significant facial features for depression severity assessment. The results found that the high severity group showed a decrease in specific action units and an increase in others, and made more intense facial movements for negative stimuli than that for positive and neutral stimuli. Hopefully, this study could benefit the interpretation of emotional functioning in depression and provide a reference for the evaluation of depression severity based on facial characteristics.
Bochao Zou, Xiangwen Lyu, Huimin Ma 0001
BIBM1
2021 Video-Based Physiological Measurement Using 3D Central Difference Convolution Attention Network
abstract
Remote photoplethysmography (rPPG) is a non-contact method to measure physiological signals, such as heart rate (HR) and respiratory rate (RR), from facial videos. In this paper, we constructed a central difference convolutional attention network with Huber loss to perform more robust remote physiological signal measurements. The proposed method consists of two key parts:1) Using central difference convolution to enhance the spatiotemporal representation, which can capture rich physiological related temporal context by gathering time difference information 2) Using Huber loss as the loss function, the gradient can be smoothly reduced as the loss value between the rPPG and ground truth PPG signal is closer to the minimum. Through experiments on multiple public datasets and cross-dataset evaluation, the good performance and robustness of the rPPG measurement network based on central difference convolution are verified.
Bochao Zou, Abdelkader Nasreddine Belkacem, Chao Chen 0046
IJCB2
2019 Real-Time Micro-expression Detection in Unlabeled Long Videos Using Optical Flow and LSTM Neural Network
Zi Tian, Xiangwen Lyu, Quande Wang, Bochao Zou, Haiyong Xie 0001
CAIP (1)5
2019 Action Units recognition based on Deep Spatial-Convolutional and Multi-label Residual network
Tongqiang Yi, Bochao Zou, Xiangwen Lyu
Neurocomputing5