VLDB 2026 Research / reviewers in the wild / expert
Caifeng Shan
dblp:98/428
· DBLP profile ↗
133ranked-venue papers
16as first author
80since 2021 · last 2026
0000-0002-2131-1671ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 65 · 11 first-author · 30 since 2021Artificial intelligence and machine learning · 52 · 10 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 1 first-author · 25 since 2021Security and privacy · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Delphi: A Neuro-Symbolic Framework for Individualized, Safe and Interpretable Treatment RecommendationabstractClinical reinforcement learning (RL) holds promise for treatment recommendation but remains hindered by black-box decision processes, limited safety guarantees, and lack of individualized reasoning. We introduce Delphi Engine, the first fully trainable neuro-symbolic causal RL framework for dynamic treatment planning, designed to answer three core clinical questions in real time: Why this action? Why is it safe? Why for this patient? Specifically, Delphi integrates: (1) causality-aware state modeling using discretized physiological variables and subtype-specific causal graphs; (2) adaptive symbolic rule constraints, combining clinical guidelines and behavior-derived rules into soft differentiable logic; and (3) interpretable decision fusion, where actions are selected based on joint neural-symbolic Q-values and explained via structured LLM-based justifications. We evaluate Delphi on the MIMIC-III sepsis cohort using both standard off-policy evaluations (WIS↑1.47, DR↑1.29, RMSE↓0.207) and the first blinded physician evaluation of an explainable RL system in healthcare. Delphi consistently outperforms historical physicians' treatments in safety (+10.4%), understandability (+8.9%), and adoption rate (+5.75%) across six clinical axes. These results highlight Delphi’s potential as a safe, interpretable, and patient-specific AI assistant for critical care medicine. Muchan Tao, Yuqi Fang, Caifeng Shan, Tieniu Tan |
AAAI | 4 |
| 2026 | InstaFace: Identity-Preserving Facial Editing with Single Image Inference
MD Wahiduzzaman Khan, Mingshan Jia, En Yu, Caifeng Shan, Kaska Musial-Gabrys |
FG | 5 |
| 2026 | Ivy-Fake: A Unified Explainable Framework and Benchmark for Image and Video AIGC DetectionabstractThe rapid development of Artificial Intelligence Generated Content (AIGC) techniques has enabled the creation of high-quality synthetic content, but it also raises significant security concerns. Current detection methods face two major limitations: (1) the lack of multidimensional explainable datasets for generated images and videos. Existing open-source datasets (e.g., WildFake, GenVideo) rely on oversimplified binary annotations, which restrict the explainability and trustworthiness of trained detectors. (2) Prior MLLM-based forgery detectors (e.g., FakeVLM) exhibit insufficiently fine-grained interpretability in their step-by-step reasoning, which hinders reliable localization and explanation. To address these challenges, we introduce Ivy-Fake, the first large-scale multimodal benchmark for fake image and video detection. It consists of over 106K richly annotated training samples (images and videos) and 5,000 manually verified evaluation examples, sourced from multiple generative models and real-world datasets through a carefully designed pipeline to ensure both diversity and quality. Furthermore, we propose Ivy-xDetector, a multimodel large language model (MLLM) based on reinforcement fine-tuning (RFT), capable of producing explainable reasoning chains and achieving robust performance across multiple fake image and video detection benchmarks. Changjiang Jiang, Fengchang Yu, Wei Peng 0009, Xinbin Yuan, Yifei Bi, Zian Zhou, Chenyang Si, Caifeng Shan |
ICMR | 11 |
| 2026 | CAS-AIR-3D: A Large-scale Low-quality Multi-modal Face Database
Qi Li 0005, Xiaoxiao Dong, Weining Wang 0001, Zhenan Sun, Tieniu Tan, Caifeng Shan |
Int. J. Comput. Vis. | 6 |
| 2026 | Low-light light field image enhancement based on illumination-guided implicit gradient representation
Deyang Liu, Xiaofei Zhou 0003, Ping An 0001, Caifeng Shan, Hongbin Zha |
Neurocomputing | 5 |
| 2026 | Reshape and rotate: Adaptive weight reshaping and fine-grained rotation for ultra-low-bit diffusion transformers quantization
Lianwei Yang, Haokun Lin, Caifeng Shan, Zhenan Sun, Qingyi Gu |
Neurocomputing | 4 |
| 2026 | Learning Knowledge-Based Prompts for Robust 3D Mask Presentation Attack Detectionabstract3D mask presentation attack detection is crucial for protecting face recognition systems against the rising threat of 3D mask attacks. While most existing methods utilize multimodal features or remote photoplethysmography (rPPG) signals to distinguish between real faces and 3D masks, they face significant challenges, such as the high costs associated with multimodal sensors and limited generalization ability. Detection-related text descriptions offer concise, universal information and are cost-effective to obtain. However, the potential of vision-language multimodal features for 3D mask presentation attack detection remains unexplored. In this paper, we propose a novel knowledge-based prompt learning framework to explore the strong generalization capability of vision-language models for 3D mask presentation attack detection. Specifically, our approach incorporates entities and triples from knowledge graphs into the prompt learning process, generating fine-grained, task-specific explicit prompts that effectively harness the knowledge embedded in pre-trained vision-language models. Furthermore, considering different input images may emphasize distinct knowledge graph elements, we introduce a visual-specific knowledge filter based on an attention mechanism to refine relevant elements according to the visual context. Additionally, we leverage causal graph theory insights into the prompt learning process to further enhance the generalization ability of our method. During training, a spurious correlation elimination paradigm is employed, which removes category-irrelevant local image patches using guidance from knowledge-based text features, fostering the learning of generalized causal prompts that align with category-relevant local patches. Experimental results demonstrate that the proposed method achieves state-of-the-art intra- and cross-scenario detection performance on benchmark datasets. Fangling Jiang, Qi Li 0005, Weining Wang 0001, Caifeng Shan, Zhenan Sun, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Silhouette generation from identity-level skeleton synthesis for improved gait recognition
Yang Wang 0103, Caifeng Shan, Yan Huang 0023, Liang Wang 0001 |
Pattern Recognit. | 3 |
| 2026 | Global and local collaborative learning for no-reference omnidirectional image quality assessment
Deyang Liu, Lifei Wan, Xiaofei Zhou 0003, Caifeng Shan |
Signal Process. Image Commun. | 5 |
| 2026 | Degradation-Aware Blind Light-Field Image Quality Assessment With Linear AttentionabstractBlind Light-Field Image Quality Assessment (LFIQA) is challenging, as degradations are microlens-dependent and spatially non-uniform, while perceptual quality relies on both spatial fidelity and angular consistency. However, many existing methods either assume globally stationary distortions or adopt global pooling or self-attention, which can be biased by locally corrupted lenslets and become computationally prohibitive when modeling long-range spatial-angular dependencies. Therefore, in this letter, we present a degradation-aware framework that first predicts a microlens reliability map to make quality inference robust to spatially non-uniform, lenslet-varying corruption. It then extracts spatial and angular features, applies a shared Receptance Weighted Key Value (RWKV) module for linear-time long-range context, and fuses them to predict perceptual quality. Experiments show a higher correlation with subjective ratings with competitive efficiency. Youzhi Zhang 0004, Jianyu Qian, Deyang Liu, Xiaofei Zhou 0003, Hongbin Zha, Caifeng Shan |
IEEE Signal Process. Lett. | 6 |
| 2026 | Efficient Detection Transformer Based on Learnable Query Queuing and Encoder Self-DistillationabstractDetection Transformers (DETRs) have achieved notable performance and outperformed traditional deep neural network models in object detection. However, the self-attention of encoder tokens incurs significant computational overhead. To improve computational efficiency and maintain detection precision simultaneously, we propose an efficient detection transformer model based on learnable query queuing and encoder self-distillation, QD-DETR. Specifically, we first design a query queuing mechanism to select object-relevant queries with the guidance of supervising signals and spatial features. Second, we construct the query representation variational information bottleneck module to optimize the query queuing via mutual information regularization. Lastly, we develop an encoder feature self-distillation method to compensate for the information loss by directly distilling the encoder output token sequences. We conducted experiments on multiple datasets and verified the accuracy and efficiency of QD-DETR in comparison with mainstream detection models. The experiment results demonstrate that QD-DETR reduces computational complexity by 46% and increases inference speed by 69% while maintaining high accuracy in complex object detection scenarios. Fangyu Li 0002, Mohan Niu, Caifeng Shan, Honggui Han |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Learning Implicit and Detail-Enhanced Network for Light Field Image Spatial-Angular Super-ResolutionabstractLight field (LF) imaging holds immense promise for applications such as post-capture refocusing and virtual reality. However, its inherent spatial-angular trade-off significantly limits both spatial and angular resolution, restricting its practicality in real-world scenarios. To address these limitations, spatial-angular super-resolution methods have been proposed to simultaneously enhance both dimensions. Yet, existing methods struggle to fully exploit the intertwined spatial-angular correlations and fail to effectively handle sparsely sampled LFs with low spatial resolution, often leading to cumulative errors during reconstruction. In this paper, we propose an Implicit and Detail-Enhanced Network (IDNet) to overcome these challenges. Our IDNet employs 3D convolution for the joint extraction of spatial and angular information, leveraging their interdependencies for more effective LF reconstruction. Additionally, we introduce an implicit detail restoration module that enhances features while encoding positional information to refine fine details. To overcome the limitations of sparse spatial and angular information on high-detail reconstruction and angular consistency in low-resolution LFs, we design a multi-representation enhancement block. This block enhances features by learning pixel differences across multiple directions in diverse representations, effectively capturing intricate details and complex correlations. Thanks to these designs, our IDNet reconstructs novel views with finer details, effectively learns occlusion relationships, and ensures geometric consistency. Experimental results on benchmark datasets demonstrate its superior quantitative and qualitative performance. The code is publicly available at https://github.com/ldyorchid/IDNet. Deyang Liu, Shizheng Li, Xiaofei Zhou 0003, Zeyu Xiao 0002, Caifeng Shan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | MambaPTP: Exploring the Potential of Mamba for Pedestrian Trajectory PredictionabstractPedestrian Trajectory Prediction (PTP) aims to predict the future trajectory of pedestrians based on a historical trajectory. Transformer-based approaches have demonstrated unparalleled performance for PTP tasks, encoding long-term temporal dependencies and heterogeneous spatial interactions of pedestrians. However, Transformer often involves redundant information and noisy interactions from irrelevant regions by considering all available trajectory features. Recently, the structured state space model, Mamba has been proposed, which captures long-range dependency in sequences with a selective mechanism to filter out redundant information. To further tap into the potential of the novel Mamba architecture for the PTP task, in this paper, we presentMambaPTP, which predicts future trajectories based purely on Mamba mechanisms, to mitigate the noisy interactions of irrelevant trajectory features and avoid repetitive trajectory modeling, while maintaining high-performance trajectory prediction. Specifically, we propose a new Bidirectional Gating Mamba (BGM) module with bidirectional state space models, which leverages the sparse gate mechanism to select informative temporal patterns and spatial interactions. Moreover, we design a Bidirectional Trajectory Alignment (BTA) module towards aligning the predicted trajectory to the ground truth, ensuring that the model to learn the effective sparse feature representation of trajectories. We conduct extensive experiments on several mainstream pedestrian trajectory prediction datasets. The results demonstrate that the proposed MambaPTP achieves competitive performance compared to advanced Transformer-based models. We hope this paper can further inspire research in Mamba for the PTP task, leading to a tighter integration of the Mamba and PTP communities. Shuangqing Zhang, Gangming Zhao, Fan Lyu, Songping Wang, Zhang Zhang 0001, Fang Zhao 0006, Caifeng Shan, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Fast Adversarial Training With Weak-to-Strong Spatial-Temporal Consistency in the Frequency Domain on VideosabstractAdversarial Training (AT) has been shown to significantly enhance adversarial robustness via a min-max optimization approach. However, its effectiveness in video recognition tasks is hampered by two main challenges. First, fast adversarial training for video models remains largely unexplored, which severely impedes its practical applications. Specifically, most video adversarial training methods are computationally costly, with long training times and high expenses. Second, existing methods struggle with the trade-off between clean accuracy and adversarial robustness. To address these challenges, we introduce Video Fast Adversarial Training with Weak-to-Strong consistency (VFAT-WS), the first fast adversarial training method for video data. Specifically, VFAT-WS incorporates the following key designs: First, it integrates a straightforward yet effective temporal frequency augmentation (TF-AUG), and its spatial-temporal enhanced form STF-AUG, along with Fast Gradient Sign Method (FGSM) to boost training efficiency and robustness. Second, it devises a weak-to-strong spatial-temporal consistency regularization, which seamlessly integrates the simple TF-AUG and the more complex STF-AUG. Leveraging the consistency regularization, it steers the learning process from simple to complex augmentations. Both of them work together to achieve a better trade-off between clean accuracy and robustness. Extensive experiments on UCF-101 and HMDB-51 with both CNN and Transformer-based models demonstrate that VFAT-WS achieves great improvements in adversarial robustness and corruption robustness, while accelerating training by nearly 490%. Songping Wang, Yueming Lyu, Xiantao Hu, Ziwen He, Wei Wang 0025, Caifeng Shan, Liang Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2026 | Probabilistic-Based Learning for Joint Light Field Image Compression and Enhancement Under Low-Light ConditionsabstractLight field (LF) imaging has attracted increasing research interest in challenging illumination conditions due to its ability to provide rich spatial and angular cues. However, such data present dual challenges: 1) the inherent multi-view structure introduces substantial data redundancy, creating high demands for efficient compression; 2) the insufficient illumination leads to severe quality degradation, which weakens inter-view consistency and visual perception. To address these coupled factors, we propose a Probabilistic-based learning for joint LF image compression and enhancement under low-light conditions (PrL-LFCE). The framework unifies structure-aware compression and feature enhancement mechanisms by introducing learnable probabilistic modeling into both feature coupling and latent distribution estimation to adaptively handle the uncertainty induced by illumination degradation and compression-related information loss. Specifically, we design a probability-based multi-directional feature coupling module that dynamically balances structural preservation and redundancy reduction across multiple directionally arranged sub-aperture images. Moreover, we introduce a swin-gated enhancement module that suppresses noise and highlights structurally salient regions in compression-aware feature representations through attention-guided gating. Extensive experiments show that PrL-LFCE consistently outperforms state-of-the-art methods, achieving at least 34.86% bitrate savings while maintaining excellent visual quality, demonstrating a strong joint compression and enhancement capability. Deyang Liu, Jimin Wang, Mounir Kaaniche, Xiaofei Zhou 0003, Gangyi Jiang, Caifeng Shan |
IEEE Trans. Image Process. | 7 |
| 2026 | Camera-Based Dual-Wavelength Defocused Speckle Imaging for Multi-Point Seismocardiographic Motion MeasurementabstractContinuous monitoring of cardiac activity is crucial for detecting anomalies such as heart failure and coronary artery disease, and it can alleviate the burden of cardiovascular disease on healthcare systems. This study introduces a novel concept for contactless monitoring of multi-point cardiac motion using dual-wavelength defocused speckle imaging (DW-DSI). A prototype system was developed to measure multi-point seismocardiography (MP-SCG) signals from the atrial and ventricular regions. In addition, blood pressure (BP) monitoring was demonstrated as a proof of concept using the time delay between atrial and ventricular motion signals. An experiment involving 19 subjects with ice water stimulation protocol demonstrated that the performance of BP estimation using time delay features of MP-SCG is comparable to BP estimated from ECG-PPG derived pulse arrival time. The results showed that the best performance was achieved using the correlation features extracted from MP-SCG, such as time delay information and heart rate, in combination with an artificial neural network model. The mean absolute error for systolic/diastolic/mean BP are 6.954 mmHg, 5.368 mmHg and 5.415 mmHg, with Pearson correlation coefficient of 0.639, 0.559, and 0.517. This demonstrates the potential of the camera-based DW-DSI system for measuring MP-SCG and the feasibility towards continuous BP monitoring. Caifeng Shan, Wenjin Wang 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | Facial Privacy Protection for Remote PhotoplethysmographyabstractRemote photoplethysmography (rPPG) has emerged as a crucial technology for contactless health monitoring, providing a convenient and non-invasive method to measure physiological signals from skin videos. Because face videos are commonly used for rPPG measurements, privacy concerns arise due to the inherent sensitivity of facial biometric data. Concerns about privacy breaches in facial video recordings have hindered telemedicine advancements and limited the creation of large-scale medical datasets, restricting the development of rPPG-based technologies. Additionally, the necessity to transmit and store rPPG videos in these applications necessitates video compression as an indispensable step. However, existing facial privacy protection techniques and video compression methods tend to degrade the rPPG signal in videos. To address these challenges, this study proposes a straightforward yet effective face anonymization module-a plug-and-play component employing spatial pixel redistribution algorithms to achieve: 1) eliminating identifiable biometric features while preserving the physiological information; 2) facilitating video compression by a macroblock reassembly strategy based on chromaticity clustering. Experiments on three rPPG datasets illustrate that the proposed method preserves physiological information in anonymized videos while effectively facilitating video compression. Jieying Wang, Caifeng Shan, Shuwang Zhou, Minglei Shu |
IEEE J. Biomed. Health Informatics | 2 |
| 2026 | PSO-HEAD: Pseudo-Supervision Guided Spatial Optimization for View-Consistent 3D Full-Head ReconstructionabstractGenerative 3D human head reconstruction in \(360^{\circ}\) is attracting increasing attention because of its flexibility in downstream animation applications. Existing generative 3D head synthesis approaches are primarily limited to near-frontal face priors, which cause distorted artifacts at large view angles. In this article, we introduce a novel Pseudo-Supervision Guided Spatial Optimization (PSO-HEAD) framework that reconstructs 3D view-consistent full-head through explicitly introducing pseudo-label of back-head supervision for spatial texture and geometric optimization. Particularly, our PSO-HEAD introduces two key improvements, i.e., Pseudo-Supervision Augmented Inversion (PSA-Inversion) and Full-Head Aware Generative Enhancement (FAGE). PSA-Inversion augments plausible invisible back-head as pseudo-supervision to optimize the view-hallucinated latent code conditioned on the augmented camera poses via GAN inversion, enforcing 3D spatial consistency across both visible and invisible regions. Furthermore, FAGE fine-tunes the 3D GAN on a proposed auxiliary FK-Enhance dataset deriving from either generated or real-world high-quality back-head images, which therefore improves the generalization of our PSO-HEAD to diverse hairstyles or underrepresented regions. Benefiting from the improvements, our PSO-HEAD enables efficient \(360^{\circ}\) view-consistent full-head generation from single input images, particularly improving reconstruction fidelity of unobserved regions, which quantitatively and qualitatively outperforms the state-of-the-art methods. Peng Zhang 0057, Chongxin Liang, Caifeng Shan |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisabstractIn the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The potential of MLLMs to process sequential visual data is still insufficiently explored, highlighting the lack of a comprehensive, high-quality assessment of their performance. In this paper, we introduce Video-MME, the first-ever full-spectrum, Multi-Modal Evaluation benchmark of MLLMs in Video analysis. Our work distinguishes from existing benchmarks through four key features: 1) Diversity in video types, spanning 6 primary visual domains with 30 subfields to ensure broad scenario generalizability; 2) Duration in temporal dimension, encompassing both short-, medium-, and long-term videos, ranging from 11 seconds to 1 hour, for robust contextual dynamics; 3) Breadth in data modalities, integrating multi-modal inputs besides video frames, including subtitles and audios, to unveil the all-round capabilities of MLLMs; 4) Quality in annotations, utilizing rigorous manual labeling by expert annotators to facilitate precise and reliable model assessment. With Video-MME, we extensively evaluate various state-of-the-art MLLMs, and reveal that Gemini 1.5 Pro is the best-performing commercial model, significantly outperforming the open-source models with an average accuracy of 75%, compared to 71.9% for GPT-4o. The results also demonstrate that Video-MME is a universal benchmark that applies to both image and video MLLMs. Further analysis indicates that subtitle and audio information could significantly enhance video understanding. Besides, a decline in MLLM performance is observed as video duration increases for all models. Our dataset along with these findings underscores the need for further improvements in handling longer sequences and multi-modal data, shedding light on future MLLM development. Project page: https://video-mme.github.io. Chaoyou Fu, Yuhan Dai, Yongdong Luo, Shuhuai Ren, Renrui Zhang, Yunhang Shen, Mengdan Zhang, Peixian Chen, Shaohui Lin, Sirui Zhao, Ke Li 0015, Tong Xu 0001, Xiawu Zheng, Enhong Chen, Caifeng Shan, Ran He 0001, Xing Sun 0001 |
CVPR | 19 |
| 2025 | Region-Aware Compositional Context Prompting for Zero-Shot Anomaly DetectionabstractZero-shot anomaly detection (ZSAD) aims to identify anomalies of unseen classes without requiring samples from those classes. Existing methods typically rely on pre-trained visual language models, such as CLIP, to detect anomalies by designing or learning generic text prompts and computing similarities with image features, which often fail to address the complexity and novelty of anomaly patterns, especially when the target domain exhibits significant differences from the source domain. To address the problems, we propose Region-aware Compositional Context Prompting (ReCo-CoP) for ZSAD, which dynamically generates contextual prompts by integrating both global and local visual information. Specifically, we introduce a Compositional Context Prompting (CCP) module that incorporates global visual features into the context through a set of basis vectors shared among images, and a Regional Context Prompting (RCP) module that optimizes the context based on image patch features, thereby enhancing the model’s ability to perceive local abnormal regions. Additionally, we combine dynamically generated prompts with static generic prompts to prevent the model from losing the essential general knowledge. Extensive experiments on 12 datasets from industrial and medical domains demonstrate the superior zero-shot detection performance of our model. The code is available at https://github.com/WenDongyp/ReCoCoP Guanglei Chu, Guosen Xie, Caifeng Shan, Fang Zhao 0006 |
ECAI | 5 |
| 2025 | Uncertainty-Aware Multimodal MRI Fusion for HIV-Associated Asymptomatic Neurocognitive Impairment Prediction
Zige Chen, Wei Wang 0411, Zhongkai Zhou, Yuqi Fang, Caifeng Shan |
MICCAI (15) | 7 |
| 2025 | Query-Level Alignment for End-to-End Lesion Detection with Human Gaze
Yan Kong, Zhixiang Peng, Yonghao Li, Jiangdong Cai, Sheng Wang 0014, Qian Wang 0001, Yuqi Fang, Caifeng Shan |
MICCAI (13) | 9 |
| 2025 | CLIP-DSA: Textual Knowledge-Guided Cerebrovascular Diseases Recognition in Multi-view Digital Subtraction Angiography
Qihang Xie, Dan Zhang 0026, Ruisheng Su, Caifeng Shan, Jiong Zhang 0004 |
MICCAI (6) | 6 |
| 2025 | AnomalyControl: Highly-Aligned Anomalous Image Generation with Controlled Diffusion ModelabstractIn industrial scenarios, diverse anomalous images are difficult to acquire, significantly limiting the performance of industrial anomaly detection methods. Automatically generating anomalous images for anomaly detection has the potential to solve the above problem. However, existing anomaly generation models are still not satisfactory regarding the authenticity and controllability of anomaly generation. In this paper, we propose a controlled anomaly generation model named AnomalyControl to generate realistic anomalous images aligned highly with both text prompts and anomaly masks. First, we introduce a CLIP-guided anomaly prompt generator that leverages a CLIP text encoder to find anomaly text prompts most aligned with real anomalous images. Secondly, we propose an anomaly appearance and shape decoupling mechanism, which designs an embedding similarity loss to enforce the alignment between the anomaly text prompt and anomalies generated with different shapes at the same location, making the appearance of generated anomalies better maintain semantic consistency when the anomaly shape changes. Then, a training-free local control enhancement strategy is employed to provide stronger control intensity to anomaly regions during inference for finer alignment with anomaly masks. Finally, a hard sample generation module is proposed to create anomalous samples with subtle shapes and imperceptible anomaly appearances, enabling the downstream anomaly detection model to focus on learning low-saliency anomaly features. Extensive experiments demonstrate that anomalous images generated by our model outperform the state-of-the-art anomaly generation methods in terms of authenticity and consistency, and can significantly improve the performance of downstream anomaly detection tasks, especially anomaly localization. Yuanyi Duan, Qinlong Wu, Guosen Xie, Fang Zhao 0006, Caifeng Shan |
ACM Multimedia | 6 |
| 2025 | MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language ModelsabstractMultimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to fully reflect the performance of MLLM, lacking a comprehensive evaluation. In this paper, we fill in this blank, presenting the first comprehensive MLLM Evaluation benchmark MME. It measures both perception and cognition abilities on a total of 14 subtasks. In order to avoid data leakage that may arise from direct use of public datasets for evaluation, the annotations of instruction-answer pairs are all manually designed. The concise instruction design allows us to fairly compare MLLMs, instead of struggling in prompt engineering. Besides, with such an instruction, we can also easily carry out quantitative statistics. A total of 30 advanced MLLMs are comprehensively evaluated on our MME, which not only suggests that existing MLLMs still have a large room for improvement, but also reveals the potential directions for the subsequent model optimization. The data are released at the project page: https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation. Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Jinrui Yang, Xiawu Zheng, Ke Li 0015, Xing Sun 0001, Yunsheng Wu, Rongrong Ji, Caifeng Shan, Ran He 0001 |
NeurIPS | 13 |
| 2025 | VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionabstractRecent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction. However, speech plays a crucial role in multimodal dialogue systems, and implementing high-performance in both vision and speech tasks remains a challenge due to the fundamental modality differences. In this paper, we propose a carefully designed multi-stage training methodology that progressively trains LLM to understand both visual and speech information, ultimately enabling fluent vision and speech interaction. Our approach not only preserves strong vision-language capacity, but also enables efficient speech-to-speech dialogue capabilities without separate ASR and TTS modules, significantly accelerating multimodal end-to-end response speed. By comparing against state-of-the-art counterparts across benchmarks for image, video, and speech, we demonstrate that our omni model is equipped with both strong visual and speech capabilities, making omni understanding and interaction. Chaoyou Fu, Haojia Lin, Yifan Zhang 0004, Yunhang Shen, Haoyu Cao 0001, Zuwei Long, Heting Gao, Ke Li 0015, Xiawu Zheng, Rongrong Ji, Xing Sun 0001, Caifeng Shan, Ran He 0001 |
NeurIPS | 15 |
| 2025 | GOOD: Training-Free Guided Diffusion Sampling for Out-of-Distribution DetectionabstractRecent advancements have explored text-to-image diffusion models for synthesizing out-of-distribution (OOD) samples, substantially enhancing the performance of OOD detection. However, existing approaches typically rely on perturbing text-conditioned embeddings, resulting in semantic instability and insufficient shift diversity, which limit generalization to realistic OOD. To address these challenges, we propose GOOD, a novel and flexible framework that directly guides diffusion sampling trajectories towards OOD regions using off-the-shelf in-distribution (ID) classifiers. GOOD incorporates dual-level guidance: (1) Image-level guidance based on the gradient of log partition to reduce input likelihood, drives samples toward low-density regions in pixel space. (2) Feature-level guidance, derived from k-NN distance in the classifier’s latent space, promotes sampling in feature-sparse regions. Hence, this dual-guidance design enables more controllable and diverse OOD sample generation. Additionally, we introduce a unified OOD score that adaptively combines image and feature discrepancies, enhancing detection robustness. We perform thorough quantitative and qualitative analyses to evaluate the effectiveness of GOOD, demonstrating that training with samples generated by GOOD can notably enhance OOD detection performance. Jiyao Liu, Yueming Lyu, Jianxiong Gao, Weichen Yu, Ningsheng Xu, Liang Wang 0001, Caifeng Shan, Ziwei Liu 0002, Chenyang Si |
NeurIPS | 9 |
| 2025 | Gait recognition via View-aware Part-wise Attention and Multi-scale Dilated Temporal Extractor
Xu Song, Yang Wang 0103, Yan Huang 0023, Caifeng Shan |
Image Vis. Comput. | 4 |
| 2025 | Effective multi-view representation learning for single-view attributed graph clustering
Heng Liu 0002, Weizhi Zhao, Zhou Bao, Mingquan Ye, Caifeng Shan |
Knowl. Based Syst. | 5 |
| 2025 | Video-based adaptive respiratory rate monitoring for clinical applications
Zhiqin Zhou, Caifeng Shan |
Mach. Vis. Appl. | 2 |
| 2025 | L3FMamba: Low-Light Light Field Image Enhancement With Prior-Injected State Space ModelsabstractIn this paper, we address the problem of low-light light field (LF) image enhancement, where spatial details and angular coherence are severely degraded due to noise and insufficient illumination. Existing methods often rely on local aggregation or naive view stacking, which fail to capture global illumination and long-range spatial-angular correlations. To overcome these limitations, we propose L3FMamba, a lightweight enhancement method that integrates Retinex and Atmospheric Scattering models with dark, bright, and average channel priors for robust illumination decomposition. Moreover, we incorporate a state space model to capture non-local spatial-angular dependencies, enabling effective propagation of global context across views. By combining physics-inspired priors with structured modeling, L3FMamba achieves accurate illumination correction and fine-detail preservation with minimal parameters. Experiments show that L3FMamba outperforms the state-of-the-art in quality. Deyang Liu, Shizheng Li, Zeyu Xiao 0002, Ping An 0001, Caifeng Shan |
IEEE Signal Process. Lett. | 5 |
| 2025 | Bi-Directional and Triangular Circulation Fusion Neural Networks for Small Object DetectionabstractDeep learning-driven object detection models are capable of accurately identifying and localizing objects. However, small objects contain limited information relative to global features, resulting in the fact that detection models often do not learn small object features adequately. To enhance the precision in detecting small objects, we propose a bi-directional and triangular circulation fusion neural network (BTFN). First, to selectively strengthen the position features of small objects, we propose a feature circulation extraction module composed of a bi-directional triangular densely nested convolutional network (BTF), thus achieving repetitive multi-layer feature fusion. Second, to fill up the semantic gaps between different scales of features, we design a mixed dual attention module (MDA) in the bi-directional triangular densely nested network. Third, to mitigate the lost information in the neural networks with deep layers as well as improve the inference time, we design a re-parameterization bi-directional composite feature fusion module (Rep-BFM) that fuses the features of multiple scales. The proposed model is evaluated extensively on the MS COCO, Tsinghua-Tencent 100k, and Haier dismantled parts of used home appliances datasets. The experiment results show that the proposed model improves the AP on MS COCO by 4%, especially the APS of small objects is improved by 7.7% compared with SOTA models. Fangyu Li 0002, Junzhu Duan, Qiyu Zhang, Caifeng Shan, Honggui Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Corruption-Invariant Person Re-Identification via Coarse-to-Fine Feature AlignmentabstractCorruption-invariant Person Re-identification (CI-ReID) aims to build robust identity correspondence across non-overlapped cameras even when severe image corruptions occur. It is challenging as those corruptions contaminate intrinsic pedestrian characteristics and cause semantic misalignment in feature space. To address this issue, this paper proposes a coarse-to-fine semantic alignment framework that learns corruption-invariant pedestrian features for re-identification from the perspective of multi-modal feature alignment. In this framework, a Coarse-to-Fine Feature Alignment Transformer (CFAT) is introduced to extract and align features of pedestrian images with different corruptions. Specifically, the CFAT aligns features of corrupted samples to that of the corresponding clean samples in a knowledge distillation manner in the coarse alignment stage, i.e., a teacher network distils identity-related semantics from clean samples and supervises the student network learning semantic-consistent features from corrupted samples. To avoid information loss of the strict alignment, we propose to integrate a Bridge Feature Generation (BFG) module into CFAT to construct meaningful latent structures among modalities in the fine alignment stage. This enables seamless alignment of the same identity between corrupted and clean modalities, leading to better re-identification performance. To evaluate the effectiveness of the proposed method, extensive experiments are conducted on three public benchmark datasets, i.e., Market-1501, CUHK-03, and MSMT-17. The experimental results demonstrate our CFAT outputs state-of-the-arts with a large margin in various corrupted scenes. Peng Zhang 0057, Caifeng Shan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Deep Sparse-to-Dense Inbetweening for Multi-View Light FieldsabstractLight field (LF) imaging, which captures both intensity and directional information of light rays, extends the capabilities of traditional imaging techniques. In this paper, we introduce a task in the field of LF imaging, sparse-to-dense inbetweening, which focuses on generating dense novel views from sparse multi-view LFs. By synthesizing intermediate views from sparse inputs, this task enhances LF view synthesis through filling in interperspective gaps within an expanded field of view and increasing data robustness by leveraging complementary information between light rays from different perspectives, which are limited by non-robust single-view synthesis and the inability to handle sparse inputs effectively. To address these challenges, we construct a high-quality multi-view LF dataset, consisting of 60 indoor scenes and 59 outdoor scenes. Building upon this dataset, we propose a baseline method. Specifically, we introduce an adaptive alignment module to dynamically align information by capturing relative displacements. Next, we explore angular consistency and hierarchical information using a multi-level feature decoupling module. Finally, a multi-level feature refinement module is applied to enhance features and facilitate reconstruction. Additionally, we introduce a universally applicable artifact-aware loss function to effectively suppress visual artifacts. Experimental results demonstrate that our method outperforms existing approaches, establishing a benchmark for sparse-to-dense inbetweening. The code is available at https://github.com/Starmao1/MutiLF. Zeyu Xiao 0002, Ping An 0001, Deyang Liu, Caifeng Shan |
IEEE Trans. Image Process. | 5 |
| 2025 | Cardiac 3D Motion Reconstruction Using Dual-Camera Defocused Speckle Imaging With Multi-Scale AmplificationabstractCardiovascular diseases are one of the leading causes of death worldwide. Accurately capturing and analyzing the multidimensional dynamics of cardiac motion is crucial for early diagnosis and rehabilitation assessment. This study introduces a novel concept for non-contact cardiac linear vibration (SCG) and rotational components (GCGx and GCGy) decoupling and reconstruction by integrating speckle motion signals captured from two cameras with different defocus levels. The intention is to overcome the motion coupling issues inherent in single-camera imaging and improve the accuracy in characterizing the cardiac complex 3D mechanical behavior. Using a sternum-mounted inertial sensor as the reference, experiments were conducted on 42 subjects in laboratory and intensive care unit settings. The results show that the reconstructed cardiac 3D motion signals exhibit greater waveform similarity to the reference signal than the raw speckle motion signal from a single camera, with similarity indices above 87.471%. In addition, with an 8 ms tolerance error, the localization accuracy of 6 key biomarkers (aortic valve opening/closing (AO/AC), mitral valve opening/closing (MO/MC), the biomarkers corresponding to the AO event in GCGy and the MC event in GCGx) are 73.080%, 99.998%, 85.587%, 86.617%, 99.683% and 77.301%, respectively. These results also outperform those obtained from the raw speckle motion signal. These findings validate the rationale and effectiveness of using dual-camera imaging with different defocus levels to reconstruct SCG, GCGx, and GCGy, offering a promising approach for accurately capturing complex cardiac 3D motion and improving cardiac function assessment. Haimiao Mo, Yingen Zhu, Yifei Da, Caifeng Shan, Wenjin Wang 0002 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | $\text{MR}^{2}$-Net: Retinal OCTA Image Stitching via Multi-Scale Representation Learning and Dynamic Location GuidanceabstractOptical coherence tomography angiography (OCTA) plays a crucial role in quantifying and analyzing retinal vascular diseases. However, the limited field of view (FOV) inherent in most commercial OCTA imaging systems poses a significant challenge for clinicians, restricting the possibility to analyze larger retinal regions of high resolution. Automatic stitching of OCTA scans in adjacent regions may provide a promising solution to extend the region of interest. However, commonly-used stitching algorithms face difficulties in achieving effective alignment due to noise, artifacts and dense vasculature present in OCTA images. To address these challenges, we propose a novel retinal OCTA image stitching network, named -Net, which integrates multi-scale representation learning and dynamic location guidance. In the first stage, an image registration network with a progressive multi-resolution feature fusion is proposed to derive deep semantic information effectively. Additionally, we introduce a dynamic guidance strategy to locate the foveal avascular zone (FAZ) and constrain registration errors in overlapping vascular regions. In the second stage, an image fusion network based on multiple mask constraints and adjacent image aggregation (AIA) strategies is developed to further eliminate the artifacts in the overlapping areas of stitched images, thereby achieving precise vessel alignment. To validate the effectiveness of our method, we conduct a series of experiments on two delicately constructed datasets, i.e., OPTOVUE-OCTA and SVision-OCTA. Experimental results demonstrate that our method outperforms other image stitching methods and effectively generates high-quality wide-field OCTA images, achieving a structural similarity index (SSIM) score of 0.8264 and 0.8014 on the two datasets, respectively. Haiting Mao, Yuhui Ma, Dan Zhang 0026, Yanda Meng, Shaodong Ma, Yuchuan Qiao, Huazhu Fu, Caifeng Shan, Da Chen 0002, Yitian Zhao, Jiong Zhang 0004 |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | Physiological Information Preserving Video Compression for rPPGabstractRemote photoplethysmography (rPPG) has recently attracted much attention due to its non-contact measurement convenience and great potential in health care and computer vision applications. Early rPPG studies were mostly developed on self-collected uncompressed video data, which limited their application in scenarios that require long-distance real-time video transmission, and also hindered the generation of large-scale publicly available benchmark datasets. In recent years, with the popularization of high-definition video and the rise of telemedicine, the pressure of storage and real-time video transmission under limited bandwidth have made the compression of rPPG video inevitable. However, video compression can adversely affect rPPG measurements. This is due to the fact that conventional video compression algorithms are not specifically proposed to preserve physiological signals. Based on this, we propose a video compression scheme specifically designed for rPPG application. The proposed approach consists of three main strategies: 1) facial ROI-based computational resource reallocation; 2) rPPG signal preserving bit resource reallocation; and 3) temporal domain up- and down-sampling coding. UBFC-rPPG, ECG-Fitness, and a self-collected dataset are used to evaluate the performance of the proposed method. The results demonstrate that the proposed method can preserve almost all physiological information after compressing the original video to 1/60 of its original size. The proposed method is expected to promote the development of telemedicine and deep learning techniques relying on large-scale datasets in the field of rPPG measurement. Jieying Wang, Caifeng Shan, Shuwang Zhou, Minglei Shu |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | DSCA: A Digital Subtraction Angiography Sequence Dataset and Spatio-Temporal Model for Cerebral Artery SegmentationabstractCerebrovascular diseases (CVDs) remain a leading cause of global disability and mortality. Digital Subtraction Angiography (DSA) sequences, recognized as the gold standard for diagnosing CVDs, can clearly visualize the dynamic flow and reveal pathological conditions within the cerebrovasculature. Therefore, precise segmentation of cerebral arteries (CAs) and classification between their main trunks and branches are crucial for physicians to accurately quantify diseases. However, achieving accurate CA segmentation in DSA sequences remains a challenging task due to small vessels with low contrast, and ambiguity between vessels and residual skull structures. Moreover, the lack of publicly available datasets limits exploration in the field. In this paper, we introduce a DSA Sequence-based Cerebral Artery segmentation dataset (DSCA), the publicly accessible dataset designed specifically for pixel-level semantic segmentation of CAs. Additionally, we propose DSANet, a spatio-temporal network for CA segmentation in DSA sequences. Unlike existing DSA segmentation methods that focus only on a single frame, the proposed DSANet introduces a separate temporal encoding branch to capture dynamic vessel details across multiple frames. To enhance small vessel segmentation and improve vessel connectivity, we design a novel TemporalFormer module to capture global context and correlations among sequential frames. Furthermore, we develop a Spatio-Temporal Fusion (STF) module to effectively integrate spatial and temporal features from the encoder. Extensive experiments demonstrate that DSANet outperforms other state-of-the-art methods in CA segmentation, achieving a Dice of 0.9033. Jiong Zhang 0004, Qihang Xie, Lei Mou, Dan Zhang 0026, Da Chen 0002, Caifeng Shan, Yitian Zhao, Ruisheng Su, Mengguo Guo |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Biphasic Face Photo-Sketch Synthesis via Semantic-Driven Generative Adversarial Network With Graph Representation LearningabstractBiphasic face photo-sketch synthesis has significant practical value in wide-ranging fields such as digital entertainment and law enforcement. Previous approaches directly generate the photo-sketch in a global view, they always suffer from the low quality of sketches and complex photograph variations, leading to unnatural and low-fidelity results. In this article, we propose a novel semantic-driven generative adversarial network to address the above issues, cooperating with graph representation learning. Considering that human faces have distinct spatial structures, we first inject class-wise semantic layouts into the generator to provide style-based spatial information for synthesized face photographs and sketches. In addition, to enhance the authenticity of details in generated faces, we construct two types of representational graphs via semantic parsing maps upon input faces, dubbed the intraclass semantic graph (IASG) and the interclass structure graph (IRSG). Specifically, the IASG effectively models the intraclass semantic correlations of each facial semantic component, thus producing realistic facial details. To preserve the generated faces being more structure-coordinated, the IRSG models interclass structural relations among every facial component by graph representation learning. To further enhance the perceptual quality of synthesized images, we present a biphasic interactive cycle training strategy by fully taking advantage of the multilevel feature consistency between the photograph and sketch. Extensive experiments demonstrate that our method outperforms the state-of-the-art competitors on the CUHK Face Sketch (CUFS) and CUHK Face Sketch FERET (CUFSF) datasets. Xingqun Qi, Muyi Sun, Zijian Wang 0009, Jiaming Liu 0003, Qi Li 0005, Fang Zhao 0006, Shanghang Zhang, Caifeng Shan |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | MPMNet: Modal Prior Mutual-Support Network for Age-Related Macular Degeneration Classification
Huaying Hao, Dan Zhang 0026, Huazhu Fu, Caifeng Shan, Yitian Zhao, Jiong Zhang 0004 |
MICCAI (1) | 6 |
| 2024 | FedGMKD: An Efficient Prototype Federated Learning Framework through Knowledge Distillation and Discrepancy-Aware AggregationabstractFederated Learning (FL) faces significant challenges due to data heterogeneity across distributed clients. To address this, we propose FedGMKD, a novel framework that combines knowledge distillation and differential aggregation for efficient prototype-based personalized FL without the need for public datasets or server-side generative models. FedGMKD introduces Cluster Knowledge Fusion, utilizing Gaussian Mixture Models to generate prototype features and soft predictions on the client side, enabling effective knowledge distillation while preserving data privacy. Additionally, we implement a Discrepancy-Aware Aggregation Technique that weights client contributions based on data quality and quantity, enhancing the global model's generalization across diverse client distributions. Theoretical analysis confirms the convergence of FedGMKD. Extensive experiments on benchmark datasets, including SVHN, CIFAR-10, and CIFAR-100, demonstrate that FedGMKD outperforms state-of-the-art methods, significantly improving both local and global accuracy in non-IID data settings. Jianqiao Zhang 0002, Caifeng Shan, Jungong Han |
NeurIPS | 2 |
| 2024 | Enhancing the utilization of uncertain pixels in semi-supervised semantic segmentation
Xingfang Chang, Changrui Chen, Caifeng Shan |
Neurocomputing | 3 |
| 2024 | Vision transformers are active learners for image copy detectionabstractImage Copy Detection (ICD) is developed to identify and track duplicated or manipulated images. The majority of existing methods rely on Convolutional Neural Networks (CNNs) and are trained using unsupervised learning techniques, which leads to subpar performance. We discover that by carefully designing the training process, Vision Transformer (ViT) backbones yield superior results. Specifically, directly training a ViT for ICD often leads to overfitting on the training images, which in turn results in poor generalization to unseen (test) images. Consequently, we initially train a CNN (such as ResNet-50), and during the ViT training, the distances between the features of CNN and ViT are regularized. We also incorporate an active learning method to further enhance performance. Notably, due to the visual discrepancy between auto-generated transformations and those used in the query set, we incorporate a small number (approximately 0.5% of unlabeled training images) of manually produced and labeled positive pairs. Training models on these pairs results in a significant performance boost though with little cost. Experimental findings demonstrate the effectiveness of our approach, and our method achieves state-of-the-art performance. Our code is available at: https://github.com/WangWenhao0716/ViT4ICD. Zhentao Tan, Caifeng Shan |
Neurocomputing | 3 |
| 2024 | Camera-based physiological measurement: Recent advances and future prospects
Jieying Wang, Caifeng Shan, Zongshen Hou |
Neurocomputing | 2 |
| 2024 | BSANet: Boundary-aware and scale-aggregation networks for CMR image segmentation
Dan Zhang 0026, Chenggang Lu, Tao Tan 0002, Behdad Dashtbozorg, Xi Long 0001, Xiayu Xu, Jiong Zhang 0004, Caifeng Shan |
Neurocomputing | 8 |
| 2024 | Boosting Micro-Expression Recognition via Self-Expression Reconstruction and Memory Contrastive LearningabstractMicro-expression (ME) is an instinctive reaction that is not controlled by thoughts. It reveals one's inner feelings, which is significant in sentiment analysis and lie detection. Since micro-expression is expressed as subtle facial changes within particular facial action units, learning discriminative and generalized features for Micro-expression Recognition (MER) is challenging. To achieve the purpose, this paper proposes a novel MER framework that simultaneously integrates supervised Prototype-based Memory Contrastive Learning (PMCL) for discriminative feature mining and adds Self-expression Reconstruction (SER) as an auxiliary task and regularization for better generalization. In particular, the proposed SER module is forced as a regularization by reconstructing input ME from the randomly dropped patch- wise features in the bottleneck. And, the PMCL module globally compares historical and current cluster agents learned from training instances to enhance intra-class compactness and inter-class separability. Extensive experiments are conducted on three benchmarks, e.g., SMIC, CASME II, and SAMM, under evaluation criteria of both Composite Database Evaluation (CDE) and Single Database Evaluation (SDE) protocols. The results show our method surpasses other state-of-the-art approaches under various evaluation metrics, achieving overall 86.30% unweighed F1-score and 88.30% unweighed average recall on the composite dataset. Furthermore, the ablation studies verify the effectiveness of our SER for better generalization and PMCL for better discrimination in learning feature representation from limited micro-expression samples. Yongtang Bao, Peng Zhang 0057, Caifeng Shan, Xianye Ben |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Spatial Attention-Guided Light Field Salient Object Detection Network With Implicit Neural RepresentationabstractRecently, many Light Field Salient Object Detection (LF SOD) methods have been proposed. However, guaranteeing the integrality and recovering more high-frequency details of the generated salient object map still remain challenging. To this end, we propose a spatial attention-guided LF SOD network with implicit neural representation to further improve LF SOD performance. We adopt an encoder-decoder structure for model construction. In order to ensure the completeness of the generated salient object map, a multi-modal and multi-scale feature fusion module is designed in the encoder part to refine the salient regions within all-in-focus image and aggregate the focal stack and all-in-focus image in spatial attention-guided manner. In order to recover more high-frequency details of the obtained salient object map, an implicit detail restoration module is proposed in the decoder part. In virtue of implicit neural representation, we convert the detail restoration problem into a functional mapping problem. By further integrating the self-attention mechanism, the derived saliency map can be depicted at a more refined level. Comprehensive experimental results demonstrate the superiority of the proposed method. Ablation studies and visual comparisons further validate that the proposed method can guarantee the integrality and recover more high-frequency detail information of the obtained saliency map. The code is publicly available athttps://github.com/ldyorchid/LFSOD-Net. Xin Zheng 0006, Zhengqu Li, Deyang Liu, Xiaofei Zhou 0003, Caifeng Shan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Gait Attribute Recognition: A New Benchmark for Learning Richer Attributes From Human Gait PatternsabstractCompared to gait recognition, Gait Attribute Recognition (GAR) is a seldom-investigated problem. However, since gait attribute recognition can provide richer and finer semantic descriptions, it is an indispensable part of building intelligent gait analysis systems. Nonetheless, the types of attributes considered in the existing datasets are very limited. This paper contributes a new benchmark dataset for gait attribute recognition named Multi-Attribute Gait (MA-Gait). Our MA-Gait contains 95 subjects recorded from 12 camera views, resulting in more than 13000 sequences, with 16 attributes labeled, including six attributes that have never been considered in the literature. Moreover, we propose a Multi-Scale Motion Encoder (MSME) to extract robust motion features, and an Attribute-Guided Feature Selection Module (AGFSM) to adaptively capture the most discriminative attribute features from static appearance features and dynamic motion features for different attributes. Our method achieves the best GAR accuracy on the new dataset. Comprehensive experiments show the effectiveness of the proposed method through both quantitative and qualitative evaluations. Xu Song, Saihui Hou, Yan Huang 0023, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Caifeng Shan |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | Meta Clothing Status Calibration for Long-Term Person Re-IdentificationabstractRecent studies have seen significant advancements in the field of long-term person re-identification (LT-reID) through the use of clothing-irrelevant or insensitive features. This work takes the field a step further by addressing a previously unexplored issue, the Clothing Status Distribution Shift (CSDS). CSDS refers to the differing ratios of samples with clothing changes to those without clothing changes between the training and test sets, leading to a decline in LT-reID performance. We establish a connection between the performance of LT-reID and CSDS, and argue that addressing CSDS can improve LT-reID performance. To that end, we propose a novel framework called Meta Clothing Status Calibration (MCSC), which uses meta-learning to optimize the LT-reID model. Specifically, MCSC simulates CSDS between meta-train and meta-test with meta-optimization objectives, optimizing the LT-reID model and making it robust to CSDS. This framework is designed to prevent overfitting and improve the generalization ability of the LT-reID model in the presence of CSDS. Comprehensive evaluations on seven datasets demonstrate that the proposed MCSC framework effectively handles CSDS and improves current state-of-the-art LT-reID methods on several LT-reID benchmarks. Yan Huang 0023, Qiang Wu 0001, Zhang Zhang 0001, Caifeng Shan, Yan Huang 0008, Yi Zhong 0002, Liang Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Camera-Based Seismocardiogram for Heart Rate Variability MonitoringabstractHeart rate variability (HRV) is a crucial metric that quantifies the variation between consecutive heartbeats, serving as a significant indicator of autonomic nervous system (ANS) activity. It has found widespread applications in clinical diagnosis, treatment, and prevention of cardiovascular diseases. In this study, we proposed an optical model for defocused speckle imaging, to simultaneously incorporate out-of-plane translation and rotation-induced motion for highly-sensitive non-contact seismocardiogram (SCG) measurement. Using electrocardiogram (ECG) signals as the gold standard, we evaluated the performance of photoplethysmogram (PPG) signals and speckle-based SCG signals in assessing HRV. The results indicated that the HRV parameters measured from SCG signals extracted from laser speckle videos showed higher consistency with the results obtained from the ECG signals compared to PPG signals. Additionally, we confirmed that even when clothing obstructed the measurement site, the efficacy of SCG signals extracted from the motion of laser speckle patterns persisted in assessing the HRV levels. This demonstrates the robustness of camera-based non-contact SCG in monitoring HRV, highlighting its potential as a reliable, non-contact alternative to traditional contact-PPG sensors. Dongfang Yu, Hongzhou Lu, Caifeng Shan, Wenjin Wang 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Exploring Generalizable Distillation for Efficient Medical Image SegmentationabstractEfficient medical image segmentation aims to provide accurate pixel-wise predictions with a lightweight implementation framework. However, existing lightweight networks generally overlook the generalizability of the cross-domain medical segmentation tasks. In this paper, we propose Generalizable Knowledge Distillation (GKD), a novel framework for enhancing the performance of lightweight networks on cross-domain medical segmentation by generalizable knowledge distillation from powerful teacher networks. Considering the domain gaps between different medical datasets, we propose the Model-Specific Alignment Networks (MSAN) to obtain the domain-invariant representations. Meanwhile, a customized Alignment Consistency Training (ACT) strategy is designed to promote the MSAN training. Based on the domain-invariant vectors in MSAN, we propose two generalizable distillation schemes, Dual Contrastive Graph Distillation (DCGD) and Domain-Invariant Cross Distillation (DICD). In DCGD, two implicit contrastive graphs are designed to model the intra-coupling and inter-coupling semantic correlations. Then, in DICD, the domain-invariant semantic vectors are reconstructed from two networks (i.e., teacher and student) with a crossover manner to achieve simultaneous generalization of lightweight networks, hierarchically. Moreover, a metric named Fréchet Semantic Distance (FSD) is tailored to verify the effectiveness of the regularized domain-invariant features. Extensive experiments conducted on the Liver, Retinal Vessel and Colonoscopy segmentation datasets demonstrate the superiority of our method, in terms of performance and generalization ability on lightweight networks. Xingqun Qi, Zhuojie Wu, Wenxuan Zou, Yifan Gao 0003, Muyi Sun, Shanghang Zhang, Caifeng Shan, Zhenan Sun |
IEEE J. Biomed. Health Informatics | 8 |
| 2024 | Multi-Modal Multi-Slice Cooperative Dual-Domain Cascaded De-Aliasing Network for MR Imaging ReconstructionabstractRecent advancements in Magnetic Resonance Imaging (MRI) reconstruction techniques aim to accelerate the imaging process. However, these methods still face two key limitations. Firstly, although the same location consistently provides anatomical information across different modalities, such as organs, tissues, or lesions, previous studies have predominantly relied on single-modality information, overlooking the potential advantages of incorporating complementary data from other modalities. Secondly, while adjacent MRI slices often capture the same location or organ with similar anatomical structures, only a few methods consider the information from neighboring slices during the reconstruction process. To address these challenges, we propose aMulti-modalMulti-slice cooperativeDual-domain cascaded de-alising network for MR imagingReconstruction (MMDR). Specifically, we design a multi-slice and multi-modal feature fusion network based on 3D convolution and swin transformer that efficiently extracts multi-modal features from MRI. Then, a dual domain cascaded recurrent network through dense-blocks with large receptive fields for fast MRI reconstruction is explored. Extensive experiments on the IXI datasets were carried out to evaluate the proposed method's robustness across varying network structures, under-sampling rates, and sampling patterns. MMDR demonstrates promising performance across both qualitative and quantitative metrics, particularly with a competitive PSNE of 42.15 and an SSIM of 0.984 for T2 reconstruction using 30% T2WI and PDWI, as well as achieving a PSNE of 39.92 and an SSIM of 0.966 for PDWI reconstruction with 30% PDWI and T2WI. Xuebin Sun, Yanwei Pang, Caifeng Shan, Shing Shin Cheng |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Living-Skin Detection Based on Spatio-Temporal Analysis of Structured Light PatternabstractLiving-skin detection is an important step for imaging photoplethysmography and biometric anti-spoofing. In this paper, we propose a new approach that exploits spatio-temporal characteristics of structured light patterns projected on the skin surface for living-skin detection. We observed that due to the interactions between laser photons and tissues inside a multi-layer skin structure, the frequency-domain sharpness feature of laser spots on skin and non-skin surfaces exhibits clear difference. Additionally, the subtle physiological motion of living-skin causes laser interference, leading to brightness fluctuations of laser spots projected on the skin surface. Based on these two observations, we designed a new living-skin detection algorithm to distinguish skin from non-skin using spatio-temporal features of structured laser spots. Experiments in the dark chamber and Neonatal Intensive Care Unit (NICU) demonstrated that the proposed setup and method performed well, achieving a precision of 85.32%, recall of 83.87%, and F1-score of 83.03% averaged over these two scenes. Compared to the approach that only leverages the property of multilayer skin structure, the hybrid approach obtains an averaged improvement of 8.18% in precision, 3.93% in recall, and 8.64% in F1-score. These results validate the efficacy of using frequency domain sharpness and brightness fluctuations to augment the features of living-skin tissues irradiated by structured light, providing a solid basis for structured light based physiological imaging. Chuchu Liao, Liping Pan, Hongzhou Lu, Caifeng Shan, Wenjin Wang 0002 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Guest Editorial Camera-Based Health Monitoring in Real-World ScenariosabstractAt Present, cameras are increasingly used to measure physiological signals from human face and body for contactless health monitoring, thereby eliminating mechanical contact with the skin that are common in wearable sensors. This is an emerging research direction developing rapidly in the last decade and which is now gradually maturing into products for patient monitoring. Advancements in biomedical optics, physiological measurement, computer vision and artificial intelligence (AI) enabled various camera-based measurements, including vital signs like heart rate (HR), respiration rate (RR), oxygen saturation (SpO2), blood pressure (BP), and physiological markers that have diagnostic capabilities, such as the detection of arrhythmia, atrial fibrillation, apnea, hypertension, etc. Image and video analysis also permit the measurement of human semantics, context and behaviours that provide new insights into health informatics (e.g., facial analysis and body actigraphy for the assessment of patient delirium), which is a unique advantage of camera sensors as compared to biomedical sensors, like e.g., photoplethysmography (PPG) and electrocardiogram (ECG). Camera-based health monitoring will bring a rich set of compelling healthcare applications that directly improve upon contact-based monitoring solutions in various scenarios like clinical units including e.g., the intensive care unit (ICU), the neonatal ICU (NICU) or sleep centers, and assisted-living homes (e.g., elderly homes or confinement centers), improving patient care experience and people's quality of life. Wenjin Wang 0002, Caifeng Shan, Steffen Leonhardt, Ramakrishna Mukkamala, Ewa Nowara |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Pedestrian Attribute Recognition via Spatio-temporal Relationship Learning for Visual SurveillanceabstractPedestrian attribute recognition (PAR) aims at predicting the visual attributes of a pedestrian image. PAR has been used as soft biometrics for visual surveillance and IoT security. Most of the current PAR methods are developed based on discrete images. However, it is challenging for the image-based method to handle the occlusion and action-related attributes in real-world applications. Recently, video-based PAR has attracted much attention in order to exploit the temporal cues in the video sequences for better PAR. Unfortunately, existing methods usually ignore the correlations among different attributes and the relations between attributes and spatio regions. To address this problem, we propose a novel method for video-based PAR by exploring the relationships among different attributes in both the spatio and temporal domains. More specifically, a spatio-temporal saliency module (STSM) is introduced to capture the key visual patterns from the video sequences, and a module for spatio-temporal attribute relationship learning (STARL) is proposed to mine the correlations among these patterns. Meanwhile, a large-scale benchmark for video-based PAR, RAP-Video, is built by extending the image-based dataset RAP-2, which contains 83,216 tracklets with 25 scenes. To the best of our knowledge, this is the largest dataset for video-based PAR. Extensive experiments are performed on the proposed benchmark as well as on MARS Attribute and DukeMTMC-Video Attribute. The superior performance demonstrates the effectiveness of the proposed method. Da Li 0003, Zhang Zhang 0001, Peng Zhang 0057, Caifeng Shan, Jungong Han |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Wavelength-Dependency of PPG Morphological Features for Camera-Based Blood Pressure EstimationabstractBlood pressure (BP) monitoring has become a daily necessity for well-being management. Camera-based BP monitoring has attracted much attention due to its comfort and convenience. It extracts photoplethysmographic (PPG) signals from the facial skin remotely and calculates morphological features for BP estimation. Although green light with strong pulsatility is commonly used, its penetration depth in skin tissues is shorter than that of the near infrared (NIR) light, indicating that the NIR light may contain more distinct PPG features (i.e., dicrotic wave) that could be more beneficial for BP estimation. In this study, we investigated the wavelength-dependency of the camera-PPG waveform for BP measurement. In particular, two morphological features (K-value and Augmentation Index) from the narrow-band green (550nm), red (660nm), and NIR (850nm) channels were measured and compared for BP calibration. Experimental data were collected from 20 healthy adult subjects using the ice water stimulation protocol. The results show that the camera-PPG waveform features have clear wavelength-dependency. The NIR channel that has deeper skin penetrability (than the green channel) but higher pulsatility (than the red channel) gives more descriptive morphological features related to BP. The conclusions drawn from this study inspire the selection of wavelength for PPG morphologv-based BP measurement. Guanghang Liao, Hongzhou Lu, Caifeng Shan, Wenjin Wang 0002 |
HealthCom | 4 |
| 2023 | Benchmark of Physiological Model Based and Deep Learning Based Remote Photoplethysmography in Automotive ApplicationsabstractRemote photoplethysmography (rPPG) can be used to monitor driver’s cardio-respiratory functions in automotive for improving the safety of driving. To understand the challenges of rPPG in this application, we created a benchmark of latest rPPG algorithms based on the MR-NIRP Car dataset, selecting the representative methods from both the physiological model based (PBV and DIS) and deep learning based (Supervised Learning and Contrastive Learning) approaches. The experimental results show that the physiological model based methods are generally more robust in this challenging scenario with vigorous motions and dynamic lighting changes, typically DIS outperforms others, with an average MAE of 6.5 bpm on RGB videos and 15.9 bpm on NIR videos. The benchmark indicates that upgrading the single wavelength NIR setup to multi-wavelength is the essential step towards robust heart-rate monitoring in automotive. Xuezhi Yang, Hongzhou Lu, Caifeng Shan, Wenjin Wang 0002 |
ICASSP | 4 |
| 2023 | Dual-focus transfer network for zero-shot learning
Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001, Tieniu Tan |
Neurocomputing | 3 |
| 2023 | Mask-guided modality difference reduction network for RGB-T semantic segmentation
Wenli Liang, Yuanjian Yang, Fangyu Li 0002, Xi Long 0001, Caifeng Shan |
Neurocomputing | 5 |
| 2023 | Constructing Stronger and Faster Baselines for Skeleton-Based Action RecognitionabstractOne essential problem in skeleton-based action recognition is how to extract discriminative features over all skeleton joints. However, the complexity of the recent State-Of-The-Art (SOTA) models for this task tends to be exceedingly sophisticated and over-parameterized. The low efficiency in model training and inference has increased the validation costs of model architectures in large-scale datasets. To address the above issue, recent advanced separable convolutional layers are embedded into an early fused Multiple Input Branches (MIB) network, constructing an efficient Graph Convolutional Network (GCN) baseline for skeleton-based action recognition. In addition, based on such the baseline, we design a compound scaling strategy to expand the model's width and depth synchronously, and eventually obtain a family of efficient GCN baselines with high accuracies and small amounts of trainable parameters, termed EfficientGCN-Bx, where "x" denotes the scaling coefficient. On two large-scale datasets, i.e., NTU RGB+D 60 and 120, the proposed EfficientGCN-B4 baseline outperforms other SOTA methods, e.g., achieving 92.1% accuracy on the cross-subject benchmark of NTU 60 dataset, while being 5.82× smaller and 5.85× faster than MS-G3D, which is one of the SOTA methods. The source code in PyTorch version and the pretrained models are available at https://github.com/yfsong0709/EfficientGCNv1. Yi-Fan Song, Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | SiamON: Siamese Occlusion-Aware Network for Visual TrackingabstractOcclusion has been proven to be one of the most challenging factors faced by most visual trackers. There are mainly two difficulties, the first one is that the number of occlusion samples are very limited even though collecting a large-scale training data set, and another one is how to correctly learn the features of the target when comes to occlusion situations. In this paper, we tried to solve these two problems together in our proposed model. To this end, we propose a novel Siamese Occlusion-aware Network (SiamON) for high-performance visual tracking. In particular, we predefine some soft-masks to solve the problem of fewer occlusion samples, which perceive patterns of occlusion contents at different locations and take these masks as the conditions to guide occlusion-aware feature learning. Meanwhile, we propose a target-aware attention mechanism allows the model to pay more attention to the target and further weaken the impact of occlusion. Extensive experiments on several popular benchmarks show that our tracking method exceeds many state-of-the-art trackers especially in the presence of occlusion and meets the requirements of real-time. Chao Fan 0001, Hongyuan Yu, Yan Huang 0008, Caifeng Shan, Liang Wang 0001, Chenglong Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Incremental Pedestrian Attribute Recognition via Dual Uncertainty-Aware Pseudo-LabelingabstractIncremental pedestrian attribute recognition (IncPAR) aims to learn novel person attributes continuously and avoid the catastrophic forgetting, which is an essential problem for image forensic and security applications, e.g., suspect search. Different from the conventional continual learning for visual classification, we formulate the IncPAR as a problem of multi-label continual learning with incomplete labels (MCL-IL), where the training samples in a novel task are annotated with only a few categories of interest but may implicitly contain other attributes of previous tasks. The incomplete label assignments is a challenging and frequently-encountered issue in real-world multi-label classification applications due to a number of reasons, e.g., incomplete data collection, moderate budget for annotations, etc. To tackle the MCL-IL problem, we propose a self-training based approach via dual uncertainty-aware pseudo-labeling (DUAPL) to transfer the knowledge learned in previous tasks to novel tasks. Specially, both kinds of uncertainties, i.e., aleatoric uncertainty and epistemic uncertainty, are modeled to mitigate the negative influences of noisy pseudo labels induced by low quality samples and immature models learned by inadequate training in early tasks. Based on the DUAPL, more reliable supervision signals can be estimated to prevent the model evolution from forgetting attributes seen in previous tasks. For standard evaluations of MCL-IL methods, two benchmarks on IncPAR, termed RAP-CL and PETA-CL, are constructed by re-organizing public human attribute datasets. Extensive experiments have been performed on these benchmarks to compare the proposed method with multiple baselines. The superior performance in terms of both recognition accuracies and forgetting ratios demonstrate the effectiveness of the proposed DUAPL for IncPAR. Da Li 0003, Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Progressive Sub-Domain Information Mining for Single-Source Generalizable Gait RecognitionabstractRecent years have witnessed the deployment of fully supervised gait recognition. However, due to domain diversity, gait recognition models designed under the fully supervised condition suffer from poor generalization in unseen domains. How to improve the generalization ability of gait recognition models and enhance their performance on unseen domains is still unexplored in existing gait recognition approaches. This paper investigate the generalizable gait recognition problem and proposes a Progressive Sub-domain Information Mining (PSIM) framework for single-source generalizable gait recognition. During training, PSIM can mine sub-domain information from a single large-scale source domain by differentiating gait features extracted from different people through unsupervised clustering. Then, domain information mitigation loss and domain homogenization loss are introduced to regularize those gait features to be domain insensitive. The above procedures is conducted iteratively until the model converges. Our PSIM framework is model-agnostic, which can directly improve the generalization ability of state-of-the-art gait recognition models without bringing too much complexity in model design. In experiments, our model-agnostic PSIM framework is adopted on several gait recognition models to show its effectiveness in boosting gait recognition performance for the single-source generalizable gait recognition task. Yang Wang 0103, Yan Huang 0023, Caifeng Shan, Liang Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Graph Flow: Cross-Layer Graph Flow Distillation for Dual Efficient Medical Image SegmentationabstractWith the development of deep convolutional neural networks, medical image segmentation has achieved a series of breakthroughs in recent years. However, high-performance convolutional neural networks always mean numerous parameters and high computation costs, which will hinder the applications in resource-limited medical scenarios. Meanwhile, the scarceness of large-scale annotated medical image datasets further impedes the application of high-performance networks. To tackle these problems, we propose Graph Flow, a comprehensive knowledge distillation framework, for both network-efficiency and annotation-efficiency medical image segmentation. Specifically, the Graph Flow Distillation transfers the essence of cross-layer variations from a well-trained cumbersome teacher network to a non-trained compact student network. In addition, an unsupervised Paraphraser Module is integrated to purify the knowledge of the teacher, which is also beneficial for the training stabilization. Furthermore, we build a unified distillation framework by integrating the adversarial distillation and the vanilla logits distillation, which can further refine the final predictions of the compact network. With different teacher networks (traditional convolutional architecture or prevalent transformer architecture) and student networks, we conduct extensive experiments on four medical image datasets with different modalities (Gastric Cancer, Synapse, BUSI, and CVC-ClinicDB). We demonstrate the prominent ability of our method on these datasets, which achieves competitive performances. Moreover, we demonstrate the effectiveness of our Graph Flow through a novel semi-supervised paradigm for dual efficient medical image segmentation. Our code will be available at Graph Flow. Wenxuan Zou, Xingqun Qi, Muyi Sun, Zhenan Sun, Caifeng Shan |
IEEE Trans. Medical Imaging | 6 |
| 2022 | Dual-branch self-attention network for pedestrian attribute recognition
Zhang Zhang 0001, Da Li 0003, Peng Zhang 0057, Caifeng Shan |
Pattern Recognit. Lett. | 5 |
| 2022 | Distilled light GaitSet: Towards scalable gait recognition
Xu Song, Yan Huang 0008, Caifeng Shan, Jilong Wang 0010 |
Pattern Recognit. Lett. | 3 |
| 2022 | Progressive Residual Learning With Memory Upgrade for Ultrasound Image Blind Super-ResolutionabstractFor clinical medical diagnosis and treatment, image super-resolution (SR) technology will be helpful to improve the ultrasonic imaging quality so as to enhance the accuracy of disease diagnosis. However, due to the differences of sensing devices or transmission media, the resolution degradation process of ultrasound imaging in real scenes is uncontrollable, especially when the blur kernel is usually unknown. This issue makes current end-to-end SR networks poor performance when applied to ultrasonic images. Aiming to achieve effective SR in real ultrasound medical scenes, in this work, we propose a blind deep SR method based on progressive residual learning and memory upgrade. Specifically, we estimate the accurate blur kernel from the spatial attention map block of low resolution (LR) ultrasound image through a multi-label classification network, then we construct three modules-up- sampling (US) module, residual learning (RL) model and memory upgrading (MU) model for ultrasound image blind SR. The US module is designed to upscale the input information and the up-sampled residual result will be used for SR reconstruction. The RL module is employed to approximate the original LR and continuously generate the updated residual and feed it to the next US module. The last MU module can store all progressively learned residuals, which offers increased interactions between the US and RL modules, augmenting the details recovery. Extensive experiments and evaluations on the benchmark CCA-US and US-CASE datasets demonstrate the proposed approach achieves better performance against the state-of-the-art methods. Heng Liu 0002, Jianyong Liu, Caifeng Shan |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Guest Editorial Emerging Challenges for Deep LearningabstractThe papers in this special section focus on the emerging challenges for deep learningin the biomedical industry. Due to the proliferation of biomedical imaging modalities such as Photoacoustic Tomography, Computed Tomography (CT), Optical Microscopy and Tomography, Single Photon Emission Computed Tomography (SPECT), Magnetic Resonance (MR) Imaging, Ultrasound, Positron Emission Tomography (PET), Magnetic Particle Imaging, EE/MEG, Electron Tomography, and Atomic Force Microscopy, massive amounts of biomedical and health informatics data are being generated on a daily basis. How can we utilize such big data to build better health profiles and predictive models so that we can better diagnose and treat diseases and provide a better life for humans? In the past years, many successful learning methods such as deep learning were proposed to answer this crucial question, which has social, economic, as well as legal implications. Shuihua Wang, Zhengchao Dong, Zheng Zhang 0006, Yuankai Huo, M. Emre Celebi 0001, Caifeng Shan |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Prior Guided Transformer for Accurate Radiology Reports GenerationabstractIn this paper, we propose a prior guided transformer for accurate radiology reports generation. In the encoder part, a radiograph is firstly represented by a set of patch features, which is obtained through a convolutional neural network and a traditional transformer encoder. Then an Additive Gaussian model is applied to represent the prior knowledge based on unsupervised clustering and sparse attention. In the decoder part, prior embeddings are acquired by probabilistically sampling from the radiograph prior. Then the visual features, language embeddings, and prior embeddings are fused by our proposed Prior Guided Attention to generate accurate radiology reports. Experiment results show that our method achieves better performance than state-of-the-art methods on two public radiology datasets, which proves the effectiveness of our prior guided transformer. Mingtao Pei, Caifeng Shan, Zhaoxing Tian |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Medical Instrument Segmentation in 3D US by Hybrid Constrained Semi-Supervised LearningabstractMedical instrument segmentation in 3D ultrasound is essential for image-guided intervention. However, to train a successful deep neural network for instrument segmentation, a large number of labeled images are required, which is expensive and time-consuming to obtain. In this article, we propose a semi-supervised learning (SSL) framework for instrument segmentation in 3D US, which requires much less annotation effort than the existing methods. To achieve the SSL learning, a Dual-UNet is proposed to segment the instrument. The Dual-UNet leverages unlabeled data using a novel hybrid loss function, consisting of uncertainty and contextual constraints. Specifically, the uncertainty constraints leverage the uncertainty estimation of the predictions of the UNet, and therefore improve the unlabeled information for SSL training. In addition, contextual constraints exploit the contextual information of the training images, which are used as the complementary information for voxel-wise uncertainty estimation. Extensive experiments on multiple ex-vivo and in-vivo datasets show that our proposed method achieves Dice score of about 68.6%-69.1% and the inference time of about 1 sec. per volume. These results are better than the state-of-the-art SSL methods and the inference time is comparable to the supervised approaches. Hongxu Yang, Caifeng Shan, R. Arthur Bouwman, Lukas R. C. Dekker, Alexander F. Kolen, Peter H. N. de With |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | CoCo DistillNet: a Cross-layer Correlation Distillation Network for Pathological Gastric Cancer SegmentationabstractIn recent years, deep convolutional neural networks have made significant advances in pathology image segmentation. However, pathology image segmentation encounters with a dilemma in which the higher-performance networks generally require more computational resources and storage. This phenomenon limits the employment of high-accuracy networks in real scenes due to the inherent high-resolution of pathological images. To tackle this problem, we propose CoCo DistillNet, a novel Cross-layer Correlation (CoCo) knowledge distillation network for pathological gastric cancer segmentation. Knowledge distillation, a general technique which aims at improving the performance of a compact network through knowledge transfer from a cumbersome network. Concretely, our CoCo DistillNet models the correlations of channel-mixed spatial similarity between different layers and then transfers this knowledge from a pre-trained cumbersome teacher network to a non-trained compact student network. In addition, we also utilize the adversarial learning strategy to further prompt the distilling procedure which is called Adversarial Distillation (AD). Furthermore, to stabilize our training procedure, we make the use of the unsupervised Paraphraser Module (PM) to boost the knowledge paraphrase in the teacher network. As a result, extensive experiments conducted on the Gastric Cancer Segmentation Dataset demonstrate the prominent ability of CoCo DistillNet which achieves state-of-the-art performance. Wenxuan Zou, Xingqun Qi, Zhuojie Wu, Zijian Wang 0009, Muyi Sun, Caifeng Shan |
BIBM | 6 |
| 2021 | CAS-AIR-3D Face: A Low-Quality, Multi-Modal and Multi-Pose 3D Face DatabaseabstractBenefiting from deep learning with large scale face databases, 2D face recognition has made significant progress in recent years. However, it still highly depends on lighting conditions and human poses, and suffers from face spoofing problem. In contrast, 3D face recognition reveals a new path that can overcome the previous limitations of 2D face recognition. One of the most important problems for 3D face recognition is to construct a suitable database, which can be exploited to train different 3D face recognition algorithms. In this work, we propose a new database, CAS-AIR-3D Face, for low-quality 3D face recognition. It includes 24713 videos from 3093 individuals, which is captured by Intel RealSense SR305. The database contains three modalities: color, depth and near infrared, and is rich in pose, expression, occlusion and distance variations. To the best of our konwledge, CAS-AIR-3D Face is the largest low-quality 3D face database in terms of the number of individuals and the sample variations. Moreover, we preprocess the data via a sophisticated face alignment method, and Point Cloud Spherical Cropping Method (SCM) is leveraged to remove the background noise in the depth images. Finally, an evaluation protocol is designed for fair comparison, and extensive experiments are conducted with different backbone networks to provide different baselines on this database. Qi Li 0005, Xiaoxiao Dong, Weining Wang 0001, Caifeng Shan |
IJCB | 4 |
| 2021 | Static and Dynamic Features Analysis from Human Skeletons for Gait RecognitionabstractGait recognition is an effective way to identify a person due to its non-contact and long-distance acquisition. In addition, the length of human limbs and the motion pattern of human from human skeletons have been proved to be effective features for gait recognition. However, the length of human limbs and motion pattern are calculated through human prior knowledge, more important or detailed information may be missing. Our method proposes to obtain the dynamic information and static information from human skeletons through disentanglement learning. In the experiments, it has been shown that the features extracted by our method are effective. Ziqiong Li, Shiqi Yu 0001, Edel B. García Reyes, Caifeng Shan, Yan-Ran Li 0001 |
IJCB | 4 |
| 2021 | Face Sketch Synthesis via Semantic-Driven Generative Adversarial NetworkabstractFace sketch synthesis has made significant progress with the development of deep neural networks in these years. The delicate depiction of sketch portraits facilitates a wide range of applications like digital entertainment and law enforcement. However, accurate and realistic face sketch generation is still a challenging task due to the illumination variations and complex backgrounds in the real scenes. To tackle these challenges, we propose a novel Semantic-Driven Generative Adversarial Network (SDGAN) which embeds global structure-level style injection and local class-level knowledge re-weighting. Specifically, we conduct facial saliency detection on the input face photos to provide overall facial texture structure, which could be used as a global type of prior information. In addition, we exploit face parsing layouts as the semantic-level spatial prior to enforce globally structural style injection in the generator of SDGAN. Furthermore, to enhance the realistic effect of the details, we propose a novel Adaptive Re-weighting Loss (ARLoss) which dedicates to balance the contributions of different semantic classes. Experimentally, our extensive experiments on CUFS and CUFSF datasets show that our proposed algorithm achieves state-of-the-art performance. Xingqun Qi, Muyi Sun, Weining Wang 0001, Xiaoxiao Dong, Qi Li 0005, Caifeng Shan |
IJCB | 6 |
| 2021 | Mask-guided contrastive attention and two-stream metric co-learning for person Re-identification
Chunfeng Song, Caifeng Shan, Yan Huang 0008, Liang Wang 0001 |
Neurocomputing | 2 |
| 2021 | Efficient and Robust Instrument Segmentation in 3D Ultrasound Using Patch-of-Interest-FuseNet with Hybrid LossabstractInstrument segmentation plays a vital role in 3D ultrasound (US) guided cardiac intervention. Efficient and accurate segmentation during the operation is highly desired since it can facilitate the operation, reduce the operational complexity, and therefore improve the outcome. Nevertheless, current image-based instrument segmentation methods are not efficient nor accurate enough for clinical usage. Lately, fully convolutional neural networks (FCNs), including 2D and 3D FCNs, have been used in different volumetric segmentation tasks. However, 2D FCN cannot exploit the 3D contextual information in the volumetric data, while 3D FCN requires high computation cost and a large amount of training data. Moreover, with limited computation resources, 3D FCN is commonly applied with a patch-based strategy, which is therefore not efficient for clinical applications. To address these, we propose a POI-FuseNet, which consists of a patch-of-interest (POI) selector and a FuseNet. The POI selector can efficiently select the interested regions containing the instrument, while FuseNet can make use of 2D and 3D FCN features to hierarchically exploit contextual information. Furthermore, we propose a hybrid loss function, which consists of a contextual loss and a class-balanced focal loss, to improve the segmentation performance of the network. With the collected challenging ex-vivo dataset on RF-ablation catheter, our method achieved a Dice score of 70.5%, superior to the state-of-the-art methods. In addition, based on the pre-trained model from ex-vivo dataset, our method can be adapted to the in-vivo dataset on guidewire and achieves a Dice score of 66.5% for a different cardiac operation. More crucially, with POI-based strategy, segmentation efficiency is reduced to around 1.3 seconds per volume, which shows the proposed method is promising for clinical use. Hongxu Yang, Caifeng Shan, R. Arthur Bouwman, Alexander F. Kolen, Peter H. N. de With |
Medical Image Anal. | 2 |
| 2021 | Richly Activated Graph Convolutional Network for Robust Skeleton-Based Action RecognitionabstractCurrent methods for skeleton-based human action recognition usually work with complete skeletons. However, in real scenarios, it is inevitable to capture incomplete or noisy skeletons, which could significantly deteriorate the performance of current methods when some informative joints are occluded or disturbed. To improve the robustness of action recognition models, a multi-stream graph convolutional network (GCN) is proposed to explore sufficient discriminative features spreading over all skeleton joints, so that the distributed redundant representation reduces the sensitivity of the action models to non-standard skeletons. Concretely, the backbone GCN is extended by a series of ordered streams which is responsible for learning discriminative features from the joints less activated by preceding streams. Here, the activation degrees of skeleton joints of each GCN stream are measured by the class activation maps (CAM), and only the information from the unactivated joints will be passed to the next stream, by which rich features over all active joints are obtained. Thus, the proposed method is termed richly activated GCN (RA-GCN). Compared to the state-of-the-art (SOTA) methods, the RA-GCN achieves comparable performance on the standard NTU RGB+D 60 and 120 datasets. More crucially, on the synthetic occlusion and jittering datasets, the performance deterioration due to the occluded and disturbed joints can be significantly alleviated by utilizing the proposed RA-GCN. Yi-Fan Song, Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Guest Editorial: Camera-Based Monitoring for Pervasive Healthcare InformaticsabstractThe papers in this special section focus on camera-based monitoring for pervasive healthcare informatics. Measuring physiological signals from the human face and body using video cameras is an emerging research topic that has grown rapidly in the last decade. Remote cameras (in both visible and infrared wavelengths) can be used to measure vital signs from a human body based on skin optics or body movements thereby avoiding mechanical contact with the skin. Camera-based health monitoring will bring a rich set of compelling healthcare applications that directly improve upon contact-based monitoring solutions and impact people’s care experience and quality of life in various scenarios, such as in hospital care units, sleep/senior centers, assisted-living homes, telemedicine and e-health, baby/elderly care at home, fitness and sports, driver monitoring in automotive applications, cardiac/ respiratory gating for MRI/CT, AR/VR based therapy and clinical training, e Wenjin Wang 0002, Steffen Leonhardt, Lionel Tarassenko, Caifeng Shan, Daniel McDuff |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | CANet: Context Aware Network for Brain Glioma SegmentationabstractAutomated segmentation of brain glioma plays an active role in diagnosis decision, progression monitoring and surgery planning. Based on deep neural networks, previous studies have shown promising technologies for brain glioma segmentation. However, these approaches lack powerful strategies to incorporate contextual information of tumor cells and their surrounding, which has been proven as a fundamental cue to deal with local ambiguity. In this work, we propose a novel approach named Context-Aware Network (CANet) for brain glioma segmentation. CANet captures high dimensional and discriminative features with contexts from both the convolutional space and feature interaction graphs. We further propose context guided attentive conditional random fields which can selectively aggregate features. We evaluate our method using publicly accessible brain glioma segmentation datasets BRATS2017, BRATS2018 and BRATS2019. The experimental results show that the proposed algorithm has better or competitive performance against several State-of-The-Art approaches under different segmentation metrics on the training and validation sets. Long Chen 0019, Feixiang Zhou, Zheheng Jiang, Qianni Zhang, Yinhai Wang, Caifeng Shan, Ling Li 0010, Huiyu Zhou 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2021 | Meta-USR: A Unified Super-Resolution Network for Multiple Degradation ParametersabstractRecent research on single image super-resolution (SISR) has achieved great success due to the development of deep convolutional neural networks. However, most existing SISR methods merely focus on super-resolution of a single fixed integer scale factor. This simplified assumption does not meet the complex conditions for real-world images which often suffer from various blur kernels or various levels of noise. More importantly, previous methods lack the ability to cope with arbitrary degradation parameters (scale factors, blur kernels, and noise levels) with a single model. A few methods can handle multiple degradation factors, e.g., noninteger scale factors, blurring, and noise, simultaneously within a single SISR model. In this work, we propose a simple yet powerful method termed meta-USR which is the first unified super-resolution network for arbitrary degradation parameters with meta-learning. In Meta-USR, a meta-restoration module (MRM) is proposed to enhance the traditional upscale module with the capability to adaptively predict the weights of the convolution filters for various combinations of degradation parameters. Thus, the MRM can not only upscale the feature maps with arbitrary scale factors but also restore the SR image with different blur kernels and noise levels. Moreover, the lightweight MRM can be placed at the end of the network, which makes it very efficient for iteratively/repeatedly searching the various degradation factors. We evaluate the proposed method through extensive experiments on several widely used benchmark data sets on SISR. The qualitative and quantitative experimental results show the superiority of our Meta-USR. Xuecai Hu, Zhang Zhang 0001, Caifeng Shan, Zilei Wang, Liang Wang 0001, Tieniu Tan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Deep Q-Network-Driven Catheter Segmentation in 3D US by Hybrid Constrained Semi-supervised Learning and Dual-UNet
Hongxu Yang, Caifeng Shan, Alexander F. Kolen, Peter H. N. de With |
MICCAI (1) | 2 |
| 2020 | Stronger, Faster and More Explainable: A Graph Convolutional Baseline for Skeleton-based Action RecognitionabstractOne essential problem in skeleton-based action recognition is how to extract discriminative features over all skeleton joints. However, the complexity of the State-Of-The-Art (SOTA) models of this task tends to be exceedingly sophisticated and over-parameterized, where the low efficiency in model training and inference has obstructed the development in the field, especially for large-scale action datasets. In this work, we propose an efficient but strong baseline based on Graph Convolutional Network (GCN), where three main improvements are aggregated, i.e., early fused Multiple Input Branches (MIB), Residual GCN (ResGCN) with bottleneck structure and Part-wise Attention (PartAtt) block. Firstly, an MIB is designed to enrich informative skeleton features and remain compact representations at an early fusion stage. Then, inspired by the success of the ResNet architecture in Convolutional Neural Network (CNN), a ResGCN module is introduced in GCN to alleviate computational costs and reduce learning difficulties in model training while maintain the model accuracy. Finally, a PartAtt block is proposed to discover the most essential body parts over a whole action sequence and obtain more explainable representations for different skeleton action sequences. Extensive experiments on two large-scale datasets, i.e., NTU RGB+D 60 and 120, validate that the proposed baseline slightly outperforms other SOTA models and meanwhile requires much fewer parameters during training and inference procedures, e.g., at most 34 times less than DGNN, which is one of the best SOTA methods. Yi-Fan Song, Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001 |
ACM Multimedia | 3 |
| 2020 | Kinematic skeleton graph augmented network for human parsing
Jinde Liu, Zhang Zhang 0001, Caifeng Shan, Tieniu Tan |
Neurocomputing | 3 |
| 2020 | Deep Unbiased Embedding Transfer for Zero-Shot LearningabstractZero-shot learning aims to recognize objects which do not appear in the training dataset. Previous prevalent mapping-based zero-shot learning methods suffer from the projection domain shift problem due to the lack of image classes in the training stage. In order to alleviate the projection domain shift problem, a deep unbiased embedding transfer (DUET) model is proposed in this paper. The DUET model is composed of a deep embedding transfer (DET) module and an unseen visual feature generation (UVG) module. In the DET module, a novel combined embedding transfer net which integrates the complementary merits of the linear and nonlinear embedding mapping functions is proposed to connect the visual space and semantic space. What's more, the end-to-end joint training process is implemented to train the visual feature extractor and the combined embedding transfer net simultaneously. In the UVG module, a visual feature generator trained with a conditional generative adversarial framework is used to synthesize the visual features of the unseen classes to ease the disturbance of the projection domain shift problem. Furthermore, a quantitative index, namely the score of resistance on domain shift (ScoreRDS), is proposed to evaluate different models regarding their resistance capability on the projection domain shift problem. The experiments on five zero-shot learning benchmarks verify the effectiveness of the proposed DUET model. As demonstrated by the qualitative and quantitative analysis, the unseen class visual feature generation, the combined embedding transfer net and the end-to-end joint training process all contribute to alleviating projection domain shift in zero-shot learning. Zhang Zhang 0001, Liang Wang 0001, Caifeng Shan, Tieniu Tan |
IEEE Trans. Image Process. | 4 |
| 2020 | Deep Salient Object Detection With Contextual Information GuidanceabstractIntegration of multi-level contextual information, such as feature maps and side outputs, is crucial for Convolutional Neural Networks (CNNs) based salient object detection. However, most existing methods either simply concatenate multi-level feature maps or calculate element-wise addition of multi-level side outputs, thus failing to take full advantages of them. In this work, we propose a new strategy for guiding multi-level contextual information integration, where feature maps and side outputs across layers are fully engaged. Specifically, shallower-level feature maps are guided by the deeper-level side outputs to learn more accurate properties of the salient object. In turn, the deeper-level side outputs can be propagated to high-resolution versions with spatial details complemented by means of shallower-level feature maps. Moreover, a group convolution module is proposed with the aim to achieve high-discriminative feature maps, in which the backbone feature maps are divided into a number of groups and then the convolution is applied to the channels of backbone feature maps within each group. Eventually, the group convolution module is incorporated in the guidance module to further promote the guidance role. Experiments on three public benchmark datasets verify the effectiveness and superiority of the proposed method over the state-of-the-art methods. Yi Liu 0038, Jungong Han, Qiang Zhang 0020, Caifeng Shan |
IEEE Trans. Image Process. | 4 |
| 2020 | RGB-T Salient Object Detection via Fusing Multi-Level CNN FeaturesabstractRGB-induced salient object detection has recently witnessed substantial progress, which is attributed to the superior feature learning capability of deep convolutional neural networks (CNNs). However, such detections suffer from challenging scenarios characterized by cluttered backgrounds, low-light conditions and variations in illumination. Instead of improving RGB based saliency detection, this paper takes advantage of the complementary benefits of RGB and thermal infrared images. Specifically, we propose a novel end-to-end network for multi-modal salient object detection, which turns the challenge of RGB-T saliency detection to a CNN feature fusion problem. To this end, a backbone network (e.g., VGG-16) is first adopted to extract the coarse features from each RGB or thermal infrared image individually, and then several adjacent-depth feature combination (ADFC) modules are designed to extract multi-level refined features for each single-modal input image, considering that features captured at different depths differ in semantic information and visual details. Subsequently, a multi-branch group fusion (MGF) module is employed to capture the cross-modal features by fusing those features from ADFC modules for a RGB-T image pair at each level. Finally, a joint attention guided bi-directional message passing (JABMP) module undertakes the task of saliency prediction via integrating the multi-level fused features from MGF modules. Experimental results on several public RGB-T salient object detection datasets demonstrate the superiorities of our proposed algorithm over the state-of-the-art approaches, especially under challenging conditions, such as poor illumination, complex background and low contrast. Qiang Zhang 0020, Nianchang Huang, Dingwen Zhang, Caifeng Shan, Jungong Han |
IEEE Trans. Image Process. | 5 |
| 2020 | Guest Editorial: Deep Learning in Ultrasound ImagingabstractAmong the different imaging modalities, ultrasound is the most widespread modality for visualizing human tissue due to it being low-cost, non-ionizing, real-time with immediate feedback to the sonographer, convenient to operate, widely available and well established, with a very large number of images generated in a single setting. On the other hand, ultrasound imaging suffers from the disadvantage of being user dependent and of variable quality,which makes the automated interpretation of ultrasound images often very difficult. In recent years, algorithms in medical imaging have been significantly improved thanks to the advent of deep learning methods (including convolutional neural networks, recurrent neural networks, autoencoders, or generative adversarial networks). To address the various challenges of automatically processing and interpreting ultrasound images, deep learning techniques have been gradually applied to various types of ultrasound data (such as B-mode ultrasound, Doppler ultrasound, or contrast-enhanced ultrasound), acquired with a range of different probes, with the aim of improving image quality, for organ segmentation, device localization and tracking, for tissue characterization, and ultimately to improve disease diagnosis and therapeutic outcome. The papers in this special section seek to present and highlight the latest development on applying advanced deep learning techniques in ultrasound imaging. Caifeng Shan, Tao Tan 0002, Shandong Wu, Julia A. Schnabel |
IEEE J. Biomed. Health Informatics | 1 |
| 2019 | A Comprehensive Study on Large-Scale Person Retrieval in Real Surveillance ScenariosabstractPerson retrieval is a hot research topic due to its important application potential for public security. Though existing algorithms have achieved impressive progresses on current public datasets, it is still a challenging task in the real surveillance scenarios due to the various viewpoints, pose variations and occlusions. Moreover, few of the existing works study the problem of person retrieval on large-scale gallery set, where lots of distractions may deteriorate the retrieval results heavily. To have a deep understanding on the above challenges, we perform a comprehensive study on current state-of-the-art person retrieval algorithms with a large-scale benchmark in real surveillance scenarios. In the study, two kinds of techniques, i.e., attribute recognition and person re-identification, including eight algorithms, are evaluated at both algorithm level and system level. Here, the system-level evaluations investigate the effects of the combinations of the above algorithms with the module of person detection, where lots of distractions in person detection results pose a big challenge for person retrieval in real scenes. Extensive evaluations with large gallery sizes (up to 243k) and comprehensive analyses are presented in the study, which will guide researchers to develop more advanced algorithms in future. Da Li 0003, Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001, Tieniu Tan |
AVSS | 3 |
| 2019 | Efficient Catheter Segmentation in 3D Cardiac Ultrasound using Slice-Based FCN With Deep Supervision and F-Score LossabstractFast and accurate catheter segmentation in 3D ultrasound (US) can improve the outcome and efficiency of cardiac interventions. In this paper, we propose an efficient catheter segmentation method based on a fully convolutional neural network (FCN). The FCN is based on a pre-trained VGG-16 model, which processes the 3D US volumes slice by slice. To enhance its performance, we modify its structure by skipping connections under a deep supervision structure, which is learned with an F-score loss function. Our method can exploit more contextual information and increase the detection of catheter-like voxels. We collected a challenging ex-vivo dataset (92 3D US images) from porcine hearts with an RF-ablation catheter inside. Our experiments on this dataset show that the proposed method achieves a segmentation performance with an F2score of 65.2% with a highly efficient inference around 1.1 sec. per volume. Hongxu Yang, Caifeng Shan, Alexander F. Kolen, Peter H. N. de With |
ICIP | 2 |
| 2019 | Automated Catheter Localization in Volumetric Ultrasound Using 3D Patch-Wise U-Net with Focal Lossabstract3D ultrasound (US) imaging has become an attractive option for image-guided interventions. Fast and accurate catheter localization in 3D cardiac US can improve the outcome and efficiency of the cardiac interventions. In this paper, we propose a catheter localization method for 3D cardiac US using the patch-wise semantic segmentation with model fitting. Our 3D U-Net is trained with the focal loss of cross-entropy, which makes the network to focus more on samples that are difficult to classify. Moreover, we adopt a dense sampling strategy to overcome the extremely imbalanced catheter occupation in the 3D US data. Extensive experiments on our challenging ex-vivo dataset show that the proposed method achieves an F-1 score of 65.1% for catheter segmentation, outperforming the state-of-the-art methods. With this, our method can localize RF-ablation catheters with an average error of 1.28 mm. Hongxu Yang, Caifeng Shan, Alexander F. Kolen, Peter H. N. de With |
ICIP | 2 |
| 2019 | Transferring from ex-vivo to in-vivo: Instrument Localization in 3D Cardiac Ultrasound Using Pyramid-UNet with Hybrid Loss
Hongxu Yang, Caifeng Shan, Tao Tan 0002, Alexander F. Kolen, Peter H. N. de With |
MICCAI (5) | 2 |
| 2019 | Hyperspectral image denoising via minimizing the partial sum of singular values and superpixel segmentation
Yang Liu 0069, Caifeng Shan, Quanxue Gao, Xinbo Gao 0001, Jungong Han, Rongmei Cui |
Neurocomputing | 2 |
| 2019 | Video-based discomfort detection for infants
Yue Sun 0001, Caifeng Shan, Tao Tan 0002, Xi Long 0001, Arash Pourtaherian, Svitlana Zinger, Peter H. N. de With |
Mach. Vis. Appl. | 2 |
| 2019 | Salient object detection employing a local tree-structured low-rank representation and foreground consistency
Qiang Zhang 0020, Zhen Huo, Yi Liu 0038, Yunhui Pan, Caifeng Shan, Jungong Han |
Pattern Recognit. | 5 |
| 2018 | Multispectral Image Analysis for Patient Tissue Tracking During Complex InterventionsabstractDuring complex interventions, patient tracking is needed for optimal motion compensation, in order to guide the physician in a minimally invasive way. Nowadays, optical tracking systems are used for tracking of markers. Despite the unobtrusiveness of this technology, the approach is cumbersome because it requires manual placement of markers, which can alter during the operations due to the presence of liquids. To improve the clinical workflow, a new feature tracking algorithm is designed, involving feature detection and tracking without optical markers. The new markers are created with the multispectral imaging. Maximally stable extremal regions (MSER) and Speeded Up Robust Feature (SURF) methods are applied to design features and track natural landmarks, e.g. moles and veins. Both methods are tested and compared in accuracy with the mean shift tracking method. SURF reaches the highest accuracy of 0.257 pixels for images at 430 nm and 0.562 pixels for images at 970 nm. This study shows that incorporating multispectral imaging in the surgical scenario leads to an attractive benefit minimizing the risk of marker obstruction and displacement. Francesca Manni, Marco Mamprin, Svitlana Zinger, Caifeng Shan, Ronald Holthuizen, Peter H. N. de With |
ICIP | 4 |
| 2018 | Catheter Detection in 3D Ultrasound Using Triplanar-Based Convolutional Neural Networksabstract3D Ultrasound (US) image-based catheter detection can potentially decrease the cost on extra equipment and training. Meanwhile, accurate catheter detection enables to decrease the operation duration and improves its outcome. In this paper, we propose a catheter detection method based on convolutional neural networks (CNNs) in 3D US. Voxels in US images are classified as catheter (or not) using triplanar-based CNNs. Our proposed CNN employs two-stage training with weighted loss function, which can cope with highly imbalanced training data and improves classification accuracy. When compared to state-of-the-art handcrafted features on ex-vivo datasets, our proposed method improves the F2-score with at least 31%. Based on classified volumes, the catheters are localized with an average position error of smaller than 3 voxels in the examined datasets, indicating that catheters are always detected in noisy and low-resolution images. Hongxu Yang, Caifeng Shan, Alexander F. Kolen, Peter H. N. de With |
ICIP | 2 |
| 2014 | Special issue on background modeling for foreground detection in real-world dynamic scenes
Thierry Bouwmans, Jordi Gonzàlez 0001, Caifeng Shan, Massimo Piccardi, Larry Davis 0001 |
Mach. Vis. Appl. | 3 |
| 2013 | The third ACM international workshop on interactive multimedia on mobile and portable devices (IMMPD'13)abstractWith the mobile and portable devices become ubiquitous for people's daily life, how to design user interfaces of these products that enable natural, intuitive and fun interaction is one of the main challenges the multimedia community is facing. Following previous successful events, the third ACM International workshop on Interactive Multimedia on Mobile and Portable Devices (IMMPD'13) aims to bring together researchers from both academia and industry in domains including computer vision, audio and speech processing, machine learning, pattern recognition, communications, human-computer interaction, and media technology to share and discuss recent advances in interactive multimedia. Jiebo Luo 0001, Caifeng Shan, Ling Shao 0001, Minoru Etoh |
ACM Multimedia | 2 |
| 2013 | Semantic assessment of shopping behavior using trajectories, shopping related actions, and context information
Mirela C. Popa, Léon J. M. Rothkrantz, Caifeng Shan, Tommaso Gritti, Pascal Wiggers |
Pattern Recognit. Lett. | 3 |
| 2013 | Shopping behavior recognition using a language modeling analogy
Mirela C. Popa, Léon J. M. Rothkrantz, Pascal Wiggers, Caifeng Shan |
Pattern Recognit. Lett. | 4 |
| 2013 | Indexing of large-scale multimedia signals
Meng Wang 0001, Xinbo Gao 0001, Yi Yang 0001, Caifeng Shan |
Signal Process. | 4 |
| 2012 | Shared-bed person segmentation based on motion estimationabstractVideo-based sleep analysis is a topic with important applications, and shared-bed occurs frequently in the context of sleep. One difficulty for the shared-bed situation is to assign the movements to the correct person because they can occur in close proximity and even overlapping. To manage to achieve person segmentation in the shared-bed situation, in this paper we propose an approach to correctly segment the region of persons based on motion estimation. In our approach, considering the consistency of the motion vectors, specifically their length and angle, the adjacent blocks are clustered. The generated clusters are then assigned to a person according to temporal correlation. The occupied region of the person is updated each frame based on the assignment result of the clusters. The proposed approach tackles the segmentation issue when the two persons are close to each other or even overlap, and the accuracy of the segmentation is beyond 82% in the data set we acquired. Xuyuan Jin, Adrienne Heinrich, Caifeng Shan, Gerard de Haan |
ICIP | 3 |
| 2012 | Assessment of customers' level of interestabstractSurveillance systems in shopping malls or supermarkets are usually designated for assuring safety and detecting abnormal behavior. We used the distributed video cameras system to design digital shopping assistants which assess the behavior of customers while shopping, detect when they need assistance, and offer their support in case there is a selling opportunity. In this paper we propose a system for analyzing human behavior patterns related to products interaction, which could reveal the customer's level of interest. We extracted discriminative features for basic action detection and analyzed different statistical and spatio-temporal classification methods, which capture relations between frames, features, and basic actions. Our experiments show that it is possible to accurately recognize different shopping related actions (85.7%) and discriminate between the proposed levels of interest in (88%) of the cases. Mirela C. Popa, Léon J. M. Rothkrantz, Caifeng Shan, Pascal Wiggers |
ICIP | 3 |
| 2012 | The second ACM international workshop on interactive multimedia on mobile and Portable devicesabstractWhen mobile and portable devices become ubiquitous for people's daily life, how to design multimedia user interfaces of these products that enable natural, intuitive and fun interaction is one of the main challenges the multimedia community is facing. Following several successful events, the 2nd ACM International workshop on Interactive Multimedia on Mobile and Portable Devices (IMMPD'12) aims to bring together researchers from both academia and industry in domains including computer vision, audio and speech processing, machine learning, pattern recognition, communications, human-computer interaction, and media technology to share and discuss recent advances in interactive multimedia. Ling Shao 0001, Caifeng Shan, Minoru Etoh |
ACM Multimedia | 2 |
| 2012 | Learning-based encoding with soft assignment for age estimation under unconstrained imaging conditions
Fares Alnajar, Caifeng Shan, Theo Gevers, Jan-Mark Geusebroek |
Image Vis. Comput. | 2 |
| 2012 | Real-time robust background subtraction under rapidly changing illumination conditions
Luc P. J. Vosters, Caifeng Shan, Tommaso Gritti |
Image Vis. Comput. | 2 |
| 2012 | Learning local binary patterns for gender classification on real-world face images
Caifeng Shan |
Pattern Recognit. Lett. | 1 |
| 2012 | Smile Detection by Boosting Pixel DifferencesabstractSmile detection in face images captured in unconstrained real-world scenarios is an interesting problem with many potential applications. This paper presents an efficient approach to smile detection, in which the intensity differences between pixels in the grayscale face images are used as features. We adopt AdaBoost to choose and combine weak classifiers based on intensity differences to form a strong classifier. Experiments show that our approach has similar accuracy to the state-of-the-art method but is significantly faster. Our approach provides 85% accuracy by examining 20 pairs of pixels and 88% accuracy with 100 pairs of pixels. We match the accuracy of the Gabor-feature-based support vector machine using as few as 350 pairs of pixels. Caifeng Shan |
IEEE Trans. Image Process. | 1 |
| 2011 | Detecting Customers' Buying Events on a Real-Life Database
Mirela C. Popa, Tommaso Gritti, Léon J. M. Rothkrantz, Caifeng Shan, Pascal Wiggers |
CAIP (1) | 4 |
| 2011 | An efficient approach to smile detectionabstractSmile detection in real-life face images is an interesting problem with many potential applications. This paper presents an efficient approach to smile detection for face images captured in real-world unconstrained scenarios. In our approach, the pixel intensities in the gray-scale face image are compared, and the intensity differences are used as features. We adopt Adaboost to choose and combine intensity differences (based weak classifiers) to form a strong classifier for smile detection. With the simple features, the detection could be very fast. Our approach achieves 85% accuracy in smile detection by examining 20 pairs of pixel difference and 88% accuracy with 100 pairs of pixel comparison. We match the accuracy of Gabor features based SVM by examining as few as 350 pairs of pixel difference. Caifeng Shan |
FG | 1 |
| 2011 | ACM international workshop on interactive multimedia on mobile and portable devices (IMMPD'11)abstractWith the mobile and portable devices become ubiquitous for people's daily life, how to design user interfaces of these products that enable natural, intuitive and fun interaction is one of the main challenges the multimedia community is facing. Following several successful events, the ACM International workshop on Interactive Multimedia on Mobile and Portable Devices (IMMPD'11) aims to bring together researchers from both academia and industry in domains including computer vision, audio and speech processing, machine learning, pattern recognition, communications, human-computer interaction, and media technology to share and discuss recent advances in interactive multimedia. Jiebo Luo 0001, Caifeng Shan, Ling Shao 0001, Minoru Etoh |
ACM Multimedia | 2 |
| 2011 | Special Issue on Video Analysis on Resource-Limited SystemsabstractThe 17 papers in this special issue focus on resource-limited systems. Rama Chellappa, Andrea Cavallaro, Ying Wu 0001, Caifeng Shan, Yun Fu 0001, Kari Pulli |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2011 | Local Binary Patterns and Its Application to Facial Image Analysis: A SurveyabstractLocal binary pattern (LBP) is a nonparametric descriptor, which efficiently summarizes the local structures of images. In recent years, it has aroused increasing interest in many areas of image processing and computer vision and has shown its effectiveness in a number of applications, in particular for facial image analysis, including tasks as diverse as face detection, face recognition, facial expression analysis, and demographic classification. This paper presents a comprehensive survey of LBP methodology, including several more recent variations. As a typical application of the LBP approach, LBP-based facial image analysis is extensively reviewed, while its successful extensions, which deal with various tasks of facial image analysis, are also highlighted. Di Huang 0001, Caifeng Shan, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002 |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2010 | Gender Classification on Real-Life Faces
Caifeng Shan |
ACIVS (2) | 1 |
| 2010 | Background Subtraction under Sudden Illumination ChangesabstractRobust background subtraction under sudden illumination changes is a challenging problem. In this paper, we propose an approach to address this issue, which combines the Eigenbackground algorithm together with a statistical illumination model. The first algorithm is used to give a rough reconstruction of the input frame, while the second one improves the foreground segmentation. We introduce an online spatial likelihood model by detecting reliable background and foreground pixels. Experimental results illustrate that our approach achieves consistently higher accuracy compared to several state-of-the-art algorithms. Luc P. J. Vosters, Caifeng Shan, Tommaso Gritti |
AVSS | 2 |
| 2010 | An event-based approach to multi-modal activity modeling and recognitionabstractThe topic of human activity modeling and recognition still provides many challenges, despite receiving considerable attention. These challenges include the large number of sensors often required for accurate activity recognition, and the need for user-specific training samples. In this paper, an approach is presented for recognition of activities of daily living (ADL) using only a single camera and microphone as sensors. Scene analysis techniques are used to classify audio and video events, which are used to model a set of activities using hidden Markov models. Data was obtained through recordings of 8 participants. The events generated by scene analysis algorithms are compared to events obtained through manual annotation. In addition, several model parameter estimation techniques are compared. In a number of experiments, it is shown that if activities are fully observed these models yield a class accuracy of 97% on annotated data, and 94% on scene analysis data. Using a sliding window approach to classify activities in progress yields a class accuracy of 79% on annotated data, and 73% on scene analysis data. It is also shown that a multi-modal approach yields superior results compared to either individual modality on scene analysis data. Finally, it can be concluded the created models perform well even across participants. Marten Pijl, Steven van de Par, Caifeng Shan |
PerCom | 3 |
| 2010 | Analysis of shopping behavior based on surveillance systemabstractClosed Circuit Television systems in shopping malls could be used to monitor the shopping behavior of people. From the tracked path, features can be extracted such as the relation with the shopping area, the orientation of the head, speed of walking and direction, pauses which are supposed to be related to the interest of the shopper. Once the interest has been detected the next step is to assess the shopper's positive or negative appreciation to the focused products by analyzing the (non-verbal) behavior of the shopper. Ultimately the system goal is to assess the opportunities for selling, by detecting if a customer needs support. In this paper we present our methodology towards developing such a system consisting of participating observation, designing shopping behavioral models, assessing the associated features and analyzing the underlying technology. In order to validate our observations we made recordings in our shop lab. Next we describe the used tracking technology and the results from experiments. Mirela C. Popa, Léon J. M. Rothkrantz, Zhenke Yang, Pascal Wiggers, Ralph Braspenning, Caifeng Shan |
SMC | 6 |
| 2010 | Special Issue on Multimodal Affective InteractionabstractThe 11 papers in this special issue can be categorized into five groups: Emotional speech synthesis and recognition; affective video content analysis; facial expressions and head movements; affect analysis in small groups; and audio-visual affective corpus. Nicu Sebe, Hamid K. Aghajan, Thomas S. Huang, Nadia Magnenat-Thalmann, Caifeng Shan |
IEEE Trans. Multim. | 5 |
| 2009 | 1st ACM international workshop on interactive multimedia for consumer electronics (IMCE'09)abstractThe ACM International workshop on Interactive Multimedia for Consumer Electronics (IMCE) aims to bring together researchers from both academia and industry in domains including computer vision, machine learning, audio and speech processing, communications, artificial intelligence and media technology to share and discuss recent advances in interactive user interfaces and multimedia applications. Multimedia interaction is becoming a technology applied in many consumer electronics devices and can make user interfaces more intuitive and controllable. Multiple modalities including audio, video and haptics can be utilized and fused for media interaction. Ling Shao 0001, Caifeng Shan, Jiebo Luo 0001, Minoru Etoh |
ACM Multimedia | 2 |
| 2009 | Facial expression recognition based on Local Binary Patterns: A comprehensive study
Caifeng Shan, Shaogang Gong, Peter W. McOwan |
Image Vis. Comput. | 1 |
| 2009 | Optimal Regularization Parameter Estimation for Spectral Regression Discriminant AnalysisabstractSpectral regression discriminant analysis (SRDA) is an efficient subspace learning method proposed recently. One important unsolved issue of SRDA is how to automatically determine an appropriate regularization parameter. In this letter, we present a method to estimate the optimal regularization parameter for SRDA. We test our method in different applications including head pose estimation, face recognition, and text categorization. Our extensive experiments evidently illustrate the effectiveness and efficiency of our approach. Caifeng Shan, Gerard de Haan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Learning Discriminative LBP-Histogram Bins for Facial Expression RecognitionabstractLocal Binary Patterns (LBP) have been well exploited for facial image analysis recently. In the existing work, the LBP histograms are extracted from local facial regions, and used as a whole for the regional description. However, not all bins in the LBP histogram are necessary to be useful for facial representation. In this paper, we propose to learn discriminative LBP-Histogram (LBPH) bins for the task of facial expression recognition. Our experiments illustrate that the selected LBPH bins provide a compact and discriminative facial representation. We experimentally illustrate that it is necessary to consider multiscale LBP for representing faces, and most discriminative information is contained in uniform patterns. By adopting SVM with the selected multiscale LBPH bins, we obtain the best recognition performance of 93.1% on the Cohn-Kanade database. 1 Caifeng Shan, Tommaso Gritti |
BMVC | 1 |
| 2008 | Local features based facial expression recognition with face registration errorsabstractIn this paper, we extensively investigate local features based facial expression recognition with face registration errors, which has never been addressed before. Our contributions are three fold. Firstly, we propose and experimentally study the histogram of oriented gradients (HOG) descriptors for facial representation. Secondly, we present facial representations based on local binary patterns (LBP) and local ternary patterns (LTP) extracted from overlapping local regions. Thirdly, we quantitatively study the impact of face registration errors on facial expression recognition using different facial representations. Overall LBP with overlapping gives the best performance (92.9% recognition rate on the Cohn-Kanade database), while maintaining a compact feature vector and best robustness against face registration errors. Tommaso Gritti, Caifeng Shan, Vincent Jeanne, Ralph Braspenning |
FG | 2 |
| 2008 | Fusing gait and face cues for human gender recognition
Caifeng Shan, Shaogang Gong, Peter W. McOwan |
Neurocomputing | 1 |
| 2007 | Learning gender from human gaits and facesabstractComputer vision based gender classification is an important component in visual surveillance systems. In this paper, we investigate gender classification from human gaits in image sequences, a relatively understudied problem. Moreover, we propose to fuse gait and face for improved gender discrimination. We exploit Canonical Correlation Analysis (CCA), a powerful tool that is well suited for relating two sets of measurements, to fuse the two modalities at the feature level. Experiments demonstrate that our multimodal gender recognition system achieves the superior recognition performance of 97.2% in large datasets. Caifeng Shan, Shaogang Gong, Peter W. McOwan |
AVSS | 1 |
| 2007 | Beyond Facial Expressions: Learning Human Emotion from Body GesturesabstractVision-based human affect analysis is an interesting and challenging problem, impacting important applications in many areas. In this paper, beyond facial expressions, we investigate affective body gesture analysis in video sequences, a relatively understudied problem. Spatial-temporal features are exploited for modeling of body gestures. Moreover, we present to fuse facial expression and body gesture at the feature level using Canonical Correlation Analysis (CCA). By establishing the relationship between the two modalities, CCA derives a semantic “affect ” space. Experimental results demonstrate the effectiveness of our approaches. 1 Caifeng Shan, Shaogang Gong, Peter W. McOwan |
BMVC | 1 |
| 2007 | Capturing Correlations Among Facial Parts for Facial Expression AnalysisabstractCapturing and analyzing the correlations among facial parts are important for interpreting facial behaviors precisely. In this paper, we exploit Canonical Correlation Analysis (CCA) to model the correlations of facial parts for facial expression analysis. We propose a Matrix-based Canonical Correlation Analysis (MCCA) for better correlation analysis on 2D image or matrix data in general. Extensive experiments have shown that compared to the traditional CCA, MCCA models more accurately correlations among image data with more compact representation using much fewer canonical factors. 1 Caifeng Shan, Shaogang Gong, Peter W. McOwan |
BMVC | 1 |
| 2007 | Visual inference of human emotion and behaviourabstractWe address the problem of automatic interpretation ofnon-exaggerated human facial and body behaviours captured in video. We illustrate our approach by three examples. (1) We introduce Canonical Correlation Analysis (CCA) and Matrix Canonical Correlation Analysis (MCCA) for capturing and analyzing spatial correlations among non-adjacent facial parts for facial behaviour analysis. (2) We extend Canonical Correlation Analysis to multimodality correlation for bebaviour inference using both facial and body gestures. (3) We model temporal correlation among human movement patterns in a wider space using a mixture of Multi-Observation Hidden Markov Model for human behaviour profiling and behavioural anomaly detection. Shaogang Gong, Caifeng Shan, Tao Xiang 0002 |
ICMI | 2 |
| 2007 | Real-time hand tracking using a mean shift embedded particle filter
Caifeng Shan, Tieniu Tan, Yucheng Wei |
Pattern Recognit. | 1 |
| 2006 | Dynamic Facial Expression Recognition Using A Bayesian Temporal Manifold ModelabstractIn this paper, we propose a novel Bayesian approach to modelling tem-poral transitions of facial expressions represented in a manifold, with the aim of dynamical facial expression recognition in image sequences. A gener-alised expression manifold is derived by embedding image data into a low dimensional subspace using Supervised Locality Preserving Projections. A Bayesian temporal model is formulated to capture the dynamic facial ex-pression transition in the manifold. Our experimental results demonstrate the advantages gained from exploiting explicitly temporal information in ex-pression image sequences resulting in both superior recognition rates and improved robustness against static frame-based recognition methods. 1 Caifeng Shan, Shaogang Gong, Peter W. McOwan |
BMVC | 1 |
| 2005 | Recognizing facial expressions at low resolutionabstractThis paper focuses on recognizing facial expressions at low resolution. We introduce local binary patterns (LBP) as novel low-computation discriminative features for low-resolution facial expression recognition. Compared to Gabor wavelets, LBP features can be derived rapidly in a single scan of raw images, whilst still retaining enough facial information in a compact representation. Support vector machine (SVM) is adopted to classify facial expressions. Extensive experiments on the Cohn-Kanade database demonstrate that the LBP features are effective and efficient for facial expression recognition, and crucially perform robustly and stably over a useful range of low resolutions. Our method yields promising performance when processing compressed low-resolution video sequences from the PETS 2003 dataset. Caifeng Shan, Shaogang Gong, Peter W. McOwan |
AVSS | 1 |
| 2005 | Conditional Mutual Infomation Based Boosting for Facial Expression RecognitionabstractThis paper proposes a novel approach for facial expression recognition by boosting Local Binary Patterns (LBP) based classifiers. L ow-cost LBP features are introduced to effectively describle local fea tures of face images. A novel learning procedure, Conditional Mutual Infomation based Boosting (CMIB), is proposed. CMIB learns a sequence of weak classifie rs that maximize their mutual information about a candidate class, conditional to the response of any weak classifier already selected; a strong cl assifier is constructed by combining the learned weak classifiers using the Naive-Bayes. Extensive experiments on the Cohn-Kanade database illustrated that LBP features are effective for expression analysis, and CMIB enables much faster training than AdaBoost, and yields a classifier of improved c lassification performance. Caifeng Shan, Shaogang Gong, Peter W. McOwan |
BMVC | 1 |
| 2005 | Robust facial expression recognition using local binary patternsabstractA novel low-computation discriminative feature space is introduced for facial expression recognition capable of robust performance over a rang of image resolutions. Our approach is based on the simple local binary patterns (LBP) for representing salient micro-patterns of face images. Compared to Gabor wavelets, the LBP features can be extracted faster in a single scan through the raw image and lie in a lower dimensional space, whilst still retaining facial information efficiently. Template matching with weighted Chi square statistic and support vector machine are adopted to classify facial expressions. Extensive experiments on the Cohn-Kanade Database illustrate that the LBP features are effective and efficient for facial expression discrimination. Additionally, experiments on face images with different resolutions show that the LBP features are robust to low-resolution images, which is critical in real-world applications where only low-resolution video input is available. Caifeng Shan, Shaogang Gong, Peter W. McOwan |
ICIP (2) | 1 |