Si-Qi Liu 0003

dblp:60/9360-3 · also Siqi Liu 0003 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0002-2795-6227ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 10 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 CAM-Interacted Vision GNN for Multi-Label Medical Images
abstract
Vision Graph Neural Network (ViG) is designed to recognize different objects through graph-level processing. However, ViG constructs graphs with appearance-level neighbors and neglects the category semantic. The oversight results in the unintentional connection of patches that belong to different objects, thus affecting the distinctiveness of categories in multi-label medical image learning. Since the pixel-level annotations for images are not easily available, category-aware graphs can not be directly built. To solve this problem, we consider localizing category-specific regions using Class Activation Maps (CAMs), an effective way to highlight regions belonging to each category without requiring manual annotations. Specifically, we propose a CAM-interacted Vision GNN (CiV-GNN), in which category-aware graphs are formed to perform intra-category graph processing. CIV-GNN includes a Class-activated Patch Division (CAPD) module, which introduces CAMs as guidance for category-aware graph building. Furthermore, we develop a Multi-graph Interactive Processing (MIP) module to model the relations between category-aware graphs, promoting inter-category interaction learning. Experimental results show that CiV-GNN performs well in surgical tool localization and multi-label medical image classification. Specifically, for m2cai16-localization, CiV-GNN exhibits a 1.43% and 7.02% improvement in mAP50 and mAP50-95, respectively, compared to YOLOv8.
Jingchao Wang 0002, Baoyao Yang, Si-Qi Liu 0003, Xiaoqi Zheng, Wenbin Yao, Junxiang Chen
IEEE J. Biomed. Health Informatics3
2025 Pattern-Anchored Adaptive Prototype Learning for Gastroscopic Lesion Detection and Beyond
Xuanye Zhang, Xiaoqing Hu, Guanbin Li, Si-Qi Liu 0003, Yuanhuan Xiong
MICCAI (9)4
2025 Sparse diffusion models for multi-annotator medical image segmentation
Haofeng Li, Guanbin Li, Si-Qi Liu 0003
Knowl. Based Syst.4
2025 Highlighted Diffusion Model as Plug-In Priors for Polyp Segmentation
abstract
Automated polyp segmentation from colonoscopy images is crucial for colorectal cancer diagnosis. The accuracy of such segmentation, however, is challenged by two main factors. First, the variability in polyps' size, shape, and color, coupled with the scarcity of well-annotated data due to the need for specialized manual annotation, hampers the efficacy of existing deep learning methods. Second, concealed polyps often blend with adjacent intestinal tissues, leading to poor contrast that challenges segmentation models. Recently, diffusion models have been explored and adapted for polyp segmentation tasks. However, the significant domain gap between RGB-colonoscopy images and grayscale segmentation masks, along with the low efficiency of the diffusion generation process, hinders the practical implementation of these models. To mitigate these challenges, we introduce the Highlighted Diffusion Model Plus (HDM+), a two-stage polyp segmentation framework. This framework incorporates the Highlighted Diffusion Model (HDM) to provide explicit semantic guidance, thereby enhancing segmentation accuracy. In the initial stage, the HDM is trained using highlighted ground-truth data, which emphasizes polyp regions while suppressing the background in the images. This approach reduces the domain gap by focusing on the image itself rather than on the segmentation mask. In the subsequent second stage, we employ the highlighted features from the trained HDM's U-Net model as plug-in priors for polyp segmentation, rather than generating highlighted images, thereby increasing efficiency. Extensive experiments conducted on six polyp segmentation benchmarks demonstrate the effectiveness of our approach.
Yuncheng Jiang 0002, Shuangyi Tan, Si-Qi Liu 0003, Zhen Li 0026, Guanbin Li
IEEE J. Biomed. Health Informatics4
2025 Polarity Prompting Vision Foundation Models for Pathology Image Analysis
abstract
The sharp rise in non-alcoholic fatty liver disease (NAFLD) cases has become a major health concern in recent years. Accurately identifying tissue alteration regions is crucial for NAFLD diagnosis but challenging with small-scale pathology datasets. Recently, prompt tuning has emerged as an effective strategy for adapting vision models to small-scale data analysis. However, current prompting techniques, designed primarily for general image classification, use generic cues that are inadequate when dealing with the intricacies of pathological tissue analysis. To solve this problem, we introduce Quantitative Attribute-based Polarity Visual Prompting (Q-PoVP), a new prompting method for pathology image analysis. Q-PoVP introduces two types of measurable attributes: K-function-based spatial attributes and histogram-based morphological attributes. Both help to measure tissue conditions quantitatively. We develop a quantitative attribute-based polarity visual prompt generator that converts quantitative visual attributes into positive and negative visual prompts, facilitating a more comprehensive and nuanced interpretation of pathological images. To enhance feature discrimination, we introduce a novel orthogonal-based polarity visual prompt tuning technique that disentangles and amplifies positive visual attributes while suppressing negative ones. We extensively tested our method on three different tasks. Our task-specific prompting demonstrates superior performance in both diagnostic accuracy and interpretability compared to existing methods. This dual advantage makes it particularly valuable for clinical settings, where healthcare providers require not only reliable results but also transparent reasoning to support informed patient care decisions. Code is available at https://github.com/7LFB/Q-PoVP.
Chong Yin, Si-Qi Liu 0003, Kaiyang Zhou, Vincent Wai-Sun Wong, Pong C. Yuen
IEEE Trans. Medical Imaging2
2024 XFibrosis: Explicit Vessel-Fiber Modeling for Fibrosis Staging from Liver Pathology Images
abstract
The increasing prevalence of non-alcoholic fatty liver disease (NAFLD) has caused public concern in recent years. The high prevalence and risk of severe complications make monitoring NAFLD progression a public health priority. Fibrosis staging from liver biopsy images plays a key role in demonstrating the histological progression of NAFLD. Fibrosis mainly involves the deposition of fibers around vessels. Current deep learning-based fi-brosis staging methods learn spatial relationships between tissue patches but do not explicitly consider the relation-ships between vessels and fibers, leading to limited performance and poor interpretability. In this paper, we propose an eXplicit vessel-fiber modeling method for Fibrosis staging from liver biopsy images, namely XFibrosis. Specifically, we transform vessels and fibers into graph-structured representations, where their micro-structures are depicted by vessel-induced primal graphs andfiber-induced dual graphs, respectively. Moreover, the fiber-induced dual graphs also represent the connectivity information between vessels caused by fiber deposition. A primal-dual graph convolution module is designed to facilitate the learning of spatial relationships between vessels and fibers, allowing for the joint exploration and interaction of their micro-structures. Experiments conducted on two datasets have shown that explicitly modeling the relationship between vessels and fibers leads to improved fibrosis staging and en-hanced interpretability.
Chong Yin, Si-Qi Liu 0003, Fei Lyu 0004, Sune Darkner, Vincent Wai-Sun Wong, Pong C. Yuen
CVPR2
2024 Prompting Vision Foundation Models for Pathology Image Analysis
abstract
The rapid increase in cases of non-alcoholic fatty liver disease (NAFLD) in recent years has raised significant public concern. Accurately identifying tissue alteration regions is crucial for the diagnosis of NAFLD, but this task presents challenges in pathology image analysis, particularly with small-scale datasets. Recently, the paradigm shift from full fine-tuning to prompting in adapting vision foundation models has offered a new perspective for small-scale data analysis. However, existing prompting methods based on task-agnostic prompts are mainly developed for generic image recognition, which fall short in providing instructive cues for complex pathology images. In this paper, we propose Quantitative Attribute-based Prompting (QAP), a novel prompting method specifically for liver pathology image analysis. QAP is based on two quantitative attributes, namely K-function-based spatial attributes and histogram-based morphological attributes, which are aimed for quantitative assessment of tissue states. Moreover, a conditional prompt generator is designed to turn these instance-specific attributes into visual prompts. Extensive experiments on three diverse tasks demonstrate that our task-specific prompting method achieves better diagnostic performance as well as better interpretability. Code is available at https://github.com/7LFBIQAP.
Chong Yin, Si-Qi Liu 0003, Kaiyang Zhou, Vincent Wai-Sun Wong, Pong C. Yuen
CVPR2
2024 Bottom-Up Domain Prompt Tuning for Generalized Face Anti-spoofing
Si-Qi Liu 0003, Pong C. Yuen
ECCV (70)1
2024 CAM-Guided Translation for Unpaired Weakly-Supervised Medical Image Segmentation
abstract
Multi-modal learning has shown advantages in improving weakly-supervised medical image segmentation (WS- MIS). However, most current works are based on paired data, which is infeasible to collect in certain scenarios. Although modal translation can be used to generate paired data, it often leads to low-quality translations, such as local deformations or irrational textures, without prior knowledge. This paper proposes a discriminative-aware image translation method, which introduces class activation maps (CAMs) to localize discriminative areas, thus overcoming the lack of pixel-wise annotations in WS-MIS. In addition, we design a CAM-correlation constraint that facilitates multi-modal complementary information exchange to enhance the consistency between CAMs generated from different modalities. Experimental results show that our method outperforms recent weakly-supervised segmentation works when using unpaired multi-modal data.
Yuebin Xie, Xiaochen He, Baoyao Yang, Fei Lyu 0004, Si-Qi Liu 0003
ICME5
2024 HistoSyn: Histomorphology-Focused Pathology Image Synthesis
Chong Yin, Si-Qi Liu 0003, Vincent Wai-Sun Wong, Pong C. Yuen
MICCAI (4)2
2024 BindingSiteDTI: differential-scale binding site modelling for drug-target interaction prediction
abstract
MOTIVATION: Enhanced by contemporary computational advances, the prediction of drug-target interactions (DTIs) has become crucial in developing de novo and effective drugs. Existing deep learning approaches to DTI prediction are frequently beleaguered by a tendency to overfit specific molecular representations, which significantly impedes their predictive reliability and utility in novel drug discovery contexts. Furthermore, existing DTI networks often disregard the molecular size variance between macro molecules (targets) and micro molecules (drugs) by treating them at an equivalent scale that undermines the accurate elucidation of their interaction. RESULTS: We propose a novel DTI network with a differential-scale scheme to model the binding site for enhancing DTI prediction, which is named as BindingSiteDTI. It explicitly extracts multiscale substructures from targets with different scales of molecular size and fixed-scale substructures from drugs, facilitating the identification of structurally similar substructural tokens, and models the concealed relationships at the substructural level to construct interaction feature. Experiments conducted on popular benchmarks, including DUD-E, human, and BindingDB, shown that BindingSiteDTI contains significant improvements compared with recent DTI prediction methods. AVAILABILITY AND IMPLEMENTATION: The source code of BindingSiteDTI can be accessed at https://github.com/MagicPF/BindingSiteDTI.
Chong Yin, Si-Qi Liu 0003, Zhaoxiang Bian, Pong C. Yuen
Bioinform.3
2024 Robust Remote Photoplethysmography Estimation With Environmental Noise Disentanglement
abstract
Remote Photoplethysmography (rPPG) has been attracting increasing attention due to its potential in a wide range of application scenarios such as physical training, clinical monitoring, and face anti-spoofing. On top of conventional solutions, deep-learning approach starts to dominate in rPPG estimation and achieves top-level performance. However, most of them try to integrate preprocessing steps such as the ROI selection into an end-to-end network, which may diverge the attention and also limit the generalization in other scenarios with different input skin regions. In this work, we focus on learning the intrinsic rPPG feature and design a lightweight but effective rPPG estimation network based on spatiotemporal convolution. To further improve the robustness, on top of the basic design we propose the Noise-Disentangled DeeprPPG (ND-DeeprPPG) by disentangling the environmental noise from the raw rPPG feature with an adversarial canonical correlation analysis learning strategy. Background regions are employed as a reference to guide the noise disentangling in a self-supervised manner. Extensive experiments show that our ND-DeeprPPG not only outperforms the state-of-the-arts on heart rate estimation but also exhibits promising robustness in cross-skin-region, cross-dataset scenarios and other rPPG-based tasks.
Si-Qi Liu 0003, Pong C. Yuen
IEEE Trans. Image Process.1
2023 Dual-bridging with Adversarial Noise Generation for Domain Adaptive rPPG Estimation
abstract
The remote photoplethysmography (rPPG) technique can estimate pulse-related metrics (e.g. heart rate and respiratory rate) from facial videos and has a high potential for health monitoring. The latest deep rPPG methods can model in-distribution noise due to head motion, video compression, etc., and estimate high-quality rPPG signals under similar scenarios. However, deep rPPG models may not generalize well to the target test domain with unseen noise and distortions. In this paper, to improve the generalization ability of rPPG models, we propose a dual-bridging network to reduce the domain discrepancy by aligning intermediate domains and synthesizing the target noise in the source domain for better noise reduction. To comprehensively explore the target domain noise, we propose a novel adversarial noise generation in which the noise generator indirectly competes with the noise reducer. To further improve the robustness of the noise reducer, we propose hard noise pattern mining to encourage the generator to learn hard noise patterns contained in the target domain features. We evaluated the proposed method on three public datasets with different types of interferences. Under different crossdomain scenarios, the comprehensive results show the effectiveness of our method.
Jingda Du, Si-Qi Liu 0003, Bochao Zhang, Pong C. Yuen
CVPR2
2023 Automatic Bleeding Risk Rating System of Gastric Varices
Luyue Shi, Guanbin Li, Xiaoguang Han 0001, Si-Qi Liu 0003
MICCAI (5)8
2023 Diffusion-Based Data Augmentation for Nuclei Image Segmentation
Guanbin Li, Wei Lou, Si-Qi Liu 0003, Haofeng Li
MICCAI (8)4
2023 Self- and Semi-supervised Learning for Gastroscopic Lesion Detection
Xuanye Zhang, Kaige Yin, Si-Qi Liu 0003, Zhijie Feng, Xiaoguang Han 0001, Guanbin Li
MICCAI (5)3
2022 Learning Sparse Interpretable Features For NAS Scoring From Liver Biopsy Images
abstract
Liver biopsy images play a key role in the diagnosis of global non-alcoholic fatty liver disease (NAFLD). The NAFLD activity score (NAS) on liver biopsy images grades the amount of histological findings that reflect the progression of NAFLD. However, liver biopsy image analysis remains a challenging task due to its complex tissue structures and sparse distribution of histological findings. In this paper, we propose a sparse interpretable feature learning method (SparseX) to efficiently estimate NAS scores. First, we introduce an interpretable spatial sampling strategy based on histological features to effectively select informative tissue regions containing tissue alterations. Then, SparseX formulates the feature learning as a low-rank decomposition problem. Non-negative matrix factorization (NMF)-based attributes learning is embedded into a deep network to compress and select sparse features for a small portion of tissue alterations contributing to diagnosis. Experiments conducted on the internal Liver-NAS and public SteatosisRaw datasets show the effectiveness of the proposed method in terms of classification performance and interpretability. regions containing tissue alterations. Then, SparseX formulates the feature learning as a low-rank decomposition problem. Non-negative matrix factorization (NMF)-based attributes learning is embedded into a deep network to compress and select sparse features for a small portion of tissue alterations contributing to diagnosis. Experiments conducted on the internal Liver-NAS and public SteatosisRaw datasets show the effectiveness of the proposed method in terms of classification performance and interpretability.
Chong Yin, Si-Qi Liu 0003, Vincent Wai-Sun Wong, Pong C. Yuen
IJCAI2
2022 Learning Temporal Similarity of Remote Photoplethysmography for Fast 3D Mask Face Presentation Attack Detection
abstract
To detect 3D mask face presentation attack, remote Photoplethysmography (rPPG), a biomedical technique that measures the heartbeat signal remotely with a normal RGB camera, is adopted as a robust liveness cue. Although existing rPPG-based solutions exhibit strong performance in experiments, the required observation time is too long (10-12 seconds) to be user-friendly in real applications such as E-payment and smartphone unlock. To shorten the observation time (within 1-second), we propose a fast rPPG-based 3D mask presentation attack detection (PAD) method by analyzing the similarity of rPPG signals in the time domain. In particular, based on facial and background local rPPG signals, we design a set of temporal similarity features to investigate the robust properties of rPPG shape and phase. Following the same direction, we refine the traditional rPPG extractor into a learnable network to cooperate with our TSrPPG feature for better robustness. An effective but lightweight spatiotemporal convolution network is constructed with a self-supervised learning strategy, aiming at enhancing the consistency of genuine facial rPPG signals and reducing the correlation of rPPG signals on masked faces. Extensive experiments are conducted on 3DMAD, HKBU-MARs V1+ and V2+, and CSMAD, which totally involve 18772 short-term video slots with a large number of real-world variations, in terms of mask type, mask transmittance, lighting condition, recording device, resolution of facial region, and compression configuration. Our proposed method persists the good performance of rPPG-based solution with only 1-second observation and outperforms the state-of-the-art competitors on discriminability and generalizability. Evaluations on prints attack, display attack, and disguise attacks with transparent masks, make-up and tattoo further exhibit its potential on handling a wider variety of attacks. To our best knowledge, this is the first work that addresses the length of observation time issue of rPPG-based 3D mask PAD.
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.1
2021 Focusing on Clinically Interpretable Features: Selective Attention Regularization for Liver Biopsy Image Classification
Chong Yin, Si-Qi Liu 0003, Rui Shao 0001, Pong C. Yuen
MICCAI (5)2
2021 Multi-Channel Remote Photoplethysmography Correspondence Feature for 3D Mask Face Presentation Attack Detection
abstract
With the advancement of 3D printing technologies, 3D mask presentation attack becomes a critical challenge in face recognition. To tackle the 3D mask presentation attack detection (PAD), remote Photoplethysmography (rPPG) is employed as an intrinsic detection cue which is independent of the mask material and appearance quality. Although the effectiveness of existing rPPG-based methods has been verified, they may not be robust enough when rPPG signals are contaminated by noise. To identify the heartbeat information from the noisy raw rPPG signals, we propose a new 3D mask PAD feature, multi-channel rPPG correspondence feature (MCCFrPPG) with the global noise-aware template learning and verification framework. To further boost the discriminability, temporal variation of the rPPG signal is considered and extracted through the multi-channel time-frequency analysis scheme. This paper also extends HKBU-MARs V2 dataset with more customized high-quality masks and increases the number of videos by two times. Comprehensive experiments were performed on existing 3D mask datasets and the extended HKBU-MARs V2+, which totally covers 3 types of masks, 12 different light settings and 6 cameras. The results not only justify the effectiveness and robustness of the proposed MCCFrPPG on 3D mask attacks but also indicate its potential on handling the replay attack with camera motion and dim light.
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.1
2020 DATA-GRU: Dual-Attention Time-Aware Gated Recurrent Unit for Irregular Multivariate Time Series
abstract
Due to the discrepancy of diseases and symptoms, patients usually visit hospitals irregularly and different physiological variables are examined at each visit, producing large amounts of irregular multivariate time series (IMTS) data with missing values and varying intervals. Existing methods process IMTS into regular data so that standard machine learning models can be employed. However, time intervals are usually determined by the status of patients, while missing values are caused by changes in symptoms. Therefore, we propose a novel end-to-end Dual-Attention Time-Aware Gated Recurrent Unit (DATA-GRU) for IMTS to predict the mortality risk of patients. In particular, DATA-GRU is able to: 1) preserve the informative varying intervals by introducing a time-aware structure to directly adjust the influence of the previous status in coordination with the elapsed time, and 2) tackle missing values by proposing a novel dual-attention structure to jointly consider data-quality and medical-knowledge. A novel unreliability-aware attention mechanism is designed to handle the diversity in the reliability of different data, while a new symptom-aware attention mechanism is proposed to extract medical reasons from original clinical records. Extensive experimental results on two real-world datasets demonstrate that DATA-GRU can significantly outperform state-of-the-art methods and provide meaningful clinical interpretation.
Qingxiong Tan, Mang Ye, Baoyao Yang, Si-Qi Liu 0003, Andy Jinhua Ma, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
AAAI4
2020 A General Remote Photoplethysmography Estimator with Spatiotemporal Convolutional Network
abstract
Remote PPG (rPPG) has been attracting increasing attention due to its potential in a wide range of application scenarios such as clinical monitoring, physical training and face presentation attack detection. On top of manually designed solutions, the deep-learning approach appears in rPPG estimation and achieves top-level performance. However, most of them try to integrate the steps of both preprocessing and ROI selection into an end-to-end network, which limits the generalization in other applications that use different input skin regions. The ROI selection model learned on face videos may not adapt to skin regions from other body parts with different appearance and size. In this paper, we leave the preprocessing apart and propose a lightweight rPPG estimation network-DeeprPPG for general use. DeeprPPG is based on spatiotemporal convolutions and can be used as a well-defined module in wider application scenarios with different types of input skin. To further boost the robustness, a spatiotemporal rPPG aggregation strategy is designed to adaptively aggregate rPPG signals from multiple skin regions into the final one. Extensive experiments are conducted and the results illustrate its robustness when facing unseen skin regions and unseen scenarios.
Si-Qi Liu 0003, Pong C. Yuen
FG1
2020 Temporal Similarity Analysis of Remote Photoplethysmography for Fast 3D Mask Face Presentation Attack Detection
abstract
To tackle the 3D mask face presentation attack, remote Photoplethysmography (rPPG), a biomedical technique that can detect heartbeat signal remotely, is employed as an intrinsic liveness cue. Although existing rPPG-based methods exhibit encouraging results, they require long observation time (10-12 seconds) to identify the heartbeat information, which limits their employment in real applications such as smartphone unlock and e-payment. To shorten the observation time (within 1-second) while keeping the performance, we propose a fast rPPG-based 3D mask presentation attack detection (PAD) method by analyzing the similarity of local facial rPPG signals in the time domain. In particular, a set of temporal similarity features of facial and background local rPPG signals are designed and fused to adapt the real world variations based on rPPG shape and phase properties. For better evaluation under practical variations, we build the HKBU-MARsV2+ dataset that includes 16 masks from 2 types and 6 lighting conditions. Finally, extensive experiments are conducted on 11092 shortterm video slots from 4 datasets with a large number of real- world variations, in terms of mask type, lighting condition, camera, resolution of face region, and compression setting. Results show that the proposed TSrPPG outperforms the state-of-the-art competitors dramatically on discriminabil- ity and generalizability. To our best knowledge, this is the first work that addresses the length of observation time issue of rPPG-based 3D mask PAD.
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
WACV1
2018 Remote Photoplethysmography Correspondence Feature for 3D Mask Face Presentation Attack Detection
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
ECCV (16)1
2016 3D Mask Face Anti-spoofing with Remote Photoplethysmography
Si-Qi Liu 0003, Pong C. Yuen, Shengping Zhang, Guoying Zhao 0001
ECCV (7)1