VLDB 2026 Research / reviewers in the wild / expert
Xingbo Dong
dblp:237/0125 · also Xing-Bo Dong
· DBLP profile ↗
41ranked-venue papers
10as first author
38since 2021 · last 2026
0000-0001-9782-6068ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 5 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 19 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A triple-head network with loss-aware label assignment for object detection
Lu Leng, Xingbo Dong |
Appl. Intell. | 4 |
| 2026 | PriP: A training-free low-light image enhancement framework via content and illumination synergistic guidance
Yongzhou Liu, Xingbo Dong, Zhe Jin 0001, Wen Sha |
Comput. Graph. | 2 |
| 2026 | ToMo-UDA++: Unsupervised Domain Adaptation for Anatomical Structure Detection Using Enhanced Topology and Morphology Knowledge
Bin Pu, Jiewen Yang, Xingguo Lv, Xingbo Dong, Lei Zhao 0013, Shengli Li 0001, Kenli Li 0001, Xiaomeng Li 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | CED: CLIP-guided entropy dynamics for robust test-time adaptation in harsh visual conditions
Liwen Wang 0002, Xingbo Dong, Yen-Lung Lai, Bin Pu, Zhao Liu 0009, Qika Lin, Zhe Jin 0001 |
Pattern Recognit. | 2 |
| 2026 | EvRWKV: A Continuous Interactive RWKV Framework for Effective Event-Guided Low-Light Image EnhancementabstractEvent cameras offer significant potential for Low-light Image Enhancement (LLIE), yet existing fusion approaches are constrained by a fundamental dilemma: early fusion struggles with modality heterogeneity, while late fusion severs crucial feature correlations. To address these limitations, we propose EvRWKV, a novel framework that enables continuous cross-modal interaction through dual-domain processing, which mainly includes a Cross-RWKV Module to capture fine-grained temporal and cross-modal dependencies, and an Event Image Spectral Fusion Enhancer (EISFE) module to perform joint adaptive frequency-domain denoising and spatial-domain alignment. This continuous interaction maintains feature consistency from low-level textures to high-level semantics. Extensive experiments on the real-world SDE and SDSD datasets demonstrate that EvRWKV significantly outperforms only image-based methods by 1.79 dB and 1.85 dB in PSNR, respectively. To further validate the practical utility of our method for downstream applications, we evaluated its impact on semantic segmentation. Experiments demonstrate that images enhanced by EvRWKV lead to a significant 35.44% improvement in mIoU. WenJie Cai, Qingguo Meng, Xingbo Dong, Zhe Jin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Collaborative Coarse-to-Fine Disease Learning With Discharge Summary Awareness for EHR Event PredictionabstractDeep learning-based models have been widely used to predict electronic health record (EHR) events by exploiting diagnostic characteristics. Despite significant progress, three limitations remain: 1) effectively modeling dynamic relationships among diseases, 2) fully leveraging diagnosis code ontologies from multiple perspectives, and 3) incorporating unstructured discharge summaries. To address these challenges, we propose a coarse-to-fine disease learning framework with patient notes for EHR event prediction, tailored to capture both dynamic and static disease characteristics. First, we construct a fine-grained dynamic disease graph by removing disease weakly correlated disease pairs based on co-occurrence distributions. Second, disease embeddings are refined by integrating coarse and fine-grained information within the hierarchical structure of ICD-9-CM codes. In addition, discharge summaries are combined with auxiliary patient notes for collaborative disease learning. Finally, gated recurrent units, location-based attention, and soft attention mechanisms are utilized to further enhance embedding representations. Experiments on two real-world EHR datasets, MIMIC-III and MIMIC-IV, demonstrate that our model consistently outperforms nine baseline methods in EHR prediction. The source code can be found at https://github.com/YNU-L/CCDLD. Yan Kang 0003, Zhuolun Li, Bin Pu, Xingbo Dong, Jiewen Yang, Lei Zhao 0013, Benteng Ma, Ningshu Li, Jianguo Chen 0001, Philip S. Yu |
IEEE Trans. Cybern. | 4 |
| 2026 | ÆMMamba: An Efficient Medical Segmentation Model With Edge EnhancementabstractMedical image segmentation is critical for disease diagnosis, treatment planning, and prognosis assessment, yet the complexity and diversity of medical images pose significant challenges to accurate segmentation. While Convolutional Neural Networks capture local features and Vision Transformers excel in the global context, both struggle with efficient long-range dependency modeling. Inspired by Mamba's State Space Modeling efficiency, we propose ÆMMamba, a novel multi-scale feature extraction framework built on the Mamba backbone network. ÆMMamba integrates several innovative modules: the Efficient Fusion Bridge (EFB) module, which employs a bidirectional state-space model and attention mechanisms to fuse multi-scale features; the Edge-Aware Module (EAM), which enhances low-level edge representation using Sobel-based edge extraction; and the Boundary Sensitive Decoder (BSD), which leverages inverse attention and residual convolutional layers to handle cross-level complex boundaries. ÆMMamba achieves state-of-the-art performance across 8 medical segmentation datasets. On polyp segmentation datasets (Kvasir, ClinicDB, ColonDB, EndoScene, ETIS), it records the highest mDice and mIoU scores, outperforming methods like MADGNet and Swin-UMamba, with a standout mDice of 72.22 on ETIS, the most challenging dataset in this domain. For lung and breast segmentation, ÆMMamba surpasses competitors such as H2Former and SwinUnet, achieving Dice scores of 84.24 on BUSI and 79.83 on COVID-19 Lung. And on the LGG brain MRI dataset, ÆMMamba attains an mDice of 87.25 and an mIoU of 79.31, outperforming all compared methods. Xingbo Dong, Iman Yi Liao, Zhe Jin 0001, Zhaozhao Xu, Bin Pu |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Leveraging Anatomical Consistency for Multi-Object Detection in Ultrasound Images via Source-free Unsupervised Domain AdaptationabstractSource-free unsupervised domain adaptation aims to eliminate domain shifts when data from the source domain and annotation from the target domain are not available. The multi-object detection tasks in medical image analysis are constrained by patient privacy and extremely huge annotation consumption. Hence, Source-free UDA is considered a more practical approach for eliminating the domain gap. However, relevant research that explores this topic is a dearth. In this paper, we design an Anatomy-aware Alignment Teacher-Student learning method using topological consistency based on a mean-teacher framework for Source-free UDA in multiple medical object detection named AATS, including Unsupervised Structure Refinement (USR) and Graph-aware Morphology Alignment (GMA). To match the student and teacher at the low-level and visual features, we propose the USR via an unsupervised clustering algorithm to group organs in ultrasound images. Based on USR, we obtain a graph with organ relations on the teacher branch. While in the student branch, we acquire visual features to construct graphical space and optimize the model with graph propagation. Finally, to match the student and teacher, GMA is designed to align the teacher and student based on both topology and morphology information that is derived from prior medical knowledge. Four groups of adaptation experiments were conducted on available medical datasets, and the outcomes demonstrate that our approach not only achieves state-of-the-art performance but also provides substantial advantages over existing methods. Bin Pu, Xingguo Lv, Jiewen Yang, Xingbo Dong, Yiqun Lin, Shengli Li 0001, Kenli Li 0001, Xiaomeng Li 0001 |
AAAI | 4 |
| 2025 | Anatomical Knowledge Mining and Matching for Semi-supervised Medical Multi-structure DetectionabstractIn medical image analysis, detecting multiple structures is crucial for evaluations and diagnosis but is often limited by the lack of high-quality annotations. Semi-supervised object detection emerges as a potent methodology to enhance model performance and generalization by leveraging a vast pool of unlabeled data alongside a minimal set of labeled data. A striking observation is that both unlabelled and labeled medical images contain a priori anatomical knowledge from human screening. In this work, we introduce a novel semi-supervised approach named Semi-akmm for mining and matching anatomical knowledge in ultrasound images. We develop an Adaptive Prior Knowledge Transfer (APKT) module to mine and explore the distribution and knowledge of potential proposal boxes by proposal proportion constraint. Furthermore, within a teacher-student learning framework, we put forward an Anatomical Structure Matching (ASM) module to facilitate co-learning consistent topological prior knowledge between the student and teacher models. To our knowledge, this marks the inception of an efficient semi-supervised medical multi-structure detection model. Our experiments across five publicly available ultrasound datasets demonstrate that Semi-akmm sets a new benchmark in performance with solid results that outperform existing methods. Bin Pu, Liwen Wang 0002, Jiewen Yang, Xingbo Dong, Benteng Ma, Zhuangzhuang Chen, Lei Zhao 0013, Shengli Li 0001, Kenli Li 0001 |
AAAI | 4 |
| 2025 | Test-Time Domain Generalization via Universe Learning: A Multi-Graph Matching Approach for Medical Image SegmentationabstractDespite domain generalization (DG) has significantly addressed the performance degradation of pre-trained models caused by domain shifts, it often falls short in real-world deployment. Test-time adaptation (TTA), which adjusts a learned model using unlabeled test data, presents a promising solution. However, most existing TTA methods struggle to deliver strong performance in medical image segmentation, primarily because they overlook the crucial prior knowledge inherent to medical images. To address this challenge, we incorporate morphological information and propose a framework based on multi-graph matching. Specifically, we introduce learnable universe embeddings that integrate morphological priors during multi-source training, along with novel unsupervised test-time paradigms for domain adaptation. This approach guarantees cycle-consistency in multi-matching while enabling the model to more effectively capture the invariant priors of unseen data, significantly mitigating the effects of domain shifts. Extensive experiments demonstrate that our method outperforms other state-of-the-art approaches on two medical image segmentation benchmarks for both multi-source and single-source domain generalization tasks. The source code is available at https://github.com/Yore0/TTDG-MGM. Xingguo Lv, Xingbo Dong, Liwen Wang 0002, Jiewen Yang, Lei Zhao 0013, Bin Pu, Zhe Jin 0001, Xuejun Li 0001 |
CVPR | 2 |
| 2025 | Learning to Zoom with Anatomical Relations for Medical Structure DetectionabstractAccurate anatomical structure detection is a critical preliminary step for diagnosing diseases characterized by structural abnormalities. In clinical practice, medical experts frequently adjust the zoom level of medical images to obtain comprehensive views for diagnosis. This common interaction results in significant variations in the apparent scale of anatomical structures across different images or fields of view. However, the information embedded in these zoom-induced scale changes is often overlooked by existing detection algorithms.
In addition, human organs possess a priori, fixed topological knowledge. To overcome this limitation, we propose ZR-DETR, a zoom-aware probabilistic framework tailored for medical object detection. ZR-DETR uniquely incorporates scale-sensitive zoom embeddings, anatomical relation constraints, and a Gaussian Process-based detection head. This architecture enables the framework to jointly model semantic context, enforce anatomical plausibility, and quantify detection uncertainty. Empirical validation across three diverse medical imaging benchmarks demonstrates that ZR-DETR consistently outperforms strong baselines in both single-domain and unsupervised domain adaptation scenarios. Bin Pu, Liwen Wang 0002, Xingbo Dong, Xingguo Lv, Zhe Jin 0001 |
NeurIPS | 3 |
| 2025 | CSP-SAM: CNN-Enhanced and Self-prompting SAM for Ultrasound Anatomical Structure Segmentation
Xingbo Dong, Bocheng Liang, Bin Pu, Zhe Jin 0001 |
PRCV (13) | 2 |
| 2025 | Rethinking Contemporary Deep Learning Techniques for Error Correction in Biometric Data
Yen-Lung Lai, Xingbo Dong, Zhe Jin 0001, Wei Jia 0001, Massimo Tistarelli, Xuejun Li 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Low-light image enhancement with luminance duality
Xingguo Lv, Xingbo Dong, Jiewen Yang, Lei Zhao 0013, Bin Pu, Zhe Jin 0001 |
Knowl. Based Syst. | 2 |
| 2025 | Single source domain generalization for palm biometrics
Congcong Jia, Xingbo Dong, Yen-Lung Lai, Andrew Beng Jin Teoh, Ziyuan Yang 0001, Liwen Wang 0002, Zhe Jin 0001, Lianqiang Yang |
Pattern Recognit. | 2 |
| 2025 | IFViT: Interpretable Fixed-Length Representation for Fingerprint Matching via Vision TransformerabstractDetermining dense feature points on fingerprints used in constructing deep fixed-length representations for accurate matching, particularly at the pixel level, is of significant interest. To explore the interpretability of fingerprint matching, we propose a multi-stage interpretable fingerprint matching network, namely Interpretable Fixed-length Representation for Fingerprint Matching via Vision Transformer (IFViT), which consists of two primary modules. The first module, an interpretable dense registration module, establishes a Vision Transformer (ViT)-based Siamese Network to capture long-range dependencies and the global context in fingerprint pairs. It provides interpretable dense pixel-wise correspondences of feature points for fingerprint alignment and enhances the interpretability in the subsequent matching stage. The second module takes into account both local and global representations of the aligned fingerprint pair to achieve an interpretable fixed-length representation extraction and matching. It employs the ViTs trained in the first module with the additional fully connected layer and retrains them to simultaneously produce the discriminative fixed-length representation and interpretable dense pixel-wise correspondences of feature points. Extensive experimental results on diverse publicly available fingerprint databases demonstrate that the proposed framework not only exhibits superior performance on dense registration and matching but also significantly promotes the interpretability in deep fixed-length representations-based fingerprint matching. Honghui Chen, Xingbo Dong, Zheng Lin 0001, Iman Yi Liao, Massimo Tistarelli, Zhe Jin 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | M3-UDA: A New Benchmark for Unsupervised Domain Adaptive Fetal Cardiac Structure DetectionabstractThe anatomical structure detection of fetal cardiac views is crucial for diagnosing fetal congenital heart disease. In practice, there is a large domain gap between different hospitals' data, such as the variable data quality due to differences in acquisition equipment. In addition, accurate annotation information provided by obstetrician experts is always very costly or even unavailable. This study explores the unsupervised domain adaptive fetal cardiac structure detection issue. Existing unsupervised domain adaptive object detection (UDAOD) approaches mainly focus on detecting objects in natural scenes, such as Foggy Cityscapes, where the structural relationships of natural scenes are uncertain. Unlike all previous UDAOD scenarios, we first collected a Fetal Cardiac Structure dataset from two hospital centers, called FCS, and proposed a multi-matching UDA approach (M3-UDA), including Histogram Matching (HM), Sub-structure Matching (SM), and Global-structure Matching (GM), to better transfer the topological knowledge of anatomical structure for UDA detection in medical scenarios. HM mitigates the domain gap between the source and target caused by pixel transformation. SM fuses the different angle information of the sub-structure to obtain the local topological knowledge for bridging the domain gap of the internal sub-structure. GM is designed to align the global topological knowledge of the whole organ from the source and target domain. Extensive experiments on our collected FCS and CardiacUDA, and experimental results show that M3-UDA outperforms existing UDAOD studies significantly. Datasets and source code are available at https://github.com/xmed-lab/M3-UDA. Bin Pu, Liwen Wang 0002, Jiewen Yang, Guannan He, Xingbo Dong, Shengli Li 0001, Zhe Jin 0001, Kenli Li 0001, Xiaomeng Li 0001 |
CVPR | 5 |
| 2024 | Validating Privacy-Preserving Face Recognition Under a Minimum AssumptionabstractThe widespread use of cloud-based face recognition technology raises privacy concerns, as unauthorized access to face images can expose personal information or be exploited for fraudulent purposes. In response, privacy-preserving face recognition (PPFR) schemes have emerged to hide visual information and thwart unauthorized access. However, the validation methods employed by these schemes often rely on unrealistic assumptions, leaving doubts about their true effectiveness in safeguarding facial privacy. In this paper, we introduce a new approach to pri-vacy validation called Minimum Assumption Privacy Protection Validation (Map2 V). This is the first exploration of formulating a privacy validation method utilizing deep image priors and zeroth-order gradient estimation, with the potential to serve as a general framework for PPFR eval-uation. Building upon Map2v, we comprehensively vali-date the privacy-preserving capability of PPFRs through a combination of human and machine vision. The exper-iment results and analysis demonstrate the effectiveness and generalizability of the proposed Map2v, showcasing its superiority over native privacy validation methods from PPFR works of literature. Additionally, this work exposes privacy vulnerabilities in evaluated state-of-the-art P P FR schemes, laying the foundation for the subsequent effective proposal of countermeasures. The source code is available at https://github.com/Beauty9882/MAP2V. Hui Zhang 0039, Xingbo Dong, Yen-Lung Lai, Xingguo Lv, Zhe Jin 0001, Xuejun Li 0001 |
CVPR | 2 |
| 2024 | Face Reconstruction Transfer Attack as Out-of-Distribution Generalization
Yoon Gyo Jung, Jaewoo Park 0001, Xingbo Dong, Hojin Park, Andrew Beng Jin Teoh, Octavia I. Camps |
ECCV (75) | 3 |
| 2024 | Unsupervised Domain Adaptation for Anatomical Structure Detection in Ultrasound ImagesabstractModels trained on ultrasound images from one institution typically experience a decline in effectiveness when transferred directly to other institutions. Moreover, unlike natural images, dense and overlapped structures exist in fetus ultrasound images, making the detection of structures more challenging. Thus, to tackle this problem, we propose a new Unsupervised Domain Adaptation (UDA) method named ToMo-UDA for fetus structure detection, which consists of the Topology Knowledge Transfer (TKT) and the Morphology Knowledge Transfer (MKT) module. The TKT leverages prior knowledge of the medical anatomy of fetal as topological information, reconstructing and aligning anatomy features across source and target domains. Then, the MKT formulates a more consistent and independent morphological representation for each substructure of an organ. To evaluate the proposed ToMo-UDA for ultrasound fetal anatomical structure detection, we introduce FUSH$^2$, a new Fetal UltraSound benchmark, comprises Heart and Head images collected from Two health centers, with 16 annotated regions. Our experiments show that utilizing topological and morphological anatomy information in ToMo-UDA can greatly improve organ structure detection. This expands the potential for structure detection tasks in medical image analysis. Bin Pu, Xingguo Lv, Jiewen Yang, Guannan He, Xingbo Dong, Yiqun Lin, Shengli Li 0001, Tan Ying, Zhe Jin 0001, Kenli Li 0001, Xiaomeng Li 0001 |
ICML | 5 |
| 2024 | Learning Frequency and Structure in UDA for Medical Object Detection
Liwen Wang 0002, Guannan He, Shengli Li 0001, Bin Pu, Zhe Jin 0001, Wen Sha, Xingbo Dong |
PRCV (14) | 9 |
| 2024 | TATrack: Target-aware transformer for object tracking
Lu Leng, Xingbo Dong |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | A spatial-temporal contexts network for object tracking
Lu Leng, Xingbo Dong |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Video-based face outline recognition
Xingbo Dong, Jiewen Yang, Andrew Beng Jin Teoh, Dahai Yu 0001, Xiaomeng Li 0001, Zhe Jin 0001 |
Pattern Recognit. | 1 |
| 2024 | MFISN: Modality Fuzzy Information Separation Network for Disease ClassificationabstractMost of the previous machine learning-based models for multi-modal medical diagnosis, primarily designed for unimodal images, usually do not fully leverage the potential of multimodal medical images, leading to limited classification accuracy. These conventional methods typically focus only on the intermodality common information, neglecting the intra-modality specific information and assuming that the common information is more effective in disease diagnosis. Moreover, they do not adequately address the impact of fuzzy information between different medical imaging modalities on diagnostic results. To this end, we propose a Modality Fuzzy Information Separation Network for disease classification, which extracts both common and specific information from fuzzy information to construct a comprehensive representation of multi-modal medical images. Specifically, we extract modality invariant features as common information by explicitly modeling and maximizing loss constraints on mutual information. For specific information extraction, a constraint on feature space independence between specific and common information is imposed on each modality. Above two steps, we concatenate common information and specific information to construct a comprehensive multi-modal representation for separating fuzzy information. Finally, we purposely design a decoder network to reconstruct medical images from uni-modal specific information and common information to demonstrate the effectiveness of the modality fuzzy information separation network. We conducted a validation of the proposed method's performance in classifying cardiomegaly, pneumothorax, edema, and skin disease. The experimental results substantiate the effectiveness of our proposed approach. Fengtao Nan, Bin Pu, Yingchun Fan, Jiewen Yang, Xingbo Dong, Zhaozhao Xu, Shuihua Wang |
IEEE Trans. Fuzzy Syst. | 6 |
| 2023 | Towards Query Efficient and Generalizable Black-Box Face Reconstruction AttackabstractIn this paper, we address the black-box face reconstruction attack with two crucial requirements: query efficiency and generalizability. A practical attack must be query efficient due to limited access to the target black-box model, and the reconstructed face must be generalizable so it can be used to attack other face recognition systems. To this end, we propose a novel face reconstruction attack that optimizes the latent vector of a pre-trained StyleGAN generator. Unlike existing methods, our method is query efficient as neither training nor simultaneous updating of multiple latent vectors is required. Furthermore, we propose a simple initialization scheme that greatly enhances the generalizability of the proposed method. We demonstrate the effectiveness of our method by a thorough evaluation on LFW and CFP-FP datasets across multiple state-of-the-art face recognition models. Project Code: github.com/1ho0jin1/Black-box-Face-Reconstruction. Hojin Park, Jaewoo Park 0001, Xingbo Dong, Andrew Beng Jin Teoh |
ICIP | 3 |
| 2023 | Minimum Assumption Reconstruction Attacks: Rise of Security and Privacy Threats Against Face Recognition
Hojin Park, Xingbo Dong, Yen-Lung Lai, Hui Zhang 0039, Andrew Beng Jin Teoh, Zhe Jin 0001 |
PRCV (5) | 3 |
| 2023 | L2DM: A Diffusion Model for Low-Light Image Enhancement
Xingguo Lv, Xingbo Dong, Zhe Jin 0001, Hui Zhang 0039, Siyi Song, Xuejun Li 0001 |
PRCV (11) | 2 |
| 2023 | A Video Face Recognition Leveraging Temporal Information Based on Vision Transformer
Hui Zhang 0039, Jiewen Yang, Xingbo Dong, Xingguo Lv, Wei Jia 0001, Zhe Jin 0001, Xuejun Li 0001 |
PRCV (5) | 3 |
| 2023 | Reconstruct face from features based on genetic algorithm using GAN generator as a distribution constraint
Xingbo Dong, Zhihui Miao, Zhe Jin 0001, Zhenhua Guo 0001, Andrew Beng Jin Teoh |
Comput. Secur. | 1 |
| 2023 | Breaking Free From Entropy's Shackles: Cosine Distance-Sensitive Error Correction for Reliable Biometric CryptographyabstractBiometric cryptosystems present a promising avenue for secure authentication; however, the efficiency and security of such systems can be hindered by errors in biometric data. To address this challenge, existing systems employ error-correction codes, but often fail to consider the distribution of biometric sources, potentially leading to an underestimation of the system’s security. In response to this issue, we propose a novel algorithm pair, designated as ENCODE and DECODE, which facilitates direct codeword generation from biometric samples. Our approach accounts for the distribution of biometric sources, thereby providing a more accurate estimation of system security compared to traditional methods. Our proposed algorithm pair generates codewords that maintain interpretability and are sensitive to the cosine distance between original biometric samples. This similarity metric is particularly well-suited for high-dimensional data analysis and enables a precise assessment of system performance. We have rigorously established the correctness of our algorithm pair, and empirical results illustrate its efficacy in tolerating distance between codewords while preserving accuracy in cosine distance-sensitive contexts. This approach has the potential to significantly improve the efficiency and security of biometric cryptosystems, rendering them more appropriate for daily cryptographic applications. Yen-Lung Lai, Xingbo Dong, Zhe Jin 0001, Massimo Tistarelli, Wun-She Yap, Bok-Min Goi |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Abandoning the Bayer-Filter to See in the DarkabstractLow-light image enhancement, a pervasive but challenging problem, plays a central role in enhancing the visibility of an image captured in a poor illumination environment. Due to the fact that not all photons can pass the Bayer-Filter on the sensor of the color camera, in this work, we first present a De-Bayer-Filter simulator based on deep neural networks to generate a monochrome raw image from the colored raw image. Next, a fully convolutional network is proposed to achieve the low-light image enhancement by fusing colored raw data with synthesized monochrome data. Channel-wise attention is also introduced to the fusion process to establish a complementary interaction between features from colored and monochrome raw images. To train the convolutional networks, we propose a dataset with monochrome and color raw pairs named Mono-Colored Raw paired dataset (MCR) collected by using a monochrome camera without Bayer-Filter and a color camera with Bayer-Filter. The proposed pipeline takes advantages of the fusion of the virtual monochrome and the color raw images, and our extensive experiments indicate that significant improvement can be achieved by leveraging raw sensor data and data-driven learning. The project is available at https://github.com/TCL-AILab/Abandon_Bayer-Filter_See_in_the_Dark. Xingbo Dong, Wanyan Xu 0001, Zhihui Miao, Jiewen Yang, Zhe Jin 0001, Andrew Beng Jin Teoh |
CVPR | 1 |
| 2022 | Recurring the Transformer for Video Action RecognitionabstractExisting video understanding approaches, such as 3D convolutional neural networks and Transformer-Based methods, usually process the videos in a clip-wise manner; hence huge GPU memory is needed and fixed-length video clips are usually required. To alleviate those issues, we introduce a novel Recurrent Vision Transformer (RViT) framework based on spatial-temporal representation learning to achieve the video action recognition task. Specifically, the proposed RViT is equipped with an attention gate to build interaction between current frame input and previous hidden state, thus aggregating the global level interframe features through the hidden state temporally. RViT is executed recurrently to process a video by giving the current frame and previous hidden state. The RViT can capture both spatial and temporal features because of the attention gate and recurrent execution. Besides, the proposed RViT can work on variant-length video clips properly without requiring large GPU memory thanks to the frame by frame processing flow. Our experiment results demonstrate that RViT can achieve state-of-the-art performance on various datasets for the video recognition task. Specifically, RViT can achieve a top-1 accuracy of 81.5% on Kinetics-400, 92.31% on Jester, 67.9% on Something-Something-V2, and an mAP accuracy of 66.1% on Charades. Jiewen Yang, Xingbo Dong, Liujun Liu, Dahai Yu 0001 |
CVPR | 2 |
| 2022 | Co-Learning to Hash Palm Biometrics for Flexible IoT DeploymentabstractSecurity enhancement via trustworthy identity authentication in Internet of Things (IoT) has soared recently. Biometrics offers a promising remedy to improve the security and utility of IoT and play a role in securing a variety of low-power and limited computing capability IoT devices to address identity management challenges. This article proposes an IoT-compliant co-learned biometric hashing network derived from palm print and palm vein dubbed PalmCohashNet. The PalmCohashNet comprises two hashing subnetworks, one for each palm modality, and is trained collaboratively to generate shared hash codes for respective modality (co-hash codes). A cross-modality hashing (CMH) loss is devised to encourage co-hash codes of palm vein and palm print from the same identity to be adjacent and consistent; meanwhile, pull the co-hash codes of each identity to a preassigned identity-specific hash centroid that is shared by both palm modalities. Two palm-based co-hash codes of a person can be generated simultaneously for deployment. The binary co-hash code is IoT compliant attributed to its highly compact form for storage and fast matching. A trained PalmCohashNet can be flexibly deployed under four operation modes: single-modality matching (print versus print or vein versus vein), multimodality matching where both print and vein are utilized, and cross-modality matching (print versus vein) depending on the IoT service context. Our empirical results on four publicly available palm databases show that the proposed method consistently outperforms state-of-the-art methods. Xingbo Dong, Muhammad Khurram Khan, Lu Leng, Andrew Beng Jin Teoh |
IEEE Internet Things J. | 1 |
| 2022 | Deep rank hashing network for cancellable face identification
Xingbo Dong, Sangrae Cho, Youngsam Kim, Soohyung Kim, Andrew Beng Jin Teoh |
Pattern Recognit. | 1 |
| 2022 | RawFormer: An Efficient Vision Transformer for Low-Light RAW Image EnhancementabstractLow-light image enhancement plays a central role in various downstream computer vision tasks. Vision Transformers (ViTs) have recently been adapted for low-level image processing and have achieved a promising performance. However, ViTs process images in a window- or patch-based manner, compromising their computational efficiency and long-range dependency. Additionally, existing ViTs process RGB images instead of RAW data from sensors, which is sub-optimal when it comes to utilizing the rich information from RAW data. We propose a fully endto-end Conv-Transformer-based model, RawFormer, to directly utilize RAW data for low-light image enhancement. RawFormer has a structure similar to that of U-Net, but it is integrated with a thoughtfully designed Conv-Transformer Fusing (CTF) block. The CTF block combines local attention and transposed selfattention mechanisms in one module and reduces the computational overhead by adopting a transposed self-attention operation. Experiments demonstrate that RawFormer outperforms state-ofthe-art models by a significant margin on low-light RAW image enhancement tasks. Wanyan Xu 0001, Xingbo Dong, Andrew Beng Jin Teoh, Zhixian Lin |
IEEE Signal Process. Lett. | 2 |
| 2021 | BioCanCrypto: An LDPC Coded Bio-Cryptosystem on Fingerprint Cancellable TemplateabstractBiometrics as a means of personal authentication has demonstrated strong viability in the past decade. However, directly deriving a unique cryptographic key from biometric data is a non-trivial task due to the fact that biometric data is usually noisy and presents large intra-class variations. Moreover, biometric data is permanently associated with the user, which leads to security and privacy issues. Cancellable biometrics and bio-cryptosystem are two main branches to address those issues, yet both approaches fall short in terms of accuracy performance, security, and privacy. In this paper, we propose a Bio-Crypto system on fingerprint Cancellable template (Bio-CanCrypto), which bridges cancellable biometrics and bio-cryptosystem to achieve a middle-ground for alleviating the limitations of both. Specifically, a cancellable transformation is applied on a fixed-length fingerprint feature vector to generate cancellable templates. Next, an LDPC coding mechanism is introduced into a reusable fuzzy extractor scheme and used to extract the stable cryptographic key from the generated cancellable templates. The proposed system can achieve both cancellability and reusability in one scheme. Experiments are conducted on a public fingerprint dataset, i.e., FVC2002. The results demonstrate that the proposed LDPC coded reusable fuzzy extractor is effective and promising. Xingbo Dong, Zhe Jin 0001, Leshan Zhao, Zhenhua Guo 0001 |
IJCB | 1 |
| 2021 | Secure Chaff-less Fuzzy Vault for Face Identification SystemsabstractBiometric cryptosystems such as fuzzy vaults represent one of the most popular approaches for secret and biometric template protection. However, they are solely designed for biometric verification, where the user is required to input both identity credentials and biometrics. Several practical questions related to the implementation of biometric cryptosystems remain open, especially in regard to biometric template protection. In this article, we propose a face cryptosystem for identification (FCI) in which only biometric input is needed. Our FCI is composed of a one-to-N search subsystem for template protection and a one-to-one match chaff-less fuzzy vault (CFV) subsystem for secret protection. The first subsystem stores N facial features, which are protected by index-of-maximum (IoM) hashing, enhanced by a fusion module for search accuracy. When a face image of the user is presented, the subsystem returns the top k matching scores and activates the corresponding vaults in the CFV subsystem. Then, one-to-one matching is applied to the k vaults based on the probe face, and the identifier or secret associated with the user is retrieved from the correct matched vault. We demonstrate that coupling between the IoM hashing and the CFV resolves several practical issues related to fuzzy vault schemes. The FCI system is evaluated on three large-scale public unconstrained face datasets (LFW, VGG2, and IJB-C) in terms of its accuracy, computation cost, template protection criteria, and security. Xingbo Dong, Soohyong Kim, Zhe Jin 0001, Jung Yeon Hwang, Sangrae Cho, Andrew Beng Jin Teoh |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Cross-spectrum Face Recognition Using Subspace Projection HashingabstractCross-spectrum face recognition, e.g. visible to thermal matching, remains a challenging task due to the large variation originated from different domains. This paper proposed a subspace projection hashing (SPH) to enable the cross-spectrum face recognition task. The intrinsic idea behind SPH is to project the features from different domains onto a common subspace, where matching the faces from different domains can be accomplished. Notably, we proposed a new loss function that can (i) preserve both inter-domain and intra-domain similarity; (ii) regularize a scaled-up pairwise distance between hashed codes, to optimize projection matrix. Three datasets, Wiki, EURECOM VIS-TH paired face and TDFace are adopted to evaluate the proposed SPH. The experimental results indicate that the proposed SPH outperforms the original linear subspace ranking hashing (LSRH) in the benchmark dataset (Wiki) and demonstrates a reasonably good performance for visible-thermal, visible-near-infrared face recognition, therefore suggests the feasibility and effectiveness of the proposed SPH. Hanrui Wang 0003, Xingbo Dong, Zhe Jin 0001, Jean-Luc Dugelay, Massimo Tistarelli |
ICPR | 2 |
| 2020 | Open-set face identification with index-of-max hashing by learning
Xingbo Dong, Soohyung Kim, Zhe Jin 0001, Jung Yeon Hwang, Sangrae Cho, Andrew Beng Jin Teoh |
Pattern Recognit. | 1 |
| 2019 | A Secure Visual-thermal Fused Face Recognition System Based on Non-Linear HashingabstractIn this paper, we propose a secure visual-thermal fused face recognition system using non-linear hashing. To extract features from both thermal and visible facial images, a deep neural network model pre-trained by visible images, namely InsightFace, is utilized in extracting deep features from both thermal and visible images. Next, we investigate into the effectiveness of using nonlinear hashing in protecting deep features extracted from both thermal and visible face images. To further boost the accuracy performance of the facial recognition system under unfavorable environment, feature- and score-level fusion of thermal and visible images for face matching are studied. The performance of different application scenarios are tested on the EURECOM VIS-TH face dataset. Experiment results suggest that: 1) feature- and score-level fusion techniques are effective in achieving higher accuracy under unfavorable situation; 2) non-linear hashing offers additional layer of protection, namely, privacy preservation, to face image. We also found that the deep model trained by using visible images is applicable to thermal images for feature extraction, which is particularly useful because there is no large thermal dataset available to train deep neural network. Xingbo Dong, Koksheik Wong, Zhe Jin 0001, Jean-Luc Dugelay |
MMSP | 1 |