VLDB 2026 Research / reviewers in the wild / expert
Sijie Niu
dblp:152/0029
· DBLP profile ↗
62ranked-venue papers
4as first author
50since 2021 · last 2026
0000-0002-1401-9859ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 3 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 12 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Entropy-aware Mutual Student Co-training for Source-free Domain Adaptation in Medical Image SegmentationabstractUnsupervised Domain Adaptation (UDA) has shown remarkable success in medical image segmentation but its practical deployment is often constrained by strict privacy regulations that prohibit access to source-domain data. Source-Free Domain Adaptation (SFDA) addresses this limitation by enabling knowledge transfer from pre-trained source models without requiring private source data. Existing SFDA methods for medical image segmentation mainly rely on self-training with pseudo-labels generated by the source model. However, severe domain discrepancies introduce substantial pseudo-label noise, while non-uniform domain shifts across target samples lead to heterogeneous data distributions, rendering uniform adaptation strategies suboptimal. To address these challenges, we propose Entropy-aware Mutual Student Co-training (EMSC), a novel SFDA framework comprising one teacher model and two student models. An Entropy-Guided Difficulty Assessment (EGDA) module partitions target samples into source-similar and source-dissimilar subsets based on predictive uncertainty. Each subset is handled by an independent student branch equipped with a Subset-specific Adaptive Injection (SAI) module, which injects semantic anchors for source-similar samples and structural noise for source-dissimilar samples to enable differentiated adaptation and suppress pseudo-label noise. Extensive experiments on two widely used medical image segmentation benchmarks demonstrate that EMSC consistently outperforms state-of-the-art SFDA methods across multiple evaluation metrics, validating its robustness under heterogeneous domain shifts. Guangrun Chen, Litan Sun, Ming Jin 0007, Xiaofeng Qu, Sijie Niu |
ICMR | 6 |
| 2026 | Enhancing Diffusion Models Towards Anomaly-Aware Reconstructions for Medical Image Anomaly DetectionabstractABSTRACT Diffusion‐based unsupervised anomaly detection in medical images has emerged as an effective paradigm, leveraging unlabelled healthy data to precisely characterize the distribution of normal anatomy and identify a wide range of pathological abnormalities. The method reconstructs a pseudo‐healthy image from a potentially anomalous input and identifies anomalies by measuring pixel‐wise reconstruction errors. However, existing approaches often preserve anomalous regions in the reconstruction, resulting in less prominent anomaly segmentation. Additionally, their inability to accurately restore normal areas can lead to increased false positives. In this work, we propose CS‐Unet to advance this paradigm by realizing the concept of anomaly‐aware reconstruction, defined as reconstructions that are consciously devoid of anomalies while faithfully restoring normal regions. Firstly, we propose a compression‐expansion DenseNet (CompExDenseNet), which performs a dense cascade of nonlinear dimension transformations to extract compact feature representations, suppressing the reconstruction of anomalous patterns. Secondly, we design an attention gate (AG) unit to control the flow of low‐frequency information, mitigating the leakage of anomalous information. Finally, we propose a frequency‐domain adaptive residual convolution (FreAR) module that selectively enhances the most relevant frequency components to facilitate high‐fidelity restoration of normal regions. Experimental results demonstrate that CS‐Unet achieves outstanding performance in unsupervised anomaly detection, confirming its effectiveness. Wanying Wu, Xiaofeng Qu, Fenghang Zhang, Xizhan Gao, Sijie Niu |
IET Image Process. | 6 |
| 2026 | From discrete to continuous: A spatiotemporal evolution-aware adversarial diffusion framework for retinal disease progression prediction
Yuhan Zhang 0001, Sijie Niu, Songtao Yuan, Qiang Chen 0004 |
Neurocomputing | 3 |
| 2026 | MCD4SR: Multimodal collaborative denoising with modality balancing for sequential recommendation
Xin Zhang 0079, Yinzhuo Chen, Shengan Wang, Dongjing Wang, Yingjie Xia, Sijie Niu, Butian Huang, Yuyu Yin |
Knowl. Based Syst. | 7 |
| 2026 | Few-shot medical anomaly detection through centroid consultation back and test-time self-calibration
Zihan Nie, Muhao Xu, Hua Wei 0007, Sijie Niu, Yi Wan 0002, Xunbin Wei, Weiye Song |
Pattern Recognit. | 6 |
| 2026 | FGCLIP-Based Augmented Language-Driven Contrastive Clustering Network for Fine-Grained Image ClusteringabstractFine-grained image clustering (FGIC) is a highly challenging task due to the large intra-class variance, small inter-class variance, and lack of annotation, aiming at grouping images into fine-grained subcategories. Existing FGIC methods generally learn parameterized localization networks to capture key objects for better clustering performance. Despite yielding promising improvements, these methods still have limitations. First, localization networks introduce additional parameters, and not all localized regions are beneficial for clustering. Second, FGIC requires more detailed semantic descriptions, however, current methods only mine supervisory signals from images, making it difficult to meet practical demands. For addressing these limitations, this paper proposes the FGCLIP-based augmented language-driven contrastive clustering (FGCLIP-ALCC) network, which uses frozen FG-CLIP to introduce external knowledge and designs parameter-efficient text branch to accurately locate key objects and learn fine-grained text semantics. More specifically, FGCLIP-ALCC contains two components: the augmented language-driven fine-grained semantic learner (ALFSL) and the multi-modal contrastive clustering heads (MmCCH). First, the ALFSL is designed with a three-stream architecture to enhance robustness, ensure the diversity and accuracy of text descriptions generated subsequently, and utilize image branches to extract visual features. Next, the text branch in ALFSL uses an augmentation-driven diverse text generation module to generate coarse-grained text descriptions, a text fine-graining module to capture key object semantics and refine the text descriptions, and a text filtering module along with a text fusion operation to enhance the intra-class cohesion of text semantics and (to) obtain unique fine-grained text embeddings for each image. Finally, the MmCCH is used to inject text semantics into visual features and obtain clustering results. Experimental results on five fine-grained datasets and four coarse-grained datasets show that FGCLIP-ALCC outperforms state-of-the-art clustering methods on all datasets and metrics, while requiring only 1.2M additional learnable parameters. The code will be released at https://github.com/xjq425/FGCLIP-ALCC. Jiaqi Xiao, Xizhan Gao, Dong Wei 0007, Xiaofeng Qu, Sijie Niu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Semantic Augmentation Variational Autoencoder for Unsupervised Anomaly Detection in Retinal OCT ImagesabstractPerforming unsupervised anomaly detection in retinal optical coherence tomography (OCT) images involves training a model solely on anomaly-free samples and detecting anomalies during inference, which reduces the cost of collecting large-scale annotated anomalous data. However, retinal OCT images exhibit significant variations in shape, thickness, and orientation, and lesions often have similar reflectance signals as normal tissues, making anomaly localization highly challenging. Existing methods address these challenges by flattening retinal layers, normalizing thickness, or leveraging reflectance priors, but their reliance on complex pre- and post-processing introduces uncertainties and limits end-to-end clinical applicability. To overcome these issues, we propose a novel semantic augmentation variational autoencoder (SeAugVAE) for unsupervised anomaly detection in retinal OCT images. Specifically, to capture the anatomical variability of normal retinas and thereby enhance anomaly sensitivity, we introduce a self-supervised semantic data augmentation strategy that enforces dual distribution consistency in both image and feature spaces during VAE training. For precise anomaly localization, we develop structural-semantic anomaly attention maps in the inference phase to detect anomalies from both local and global perspectives, and combine them to calculate anomaly score maps as the metric for localizing anomalous regions in images. Extensive experiments on multiple publicly and privately collected Cirrus and Spectralis OCT datasets demonstrate the effectiveness of SeAugVAE in pixel-wise unsupervised anomaly detection across multiple retinal diseases. Our codes are available at https://github.com/xyzhou1121/SeAugVAE. Xueying Zhou, Sijie Niu, Xiangmin Han, Xizhan Gao, Jun Shi 0004 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Garlic Scale Bud Orientation Recognition Algorithm with YOLO and Progressive Learning Strategy
Litan Sun, Xiancheng Shen, Yongshuai Shen, Lexin Jiang, Jun Chong, Sijie Niu |
ICIC (14) | 6 |
| 2025 | CNNFormer: A CNN-Transformer Hybrid Model for Referring Image Segmentation
Kangsai Yao, Xizhan Gao, Xiaofeng Qu, Sijie Niu |
ICIC (5) | 5 |
| 2025 | CoDiff-SaK: Controllable Diffusion Model with Segment Anything Knowledge for Low-dose CT Image DenoisingabstractElectronic noise, a natural characteristic of low-dose CT (LDCT) images, poses a major challenge for achieving accurate quantitative analysis and precise diagnosis. Deep learning-based methods have gained remarkable improvements in structure-preserving noise reduction but face limitations in controlling the generation of boundaries. To address this challenge, we propose a knowledge-driven controllable LDCT denoising diffusion framework. Specifically, we employ the Segment Anything Model to obtain hierarchical prior knowledge and develop a discretized feature fusion module, both of which effectively guide the diffusion model to reconstruct the denoised images. Second, we construct a negative sample set and design a joint optimization loss to amplify the distance between the model output and the negative samples, making the denoised image closer to the ground-truth distribution. Extensive experiments conducted on two datasets demonstrate our proposed method achieves state-of-the-art performance compared with other methods, and ablation experiments also show the effectiveness of our proposed method in boundary generation. Fenghang Zhang, Xizhan Gao, Wanying Wu, Sijie Niu |
ICME | 5 |
| 2025 | Predefined-Time Synchronization Control of Fractional-Order Multi-modal Memristive Neural NetworksabstractIn this paper, the problem of predefined time synchronization in fractional-order multi-modal memristor neural networks (FOMM-MNNs) is explored. To address the problem, a new predefined time stability theorem is proposed and its sufficient condition is derived through definite integral and inequality transformation. On the basis of this theory, an effective controller is designed, and specific sufficient conditions for predefined time synchronization in FOMM-MNNs networks are further proposed. In addition, the synchronous dynamics of the FOMM-MNNs drive-response system is investigated and analyzed in detail in the time domain. Finally, the correctness and validity of the theoretical conclusions are verified by numerical simulations, and the synchronization method is applied to the field of signal encryption and decryption, thus further verifying its practical application value. Mengwei Niu, Pengxiang Fu, Hui Zhao 0009, Sijie Niu, Xizhan Gao, Mingwen Zheng, Lixiang Li 0001 |
IJCNN | 4 |
| 2025 | ASP-CLIP: Adaptive and Static Prompts Learning for Zero-shot Anomaly DetectionabstractZero-shot anomaly detection aims to identify anomalies in unseen classes using knowledge learned from seen classes. Advances in vision-language models like CLIP have demonstrated strong potential for zero-shot anomaly detection tasks. However, CLIP was originally designed to align text with global visual features, focusing on image-level semantics information, making it difficult to capture local anomaly features in pixel-level localization tasks accurately. To address this issue, we propose a CLIP-based Adaptive and Static Prompts learning method for zero-shot anomaly detection (ASP-CLIP). It combines static text prompts and adaptive category prompts to form hybrid text prompts, optimizing text representations. Static text prompts capture general information across categories, while adaptive category prompts are generated by an image adapter module that introduces category-specific information by mapping visual features to the text embedding space. Specifically, we design the Stepwise Adjustment module that improves the alignment between local visual features and text embeddings by utilizing similarity maps identified from shallow local visual features to refine deep local visual features. Experimental results validate the effectiveness of our approach, demonstrating highly competitive performance on the MVTec and VisA datasets. Buqing Zou, Xiaofeng Qu, Fenghang Zhang, Xizhan Gao, Sijie Niu |
IJCNN | 6 |
| 2025 | Learning discriminative features via deep metric learning for video-based person re-identification
Xizhan Gao, Sijie Niu, Hui Zhao 0009 |
Expert Syst. Appl. | 3 |
| 2025 | IATA: Instance-driven advancing targeted attacks with transferable pattern embedding
Litan Sun, Xiaofeng Qu, Guangrun Chen, Haokun Geng, Sijie Niu |
Neurocomputing | 6 |
| 2025 | Masked Superpixel Contrastive Subspace Clustering Network for Unsupervised Large-Scale Hyperspectral Image ClassificationabstractSubspace clustering contributes a lot to the development of unsupervised hyperspectral image (HSI) classification task due to its ability of processing high-dimensional data. Generally, subspace clustering suffers from the bottlenecks of high computational cost when dealing with large-scale HSI. Some researches address this problem by using superpixel segmentation. However, the introduction of superpixel segmentation may cause some adverse effects, including interference from non-target region samples and the oversmoothing of superpixel-level samples. To overcome these issues, we propose a masked superpixel contrastive subspace clustering (MSCSC) for large-scale HSI classification. Specifically, we leverage hyperspectral masked autoencoder rather than conventional autoencoder as the backbone to mitigate the interference of non-target region samples during feature extraction. Then, based on this backbone, we integrate contrastive learning with self-expressiveness based subspace clustering to form the superpixel-level contrastive subspace clustering network, which aims to learn the discriminative superpixel-level self-representation. Moreover, we design a novel data augmentation strategy to ensure learn global and local information, improving the distinguishability of superpixel-level features. Experiments on four popular HSI datasets validate the superiority of our proposed method, with a large accuracy improvement compared to state-of-the-art methods. Tianhao Han, Xiaofeng Qu, Xizhan Gao, Xinwang Liu 0002, Sijie Niu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Multitask Hybrid Knowledge Distillation for Unsupervised Anomaly DetectionabstractDetecting both logical and structural anomalies in an unsupervised anomaly detection task is a significant challenge due to the inherent differences between the two types of anomalies. The use of two-branch knowledge distillation to deal with these two types of anomalies separately is a generalized approach. However, existing methods often design dual branches separately, which does not effectively utilize the shared information between these two branches. Also, due to the introduction of bottleneck layers, a large amount of detailed information is often lost during the reconstruction process, resulting in many false positives. To overcome these drawbacks, we structure the student network as a multitask model to enhance its feature extraction capability, thereby improving its ability to distinguish between logical and structural anomalies, especially under the constraint of limited training data. In addition, we incorporated a self-supervised distillation loss within the logical detection branch and trained the model using a hybrid distillation approach. By leveraging the differences in features between self-distillations to detect logical anomalies, we effectively minimized the false positives that often arise from image reconstruction blurring due to feature compression in the logical branch. We conducted experiments on three well-known anomaly detection datasets to demonstrate the effectiveness of our approach. In particular, on the challenging MVTec LOCO AD dataset, our method achieved impressive results with a pixel-level sPRO of 82.9% and an image-level area under the receiver operating characteristic curve (AUROC) of 91.0%. Muhao Xu, Cuiping Zhu, Sijie Niu |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | Adaptive Anchor-Guided Representation Learning for Efficient Multi-View Subspace ClusteringabstractMulti-view Subspace Clustering (MVSC) effectively aggregating multiple data sources to promise clustering performance. Recently, various anchor-based variants have been introduced to effectively alleviate the computation complexity of MVSC. Although satisfactory advancement has been achieved, existing methods either independently learn anchor matrices and their anchor representations or learn a consensus anchor matrix and unified anchor representation, failing to capture both consistency and complementary information simultaneously. In addition, the time complexity of obtaining clustering results by applying Singular Value Decomposition (SVD) on the anchor representation matrix remains high. To tackle the above problems, we propose an Adaptive Anchor-guided Representation Learning for Efficient Multi-view Subspace Clustering (A2RL-EMVSC) framework, which integrates consensus anchors learning, anchor-guided representation learning and matrix factorization to enhance clustering performance and scalability. Technically, the proposed method learns view-specific anchor representation matrices by consensus anchors guidance, which simultaneously exploit consistency and complementary information. Moreover, by applying matrix decomposition to the view-specific anchor representation matrices, clustering results can be achieved with linear time complexity. Extensive experiments on ten challenging multi-view datasets show that the proposed method can improve the effectiveness and superiority of clustering compared with state-of-the-art methods. Xinwang Liu 0002, Tianhao Han, Xiaofeng Qu, Sijie Niu |
IEEE Trans. Image Process. | 5 |
| 2025 | Long and Recent Preference Learning With Recent-K Items Distribution for Recommender SystemabstractReinforcement learning (RL) aims to formulate the recommendation task as a Markov decision process (MDP) and trains an agent to automatically learn the optimal recommendation policy from interaction trajectories through trial-and-error and reward mechanisms. However, most existing RL-based approaches overlook the correlation between items and the dynamics of user interests implied in temporally close interactions. Therefore, in this paper, we propose a reinforcement learning method that incorporates a “recent-k items” distribution to capture users' local preferences. Specifically, we model the output layer as two distinct branches. The “recent-k items” branch, formulated with a Kullback-Leibler divergence loss, learns the recent interests of users, whereas the other branch utilizes a one-step temporal difference error to capture long-term preferences. The proposed structure is integrated into deep Q-learning and actor-critics, resulting in two enhanced methods named R$k$Q and R$k$AC, respectively. Furthermore, a novel soft inter-reward is carefully designed to enhance the proposed method, and we theoretically prove the convergence of the proposed algorithm. We perform extensive experiments on two large real-world datasets and conduct further analysis of the influences of different action sequences, time intervals, and enhancement capabilities for state-of-the-art models. The experimental results demonstrate the efficacy of our proposed methods. Yongbiao Gao, Sijie Niu, Guohua Lv, Miaogen Ling, Xin Geng 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | DIRL: Learning Discriminative ID-Related Representations for Video Visible-Infrared Person ReIDabstractThe core of Video-Based Visible-Infrared Person Re-Identification (VVI-ReID) lies in learning modal-sharing features, namely ID-related feature representations, that are often mixed within modal-invariant features. Existing methods use weight sharing networks to learn frame-level modal-invariant features, but fail to consider modality invariance when computing sequence-level features, resulting in a significant gap between modalities. Moreover, these methods do not explicitly separate the ID-related features and modal-related features, which reduces the model’s discriminative ability and leads to interference in VVI-ReID. In this article, we propose a Discriminative ID-Related Representation Learning (DIRL) network for VVI-ReID. DIRL network consists of three key components, that is Two-Stream Backbone Module (TBM), Cross-Modality Interaction Module (CIM), and Feature Decoupling Module (FDM). More specifically, the TBM is first constructed to preliminarily capture frame-level modal-invariant features. Then, the CIM is designed to interact information between modals and aggregate temporal features simultaneously, thereby obtaining sequence-level modal-invariant features. Finally, the FDM is designed to explicitly separate modal-related features from ID-related ones within the modal-invariant features, thereby leaving only discriminative ID-related representations. Through extensive benchmark experiments, our method demonstrates superior performance over state-of-the-art approaches by significant margins. Our code will be available at https://github.com/JhSearch/DIRL . Xizhan Gao, Sijie Niu, Hui Zhao 0009 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Efficient Large-Scale Pre-Trained Model Guided MR Imaging Super-ResolutionabstractSuper-resolution (SR) is a post-processing technique that can effectively improve the resolution of MR Imaging without upgrading hardware devices. Although existing single-contrast SR reconstruction algorithms and multi-contrast SR algorithms have achieved impressive results, the ill-posed nature of the SR task makes it challenging to achieve significant performance improvements by simply improving the model architecture. Recently, pre-trained models have shown great potential in low-level visual restoration tasks. In this study, we explore using the powerful representational capabilities of pre-trained models to improve the performance of MRI SR and propose a Pre-Training Guided MRI SR (PTG-SR) architecture. Firstly, an effective and simple baseline model is constructed that improves the local and global perceptual capabilities of the modules while maintaining low computational resources. Secondly, we design a Pre-Training Guided Dynamic Alignment Module (PTG-DAM) that transforms the features extracted by the pre-trained model into information related to image degradation, resulting in high-quality and fine-grained SR images. Furthermore, an improved routing attention module, termed DRAM, is proposed which captures the relationships between different paths from both local and global perspectives with fewer resources, further enhancing the representational capabilities of the model. Extensive experimental results on the IXI and BraTS2020 datasets demonstrate that PTG-SR achieves advanced performance and robustness, outperforming existing state-of-the-art (SOTA) SR approaches. Fenghang Zhang, Xizhan Gao, Sijie Niu |
BIBM | 4 |
| 2024 | Contrastive Learning with Global Representation for Face Anti-spoofing
Jiwen Dong, Xizhan Gao, Hui Zhao 0009, Jinglan Tian, Sijie Niu |
ICIC (5) | 9 |
| 2024 | A Unified Dual Attention-Guided Reverse Distillation Framework for Anomaly Detection
Cuiping Zhu, Muhao Xu, Sijie Niu |
ICIC (6) | 5 |
| 2024 | Correlation-Guided Image-to-Video Transfer Learning for Video Recognition
Xizhan Gao, Sijie Niu, Hui Zhao 0009 |
ICONIP (8) | 3 |
| 2024 | Deep Self-supervised Subspace Clustering with Triple Loss
Xiaotong Bu, Jiwen Dong, Xizhan Gao, Sijie Niu |
MMM (2) | 6 |
| 2024 | Uncertainty-weighted prototype active learning in domain adaptive semantic segmentation
Sijie Niu, Xizhan Gao, Jinping Li, Xiuli Shao |
Expert Syst. Appl. | 2 |
| 2024 | Coarse-to-fine online latent representations matching for one-stage domain adaptive semantic segmentation
Sijie Niu, Xizhan Gao, Xiuli Shao |
Pattern Recognit. | 2 |
| 2024 | A novel active contour model based on features for image segmentation
Sijie Niu |
Pattern Recognit. | 2 |
| 2024 | EDLRDPL_Net: A New Deep Dictionary Learning Network for SAR Image ClassificationabstractSAR image classification is one of the research hotspots in the field of remote sensing. The performance of SAR image classification greatly depends on the learning of features and the design of classifiers. However, traditional SAR image classification methods either focus on learning (deep) features or focus on designing discriminative classifiers, while ignoring the correlation between them, resulting in the learned features and classifiers often not matching. To address the above issues, in this paper a novel classifier, termed entropy based low-rank dictionary pair learning (ELRDPL) method is first designed, which introduces the entropy theory and low-rank constraints into the objective function of dictionary learning for increasing the discriminative ability and decreasing the space occupation. Based on the constructed classifier, an entropy based deep lowrank dictionary pair learning network (EDLRDPL Net) is then proposed, which performs joint learning of deep features and discriminative dictionaries by embedding the ELRDPL classifier into the deep neural network. Extensive experimental results on four SAR image classification datasets demonstrate the effectiveness of EDLRDPL Net. Xizhan Gao, Kang Wei 0003, Sijie Niu, Hui Zhao 0009, Jiwen Dong |
IEEE Signal Process. Lett. | 3 |
| 2024 | Joint Metric Learning-Based Class-Specific Representation for Image Set ClassificationabstractWith the rapid advances in digital imaging and communication technologies, recently image set classification has attracted significant attention and has been widely used in many real-world scenarios. As an effective technology, the class-specific representation theory-based methods have demonstrated their superior performances. However, this type of methods either only uses one gallery set to measure the gallery-to-probe set distance or ignores the inner connection between different metrics, leading to the learned distance metric lacking robustness, and is sensitive to the size of image sets. In this article, we propose a novel joint metric learning-based class-specific representation framework (JMLC), which can jointly learn the related and unrelated metrics. By iteratively modeling probe set and related or unrelated gallery sets as affine hull, we reconstruct this hull sparsely or collaboratively over another image set. With the obtained representation coefficients, the combined metric between the query set and the gallery set can then be calculated. In addition, we also derive the kernel extension of JMLC and propose two new unrelated set constituting strategies. Specifically, kernelized JMLC (KJMLC) embeds the gallery sets and probe sets into the high-dimensional Hilbert space, and in the kernel space, the data become approximately linear separable. Extensive experiments on seven benchmark databases show the superiority of the proposed methods to the state-of-the-art image set classifiers. Xizhan Gao, Sijie Niu, Dong Wei 0007, Xingrui Liu, Tingwei Wang, Fa Zhu, Jiwen Dong, Quan-Sen Sun |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Predefined-Time Event-Triggered Consensus for Nonlinear Multi-Agent Systems with Uncertain Parameter
Yafei Lu, Hui Zhao 0009, Aidi Liu, Mingwen Zheng, Sijie Niu, Xizhan Gao, Xiju Zong |
ICONIP (1) | 5 |
| 2023 | New Predefined-Time Stability Theorem and Applications to the Fuzzy Stochastic Memristive Neural Networks with Impulsive Effects
Hui Zhao 0009, Qingjie Wang, Sijie Niu, Xizhan Gao, Xiju Zong |
ICONIP (2) | 4 |
| 2023 | Predefined-Time Synchronization of Complex Networks with Disturbances by Using Sliding Mode Control
Hui Zhao 0009, Aidi Liu, Sijie Niu, Xizhan Gao, Xiju Zong |
ICONIP (7) | 4 |
| 2023 | Two-directional two-dimensional fractional-order embedding canonical correlation analysis for multi-view dimensionality reduction and set-based video recognition
Yinghui Sun, Xizhan Gao, Sijie Niu, Dong Wei 0007, Zhen Cui 0001 |
Expert Syst. Appl. | 3 |
| 2023 | A fine-to-coarse-to-fine weakly supervised framework for volumetric SD-OCT image segmentationabstractAbstract Obtaining accurate segmentation of central serous chorioretinopathy in spectral‐domain optical coherence tomography (SD‐OCT) is critical for the determination of the disease severity. Although existing methods achieve considerable segmentation results, they heavily depend on large‐scale data with high‐quality annotations. Also, the lesions bear a large shape variation across different patients, which are often difficult to encode. To address the above problems, we propose a fine‐to‐coarse‐to‐fine weakly supervised framework. Specifically, global alternate max‐avg pooling (GTP) network can be employed to locate the lesion regions accurately by using only image‐level annotations. A network module based on the GTP network and a semantic transfer module are proposed to iteratively guide the network to continuously discover and expand the target lesion regions. Then, we employ 3D grey distribution histogram to generate pseudo‐volumetric labels. Finally, a novel 3D level set loss function is proposed to perform coarse‐to‐fine volumetric segmentation. Experiments on a challenging dataset demonstrate that the performance of our proposed method is closer to those of models trained with pixel‐level supervision. Sijie Niu, Ruiwen Xing, Xizhan Gao, Yuehui Chen |
IET Comput. Vis. | 1 |
| 2023 | A novel time series prediction method based on pooling compressed sensing echo state network and its application in stock market
Hui Zhao 0009, Mingwen Zheng, Sijie Niu, Xizhan Gao, Lixiang Li 0001 |
Neural Networks | 4 |
| 2023 | An Improved Fixed-Time Stability Theorem and its Application to the Synchronization of Stochastic Impulsive Neural Networks
Qingjie Wang, Hui Zhao 0009, Aidi Liu, Sijie Niu, Xizhan Gao, Xiju Zong, Lixiang Li 0001 |
Neural Process. Lett. | 4 |
| 2023 | Exploiting Sparse Self-Representation and Particle Swarm Optimization for CNN CompressionabstractStructured pruning has received ever-increasing attention as a method for compressing convolutional neural networks. However, most existing methods directly prune the network structure according to the statistical information of the parameters. Besides, these methods differentiate the pruning rates only in each pruning stage or even use the same pruning rate across all layers, rather than using learnable parameters. In this article, we propose a network redundancy elimination approach guided by the pruned model. Our proposed method can easily tackle multiple architectures and is scalable to the deeper neural networks because of the use of joint optimization during the pruning procedure. More specifically, we first construct a sparse self-representation for the filters or neurons of the well-trained model, which is useful for analyzing the relationship among filters. Then, we employ particle swarm optimization to learn pruning rates in a layerwise manner according to the performance of the pruned model, which can determine optimal pruning rates with the best performance of the pruned model. Under this criterion, the proposed pruning approach can remove more parameters without undermining the performance of the model. Experimental results demonstrate the effectiveness of our proposed method on different datasets and different architectures. For example, it can reduce 58.1% FLOPs for ResNet50 on ImageNet with only a 1.6% top-five error increase and 44.1% FLOPs for FCN_ResNet50 on COCO2017 with a 3% error increase, outperforming most state-of-the-art methods. Sijie Niu, Kun Gao 0002, Xizhan Gao, Hui Zhao 0009, Jiwen Dong, Yuehui Chen, Dinggang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Deep Dictionary Pair Learning for SAR Image Classification
Kang Wei 0003, Jiwen Dong, Sijie Niu, Hui Zhao 0009, Xizhan Gao |
ICANN (3) | 4 |
| 2022 | Vision Transformer with Depth Auxiliary Information for Face Anti-spoofing
Shenyuan Li, Jiwen Dong, Xizhan Gao, Sijie Niu |
ICONIP (3) | 5 |
| 2022 | Class-specific representation based distance metric learning for image set classification
Xizhan Gao, Zeming Feng, Dong Wei 0007, Sijie Niu, Hui Zhao 0009, Jiwen Dong |
Knowl. Based Syst. | 4 |
| 2022 | A new predefined-time stability theorem and its application in the synchronization of memristive complex-valued BAM neural networks
Aidi Liu, Hui Zhao 0009, Qingjie Wang, Sijie Niu, Xizhan Gao, Chuan Chen 0001, Lixiang Li 0001 |
Neural Networks | 4 |
| 2022 | Rapid construction of 4D high-quality microstructural image for cement hydration using partial information registration
Lin Wang 0004, Bo Yang 0001, Sijie Niu, Sung-Kwun Oh |
Pattern Recognit. | 4 |
| 2022 | Deep Low-Rank Graph Convolutional Subspace Clustering for Hyperspectral ImageabstractDeep subspace clustering (DSC) has achieved considerable success in the classification task of hyperspectral image (HSI) without background (defined as noisy samples) compared with traditional subspace clustering methods. Unfortunately, directly applying DSC to classify land-cover on HSI datasets with background may suffer from the degradation of classification performance. In this paper, we propose an effective deep low-rank graph convolutional subspace clustering (DLR-GCSC) framework for improving the performance of land-cover classification on HSI datasets with background. Specifically, we design a joint spatial-spectral network to extract band-level and patch-level features simultaneously by combining 1D and 2D auto-encoder. Moreover, we construct a low-rank constrained fully connected layer as self-expression layer in the network to make the joint features more discriminative. To reduce the influence of noisy samples and obtain an informative affinity matrix, we recast the joint features into a non-Euclidean domain by introducing graph convolution. Finally, spectral clustering is applied on the informative affinity matrix to obtain the classification results. Experiments on three benchmark HSI datasets show that our proposed method achieves competitive classification performance to the state-of-the-art methods on both HSI data with background and without background. Tianhao Han, Sijie Niu, Xizhan Gao, Wenyue Yu, Na Cui, Jiwen Dong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | AFLLC: A Novel Active Contour Model Based on Adaptive Fractional Order Differentiation and Local-Linearly Constrained Bias Field
Yingying Han, Jiwen Dong, Fan Li 0026, Xizhan Gao, Sijie Niu |
ICONIP (3) | 6 |
| 2021 | Unsupervised Domain Adaptation with Self-selected Active Learning for Cross-domain OCT Image Segmentation
Sijie Niu, Xizhan Gao, Jiwen Dong |
ICONIP (2) | 2 |
| 2021 | Iterative registration for multi-modality retinal fundus photographs using directional vessel skeletonabstractAbstract This paper proposes an automated registration method for multi‐modality retinal fundus photographs based on the directional vessel skeleton. The main purpose is to register two retinal fundus photographs with different modalities of the same scanning region, which can provide multi‐modality information for clinicians to diagnose retinal diseases or to make a treatment decision. The directional vessel skeleton of each fundus image is first detected by bias field correction and Gabor filter. The final registered fundus photographs are then obtained by the iterative affine registration between the detected directional vessel skeletons of two photographs. In this work, four kinds of fundus photographs in the macular regions of the patient with diseases, consisting of 20 optical coherence tomography fundus images, 20 colour fundus photographs, 20 fluorescein fundus angiography images and 20 indocyanine green angiography images, are utilised to quantitatively evaluate the proposed method. The root‐mean‐square errors show an advantageous performance in both registration success rate and accuracy. Wenwen Kong, Pengxiao Zang, Sijie Niu, Dengwang Li |
IET Image Process. | 3 |
| 2021 | MFNet-LE: Multilevel fusion network with Laplacian embedding for face presentation attacks detectionabstractAbstract Face detection is playing a pivotal role for crowd counting and abnormal events detection. However, it is vulnerable to face presentation attacks by printed photos, videos, and 3D masks of real human faces. Although numerous detection techniques based on deep learning have been employed to address the problem of face presentation attacks, there are still several weaknesses in these approaches, such as high algorithm complexity and a lack of detection ability. To overcome these weaknesses, a method based on a multilevel fusion network with Laplacian embedding (MFNet‐LE) for the detection of face presentation attacks is proposed. First, a shallow network that contains just three layers was developed, which makes the model faster. Then, an optimised multilevel fusion strategy was developed to combine the input with the output of all previous layers to improve the detection ability of the method. Finally, a Laplacian embedding algorithm is introduced to maintain the inter‐class discrimination and penalise the intra‐class distance. Under the joint supervision of Laplacian loss and softmax loss, the proposed approach can obtain more discriminative features, which enhance the accuracy of attack detection. Experiments were conducted with three public databases for face presentation attacks: CASIA FASD, Idiap Replay Attack database and MSU USSA. The results demonstrate that the MFNet‐LE model can outperform the state‐of‐the‐art methods. Sijie Niu, Xiaofeng Qu, Xizhan Gao, Tingwei Wang, Jiwen Dong |
IET Image Process. | 1 |
| 2021 | Fast High-Order Sparse Subspace Clustering With Cumulative MRF for Hyperspectral ImagesabstractSparse subspace clustering (SSC), as a powerful tool in hyperspectral image (HSI) segmentation, has caused widespread attention recently. However, the existing methods are generally affected by the limitations of region consistency and time complexity. To address these issues, we propose a fast high-order SSC (FHoSSC) with the cumulative Markov random field (MRF) algorithm in this letter. By capturing high-order information, we explore the pixel-level contextual restraints within the same high-order data to preserve consistency. Also, a new regularization term is introduced by the spatial constraints among adjacent high-order data for improving segmentation accuracy. Finally, cumulative MRF, as a variant of MRF, is used to further refine the segmentation result through combining original HSI information. Experiments on real data sets demonstrate that the proposed method not only outperforms the state-of-the-art methods in segmentation accuracy but also reduces the time complexity significantly. Limei Wang, Sijie Niu, Xizhan Gao, Kun Liu 0022, Feixia Lu, Qi Diao, Jiwen Dong |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | An integrated time adaptive geographic atrophy prediction model for SD-OCT images
Yuhan Zhang 0001, Zexuan Ji, Sijie Niu, Theodore Leng, Daniel L. Rubin, Songtao Yuan, Qiang Chen 0004 |
Medical Image Anal. | 4 |
| 2021 | Multi-Site Infant Brain Segmentation Algorithms: The iSeg-2019 ChallengeabstractTo better understand early brain development in health and disorder, it is critical to accurately segment infant brain magnetic resonance (MR) images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF). Deep learning-based methods have achieved state-of-the-art performance; h owever, one of the major limitations is that the learning-based methods may suffer from the multi-site issue, that is, the models trained on a dataset from one site may not be applicable to the datasets acquired from other sites with different imaging protocols/scanners. To promote methodological development in the community, the iSeg-2019 challenge (http://iseg2019.web.unc.edu) provides a set of 6-month infant subjects from multiple sites with different protocols/scanners for the participating methods. T raining/validation subjects are from UNC (MAP) and testing subjects are from UNC/UMN (BCP), Stanford University, and Emory University. By the time of writing, there are 30 automatic segmentation methods participated in the iSeg-2019. In this article, 8 top-ranked methods were reviewed by detailing their pipelines/implementations, presenting experimental results, and evaluating performance across different sites in terms of whole brain, regions of interest, and gyral landmark curves. We further pointed out their limitations and possible directions for addressing the multi-site issue. We find that multi-site consistency is still an open issue. We hope that the multi-site dataset in the iSeg-2019 and this review article will attract more researchers to address the challenging and critical multi-site issue in practice. Yue Sun 0001, Kun Gao 0002, Zhengwang Wu, Xiaopeng Zong, Zhihao Lei, Ying Wei 0007, Jun Ma 0016, Xiaoping Yang 0001, Xue Feng 0001, Li Zhao 0001, Trung Le Phan, Jitae Shin, Tao Zhong 0002, Yu Zhang 0064, Lequan Yu, Caizi Li, Ramesh Basnet, M. Omair Ahmad, M. N. S. Swamy 0001, Wenao Ma, Qi Dou 0001, Toan Duc Bui, Camilo Bermudez, Bennett A. Landman, Ian H. Gotlib, Kathryn L. Humphreys, Sarah Shultz, Longchuan Li, Sijie Niu, Weili Lin, Valerie Jewells, Dinggang Shen, Gang Li 0001, Li Wang 0026 |
IEEE Trans. Medical Imaging | 30 |
| 2020 | Automatic retinal layer segmentation in SD-OCT images with CSC guided by spatial characteristics
Kun Gao 0002, Wenwen Kong, Sijie Niu, Dengwang Li, Yuehui Chen |
Multim. Tools Appl. | 3 |
| 2020 | MS-CAM: Multi-Scale Class Activation Maps for Weakly-Supervised Segmentation of Geographic Atrophy Lesions in SD-OCT ImagesabstractAs one of the most critical characteristics in advanced stage of non-exudative Age-related Macular Degeneration (AMD), Geographic Atrophy (GA) is one of the significant causes of sustained visual acuity loss. Automatic localization of retinal regions affected by GA is a fundamental step for clinical diagnosis. In this paper, we present a novel weakly supervised model for GA segmentation in Spectral-Domain Optical Coherence Tomography (SD-OCT) images. A novel Multi-Scale Class Activation Map (MS-CAM) is proposed to highlight the discriminatory significance regions in localization and detail descriptions. To extract available multi-scale features, we design a Scaling and UpSampling (SUS) module to balance the information content between features of different scales. To capture more discriminative features, an Attentional Fully Connected (AFC) module is proposed by introducing the attention mechanism into the fully connected operations to enhance the significant informative features and suppress less useful ones. Based on the location cues, the final GA region prediction is obtained by the projection segmentation of MS-CAM. The experimental results on two independent datasets demonstrate that the proposed weakly supervised model outperforms the conventional GA segmentation methods and can produce similar or superior accuracy comparing with fully supervised approaches. The source code has been released and is available on GitHub: https://github.com/ jizexuan/Multi-Scale-Class-Activation-Map-Tensorflow. Xiao Ma 0011, Zexuan Ji, Sijie Niu, Theodore Leng, Daniel L. Rubin, Qiang Chen 0004 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Adaptive-Guided-Coupling-Probability Level Set for Retinal Layer SegmentationabstractQuantitative assessment of retinal layer thickness in spectral domain-optical coherence tomography (SD-OCT) images is vital for clinicians to determine the degree of ophthalmic lesions. However, due to the complex retinal tissues, high-level speckle noises and low intensity constraint, how to accurately recognize the retinal layer structure still remains a challenge. To overcome this problem, this paper proposes an adaptive-guided-coupling-probability level set method for retinal layer segmentation in SD-OCT images. Specifically, based on Bayes's theorem, each voxel probability representation is composed of two probability terms in our method. The first term is constructed as neighborhood Gaussian fitting distribution to characterize intensity information for each intra-retinal layer. The second one is boundary probability map generated by combining anatomical priors and adaptive thickness information to ensure surfaces evolve within a proper range. Then, the voxel probability representation is introduced into the proposed segmentation framework based on coupling probability level set to detect layer boundaries. A total of 1792 retinal B-scan images from 4 SD-OCT cubes in healthy eyes, 5 cubes in abnormal eyes with central serous chorioretinaopathy and 5 SD-OCT cubes in abnormal eyes with age-related macular disease are used to evaluate the proposed method. The experiment demonstrates that the segmentation results obtained by the proposed method have a good consistency with ground truth, and the proposed method outperforms six methods in the layer segmentation of uneven retinal SD-OCT images. Yue Sun 0001, Sijie Niu, Xizhan Gao, Jie Su 0010, Jiwen Dong, Yuehui Chen, Li Wang 0026 |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | shallowCNN-LE: A shallow CNN with Laplacian Embedding for face anti-spoofingabstractIdentity authentication based on face recognition has been significantly improved due to the outstanding ability of face detection, thus it plays an important role in society. However, face recognition system might be deceived by malicious face spoof attacks raising risk from both safety and property. The algorithm to accurately detect face anti-spoofing in identity authentication system is becoming crucial. In this paper, a shallow convolutional neural network with laplacian embedding (shallowCNN-LE) is proposed for face anti-spoofing. Two different types of features are concatenated to accurately detect the face liveness, including depth features and dynamic texture features. First, the developed shallow CNN model contains four layers which make the model faster. Second, we integrate dynamic texture features extracted by using the dual tree complex wavelet transform (DT-CWT) with the depth features as input features to feed into the proposed model. Finally, we propose a laplacian embedding algorithm, which can maintain the inter-class discrimination and penalize the distance of intra-class. When embedding the laplacian loss with the softmax loss, the proposed method can obtain much more discriminative features, which is helpful to detect face anti-spoofing. Experimental results on public databases of CASIA FASD, Replay attack and MSU USSA database demonstrate that our proposed method outperforms the state-of-the-art methods for face anti-spoofing detection. Xiaofeng Qu, Jiwen Dong, Sijie Niu |
FG | 3 |
| 2019 | Two-Directional Two-Dimensional Kernel Canonical Correlation AnalysisabstractTwo-directional two-dimensional canonical correlation analysis ((2D)2CCA) directly seeks linear relationship between different image data sets without reshaping images into vectors. However, it fails in finding the nonlinear correlation. In this letter, a novel method named as two-directional two-dimensional kernel canonical correlation analysis is proposed, which is a nonlinear version of (2D)2CCA and is able to find the nonlinear relationship between different image data sets. Experimental results in different expressions, illumination conditions and poses show the effectiveness of the proposed method. Xizhan Gao, Sijie Niu, Quan-Sen Sun |
IEEE Signal Process. Lett. | 2 |
| 2018 | A Semantic Context Model for Automatic Image Annotation
Xin Fu 0011, Sijie Niu, Hengcai Zhang |
ICIC (2) | 3 |
| 2018 | Multi-path 3D Convolution Neural Network for Automated Geographic Atrophy Segmentation in SD-OCT Images
Rongbin Xu, Sijie Niu, Kun Gao 0002, Yuehui Chen |
ICIC (2) | 2 |
| 2018 | Structural Compact Core Tensor Dictionary Learning for Multispec-Tral Remote Sensing Image DeblurringabstractThe multispectral remote sensing image (MS-RSI) is blurred existing multispectral camera due to various hardware limitations. In this paper, we propose a novel structural compact core tensor dictionary learning (SCCTDL) model for MS-RSI deblurring. First, the multispectral patch is modeled by three-order tensor and high-order singular value decomposition is applied to the tensor. Then the task of MS-RSI deblurring is formulated as a minimum sparse core tensor estimation problem. To improve the accuracy of core tensor coding, the core tensor estimation based on the structural compact principle is introduced into the SCCTDL model to exploit abundant structural similarity in image. Experimental results suggest that our method outperforms several existing MS-RSI deblurring methods in both subjective image quality and visual perception. Leilei Geng, Xiushan Nie, Sijie Niu, Yilong Yin |
ICIP | 3 |
| 2018 | Beyond Retinal Layers: A Large Blob Detection for Subretinal Fluid Segmentation in SD-OCT Images
Zexuan Ji, Qiang Chen 0004, Sijie Niu, Wen Fan 0003, Songtao Yuan, Quan-Sen Sun |
MICCAI (2) | 4 |
| 2018 | Automated Choroidal Neovascularization Detection for Time Series SD-OCT Images
Sijie Niu, Zexuan Ji, Wen Fan 0003, Songtao Yuan, Qiang Chen 0004 |
MICCAI (2) | 2 |
| 2018 | Automated and Robust Geographic Atrophy Segmentation for Time Series SD-OCT Images
Sijie Niu, Zexuan Ji, Qiang Chen 0004 |
PRCV (1) | 2 |
| 2017 | Robust noise region-based active contour model via local similarity factor for image segmentation
Sijie Niu, Qiang Chen 0004, Luis de Sisternes, Zexuan Ji, Ze Ming Zhou, Daniel L. Rubin |
Pattern Recognit. | 1 |