VLDB 2026 Research / reviewers in the wild / expert
Lijun Guo
dblp:35/3362
· DBLP profile ↗
53ranked-venue papers
3as first author
47since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 1 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 13 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond reconstruction: Enhancing masked autoencoders with contrastive learning for video representation learning
Yawei Feng, Lijun Guo, Guitao Yu, Rong Zhang 0007, Jiangbo Qian, Chong Wang 0001, Shangce Gao |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Round versus square: Exploring the influence of facial shape in robo-advisors on consumer acceptance of investment advice
Zhongpeng Cao, Lijun Guo |
Inf. Manag. | 3 |
| 2026 | Boosting representation diversity in video transformers via segmented contrastive masked autoencoders
Yawei Feng, Lijun Guo, Guitao Yu, Rong Zhang 0007, Jiangbo Qian, Chong Wang 0001, Shangce Gao |
Neurocomputing | 2 |
| 2026 | Particle swarm optimization with problem-aware hyperparameter design for feature selection in high dimensions
Jinrui Gao, Zhenyu Lei 0002, Lijun Guo, Yirui Wang 0001, Shangce Gao |
Inf. Sci. | 4 |
| 2026 | Dual uncertainty-aware hierarchical semantic matching for text-to-image person search
Ting Tuo, Lijun Guo |
Image Vis. Comput. | 2 |
| 2025 | Soft-Evidence Fused Graph Neural Network for Cancer Driver Gene Identification across Multi-View Biological GraphsabstractIdentifying cancer driver genes (CDGs) is essential for understanding cancer mechanisms and developing targeted therapies. Graph neural networks (GNNs) have recently been employed to identify CDGs by capturing patterns in biological interaction networks. However, most GNN-based approaches rely on a single protein-protein interaction (PPI) network, ignoring complementary information from other biological networks. Some studies integrate multiple networks by aligning features with consistency constraints to learn unified gene representations for CDG identification. However, such representation-level fusion often assumes congruent gene relationships across networks, which may overlook network heterogeneity and introduce con-flicting information. To address this, we propose Soft-Evidence Fusion Graph Neural Network (SEFGNN), a novel framework for CDG identification across multiple networks at the decision level. Instead of enforcing feature-level consistency, SEFGNN treats each biological network as an independent evidence source and performs uncertainty-aware fusion at the decision level using Dempster-Shafer Theory (DST). To alleviate the risk of overconfidence from DST, we further introduce a Soft Evidence Smoothing (SES) module that improves ranking stability while preserving discriminative performance. Experiments on three cancer datasets show that SEFGNN consistently outperforms state-of-the-art baselines and exhibits strong potential in discovering novel CDGs. Bang Chen, Lijun Guo, Houli Fan |
BIBM | 2 |
| 2025 | Multi-Modal Timely Pancreatitis Severity Assessment via Hierarchical Evidential Conflictive LearningabstractAcute pancreatitis (AP) can rapidly progress to severe acute pancreatitis (SAP), which carries a high risk of mortality. Early screening of high-risk patients using computed tomography (CT) imaging and laboratory indicators is beneficial to improving clinical outcomes. To fully leverage the advantages of multi-modal data, we address pancreatitis severity assessment from multiple views and introduce multi-view categorical uncertainty quantification to enhance model reliability. Existing uncertainty-aware multi-view classification methods often assume consistency across views and indiscriminately reduce uncertainty. Nevertheless, in the context of pancreatitis, view conflicts underlying different data modalities are common as each modality provides unique pathological information. To tackle these challenges, we develop a Hierarchical Evidential Conflictive Learning (HECL) method for pancreatitis severity assessment, which estimates both the uncertainty and belief mass through subjective logic from multi-modal data and aggregates the conflicting opinions via logarithmic opinion pooling, by weighted geometric averaging of belief distributions and effectively transforming view conflicts into measurable prediction uncertainties. Additionally, HECL incorporates a hierarchical fusion strategy and introduces cross-modal pseudo-views to enhance the representation and interaction across different views. Experimental results show that HECL significantly outperforms several state-of-the-art multi-view methods. Visualization analysis reveals that the model can accurately localize pancreatic lesion regions and identify key predictive indicators, which can provide trustworthy diagnoses for pancreatitis severity assessment. Houli Fan, Lijun Guo, Xiuchao He, Bang Cheng, Yingqing Zeng, Jiang Duan, Rong Zhang 0007 |
BIBM | 2 |
| 2025 | PoseDucer: Implicit relation inducement for invisible keypoint reconstruction in real-world occluded scenes
Junneng Feng, Rong Zhang 0007, Yirui Wang 0001, Shangce Gao, Lijun Guo |
Neurocomputing | 5 |
| 2025 | Robust auxiliary modality is beneficial for video-based cloth-changing person re-identification
Youming Chen, Ting Tuo, Lijun Guo, Rong Zhang 0007, Yirui Wang 0001, Shangce Gao |
Image Vis. Comput. | 3 |
| 2025 | ActiveFreq: Integrating Active Learning and Frequency Domain Analysis for Interactive Segmentation
Lijun Guo, Qian Zhou 0001, Zidi Shi, Hua Zou 0002, Gang Ke |
Knowl. Based Syst. | 1 |
| 2025 | DSFormer: Dynamic size attention with enhanced long-range dependency modeling for artery/vein classification
Zeyuan Ju, Chouyu Chen, Lijun Guo, Zhenyu Lei 0002, Masaaki Omura, Shangce Gao |
Knowl. Based Syst. | 4 |
| 2025 | Part2Pose: Inferring Human Pose From Parts in Complex ScenesabstractMost of existing Human Pose Estimation (HPE) methods struggle to handle with challenges such as changeable poses, complex backgrounds, and occlusion encountered in complex scenes. To address these problems, a novel HPE network, called Part2Pose, is proposed in this paper. In our Part2Pose, instead of focusing on small-sized keypoints like existing HPE methods do, we first extract image features based on human body parts to expand the detection scope. This strategy enhances the robustness of the extracted features to variations and distractions in complex scenes. Then, a Transformer-based Global Part Relation Module (GPRM) and a graph convolutional network-based Local Part Relation Module (LPRM) are used to capture global and local relationships among different body parts to help infer the position of keypoints. Extensive experiments on challenging datasets, including COCO, CrowdPose and OCHuman, show that the proposed Part2Pose can surpass existing popular state-of-the-art HPE methods. The combination with lightweight networks confirms the robustness and generalizability of our Part2Pose. Rong Zhang 0007, Junneng Feng, Cun Feng, Yirui Wang 0001, Lijun Guo |
IEEE Signal Process. Lett. | 5 |
| 2025 | Event-Based Video Reconstruction Via Spatial-Temporal Heterogeneous Spiking Neural NetworkabstractEvent cameras detect per-pixel brightness changes and output asynchronous event streams with high temporal resolution, high dynamic range, and low latency. However, the unstructured nature of event streams means that humans cannot analyze and interpret them in the same way as natural images. Event-based video reconstruction is a widely used method aimed at reconstructing intuitive videos from event streams. Most reconstruction methods based on traditional artificial neural networks (ANNs) have high energy consumption, which counteracts the low-power advantage of event cameras. Spiking neural networks (SNNs) are a new generation of event-driven neural networks that encode information via discrete spikes, which leads to greater computational efficiency. Previous methods based on SNNs overlooked the asynchronous nature of event streams, leading to reconstructions that suffer from artifacts, flickering, low contrast, etc. In this work, we analyze event streams and spiking neurons and explain poor reconstruction quality. We specifically propose a novel spatial-temporal heterogeneous (STH) spiking neuron suitable for reconstructing asynchronous event streams. The STH neuron adjusts the membrane decay coefficient adaptively and has better spatiotemporal perception. In addition, we propose a temporal-frequency calibration module (TFCM) based on the Fourier transform to improve the contrast of the reconstructions. On the basis of the above proposed neuron and module, we construct two SNN-based models, referred to as the STHSNN and TFCSNN. The goal of the former is to reduce the artifacts and flickering in reconstructions, whereas the latter focuses on enhancing the contrast. The experimental results demonstrate that our models can yield reconstructions in various scenarios, achieving better quality and lower energy consumption than previous SNNs. Specifically, the TFCSNN and STHSNN achieve top-2 performance among the SNN-based models, with energy consumption reductions of 3.48 times and 12.40 times, respectively. Lijun Guo, Chong Wang 0001, Guoqi Li 0002, Jiangbo Qian |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | SpikeHCD: Spiking Transformer With Parallel Neurons and Memory-Enhanced Attention for Hyperspectral Change DetectionabstractHyperspectral change detection is a critical technology in remote sensing, widely applied in urban planning, environmental monitoring, and disaster detection. However, hyperspectral data exhibits higher spectral dimensionality compared to conventional RGB data, making existing methods struggle to balance high accuracy and low energy consumption. As the third generation of neural networks, spiking neural networks (SNNs) demonstrate the advantage of low energy efficiency, but the iterative computation process in spiking neurons significantly increases training and inference burdens when applied to hyperspectral change detection. To address these challenges, we propose a novel spiking Transformer with parallel neurons and memory-enhanced attention for hyperspectral change detection named SpikeHCD, the first SNNs specifically designed for hyperspectral change detection. SpikeHCD not only maintains low-energy advantage but also employs a probability-driven parallel spiking neurons (PPSN) to improve computational efficiency, enabling more effective application in remote sensing tasks. We further design a memory-enhanced spiking attention (MSA) module to enhance temporal modeling capability, and thoroughly extract spatial-spectral features. Additionally, a spiking difference module (SDM) is introduced to capture change features across different timesteps. Experimental results demonstrate that SpikeHCD can achieve several state-of-the-art (SOTA) results on multiple hyperspectral datasets, with faster detection, lower energy consumption, and fewer number of parameters. The codes are available at https://github.com/mzhcode/HCD_snn. Zihao Mei, Chong Wang 0001, Lijun Guo, Jiangbo Qian |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | A Deep-Learning-Based Lumbosacral Localization and Landmark Detection Network for Automatic Lumbar Stability and Spondylolisthesis Grading AssessmentabstractThe accurate detection of vertebral landmarks is crucial for clinical diagnosis and research on the lumbar stability and spondylolisthesis grading. However, the small size of vertebral landmarks in X-ray images and the morphological similarity among vertebrae complicate this detection task. Recent advances in deep learning have enhanced the spinal landmark detection. To further improve the assessment methods for the lumbar stability and spondylolisthesis grading, we propose the LSLD-Net, a novel network based on the lumbosacral localization and landmark detection. In clinical practices, the X-ray images used for evaluating the lumbar spine stability can often include the extraneous information from non-lumbosacral areas, such as the thoracic spine. The proposed LSLD-Net first extracts the lumbosacral region from the complex X-ray images and then performs the precise landmark detection within the identified area. The landmark detection stage integrates the HRNet and U-net architectures, effectively capturing the long-range contextual information, the overall structural layout, and the fine local details in lumbar X-ray images to optimize landmark detection. Additionally, we propose a Multi-Scale Attention Module that enhances the relevant features and suppresses the irrelevant ones through channel and spatial attention, thereby achieving precise landmark detection and improving the robustness of the network. The evaluations on the private and public BUU-LSPINE datasets indicate that the LSLD-Net surpasses other state-of-the-art methods in landmark detection, enhancing the accuracy and efficiency of Sagittal Displacement and Intervertebral Space Angle measurements. This performance excels in assessing the lumbar stability and spondylolisthesis grading, offering the significant support to clinicians for early quantitative diagnosis and evaluation. Rong Zhang 0007, Baolin Xu, Dongdong Xia, Lijun Guo |
BIBM | 6 |
| 2024 | Tri-Hash Progressive Sampling Neural Attenuation Field for Sparse-View CBCT ReconstructionabstractSparse-view CBCT reconstruction is essential to reduce the X-ray radiation dose in clinical CBCT imaging. However, reducing the number of views often lower the image quality. Existing neural radiance field (NeRF) technologies can achieve the high-quality 3D reconstruction and novel view generation in natural scenes under sparse-view conditions. Nevertheless, applying these technologies to the 3D reconstruction of human tissues in the medical field still has limitations. The recently proposed neural attenuation field (NAF) technique has shown progress in 3D reconstruction of human tissues in medical images. However, in the context of sparse views, there remains the issue of inferior reconstruction quality due to insufficient acquisition of spatial structural information in human tissues. To address these challenges, we propose a novel framework called the Tri-Hash Progressive Sampling Neural Attenuation Field (THP-NAF). First, we introduced an Enhanced Tri-Hash Representation mechanism that enhanced the extraction of 3D spatial information through the 2D plane mapping. This mechanism captures more contextual and spatial information and achieved an optimized balance between the image quality and generation efficiency. Additionally, to mitigate the sampling inefficiency caused by random sampling, we employed a Sobel-based adaptive point–ray sampling strategy. This strategy combined the global and local information for structure-aware ray sampling and could dynamically adjust the number of sampling points, thereby enhancing the sampling flexibility and efficiency. Our method was validated across multiple datasets, demonstrating its ability to improve image quality and its significant potential for clinical applications. Lijun Guo, Rong Zhang 0007, Wenming He, Shangce Gao |
BIBM | 2 |
| 2024 | HR-xNet: A Novel High-Resolution Network for Human Pose Estimation with Low Resource ConsumptionabstractIn our study, we are aiming to find an effective and lightweight solution for the human pose estimation task. To this end, a novel high-resolution representation network called as HR-xNet is proposed. First, HR-xNet is derived from the HRNet [20], by replacing the stem and Stage 1 of HRNet with the lightweight feature extraction head and pruning HRNet. Then, our focus is on improving the model's representation ability for changeable human poses and small targets such as human keypoints, while continuing to reduce model's resource consumption. First, the novel Lightweight Multi-Scale Dynamic Convolution (LMSD Conv) is introduced. The LMSD Conv greatly improves the learning capacity of the network by adaptively generating convolution kernels of different sizes to extract features with different receptive fields from the input. Second, low-level detailed and high-level semantic features are interacted with in the Feature Enhancement Module to relearn the lost detailed features for small keypoints and changeable human poses. On the COCO and CrowdPose datasets, our model can compete with some mainstream large networks and existing state-of-the-art lightweight methods at a low cost. Cun Feng, Lijun Guo |
FG | 3 |
| 2024 | Progressive Learning Based Knowledge Distillation for Low Resolution Cerebral Microbleed SegmentationabstractThis study aims to address key technical issues in the segmentation of Cerebral MicroBleeds (CMBs) based on Low-Resolution (LR) Magnetic Resonance Imaging (MRI) data. There are two challenges in this task. First, the CMB lesions are typically small in size and easily confused with various mimics. Second, anisotropy becomes more prominent and adverse in LR MRI sequences than HR sequences. To address these issues, we propose a Progressive Learning based Knowledge Distillation method. This method progressively transfers knowledge from HR models to their LR counterparts, thereby minimizing the occurrence of false positives attributable to noise from Super-Resolution. To further eliminate the influence of anisotropy, an encoding-enhanced network, called E2U-Net, is proposed in this paper. It can effectively capture anisotropic information and mitigates potential feature loss. The experimental results on multiple publicly accessible CMBs datasets demonstrated the superiority of our proposed approach over existing deep-learning methods. Tianxiang Xia, Rong Zhang 0007, Zhenzuo Chen, Guomin Xie, Xiping Wu, Zhongyue Lv, Lijun Guo |
ICASSP | 7 |
| 2024 | Multimodal Video Highlight Detection with Noise-Robust LearningabstractVideo highlight detection aims to select the most interesting and attractive clips from lengthy videos, which is crucial for enhancing the video editing and viewing experience on social media platforms. Existing video highlight detection methods predominantly rely on visual modality information, and underutilize the abundant multimodality of videos. Furthermore, in supervised video analysis tasks, subjective judgments during label annotation can generate uncertain noise labels that negatively impact the learning process. To address these issues, we propose a noise-robust multimodal video highlight detection approach. Our approach first enhances feature representation by incorporating multimodal representations of a video’s visual and auditory information. This allows for the extraction of complementary information from different modalities. We then implement a noise-cleaning mechanism that utilizes multiple modalities to clean noise samples. This helps to suppress the negative impact of noise samples on the learning process, ensuring that the network learns more robust features from clean samples. We evaluate our approach on two public datasets, YouTube Highlights and TVSum, and demonstrate its efficacy in mitigating the impact of noise labels, while also improving the accuracy and robustness of video highlight detection. Yinhui Jiang, Sihui Luo 0001, Lijun Guo |
IJCNN | 3 |
| 2024 | A spherical evolution algorithm with two-stage search for global optimization and real-world problems
Yirui Wang 0001, Zonghui Cai, Lijun Guo, Yang Yu 0013, Shangce Gao |
Inf. Sci. | 3 |
| 2024 | QSMT-net: A query-sensitive proposal and multi-temporal-span matching network for video grounding
Qingqing Wu 0014, Lijun Guo, Rong Zhang 0007, Jiangbo Qian, Shangce Gao |
Image Vis. Comput. | 2 |
| 2024 | MCT-VHD: Multi-modal contrastive transformer for video highlight detection
Yinhui Jiang, Sihui Luo 0001, Lijun Guo |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | IIESC-Net: Incorporating Implicit and Explicit Structural Constraints for Hip Joint Landmark Detection in pelvic X-rayabstractThe hip joint, a vital weight-bearing joint in the human body, is susceptible to various hip-related diseases. The accurate identification of anatomical landmarks in the hip joint is essential for both disease diagnosis and surgical planning. However, these landmarks are often inconspicuous in X-ray images, where irrelevant background interference increases detection difficulty. In this study, we proposed an IIESC-Net model that combined local and global features, enabling a hierarchical understanding of the implicit structural characteristics of the hip joint. Additionally, drawing from domain expertise based on the explicit physiological structure of the hip joint, we designed a Hip Morphology-Aware loss function to constrain large landmark errors through the application of high-confidence landmarks with robust distinctive identification, thereby achieving accuracy in automatic detection. Furthermore, we constructed a dataset comprising of 843 pelvic X-ray images. The experimental results demonstrated a substantial enhancement in hip joint landmark detection accuracy attributed to the proposed IIESC-Net. This innovation established state-of-the-art performance, notably excelling in attaining heightened successful detection rates under stringent error tolerance. This achievement has profound practical implications for clinical applications. Lijun Guo, Lixin Ni, Xiuchao He, Rong Zhang 0007 |
BIBM | 2 |
| 2023 | A Dual-View Fusion Network for Automatic Spinal Keypoint Detection in Biplane X-ray ImagesabstractAccurate keypoint detection in medical images of the spine is critical for the assessment, diagnosis, treatment planning, and clinical investigation of spinal deformities. However, due to severe occlusions of spinal structures in lateral X-ray images, accurate keypoint detection can be hardly achieved in lateral X-ray images based on single-view information. Thus, methods based on both the anterior-posterior (AP) and lateral (LAT) X-ray image views have been proposed to alleviate occlusion problems and achieve better keypoint detection performance. Although some progress has been made with these dual-view methods, they do not effectively exploit a priori knowledge of the spine and hence cannot adequately account for the structural correlation of the vertebrae across views. In this paper, a new dual-view fusion network (DVFNet) framework is proposed for keypoint detection in spinal X-ray images. This framework obtains structural correlations between AP and LAT views of the spine based on a priori spine knowledge represented by high-level semantic features. Meanwhile, the proposed framework combines local and global features extracted respectively by a local subnetwork and a global subnetwork. On the one hand, the local subnetwork is constructed as an enhanced codec structure based on both the AP and LAT views. This subnetwork is trained to output local features that contain both joint semantic features of the two views and independent fine-grained features of each individual view. This scheme leads to accurate keypoint estimation locally. On the other hand, the global subnetwork utilizes a self-attention mechanism to extract view-specific global features based on either the AP view or the LAT view in order to eliminate ambiguity, and reduce confusion on keypoint locations. Further, we propose a weighted feature fusion (WFF) module for adaptive fusion of the local and global features. We evaluated the DVFNet model on a private dataset and found that our proposed method achieves more accurate spinal keypoint detection compared to other state-of-the-art methods, and thus our method can provide reliable assistance to clinicians. Lijun Guo, Rong Zhang 0007, Xiuchao He |
BIBM | 2 |
| 2023 | SPC-Net: Structure-Aware Pixel-Level Contrastive Learning Network for OCTA A/V Segmentation and Differentiation
Huaying Hao, Yuhui Ma, Lijun Guo, Jiong Zhang 0004, Yitian Zhao |
CGI (1) | 4 |
| 2023 | TransGait: Multimodal-based gait recognition with set transformer
Lijun Guo, Rong Zhang 0007, Jiangbo Qian, Shangce Gao |
Appl. Intell. | 2 |
| 2023 | Blind inverse light transport using unrolling network
Wenting Yin, Xulun Ye, Lijun Guo |
Appl. Intell. | 5 |
| 2023 | An adaptive position-guided gravitational search algorithm for function optimization and image threshold segmentation
Anjing Guo, Yirui Wang 0001, Lijun Guo, Rong Zhang 0007, Yang Yu 0013, Shangce Gao |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Can relearning local representation help small networks for human pose estimation?
Dingning Xu, Lijun Guo, Rong Zhang 0007, Jiangbo Qian, Shangce Gao |
Neurocomputing | 2 |
| 2023 | Prediction of PM2.5 time series by seasonal trend decomposition-based dendritic neuron model
Zijing Yuan, Shangce Gao, Yirui Wang 0001, Chunzhi Hou, Lijun Guo |
Neural Comput. Appl. | 6 |
| 2023 | Laplacian Lp norm least squares twin support vector machine
Xijiong Xie, Feixiang Sun, Jiangbo Qian, Lijun Guo, Rong Zhang 0007, Xulun Ye, Zhijin Wang |
Pattern Recognit. | 4 |
| 2023 | Multi-Label Hashing for Dependency Relations Among Multiple ObjectivesabstractLearning hash functions have been widely applied for large-scale image retrieval. Existing methods usually use CNNs to process an entire image at once, which is efficient for single-label images but not for multi-label images. First, these methods cannot fully exploit independent features of different objects in one image, resulting in some small object features with important information being ignored. Second, the methods cannot capture different semantic information from dependency relations among objects. Third, the existing methods ignore the impacts of imbalance between hard and easy training pairs, resulting in suboptimal hash codes. To address these issues, we propose a novel deep hashing method, termed multi-label hashing for dependency relations among multiple objectives (DRMH). We first utilize an object detection network to extract object feature representations to avoid ignoring small object features and then fuse object visual features with position features and further capture dependency relations among objects using a self-attention mechanism. In addition, we design a weighted pairwise hash loss to solve the imbalance problem between hard and easy training pairs. Extensive experiments are conducted on multi-label datasets and zero-shot datasets, and the proposed DRMH outperforms many state-of-the-art hashing methods with respect to different evaluation metrics. Liangkang Peng, Jiangbo Qian, Zhengtao Xu, Lijun Guo |
IEEE Trans. Image Process. | 5 |
| 2023 | VLTENet: A Deep-Learning-Based Vertebra Localization and Tilt Estimation Network for Automatic Cobb Angle EstimationabstractScoliosis diagnosis and assessment rely upon Cobb angle estimation from X-ray images of the spine. Recently, automated scoliosis assessment has been greatly improved using deep learning methods. However, in such methods, the Cobb angle is usually predicted based on regression models that don't account for information of the spine structure. Alternatively, the Cobb angle can be estimated indirectly through landmark-detection and vertebra-segmentation, but this approach is still highly sensitive to small detection and segmentation errors. This paper proposes a novel deep-learning architecture, called the vertebra localization and tilt estimation network (VLTENet). This network boosts the Cobb angle estimation accuracy through employing vertebra localization and tilt estimation as network prediction goals. In particular, the VLTENet model innovatively combines a deep high-resolution network (HRNet) and a fully-convolutional U-Net architecture for capturing long-range contextual information, the overall structure, and local details in spinal X-ray images. A feature fusion channel attention (FFCA) module is also proposed to selectively emphasize more informative features and suppress less informative ones. In addition, a joint spine loss function (JS-Loss) is designed to account for the spine shape and other spatial constraints, so that the network focuses more on spine-related regions and ignore irrelevant background regions. Finally, we propose a new Cobb angle estimation method conforms with the clinical Cobb angle calculation guidelines, and produces accurate estimates for different types of scoliosis. Extensive experiments on the publically-available AASCE challenge dataset and on an in-house dataset demonstrated the superiority of our method for the task of automatic assessment of scoliosis. Lulin Zou, Lijun Guo, Rong Zhang 0007, Lixin Ni, Zhenzuo Chen, Xiuchao He |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Topology-Aware Learning for Semi-supervised Cross-domain Retinal Artery/Vein Classification
Jianyang Xie, Yonghuai Liu, Huaying Hao, Lijun Guo, Jiong Zhang 0004, Yitian Zhao |
CGI | 5 |
| 2022 | NerveFormer: A Cross-Sample Aggregation Network for Corneal Nerve Segmentation
Lei Mou, Shaodong Ma, Huazhu Fu, Lijun Guo, Yalin Zheng, Jiong Zhang 0004, Yitian Zhao |
MICCAI (4) | 5 |
| 2022 | LDNet: Lightweight dynamic convolution network for human pose estimation
Dingning Xu, Rong Zhang 0007, Lijun Guo, Cun Feng, Shangce Gao |
Adv. Eng. Informatics | 3 |
| 2022 | mmGaitSet: multimodal based gait recognition for countering carrying and clothing changes
Lijun Guo, Rong Zhang 0007, Xijiong Xie, Xulun Ye |
Appl. Intell. | 2 |
| 2022 | MetaCRS: unsupervised clustering of contigs with the recursive strategy of reducing metagenomic dataset's complexityabstractBACKGROUND: Metagenomics technology can directly extract microbial genetic material from the environmental samples to obtain their sequencing reads, which can be further assembled into contigs through assembly tools. Clustering methods of contigs are subsequently applied to recover complete genomes from environmental samples. The main problems with current clustering methods are that they cannot recover more high-quality genes from complex environments. Firstly, there are multiple strains under the same species, resulting in assembly of chimeras. Secondly, different strains under the same species are difficult to be classified. Thirdly, it is difficult to determine the number of strains during the clustering process. RESULTS: In view of the shortcomings of current clustering methods, we propose an unsupervised clustering method which can improve the ability to recover genes from complex environments and a new method for selecting the number of sample's strains in clustering process. The sequence composition characteristics (tetranucleotide frequency) and co-abundance are combined to train the probability model for clustering. A new recursive method that can continuously reduce the complexity of the samples is proposed to improve the ability to recover genes from complex environments. The new clustering method was tested on both simulated and real metagenomic datasets, and compared with five state-of-the-art methods including CONCOCT, Maxbin2.0, MetaBAT, MyCC and COCACOLA. In terms of the number and quality of recovered genes from metagenomic datasets, the results show that our proposed method is more effective. CONCLUSIONS: A new contigs clustering method is proposed, which can recover more high-quality genes from complex environmental samples. Zhongjun Jiang, Lijun Guo |
BMC Bioinform. | 3 |
| 2022 | Self-trained prediction model and novel anomaly score mechanism for video anomaly detection
AiBin Guo, Lijun Guo, Rong Zhang 0007, Yirui Wang 0001, Shangce Gao |
Image Vis. Comput. | 2 |
| 2022 | Symmetric uncertainty-incorporated probabilistic sequence-based ant colony optimization for feature selection in classification
Shangce Gao, Yong Zhang 0016, Lijun Guo |
Knowl. Based Syst. | 4 |
| 2022 | Band Selection for HSI Classification Using Binary Constrained OptimizationabstractHyperspectral images (HSIs) containing tens to hundreds of bands can be used in various image classification tasks. However, due to the high data redundancy of the spectral information, the acquiring and analysis of HSIs are usually relatively time-consuming and wasteful of storage space, and therefore limit the practical application of HSIs. Selecting a subset of bands without sacrificing classification accuracy is a strategy to relieve such problems. In this letter, we present an optimization-based method, which can jointly optimize the band selection (BS) and the classification network parameters for HSIs. The proposed method regards the discrete selection problem as a continuous constrained optimization problem and adaptively selects the informative band subsets for classification. Besides, the experimental results on three public datasets show that our BS method outperforms the state-of-the-art methods in terms of classification accuracy. Xueyan Tian, Chong Wang 0001, Lijun Guo |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Hierarchical attention network for attributed community detection of joint representation
Qiqi Zhao, Huifang Ma, Lijun Guo, Zhixin Li 0001 |
Neural Comput. Appl. | 3 |
| 2022 | Self-Label Refining for Unsupervised Person Re-IdentificationabstractFully unsupervised person Re-ID is a challenging task. State-of-the-art methods perform model training with the pseudo labels generated by clustering algorithms on the unlabeled dataset. However, the label noise caused by clustering limits the performance of person Re-ID tasks. To alleviate the problem, this paper proposes a Self-Label Refining network (SLRNet). It is considered that the local parts naturally mitigate the variation of intra-identity samples caused by cross-view. Thus, the self-label refining module (SLR) estimates the similarities between global and local pseudo labels with clustering consensus, and then it refines the global pseudo labels by integrating propagated local pseudo labels into global pseudo labels. Meanwhile, a symmetric ClusterNCE loss is further proposed to enhance the robustness of the network to noisy labels. Extensive experiments show that our method achieves state-of-the-art performance on three widely used person Re-ID datasets. Xiaoting Yu, Lijun Guo, Rong Zhang 0007 |
IEEE Signal Process. Lett. | 2 |
| 2022 | An Unsupervised Multi-Shot Person Re-Identification Method via Mutual Normalized Sparse Representation and Stepwise LearningabstractDue to abundant prior information and widespread applications, multi-shot based person re-identification has drawn increasing attention in recent years. In this paper, the high labeling cost and huge unlabeled data motivate us to focus on the unsupervised scenario and a unified coarse-to-fine framework is proposed, named by Mutual Normalized Sparse Representation (MNSR). Our method is an iteration procedure and each iteration involves two key steps: label estimation and metric model learning. In the former, we present a MNSR model to infer the pairwise labels of cross-camera by endowing sparse representation coefficient with the probability property. MNSR explicitly takes the mutually correlation between cameras into consideration and thus produces more accurate results. Meanwhile, we propose a probability-guided positive pairwise label prediction method to mine hard positive samples. For the latter, we learn a metric model with the estimated pairwise labels as supervision. In this procedure, we select some reliable labels for training by configuring with a stepwise learning method, rather than use all the estimated pair samples. This procedure helps to prevent the noise samples damaging the learning of discriminative metric model, especially for the initial iterations. Extensive experiments are conducted on four publicly available datasets, including PRID 2011, iLIDS-VID, SAIVT-SoftBio and MARS, and the results demonstrate the superior performance of the MNSR method in comparison with state-of-the-art unsupervised multi-shot person re-identification methods. Xiaobao Li, Qingyong Li, Wen Wang 0019, Lijun Guo |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Unsupervised Multi-shot Person Re-identification via Dynamic Bi-directional Normalized Sparse Representation
Xiaobao Li, Wen Wang 0019, Qingyong Li, Lijun Guo |
MMM (1) | 4 |
| 2021 | HLFNet: High-low Frequency Network for Person Re-IdentificationabstractPerson re-identification (re-ID) technology has attracted many scholars in the past few years. With the recent developments of deep learning technology, person re-ID has been greatly improved. However, the main chalenge of re-ID is to distinguish the detailed information in different images. Consequently, it is of significant importance to extract fine-grained features in the re-ID tasks. In the present study, a novel method, called the high-low frequency network (HLFNet), is proposed to effectively use the image information of different frequencies and focus on the detailed information between different individual images. In this regard, high frequency and low-frequency information are initially extracted from the original image, and then two backbones are applied to extract the features from the two information branches. Different frequencies of image information complement each other so that a better recognition effect can be achieved. Moreover, a local branch is utilized to extract the distinguishable local features for guiding the global feature branch in the training stage. Finally, only the extracted global feature from the trained network is required in the inference phase of re-ID. Performed experiments demonstrate that the proposed method can significantly enhance the feature representation accuracy and achieve the state-of-the-art performance on diverse benchmarks. Cen Liu, Lijun Guo, Rong Zhang 0007 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Discriminative Feature Network Based on a Hierarchical Attention Mechanism for Semantic Hippocampus SegmentationabstractThe morphological analysis of hippocampus is vital to various neurological studies including brain disorders and brain anatomy. To assist doctors in analyzing the shape and volume of the hippocampus, an accurate and automatic hippocampus segmentation method is highly demanded in the clinical practice. Given that fully convolutional networks (FCNs) have made significant contributions in biomedical image segmentation applications, we propose a notably discriminative feature network based on a hierarchical attention mechanism in hippocampal segmentation. First, considering the problem that the hippocampus is a rather small part in MR images, we design a context-aware high-level feature extraction module (CHFEM) to extract high-level features of scale invariance in the encoder stage. Further, we introduce a hierarchical attention mechanism into our segmentation framework. The mechanism is divided into three parts: a low-level feature spatial attention module (LFSAM) is developed to learn the spatial relationship between different pixels on each channel in the low-level stage of the encoder, a high-level feature channel attention module (HFCAM) is to model the semantic information relationship on different channel images in the high-level stage of the encoder, and a cross-connected attention module (CCAM) is designed in the decoder part to further suppress the noisy boundaries of hippocampus and simultaneously utilize the attentional low-level features from the encoder to better guide the high-level hippocampus edge segmentation in the decoder phase. The proposed approach achieves outstanding performance on the ADNI dataset and the Decathlon dataset compared with other semantic segmentation models and existing hippocampal segmentation approaches. Source code is available at https://github.com/LannyShi/Hippocampal-segmentation. Jiali Shi, Rong Zhang 0007, Lijun Guo, Linlin Gao, Huifang Ma |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Bayesian Adversarial Spectral Clustering With Unknown Cluster NumberabstractSpectral clustering is a popular tool in many unsupervised computer vision and machine learning tasks. Recently, due to the encouraging performance of deep neural networks, many conventional spectral clustering methods have been extended to the deep framework. Although these deep spectral clustering methods are quite powerful and effective, learning the cluster number from data is still a challenge. In this paper, we aim to tackle this problem by integrating the spectral clustering, generative adversarial network and low rank model within a unified Bayesian framework. First, we adapt the low rank method to the cluster number estimation problem. Then, an adversarial-learning-based deep clustering method is proposed and incorporated. When introducing the spectral clustering method into our model clustering procedure, a hidden space structure preservation term is proposed. Via a Bayesian framework, the structure preservation term is embedded into the generative process, which can then be used to deduce a spectral clustering in the optimization procedure. Finally, we derive a variational-inference-based method and embed it into the network optimization and learning procedure. Experiments on different datasets prove that our model has the cluster number estimation capability and show that our method can outperform many similar graph clustering methods. Xulun Ye, Jieyu Zhao 0002, Yu Chen 0067, Lijun Guo |
IEEE Trans. Image Process. | 4 |
| 2019 | Heterogeneity of synaptic input connectivity regulates spike-based neuronal avalanches
Shengdun Wu, Yangsong Zhang 0001, Yan Cui 0004, Jiakang Wang, Lijun Guo, Dezhong Yao 0001, Peng Xu 0001, Daqing Guo |
Neural Networks | 6 |
| 2019 | A Nonparametric Deep Generative Model for Multimanifold ClusteringabstractMultimanifold clustering separates data points approximately lying on a union of submanifolds into several clusters. In this paper, we propose a new nonparametric Bayesian model to handle the manifold data structure. In our framework, we first model the manifold mapping function between Euclidean space and topological space by applying a deep neural network, and then construct the corresponding generation process of multiple manifold data. To solve the posterior approximation problem, in the optimization procedure, we apply a variational auto-encoder-based optimization algorithm. Especially, as the manifold algorithm has poor performance on the real dataset where nonmanifold and manifold clusters are appearing simultaneously, we expand our proposed manifold algorithm by integrating it with the original Dirichlet process mixture model. Experimental results have been carried out to demonstrate the state-of-the-art clustering performance. Xulun Ye, Jieyu Zhao 0002, Lijun Guo |
IEEE Trans. Cybern. | 4 |
| 2017 | Unsupervised video object segmentation by spatiotemporal graphical model
Lijun Guo, Ting-Ting Cheng, Yuanjie Huang, Jieyu Zhao 0002, Rong Zhang 0007 |
Multim. Tools Appl. | 1 |
| 2015 | Video human segmentation based on multiple-cue integration
Lijun Guo, Ting-Ting Cheng, Rong Zhang 0007, Jieyu Zhao 0002 |
Signal Process. Image Commun. | 1 |
| 2013 | Unsupervised Natural Image Segmentation via Bayesian Ying-Yang Harmony Learning Theory
Shaojun Zhu, Jieyu Zhao 0002, Lijun Guo |
Neurocomputing | 3 |