EDBT 2026 Demo / reviewers in the wild / expert
Rong Zhang 0007
dblp:13/5366-7
· DBLP profile ↗
28ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0001-8019-245XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond reconstruction: Enhancing masked autoencoders with contrastive learning for video representation learning
Yawei Feng, Lijun Guo, Guitao Yu, Rong Zhang 0007, Jiangbo Qian, Chong Wang 0001, Shangce Gao |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Boosting representation diversity in video transformers via segmented contrastive masked autoencoders
Yawei Feng, Lijun Guo, Guitao Yu, Rong Zhang 0007, Jiangbo Qian, Chong Wang 0001, Shangce Gao |
Neurocomputing | 4 |
| 2025 | Multi-Modal Timely Pancreatitis Severity Assessment via Hierarchical Evidential Conflictive LearningabstractAcute pancreatitis (AP) can rapidly progress to severe acute pancreatitis (SAP), which carries a high risk of mortality. Early screening of high-risk patients using computed tomography (CT) imaging and laboratory indicators is beneficial to improving clinical outcomes. To fully leverage the advantages of multi-modal data, we address pancreatitis severity assessment from multiple views and introduce multi-view categorical uncertainty quantification to enhance model reliability. Existing uncertainty-aware multi-view classification methods often assume consistency across views and indiscriminately reduce uncertainty. Nevertheless, in the context of pancreatitis, view conflicts underlying different data modalities are common as each modality provides unique pathological information. To tackle these challenges, we develop a Hierarchical Evidential Conflictive Learning (HECL) method for pancreatitis severity assessment, which estimates both the uncertainty and belief mass through subjective logic from multi-modal data and aggregates the conflicting opinions via logarithmic opinion pooling, by weighted geometric averaging of belief distributions and effectively transforming view conflicts into measurable prediction uncertainties. Additionally, HECL incorporates a hierarchical fusion strategy and introduces cross-modal pseudo-views to enhance the representation and interaction across different views. Experimental results show that HECL significantly outperforms several state-of-the-art multi-view methods. Visualization analysis reveals that the model can accurately localize pancreatic lesion regions and identify key predictive indicators, which can provide trustworthy diagnoses for pancreatitis severity assessment. Houli Fan, Lijun Guo, Xiuchao He, Bang Cheng, Yingqing Zeng, Jiang Duan, Rong Zhang 0007 |
BIBM | 8 |
| 2025 | PoseDucer: Implicit relation inducement for invisible keypoint reconstruction in real-world occluded scenes
Junneng Feng, Rong Zhang 0007, Yirui Wang 0001, Shangce Gao, Lijun Guo |
Neurocomputing | 2 |
| 2025 | Robust auxiliary modality is beneficial for video-based cloth-changing person re-identification
Youming Chen, Ting Tuo, Lijun Guo, Rong Zhang 0007, Yirui Wang 0001, Shangce Gao |
Image Vis. Comput. | 4 |
| 2025 | Part2Pose: Inferring Human Pose From Parts in Complex ScenesabstractMost of existing Human Pose Estimation (HPE) methods struggle to handle with challenges such as changeable poses, complex backgrounds, and occlusion encountered in complex scenes. To address these problems, a novel HPE network, called Part2Pose, is proposed in this paper. In our Part2Pose, instead of focusing on small-sized keypoints like existing HPE methods do, we first extract image features based on human body parts to expand the detection scope. This strategy enhances the robustness of the extracted features to variations and distractions in complex scenes. Then, a Transformer-based Global Part Relation Module (GPRM) and a graph convolutional network-based Local Part Relation Module (LPRM) are used to capture global and local relationships among different body parts to help infer the position of keypoints. Extensive experiments on challenging datasets, including COCO, CrowdPose and OCHuman, show that the proposed Part2Pose can surpass existing popular state-of-the-art HPE methods. The combination with lightweight networks confirms the robustness and generalizability of our Part2Pose. Rong Zhang 0007, Junneng Feng, Cun Feng, Yirui Wang 0001, Lijun Guo |
IEEE Signal Process. Lett. | 1 |
| 2024 | A Deep-Learning-Based Lumbosacral Localization and Landmark Detection Network for Automatic Lumbar Stability and Spondylolisthesis Grading AssessmentabstractThe accurate detection of vertebral landmarks is crucial for clinical diagnosis and research on the lumbar stability and spondylolisthesis grading. However, the small size of vertebral landmarks in X-ray images and the morphological similarity among vertebrae complicate this detection task. Recent advances in deep learning have enhanced the spinal landmark detection. To further improve the assessment methods for the lumbar stability and spondylolisthesis grading, we propose the LSLD-Net, a novel network based on the lumbosacral localization and landmark detection. In clinical practices, the X-ray images used for evaluating the lumbar spine stability can often include the extraneous information from non-lumbosacral areas, such as the thoracic spine. The proposed LSLD-Net first extracts the lumbosacral region from the complex X-ray images and then performs the precise landmark detection within the identified area. The landmark detection stage integrates the HRNet and U-net architectures, effectively capturing the long-range contextual information, the overall structural layout, and the fine local details in lumbar X-ray images to optimize landmark detection. Additionally, we propose a Multi-Scale Attention Module that enhances the relevant features and suppresses the irrelevant ones through channel and spatial attention, thereby achieving precise landmark detection and improving the robustness of the network. The evaluations on the private and public BUU-LSPINE datasets indicate that the LSLD-Net surpasses other state-of-the-art methods in landmark detection, enhancing the accuracy and efficiency of Sagittal Displacement and Intervertebral Space Angle measurements. This performance excels in assessing the lumbar stability and spondylolisthesis grading, offering the significant support to clinicians for early quantitative diagnosis and evaluation. Rong Zhang 0007, Baolin Xu, Dongdong Xia, Lijun Guo |
BIBM | 2 |
| 2024 | Tri-Hash Progressive Sampling Neural Attenuation Field for Sparse-View CBCT ReconstructionabstractSparse-view CBCT reconstruction is essential to reduce the X-ray radiation dose in clinical CBCT imaging. However, reducing the number of views often lower the image quality. Existing neural radiance field (NeRF) technologies can achieve the high-quality 3D reconstruction and novel view generation in natural scenes under sparse-view conditions. Nevertheless, applying these technologies to the 3D reconstruction of human tissues in the medical field still has limitations. The recently proposed neural attenuation field (NAF) technique has shown progress in 3D reconstruction of human tissues in medical images. However, in the context of sparse views, there remains the issue of inferior reconstruction quality due to insufficient acquisition of spatial structural information in human tissues. To address these challenges, we propose a novel framework called the Tri-Hash Progressive Sampling Neural Attenuation Field (THP-NAF). First, we introduced an Enhanced Tri-Hash Representation mechanism that enhanced the extraction of 3D spatial information through the 2D plane mapping. This mechanism captures more contextual and spatial information and achieved an optimized balance between the image quality and generation efficiency. Additionally, to mitigate the sampling inefficiency caused by random sampling, we employed a Sobel-based adaptive point–ray sampling strategy. This strategy combined the global and local information for structure-aware ray sampling and could dynamically adjust the number of sampling points, thereby enhancing the sampling flexibility and efficiency. Our method was validated across multiple datasets, demonstrating its ability to improve image quality and its significant potential for clinical applications. Lijun Guo, Rong Zhang 0007, Wenming He, Shangce Gao |
BIBM | 3 |
| 2024 | Progressive Learning Based Knowledge Distillation for Low Resolution Cerebral Microbleed SegmentationabstractThis study aims to address key technical issues in the segmentation of Cerebral MicroBleeds (CMBs) based on Low-Resolution (LR) Magnetic Resonance Imaging (MRI) data. There are two challenges in this task. First, the CMB lesions are typically small in size and easily confused with various mimics. Second, anisotropy becomes more prominent and adverse in LR MRI sequences than HR sequences. To address these issues, we propose a Progressive Learning based Knowledge Distillation method. This method progressively transfers knowledge from HR models to their LR counterparts, thereby minimizing the occurrence of false positives attributable to noise from Super-Resolution. To further eliminate the influence of anisotropy, an encoding-enhanced network, called E2U-Net, is proposed in this paper. It can effectively capture anisotropic information and mitigates potential feature loss. The experimental results on multiple publicly accessible CMBs datasets demonstrated the superiority of our proposed approach over existing deep-learning methods. Tianxiang Xia, Rong Zhang 0007, Zhenzuo Chen, Guomin Xie, Xiping Wu, Zhongyue Lv, Lijun Guo |
ICASSP | 2 |
| 2024 | QSMT-net: A query-sensitive proposal and multi-temporal-span matching network for video grounding
Qingqing Wu 0014, Lijun Guo, Rong Zhang 0007, Jiangbo Qian, Shangce Gao |
Image Vis. Comput. | 3 |
| 2023 | IIESC-Net: Incorporating Implicit and Explicit Structural Constraints for Hip Joint Landmark Detection in pelvic X-rayabstractThe hip joint, a vital weight-bearing joint in the human body, is susceptible to various hip-related diseases. The accurate identification of anatomical landmarks in the hip joint is essential for both disease diagnosis and surgical planning. However, these landmarks are often inconspicuous in X-ray images, where irrelevant background interference increases detection difficulty. In this study, we proposed an IIESC-Net model that combined local and global features, enabling a hierarchical understanding of the implicit structural characteristics of the hip joint. Additionally, drawing from domain expertise based on the explicit physiological structure of the hip joint, we designed a Hip Morphology-Aware loss function to constrain large landmark errors through the application of high-confidence landmarks with robust distinctive identification, thereby achieving accuracy in automatic detection. Furthermore, we constructed a dataset comprising of 843 pelvic X-ray images. The experimental results demonstrated a substantial enhancement in hip joint landmark detection accuracy attributed to the proposed IIESC-Net. This innovation established state-of-the-art performance, notably excelling in attaining heightened successful detection rates under stringent error tolerance. This achievement has profound practical implications for clinical applications. Lijun Guo, Lixin Ni, Xiuchao He, Rong Zhang 0007 |
BIBM | 6 |
| 2023 | A Dual-View Fusion Network for Automatic Spinal Keypoint Detection in Biplane X-ray ImagesabstractAccurate keypoint detection in medical images of the spine is critical for the assessment, diagnosis, treatment planning, and clinical investigation of spinal deformities. However, due to severe occlusions of spinal structures in lateral X-ray images, accurate keypoint detection can be hardly achieved in lateral X-ray images based on single-view information. Thus, methods based on both the anterior-posterior (AP) and lateral (LAT) X-ray image views have been proposed to alleviate occlusion problems and achieve better keypoint detection performance. Although some progress has been made with these dual-view methods, they do not effectively exploit a priori knowledge of the spine and hence cannot adequately account for the structural correlation of the vertebrae across views. In this paper, a new dual-view fusion network (DVFNet) framework is proposed for keypoint detection in spinal X-ray images. This framework obtains structural correlations between AP and LAT views of the spine based on a priori spine knowledge represented by high-level semantic features. Meanwhile, the proposed framework combines local and global features extracted respectively by a local subnetwork and a global subnetwork. On the one hand, the local subnetwork is constructed as an enhanced codec structure based on both the AP and LAT views. This subnetwork is trained to output local features that contain both joint semantic features of the two views and independent fine-grained features of each individual view. This scheme leads to accurate keypoint estimation locally. On the other hand, the global subnetwork utilizes a self-attention mechanism to extract view-specific global features based on either the AP view or the LAT view in order to eliminate ambiguity, and reduce confusion on keypoint locations. Further, we propose a weighted feature fusion (WFF) module for adaptive fusion of the local and global features. We evaluated the DVFNet model on a private dataset and found that our proposed method achieves more accurate spinal keypoint detection compared to other state-of-the-art methods, and thus our method can provide reliable assistance to clinicians. Lijun Guo, Rong Zhang 0007, Xiuchao He |
BIBM | 3 |
| 2023 | TransGait: Multimodal-based gait recognition with set transformer
Lijun Guo, Rong Zhang 0007, Jiangbo Qian, Shangce Gao |
Appl. Intell. | 3 |
| 2023 | An adaptive position-guided gravitational search algorithm for function optimization and image threshold segmentation
Anjing Guo, Yirui Wang 0001, Lijun Guo, Rong Zhang 0007, Yang Yu 0013, Shangce Gao |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Can relearning local representation help small networks for human pose estimation?
Dingning Xu, Lijun Guo, Rong Zhang 0007, Jiangbo Qian, Shangce Gao |
Neurocomputing | 3 |
| 2023 | Laplacian Lp norm least squares twin support vector machine
Xijiong Xie, Feixiang Sun, Jiangbo Qian, Lijun Guo, Rong Zhang 0007, Xulun Ye, Zhijin Wang |
Pattern Recognit. | 5 |
| 2023 | VLTENet: A Deep-Learning-Based Vertebra Localization and Tilt Estimation Network for Automatic Cobb Angle EstimationabstractScoliosis diagnosis and assessment rely upon Cobb angle estimation from X-ray images of the spine. Recently, automated scoliosis assessment has been greatly improved using deep learning methods. However, in such methods, the Cobb angle is usually predicted based on regression models that don't account for information of the spine structure. Alternatively, the Cobb angle can be estimated indirectly through landmark-detection and vertebra-segmentation, but this approach is still highly sensitive to small detection and segmentation errors. This paper proposes a novel deep-learning architecture, called the vertebra localization and tilt estimation network (VLTENet). This network boosts the Cobb angle estimation accuracy through employing vertebra localization and tilt estimation as network prediction goals. In particular, the VLTENet model innovatively combines a deep high-resolution network (HRNet) and a fully-convolutional U-Net architecture for capturing long-range contextual information, the overall structure, and local details in spinal X-ray images. A feature fusion channel attention (FFCA) module is also proposed to selectively emphasize more informative features and suppress less informative ones. In addition, a joint spine loss function (JS-Loss) is designed to account for the spine shape and other spatial constraints, so that the network focuses more on spine-related regions and ignore irrelevant background regions. Finally, we propose a new Cobb angle estimation method conforms with the clinical Cobb angle calculation guidelines, and produces accurate estimates for different types of scoliosis. Extensive experiments on the publically-available AASCE challenge dataset and on an in-house dataset demonstrated the superiority of our method for the task of automatic assessment of scoliosis. Lulin Zou, Lijun Guo, Rong Zhang 0007, Lixin Ni, Zhenzuo Chen, Xiuchao He |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | LDNet: Lightweight dynamic convolution network for human pose estimation
Dingning Xu, Rong Zhang 0007, Lijun Guo, Cun Feng, Shangce Gao |
Adv. Eng. Informatics | 2 |
| 2022 | mmGaitSet: multimodal based gait recognition for countering carrying and clothing changes
Lijun Guo, Rong Zhang 0007, Xijiong Xie, Xulun Ye |
Appl. Intell. | 3 |
| 2022 | Self-trained prediction model and novel anomaly score mechanism for video anomaly detection
AiBin Guo, Lijun Guo, Rong Zhang 0007, Yirui Wang 0001, Shangce Gao |
Image Vis. Comput. | 3 |
| 2022 | Self-Label Refining for Unsupervised Person Re-IdentificationabstractFully unsupervised person Re-ID is a challenging task. State-of-the-art methods perform model training with the pseudo labels generated by clustering algorithms on the unlabeled dataset. However, the label noise caused by clustering limits the performance of person Re-ID tasks. To alleviate the problem, this paper proposes a Self-Label Refining network (SLRNet). It is considered that the local parts naturally mitigate the variation of intra-identity samples caused by cross-view. Thus, the self-label refining module (SLR) estimates the similarities between global and local pseudo labels with clustering consensus, and then it refines the global pseudo labels by integrating propagated local pseudo labels into global pseudo labels. Meanwhile, a symmetric ClusterNCE loss is further proposed to enhance the robustness of the network to noisy labels. Extensive experiments show that our method achieves state-of-the-art performance on three widely used person Re-ID datasets. Xiaoting Yu, Lijun Guo, Rong Zhang 0007 |
IEEE Signal Process. Lett. | 3 |
| 2021 | HLFNet: High-low Frequency Network for Person Re-IdentificationabstractPerson re-identification (re-ID) technology has attracted many scholars in the past few years. With the recent developments of deep learning technology, person re-ID has been greatly improved. However, the main chalenge of re-ID is to distinguish the detailed information in different images. Consequently, it is of significant importance to extract fine-grained features in the re-ID tasks. In the present study, a novel method, called the high-low frequency network (HLFNet), is proposed to effectively use the image information of different frequencies and focus on the detailed information between different individual images. In this regard, high frequency and low-frequency information are initially extracted from the original image, and then two backbones are applied to extract the features from the two information branches. Different frequencies of image information complement each other so that a better recognition effect can be achieved. Moreover, a local branch is utilized to extract the distinguishable local features for guiding the global feature branch in the training stage. Finally, only the extracted global feature from the trained network is required in the inference phase of re-ID. Performed experiments demonstrate that the proposed method can significantly enhance the feature representation accuracy and achieve the state-of-the-art performance on diverse benchmarks. Cen Liu, Lijun Guo, Rong Zhang 0007 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Discriminative Feature Network Based on a Hierarchical Attention Mechanism for Semantic Hippocampus SegmentationabstractThe morphological analysis of hippocampus is vital to various neurological studies including brain disorders and brain anatomy. To assist doctors in analyzing the shape and volume of the hippocampus, an accurate and automatic hippocampus segmentation method is highly demanded in the clinical practice. Given that fully convolutional networks (FCNs) have made significant contributions in biomedical image segmentation applications, we propose a notably discriminative feature network based on a hierarchical attention mechanism in hippocampal segmentation. First, considering the problem that the hippocampus is a rather small part in MR images, we design a context-aware high-level feature extraction module (CHFEM) to extract high-level features of scale invariance in the encoder stage. Further, we introduce a hierarchical attention mechanism into our segmentation framework. The mechanism is divided into three parts: a low-level feature spatial attention module (LFSAM) is developed to learn the spatial relationship between different pixels on each channel in the low-level stage of the encoder, a high-level feature channel attention module (HFCAM) is to model the semantic information relationship on different channel images in the high-level stage of the encoder, and a cross-connected attention module (CCAM) is designed in the decoder part to further suppress the noisy boundaries of hippocampus and simultaneously utilize the attentional low-level features from the encoder to better guide the high-level hippocampus edge segmentation in the decoder phase. The proposed approach achieves outstanding performance on the ADNI dataset and the Decathlon dataset compared with other semantic segmentation models and existing hippocampal segmentation approaches. Source code is available at https://github.com/LannyShi/Hippocampal-segmentation. Jiali Shi, Rong Zhang 0007, Lijun Guo, Linlin Gao, Huifang Ma |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Unsupervised video object segmentation by spatiotemporal graphical model
Lijun Guo, Ting-Ting Cheng, Yuanjie Huang, Jieyu Zhao 0002, Rong Zhang 0007 |
Multim. Tools Appl. | 5 |
| 2017 | Fast algorithm for 2D fragment assembly based on partial EMD
Shuang-Min Chen, Zhenyu Shu, Shi-Qing Xin, Jieyu Zhao 0002, Guang Jin, Rong Zhang 0007, Jürgen Beyerer |
Vis. Comput. | 7 |
| 2015 | In-camera JPEG compression detection for doubly compressed images
Rong Zhang 0007, Rangding Wang |
Multim. Tools Appl. | 1 |
| 2015 | Video human segmentation based on multiple-cue integration
Lijun Guo, Ting-Ting Cheng, Rong Zhang 0007, Jieyu Zhao 0002 |
Signal Process. Image Commun. | 4 |
| 2011 | Distinguishing Photographic Images and Photorealistic Computer Graphics Using Visual Vocabulary on Local Image Edges
Rong Zhang 0007, Rangding Wang, Tian-Tsong Ng |
IWDW | 1 |