Zongyi Xu

dblp:125/3642 · DBLP profile ↗
← Back
20ranked-venue papers
12as first author
14since 2021 · last 2026
0000-0002-8267-6834ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Computer networks · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-Cancer
abstract
Single-cell RNA sequencing (scRNA-seq) is essential for decoding tumor heterogeneity. However, pan-cancer research still faces two key challenges: learning discriminative and efficient single-cell representations, and establishing a comprehensive evaluation benchmark. In this paper, we introduce \algoname, a lightweight hybrid neural network that combines the strengths of Transformers and state-space models to achieve a balance between performance and efficiency. \algoname consists of a front-end local-context encoder with shared self-attention layers to capture complex, order-independent gene interactions; and a back-end global sequential feature decoder that efficiently integrates global context using a linear-time state-space model. This modular design preserves the expressive power of Transformers while leveraging the scalability of Mamba to enable transcriptome modeling, effectively capturing both local and global regulatory signals. To enable robust evaluation, we also construct a large-scale pan-cancer single-cell benchmark, \algoname Bench, containing over 3.5 million high-quality cells across 33 cancer subtypes, curated through a rigorous preprocessing pipeline. Experimental results show that \algoname outperforms state-of-the-art models on our pan-cancer benchmark (+4.0\%) and across multiple public tasks, including cell type annotation (+7.4\%), batch integration (+4.0\%) and multi-omics integration (+3.1\%).
Xiaoshui Huang, Tianlin Zhu, Yifan Zuo 0001, Xue Xia 0005, Zonghan Wu, Jiebin Yan, Dingli Hua, Zongyi Xu, Yuming Fang 0001, Jian Zhang 0002
AAAI8
2026 Non-target divergence hypothesis: Toward understanding modality differences in cross-modal knowledge distillation
Zongyi Xu, Xiaoshui Huang, Shanshan Zhao 0001, Xinbo Gao 0001
Neural Networks2
2026 HiViTrack: Hierarchical vision transformer with efficient target-prompt update for visual object tracking
Bailian Xie, Zongyi Xu, Weisheng Li 0001, Xinbo Gao 0001
Pattern Recognit.5
2026 GarmentRec: Towards Individual Garment Reconstruction From a Monocular Human Image
abstract
Reconstructing high-quality garment models from monocular images is important as it provides a practical and effective solution for human digitization and virtual try-on etc. Recent implicit function-based garment reconstruction methods recover free-form geometry but struggle to reconstruct individual garment meshes from human images and tend to produce disembodied limbs or degenerate shapes for novel views. In contrast, explicit parametric garment template models can be utilised to construct separate meshes and constrain the shape reconstruction robustly. However, this limits the reconstruction of garment details and shape variations, such as the wrinkles and pockets etc. To address this problem, in this paper, we introduce a novel explicit garment template that is designed for both closed and open garment topology. Powered by our new garment template, we further propose a detailed garment reconstruction method based on a monocular view that can process both the closed and open types for shape recovery. To capture those challenging parts with unknown geometry and topology, we predict displacement maps on the parameterization domain for the target garment from the monocular image and elaborate it to the 3D garment surface via the UV coordinates, achieving realistic details on the 3D garment shape. Extensive experiments demonstrate the accuracy and robustness of our method and show that realistic details like garment wrinkles and pockets can be faithfully recovered in an explicit way. The code and dataset are available at https://github.com/worryDes/GarmentRec.
Zongyi Xu, Shiyang Cheng 0001, Wang Fei, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2025 ALCReg: Active Label Correction for Partial Point Cloud Registration
abstract
Deep point cloud registration methods encounter challenges due to partial overlaps and are heavily reliant on labeled data. In this paper, we propose ALCReg, an active label correction method for partial point cloud registration learning. ALCReg utilises a multimodal approach to generate pseudo labels, mitigating the cold-start issue in active learning. To ensure the diversity and representativeness of selected samples, we propose an inlier ratio based query strategy for manual correction. Furthermore, an innovative self-correction mechanism based on consistency is introduced, allowing the model to refine pseudo labels autonomously and further improve model performance. Experimental results on the 3DMatch and 3DLoMatch datasets demonstrate that ALCReg achieves comparable performance with the fully-supervised registration methods, even with only 5% of labeled samples, making it the first active learning method tailored for partial point cloud registration. Code is available at https://github.com/Jiang0903/ALCReg.
Zongyi Xu, Xinqi Jiang, Shanshan Zhao 0001, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001
ICME1
2025 S2Reg: Structure-semantics collaborative point cloud registration
Zongyi Xu, Xinqi Jiang, Shiyang Cheng 0001, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001
Pattern Recognit.1
2025 CF3d: Category fused 3D point cloud retrieval
Zongyi Xu, Ruicheng Zhang, Zuo Li, Shiyang Cheng 0001, Huiyu Zhou 0001, Weisheng Li 0001, Xinbo Gao 0001
Signal Process.1
2025 Weakly Supervised LiDAR Semantic Segmentation via Scatter Image Annotation
abstract
Weakly supervised LiDAR semantic segmentation has made significant strides with limited labeled data. However, most existing methods focus on the network training under weak supervision, while efficient annotation strategies remain largely unexplored. To tackle this gap, we implement LiDAR semantic segmentation using scatter image annotation, effectively integrating an efficient annotation strategy with network training. Specifically, we propose employing scatter images to annotate LiDAR point clouds, combining a pre-trained optical flow estimation network with a foundational image segmentation model to rapidly propagate manual annotations into dense labels for both images and point clouds. Moreover, we propose ScatterNet, a network that includes three pivotal strategies to reduce the performance gap caused by such annotations. First, it utilizes dense semantic labels as supervision for the image branch, alleviating the modality imbalance between point clouds and images. Second, an intermediate fusion branch is proposed to obtain multimodal texture and structural features. Finally, a perception consistency loss is introduced to determine which information needs to be fused and which needs to be discarded during the fusion process. Extensive experiments on the nuScenes and SemanticKITTI datasets demonstrate that our method requires less than 0.02% of the labeled points to achieve over 95% of the performance of fully-supervised methods. Notably, our labeled points are only 5% of those used in the most advanced weakly supervised methods.
Zongyi Xu, Xiaoshui Huang, Shanshan Zhao 0001, Xinqi Jiang, Xinbo Gao 0001
IEEE Trans. Multim.2
2024 Retrieval-and-alignment based large-scale indoor point cloud semantic segmentation
Zongyi Xu, Xiaoshui Huang, Yangfu Wang, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001
Sci. China Inf. Sci.1
2024 IGReg: Image-Geometry-Assisted Point Cloud Registration via Selective Correlation Fusion
abstract
Point cloud registration suffers from repeated patterns and low geometric structures in indoor scenes. The recent transformer utilises attention mechanism to capture the global correlations in feature space and improves the registration performance. However, for indoor scenarios, global correlation loses its advantages as it cannot distinguish real useful features and noise. To address this problem, we propose an image-geometry-assisted point cloud registration method by integrating image information into point features and selectively fusing the geometric consistency with respect to reliable salient areas. Firstly, an Intra-Image-Geometry fusion module is proposed to integrate the texture and structure information into the point feature space by the cross-attention mechanism. Initial corresponding superpoints are acquired as salient anchors in the source and target. Then, a selective correlation fusion module is designed to embed the correlations between the salient anchors and points. During training, the saliency location and selective correlation fusion modules exchange information iteratively to identify the most reliable salient anchors and achieve effective feature fusion. The obtained distinctive point cloud features allow for accurate correspondence matching, leading to the success of indoor point cloud registration. Extensive experiments are conducted on 3DMatch and 3DLoMatch datasets to demonstrate the outstanding performance of the proposed approach compared to the state-of-the-art, particularly in those geometrically challenging cases such as repetitive patterns and low-geometry regions.
Zongyi Xu, Xinqi Jiang, Changjun Gu, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Multim.1
2023 Hierarchical Point-based Active Learning for Semi-supervised Point Cloud Semantic Segmentation
abstract
Impressive performance on point cloud semantic segmentation has been achieved by fully-supervised methods with large amounts of labelled data. As it is labour-intensive to acquire large-scale point cloud data with point-wise labels, many attempts have been made to explore learning 3D point cloud segmentation with limited annotations. Active learning is one of the effective strategies to achieve this purpose but is still under-explored. The most recent methods of this kind measure the uncertainty of each pre-divided region for manual labelling but they suffer from redundant information and require additional efforts for region division. This paper aims at addressing this issue by developing a hierarchical point-based active learning strategy. Specifically, we measure the uncertainty for each point by a hierarchical minimum margin uncertainty module which considers the contextual information at multiple levels. Then, a feature-distance suppression strategy is designed to select important and representative points for manual labelling. Besides, to better exploit the unlabelled data, we build a semi-supervised segmentation framework based on our active strategy. Extensive experiments on the S3DIS and ScanNetV2 datasets demonstrate that the proposed framework achieves 96.5% and 100% performance of fully-supervised baseline with only 0.07% and 0.1% training data, respectively, outperforming the state-of-the-art weakly-supervised and active learning methods. The code will be available at https://github.com/SmiletoE/HPAL.
Zongyi Xu, Shanshan Zhao 0001, Qianni Zhang, Xinbo Gao 0001
ICCV1
2022 Anomaly Warning: Learning and Memorizing Future Semantic Patterns for Unsupervised Ex-ante Potential Anomaly Prediction
abstract
Existing video anomaly detection methods typically utilize reconstruction or prediction error to detect anomalies in the current frame. However, these methods cannot predict ex-ante potential anomalies in future frames, which is imperative in real scenes. Inspired by the ex-ante prediction ability of humans, we propose an unsupervised Ex-ante Potential Anomaly Prediction Network (EPAP-Net), which learns to build a semantic pool to memorize the normal semantic patterns of future frames for indirect anomaly prediction. At the training time, the memorized patterns are encouraged to be discriminated through our Semantic Pool Building Module (SPBM) with the novel padding and updating strategies. Moreover, we present a novel Semantic Similarity Loss (SSLoss) at the feature level to maximize the semantic consistency of memorized items and corresponding future frames. Specially, to enhance the value of our work, we design a Multiple Frames Prediction module (MFP) to achieve anomaly prediction in future multiple frames. At the test time, we utilize the trained semantic pool instead of ground truth to evaluate the anomalies of future frames. Besides, to obtain better feature representations for our task, we introduce a novel Channel-selected Shift Encoder (CSE), which shifts channels along the temporal dimension between the input frames to capture motion information without generating redundant features. Experimental results demonstrate that the proposed EPAP-Net can effectively predict the potential anomalies in future frames and exhibit superior or competitive performance on video anomaly detection.
Jiaxu Leng, Mingpi Tan, Xinbo Gao 0001, Wen Lu 0004, Zongyi Xu
ACM Multimedia5
2022 Robust real-world point cloud registration by inlier detection
Xiaoshui Huang, Yangfu Wang, Sheng Li 0020, Guofeng Mei, Zongyi Xu, Yucheng Wang 0003, Jian Zhang 0002, Mohammed Bennamoun
Comput. Vis. Image Underst.5
2021 Building High-Fidelity Human Body Models From User-Generated Data
abstract
We propose a key point-based approach, refers to asKPhub-PC, to estimate high-fidelity human body models from low-quality point clouds acquired with an affordable 3D scanner and a variationKPhub-Ithat can achieve the same purpose based on low-resolution single images taken by smartphones. In KPhub-PC, a sparse set of key points is annotated to guide the deformation of a parametric 3D human body model SMPL and then a high-fidelity human body model that can explain the target point clouds is built. Besides building 3D human body models from point clouds, KPhub-I is designed to estimate accurate 3D human body models from single 2D images. The SMPL model is fitted to 2D joints and the boundary of the human body which are detected using CNN based methods automatically. Considering that people are in stable poses most of the time, a stable pose prior is defined from CMU motion capture dataset for further improving accuracy. Extensive experiments demonstrate that in both types of user-generated data, the proposed approaches can build believable and animatable human body models robustly. Our approach outperforms the state-of-the-arts in the accuracy of both human body shape and pose estimation.
Zongyi Xu, Yindi Zhu, Huiyu Zhou 0001, Qianni Zhang
IEEE Trans. Multim.1
2020 DNRTI: A Large-scale Dataset for Named Entity Recognition in Threat Intelligence
abstract
Named entity recognition is an important and challenging problem in Natural language processing. Although the past decade has witnessed major advances in entity recognition in many fields, such successes have been slow to network security field, not only because of the data in the network security field is very professional, but also due to the sensitive information in the data. To advance named entity recognition research in network security field, we introduce a large-scale Dataset for Named Entity Recognition in Threat Intelligence (DNRTI). To this end, we collect more than 300 pieces of threat intelligence. The data in DNRTI is all annotated by experts in threat intelligence interpretation using 13 object categories. The fully annotated DNRTI contains 175220 words. To build a baseline for named entity recognition in the threat intelligence field, we evaluate some deep learning model on DNRTI. Experiments demonstrate that DNRTI well represents the key information in threat intelligence and are quite challenging.
Xuren Wang, Xinpei Liu, Shengqin Ao, Zhengwei Jiang, Zongyi Xu, Zihan Xiong, Mengbo Xiong
TrustCom6
2018 Region Based User-Generated Human Body Scan Registration
abstract
We present a region-based registration method to robustly register low-quality human body scans that are acquired with cost-effective devices accessible to general users. When traditional closest point based registration approaches are performed on these noisy scan data, it is easy to fall into local minimum. To address this problem, we learn prior knowledge of body shape from publicly available dataset and combine it with the Iterative Closest Point (ICP) algorithm. Firstly, sparse markers are used to change pose of template, making it perform in the same way as target scans do. In the registration stage, the holistic shape model for the basic figure of human and a set of local shape models for describing the details of each human body part are trained. We fit the holistic model roughly to the target mesh. To capture more body details, we combine local shape models with the non-rigid ICP method to deform the template part-by-part. Extensive experiments over data scanned using devices from professional to low-cost types verify that our approach is both accurate and robust to incomplete and noisy data.
Zongyi Xu, Qianni Zhang
ICME1
2018 Multilevel active registration for kinect human body scans: from low quality to high quality
abstract
Registration of 3D human body has been a challenging research topic for over decades. Most of the traditional human body registration methods require manual assistance, or other auxiliary information such as texture and markers. The majority of these methods are tailored for high-quality scans from expensive scanners. Following the introduction of the low-quality scans from cost-effective devices such as Kinect, the 3D data capturing of human body becomes more convenient and easier. However, due to the inevitable holes, noises and outliers in the low-quality scan, the registration of human body becomes even more challenging. To address this problem, we propose a fully automatic active registration method which deforms a high-resolution template mesh to match the low-quality human body scans. Our registration method operates on two levels of statistical shape models: (1) the first level is a holistic body shape model that defines the basic figure of human; (2) the second level includes a set of shape models for every body part, aiming at capturing more body details. Our fitting procedure follows a coarse-to-fine approach that is robust and efficient. Experiments show that our method is comparable with the state-of-the-art methods for high-quality meshes in terms of accuracy and it outperforms them in the case of low-quality scans where noises, holes and obscure parts are prevalent.
Zongyi Xu, Qianni Zhang, Shiyang Cheng 0001
Multim. Syst.1
2016 Symmetry-Aware Human Shape Correspondence Using Skeleton
Zongyi Xu, Qianni Zhang
MMM (1)1
2016 Decomposition and matching: Towards efficient automatic Chinese character stroke extraction
abstract
In this paper, given images of Chinese characters, we present an automatic stroke extraction system, which consists of character decomposition, shape matching, cross area extraction and stroke segment combination. First, a character is decomposed into isolated stroke structures according to the connection. Then, we extract shape contexts of stroke structures and find the matched counterparts in a standard database by shape matching. Cross points are computed, and cross areas are extracted according to a proposed adaptive cross area extraction method based on point-to-boundary orientation distance. For a shape structure with the matched structure, we optimize cross point set and combine its stroke segments according to the correct cross point set and combination way of its matched counterpart. For those without matched structures, we propose an angle based stroke segment combination method to combine the segments into a complete stroke. Experimental results indicate that the proposed system achieves high accuracy and demonstrates prominent augmentation on efficiency.
Zongyi Xu, Qianni Zhang, Ebroul Izquierdo
VCIP1
2012 Let Me Take Care of Myself: A Vehicle Self-Gratification System Using Vehicular Sensors and Mobile Phones
abstract
If an automobile can recognize whether a place is refuel station, parking lot or maintenance shop automatically, and it can tell other automobiles where the place is, the automobile will be able to fulfill its basic needs like refueling, parking and maintaining without any manual intervention. People will be freed from asking passengers or searching in the map of Geographic Information System (GIS) about whereto refuel, park or maintain. In this paper, we propose a novel vehicle self-gratification system using vehicular sensors and mobile phones to achieve the aforementioned assumption. Also, through statistics analysis, some new and useful information like popularity of a maintenance shop and probability of an available parking space at a certain time can be generated and guide people to make better choices. Experiments in real driving environment show that our system can accurately recognize the locations of fuel stations, parking lots and maintenance shops, and the information of popularity and probability make people's choices not random but good.
Hai-gang Gong, Kexiong Zeng, Jinchuan Tang, Nianbo Liu, Zongyi Xu, Bang Liu 0001
MSN6