EDBT 2026 Demo / reviewers in the wild / expert
Jiucheng Xie
dblp:242/8286
· DBLP profile ↗
23ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0003-2336-8521ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Security and privacy · 3 · 3 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HazeRes-DFDet: Haze-Resilient Depth-Frequency Detector for foggy drone images
Guangwei Gao, Jiucheng Xie |
Pattern Recognit. | 5 |
| 2026 | Cellular Aggregation Graph Convolutional Network for Point Cloud Quality AssessmentabstractPoint cloud quality assessment (PCQA) is a challenging task due to the inherently disordered nature of points. Existing point-based methods, such as sparse convolution and PointNet, are limited by local spatial modeling and structural feature extraction. Although 3D graph convolutional networks (GCNs) offer advantages in capturing local structural features through explicit geometric modeling and deformable kernels, their scalability is hindered by the high memory consumption associated with storing neighborhood matrices, particularly for large-scale point clouds. In this paper, to better extract hierarchical structural information and maintain efficiency in computational memory, we propose a novel point-based no-reference PCQA method, namely cellular aggregation network (CANet). The method effectively and efficiently extracts the quality-aware features of large patches in a divide-and-conquer manner. Specifically, a cellular sampling (CS) module is introduced to divide large patches into smaller cells, effectively avoiding the problem of memory explosion. A cellular aggregation (CA) module is proposed to extract intra-cell features and fuse inter-cell features. Moreover, a global aggregation (GA) module is presented to extract global sketch information. Finally, a long-term fusion (LTF) module is introduced to capture long-term dependencies between the features of the CA and GA modules. Experimental results on benchmark datasets demonstrate that the proposed model achieves state-of-the-art performance. Jian Xiong 0005, Lingxia Jiang, Qiang Hu 0003, Jiucheng Xie, Hao Gao 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Effective Gaussian Management for High-Fidelity Scene ReconstructionabstractThis paper proposes an effective Gaussian management framework for high-fidelity scene reconstruction of both appearance and geometry. Unlike recent Gaussian Splatting (GS) pipelines that treat all primitives uniformly during optimization, our framework explicitly manages the attribute activation, representation and pruning of Gaussian. Specifically, our framework first introduces GauSep, a novel densification strategy that selectively activates Gaussian color or normal attributes to alleviate destructive gradient conflicts arising from dual supervision. We further propose GauRep, an adaptive Gaussian representation that dynamically adjusts spherical harmonics (SHs) orders and performs task-decoupled pruning to reduce redundancy at both the individual and global levels. To provide reliable geometric supervision for above mangement process, we additionally introduce CoRe, an regularized surface reconstruction module that distills robust normal fields from an SDF branch to the Gaussian representation through a confidence mechanism. Notably, the proposed Gaussian management is compatible with various reconstruction architectures and can be seamlessly integrated to improve performance while reducing size of the model. Extensive experiments demonstrate that our approach achieves superior or comparable performance in appearance and geometry reconstruction compared with state-of-the-art methods, while using significantly fewer parameters. Jiateng Liu, Hao Gao 0005, Jiucheng Xie, Chi-Man Pun, Jian Xiong 0005, Haolun Li 0001, Junxin Chen 0001, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Multi-scale spatio-temporal cross human motion prediction based on self-attention mechanismabstractHuman motion prediction is crucial for applications ranging from robotics to human-computer interaction. This paper introduces a novel, multi-scale, spatiotemporal cross-attention-based algorithm for human motion prediction, which effectively models long-term dependencies in motion sequences. The proposed method leverages a dual-stream spatio-temporal Transformer framework that decouples temporal and spatial features, allowing each to independently capture dynamic temporal dependencies and spatial correlations. A key innovation is the introduction of a cross-attention mechanism, which ensures consistent information exchange between the temporal and spatial streams. Additionally, a multi-scale cross-attention mechanism is employed to capture relationships across different scales. Extensive experiments on benchmark datasets such as Human3.6M, CMU-MoCap, AMASS, and 3DPW demonstrate that the proposed model outperforms state-of-the-art methods in both short-term and long-term prediction accuracy. Metrics including MPJPE, MAE, PSEnt, and PSKLD validate the model's ability to generate accurate and smooth motion trajectories. Ablation studies further confirm the critical contributions of each component, highlighting the algorithm's robustness and efficiency. This research represents a significant advancement in human motion prediction, offering precise and reliable solutions for understanding and forecasting complex motion patterns. Chuyi Gao, Haidong Hu, Jiucheng Xie |
IJCNN | 3 |
| 2025 | CANet: Cellular Aggregation Network for Point Cloud Quality AssessmentabstractThe concept of visual masking reveals that human visual perception is influenced by content and distortion information. Existing projection-based methods lose depth information and intrinsic topological structures. Due to the limitations of computational memory, the existing point-based methods tend to deal with small patches with little content information. In this paper, we propose a novel point-based no-reference quality assessment method, namely cellular aggregation network (CANet). The method effectively extracts the quality-aware features of large patches in a divide-and-conquer manner. Specifically, the cellular sampling module is used to divide large patches into smaller cells, which effectively avoids the memory explosion problem. The cellular aggregation module is proposed to obtain more content information from small cells. A global aggregation module is proposed to extract global sketch information. Furthermore, a long-term fusion module is introduced to capture long-term dependencies, which can better receive content-aware semantic features. Experimental results on benchmark databases demonstrate that CANet achieves competitive performances. Lingxia Jiang, Jian Xiong 0005, Jiucheng Xie, Hao Gao 0005 |
ISCAS | 4 |
| 2025 | Hierarchical Local Temporal Network for 2D-to-3D Human Pose EstimationabstractRecent advancements in transformer-based methods have yielded substantial success in 2D-to-3D human pose estimation. Transformer-based estimators possess inherent advantages like the global receptive field. Nevertheless, existing transformer approaches ignore the differences among local contexts, resulting in insufficient learning of local information. To address this issue, we introduce nonuniform graph convolution to extract spatial local relationships in skeletons, remedying the limitations of traditional transformers in learning human body topology effectively. Additionally, our proposed hierarchical local temporal network (HLTN) models local temporal associations across three hierarchical levels: 1) joints; 2) body-parts; and 3) poses, effectively addressing the constraint of traditional transformers in learning localized human movements. We connect these two modules in parallel with the spatial and temporal transformer to obtain better features of skeleton sequences. Furthermore, we integrate nonuniform graph convolution with spatial Transformer methods to achieve interaction between local and global features at the attention level. Through these improved methods, our network not only effectively identifies global trends but also exhibits stronger sensitivity to local variations. Compared with the latest methods, our method achieves state-of-the-art performance on multiple datasets (Human3.6M and Mpi-Inf-3DHP). Jiucheng Xie, Haolun Li 0001, Hao Gao 0005 |
IEEE Internet Things J. | 2 |
| 2025 | Lifespan age synthesis on human faces with decorrelation constraints and geometry guidance
Jiucheng Xie, Lingqing Zhang, Hao Gao 0005, Chi-Man Pun |
Pattern Recognit. Lett. | 1 |
| 2025 | Multi-Task Learning Model for V-PCC Geometry Compression Artifact RemovalabstractIn video-based point cloud compression (V-PCC), point clouds are projected as videos using a patch projection method and then compressed using video coding techniques. However, the lossy video compression and the down-sampling of occupancy maps (OMs) can lead to geometry compression artifacts, i.e., depth errors and OM errors, respectively. These errors can significantly affect the reconstruction quality of the point clouds. Existing methods can only eliminate one type of error and therefore have limited quality improvement. In this paper, to improve the quality maximally, a multi-task learning-based geometry compression artifact removal method is proposed to reduce both types of errors simultaneously. Considering the differences between the two tasks, the proposed method deals with the challenges of shared feature extraction and heterogeneous objective optimization. First, we propose a context-aware multi-task learning (CAML) model. The proposed CAML model can extract shared features that are context-aware and satisfy both tasks. Second, an improved optimization scheme is presented to train the proposed model. The improved optimization can fix the gradient imbalance of model updating. Cross-validation experiments show that the proposed method saves an average of over 45% Bjϕntegaard Delta bitrate in terms of the D2 metric. Jian Xiong 0005, Jiucheng Xie, Hui Yuan 0001, Hao Gao 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | GaussianHead: High-Fidelity Head Avatars With Learnable Gaussian DerivationabstractCreating lifelike 3D head avatars and generating compelling animations for diverse subjects remain challenging in computer vision. This paper presents GaussianHead, which models the active head based on anisotropic 3D Gaussians. Our method integrates a motion deformation field and a single-resolution tri-plane to capture the head's intricate dynamics and detailed texture. Notably, we introduce a customized derivation scheme for each 3D Gaussian, facilitating the generation of multiple "doppelgangers" through learnable parameters for precise position transformation. This approach enables efficient representation of diverse Gaussian attributes and ensures their precision. Additionally, we propose an inherited derivation strategy for newly added Gaussians to expedite training. Extensive experiments demonstrate GaussianHead's efficacy, achieving high-fidelity visual results with a remarkably compact model size ($\approx 12$≈12 MB). Our method outperforms state-of-the-art alternatives in tasks such as reconstruction, cross-identity reenactment, and novel view synthesis. Jie Wang 0137, Jiucheng Xie, Xianyan Li, Feng Xu 0005, Chi-Man Pun, Hao Gao 0005 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Geometry Compression Artifact Removal for V-PCC over a Wide Bitrate RangeabstractIn video-based point cloud compression (V-PCC), point clouds are generated as videos via patch projection to be compressed using video coding techniques. However, a large number of filled empty pixels in the videos creates a fake context, which reduces the noise prediction accuracy in compression artifact removal. Moreover, mean square error (MSE)-based trained models perform better on low-bitrates than on high-bitrates due to the unbalanced parameter updates. This paper proposes an learning-based geometry compression artifact removal for V-PCC over a wide range of bitrates. Firstly, an occupancy map-based contextual feature extraction is proposed to eliminate the interference of empty pixels on the neighboring non-empty pixels. Secondly, an incremental Peak Signal to Noise Ratio (PSNR)-based training scheme is presented to balance the error differences. Experimental results show the effectiveness of the proposed method. Jian Xiong 0005, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 4 |
| 2024 | Local Optimization Networks for Multi-View Multi-Person Human Posture EstimationabstractWith the growing applicability of multi-view multi-person 3D human pose estimation across diverse scenarios, the impact of external environmental factors and occlusion on accuracy has garnered substantial attention. In this research, we introduce a novel approach to multi-view multi-person 3D human pose estimation, leveraging a localized optimization strategy. Specifically, our method enhances the interplay of feature information from different channels and fine-tunes the optimal feature weights to capture intricate dependencies among joints. This refinement leads to improved accuracy in handling external environmental factors. Experimental evaluations were conducted on two prominent benchmark datasets, namely Campus and Shelf. The proposed method achieved a remarkable performance, with a Percentage of Correct Parts (PCP) score of 97.4% and 98.2% for the Campus and Shelf datasets, respectively. Jucheng Song, Chi-Man Pun, Haolun Li 0001, Rushi Lan, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 5 |
| 2024 | Scene flow estimation from 3D point clouds based on dual-branch implicit neural representationsabstractAbstract Recently, online optimisation‐based scene flow estimation has attracted significant attention due to its strong domain adaptivity. Although online optimisation‐based methods have made significant advances, the performance is far from satisfactory as only flow priors are considered, neglecting scene priors that are crucial for the representations of dynamic scenes. To address this problem, the authors introduce a dual‐branch MLP‐based architecture to encode implicit scene representations from a source 3D point cloud, which can additionally synthesise a target 3D point cloud. Thus, the mapping function between the source and synthesised target 3D point clouds is established as an extra implicit regulariser to capture scene priors. Moreover, their model infers both flow and scene priors in a stronger bidirectional manner. It can effectively establish spatiotemporal constraints among the synthesised, source, and target 3D point clouds. Experiments on four challenging datasets, including KITTI scene flow, FlyingThings3D, Argoverse, and nuScenes, show that our method can achieve potential and comparable results, proving its effectiveness and generality. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
IET Comput. Vis. | 3 |
| 2024 | Geometry-guided generalizable NeRF for human rendering
Jiucheng Xie, Yiqin Yao, Xun Lv, Shuliang Zhu, Yijing Guo, Hao Gao 0005 |
Multim. Tools Appl. | 1 |
| 2023 | Boosting Face Recognition Performance with Synthetic Data and Limited Real DataabstractFace recognition is one of the most precise and straightforward methods to establish individual identity, and is important in our daily life. To solve the issues of privacy, bias, and collection difficulty caused by face recognition relying heavily on collecting a huge number of real face images from the Internet, a seemingly promising idea is to employ GAN-generated synthetic faces as the training data. However, there are obvious surface gaps and domain gaps between real and synthetic face images, and cannot be replaced directly. In this paper, we attempt to boost face recognition simultaneously using synthetic data and limited real data. Specifically, we first design an augmented space for auto augmentation methods to augment synthetic images to alleviate the surface gap, then propose to disentangle the underlying style distributions through dual batch normalization layers so that both synthetic and real images can be learned jointly by convolution layers without mixing across domains. Extensive experiments demonstrate our method can achieve better results than training with large quantities of real data. Wenqing Wang 0002, Lingqing Zhang, Chi-Man Pun, Jiucheng Xie |
ICASSP | 4 |
| 2023 | Spike-Based Optical Flow Estimation Via Contrastive LearningabstractSpiking cameras have shown promising advantages for optical flow estimation in high-speed scenarios. The recent work SCFlow [1] attempts to train an optical flow model using spike frames based on a multi-scale flow reconstruction loss. However, only using the flow reconstruction loss is unable to effectively deal with the details of motion, which may lead to noise and blur in the estimated flow fields. To address this issue, we introduce a contrastive loss into spike-based optical flow estimation, which exploits both the information of positive samples and negative samples. Moreover, we propose a refinement step with flexible reception fields to effectively refine the initial flow fields. Experiments on the spiking optical flow dataset PHM demonstrate that the proposed network is effective for spike-based optical flow estimation. In addition, our method achieves competitive performance compared to recent spike-based, frame-based, and event-based methods. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 3 |
| 2023 | Cross-Modal Optical Flow Estimation via Modality Compensation and AlignmentabstractCross-modal optical flow estimation aims to predict motion fields between two frames collected from different modalities, recently attracting intensive attention. However, a substantial yet challenging problem is how to match images across a large modal discrepancy. In this paper, we propose a modality compensation module (MCM) to extract complementary features from different modalities adaptively. Moreover, a cross-modal feature alignment loss is introduced into our network, pulling the compensative features of two cross-modal frames closer and effectively reducing the modal discrepancy. The experimental results demonstrate that our method can achieve competitive performance on the cross-modal optical flow dataset CrossKITTI. Moreover, we experimentally verify that the proposed MCM and cross-modal feature alignment loss are effective for cross-modal optical flow estimation. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 3 |
| 2023 | Learning Scene Flow from 3d Point Clouds with Cross-Transformer and Global Motion CuesabstractScene flow estimation is critical for real-world vision problems such as autonomous driving and augmented reality. Due to the popularity of 3D LiDAR sensors, scene flow estimation from 3D point clouds arouses increasing attention. Existing methods usually use a flow embedding-based layer to find correspondences between point pairs. However, only using a flow embedding-based layer is not enough to model the global mutual relationship between two features due to local matching. In this paper, we introduce a cross-transformer to capture more reliable dependencies for point pairs. Moreover, a global motion-aware module is adopted to learn large displacements with a non-local approach. The experimental results demonstrate that the proposed method achieves comparable performance on public datasets and confirm the effectiveness of exploiting the cross-transformer and global motion cues for scene flow estimation. Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005 |
ICASSP | 3 |
| 2023 | Scene Flow Estimation from Point Clouds with Contrastive Loss and Dual Pseudo LabelsabstractScene flow estimation aims to extract the 3D motion vector between each surface point in two consecutive point clouds. Pseudo-label-based approaches usually exploit point-to-point relations and 3D geometry information to generate the pseudo label for self-supervised learning. However, unreasonable results are still obtained due to the unexploited information of negative samples. Moreover, previous approaches are limited by the fact that pseudo labels are only generated along the forward direction, ignoring the backward direction that has strong spatiotemporal correlations with the forward direction. In this paper, we address these issues in a simple yet effective manner. Specifically, we introduce a contrastive loss to exploit both the information of positive samples and negative samples. Furthermore, we design a dual pseudo labels generation strategy to provide a bidirectional self-supervision for scene flow estimation. Experiments on FlyingThings3D and KITTI datasets show that our method can achieve competitive performance compared to recent self-supervised methods. Mingliang Zhai, Kang Ni, Jiucheng Xie, Xuezhi Xiang, Hao Gao 0005 |
ICIP | 3 |
| 2022 | Implicit and Explicit Feature Purification for Age-Invariant Facial Representation LearningabstractThis paper presents a new method, named implicit and explicit feature purification (IEFP), for age-invariant face recognition. Facial features extracted from a face image contain the information about the identity, age, and other attributes. For age-invariant face recognition, it is important to remove the irrelevant information, and retain the identity information only, in the facial features. Through the two proposed feature purification mechanisms, our framework can produce facial-feature embeddings that preserve identity information as much as possible and are insensitive to age variations. Specifically, on the one hand, a special network module is devised to implicitly purify the original facial features obtained from a face encoder. On the other hand, to obtain purer facial feature representations for age-invariant face recognition, irrelevant information within the implicitly purified features, such as the age, is further removed. This is realized by using a regularizer, based on information theory, to explicitly minimize the correlation between identity-related features and age-related features. Comprehensive ablation studies show that these two feature purification schemes can work independently, as well as collaboratively, to achieve better performance. Extensive evaluations on several benchmark data sets show that the IEFP method is on par with those competitors learned on far more favorable training samples, and it achieves the best performance in a fair comparison. Furthermore, we provide mathematical interpretation to explain the effectiveness of our approach, and find that it tends to generate low-rank, yet high-dimensional, representations for age-invariant face recognition. Jiucheng Xie, Chi-Man Pun, Kin-Man Lam 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Action Recognition Framework in Traffic Scene for Autonomous Driving SystemabstractFor the autonomous driving system, accurately recognizing the actions of different roles in the traffic scene is the prerequisite for realizing this kind of human-vehicle information interaction. In this paper, we propose a complete framework based on 3D human pose estimation to recognize the actions of different roles on the road. The main objects recognized include traffic police, cyclists, and some passersby in need. We perform action recognition based on a dynamic adaptive graph convolutional network, which can realize the action recognition of objects based on 3D human pose. In addition to the action recognition module, we have optimized both the object detection module and the human pose estimation module in the framework so that the framework can handle multiple objects at the same time, which can be closer to the real traffic scene. To realize complex and changeable human action recognition, we built a multi-view camera system to collect responsible 3D human pose datasets containing traffic police gestures, cyclist gestures, and pedestrians’ body movements. In the experiments, compared to other state-of-the-art researches, the proposed framework can achieve comparable results with the same dataset. Satisfactory performance has also been obtained on the real data we collected, which can handle a variety of different action recognition tasks at the same time. Feiyi Xu, Feng Xu 0005, Jiucheng Xie, Chi-Man Pun, Huimin Lu 0001, Hao Gao 0005 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Deep and Ordinal Ensemble Learning for Human Age Estimation From Facial ImagesabstractSome recent work treats age estimation as an ordinal ranking task and decomposes it into multiple binary classifications. However, a theoretical defect lies in this type of methods: the ignorance of possible contradictions in individual ranking results. In this paper, we partially embrace the decomposition idea and propose the Deep and Ordinal Ensemble Learning with Two Groups Classification (DOEL2groups) for age prediction. An important advantage of our approach is that it theoretically allows the prediction even when the contradictory cases occur. The proposed method is characterized by a deep and ordinal ensemble and a two-stage aggregation strategy. Specifically, we first set up the ensemble based on Convolutional Neural Network (CNN) techniques, while the ordinal relationship is implicitly constructed among its base learners. Each base learner will classify the target face into one of two specific age groups. After achieving probability predictions of different age groups, then we make aggregation by transforming them into counting value distributions of whole age classes and getting the final age estimation from their votes. Moreover, to further improve the estimation performance, we suggest to regard the age class at the boundary of original two age groups as another age group and this modified version is named the Deep and Ordinal Ensemble Learning with Three Groups Classification (DOEL3groups). Effectiveness of this new grouping scheme is validated in theory and practice. Finally, we evaluate the proposed two ensemble methods on controlled and wild aging databases, and both of them produce competitive results. Note that the DOEL3groupsshows the state-of-the-art performance in most cases. Jiucheng Xie, Chi-Man Pun |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2019 | Automatic Medical Image Registration Based on an Integrated Method Combining Feature and Area Information
Jiucheng Xie, Chi-Man Pun, Zhaoqing Pan, Hao Gao 0005, Baoyun Wang |
Neural Process. Lett. | 1 |
| 2019 | Chronological Age Estimation Under the Guidance of Age-Related Facial AttributesabstractAlthough the researches of facial attributes' analysis have been launched for decades, the estimation of chronological age attribute remains a big challenge. Previous researchers have found that some facial attributes (e.g., gender and race attributes) have close connections with the age attribute and make age estimation under a specific condition decided by various combinations of those age-related attributes which should be more reasonable. In this paper, we propose a generic framework based on a convolutional neural network, which can consider different conditions for age estimation and jointly output age and age-related facial attributes in the end. Compared with conventional methods, it is more efficient and universal. Besides, we view age estimation as a special multi-class ordinal classification problem and use a losses combination function to optimize the predicted probability distribution of individual age classes. These operations further improve the performance of age estimation. Finally, the proposed method achieves state-of-the-art results on both controlled and wild face datasets. Jiucheng Xie, Chi-Man Pun |
IEEE Trans. Inf. Forensics Secur. | 1 |