Deying Kong

dblp:176/1065 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 GATTCR: A Graph Attention Network With Multi-Feature Fusion for Peripheral Blood TCR Repertoire Classification
abstract
T cell receptor (TCR) repertoire profiling provides a promising avenue for noninvasive disease diagnostics by capturing immune signatures directly from peripheral blood. However, the high diversity and sparsity of TCR sequences pose significant challenges for robust immune state classification. In this work, we propose GATTCR, a novel framework that integrates Graph Attention Networks (GATs) with multi-feature fusion to model complex dependencies within TCR repertoires. By representing TCRs as graph nodes and incorporating biological priors-such as sequence embeddings, structural similarity, V gene usage, and clonal frequency-GATTCR enables context-aware, structure-informed representation learning. We evaluate GATTCR across a comprehensive panel of cancer and infectious disease datasets, demonstrating consistent improvements over existing methods. Notably, GATTCR achieves AUROC gains of up to +47.9% under few-shot learning scenarios, highlighting its ability to generalize from limited labeled data. Ablation studies further confirm the critical role of graph-based modeling and immunological features in driving performance gains. Overall, GATTCR offers a scalable approach for TCR repertoire analysis and paves the way for routine, noninvasive immune monitoring in precision medicine.
Hengwei Ju, Deying Kong, Yuhao Tao, Fei Wang 0017
IEEE Trans. Comput. Biol. Bioinform.2
2025 UMAMI: Unifying Masked Autoregressive Models and Deterministic Rendering for View Synthesis
abstract
Novel view synthesis (NVS) seeks to render photorealistic, 3D‑consistent images of a scene from unseen camera poses given only a sparse set of posed views. Existing deterministic networks render observed regions quickly but blur unobserved areas, whereas stochastic diffusion‑based methods hallucinate plausible content yet incur heavy training‑ and inference‑time costs. In this paper, we propose a hybrid framework that unifies the strengths of both paradigms. A bidirectional transformer encodes multi‑view image tokens and Plücker‑ray embeddings, producing a shared latent representation. Two lightweight heads then act on this representation: (i) a feed‑forward regression head that renders pixels where geometry is well constrained, and (ii) a masked autoregressive diffusion head that completes occluded or unseen regions. The entire model is trained end‑to‑end with joint photometric and diffusion losses, without handcrafted 3D inductive biases, enabling scalability across diverse scenes. Experiments demonstrate that our method attains state‑of‑the‑art image quality while reducing rendering time by an order of magnitude compared with fully generative baselines.
Thanh-Tung Le, Deying Kong, Xiaohui Xie, Stephan Mandt
NeurIPS4
2024 Hybrid-CSR: Coupling Explicit and Implicit Reconstruction of Cortical Surface
Shanlin Sun, Pooya Khosravi, Chenyu You, Deying Kong, Xiangyi Yan, Xiaohui Xie
BMVC7
2024 Handformer2T: A Lightweight Regression-based Model for Interacting Hands Pose Estimation from A Single RGB Image
abstract
Despite its extensive range of potential applications in virtual reality and augmented reality, 3D interacting hand pose estimation from RGB image remains a very challenging problem, due to appearance confusions between keypoints of the two hands, and severe hand-hand occlusion. Due to their ability to capture long range relationships between keypoints, transformer-based methods have gained popularity in the research community. However, the existing methods usually deploy tokens at keypoint level, which inevitably results in high computational and memory complexity. In this paper, we propose a simple yet novel mechanism, i.e., hand-level tokenization, in our transformer based model, where we deploy only one token for each hand. With this novel design, we also propose a pose query enhancer module, which can refine the pose prediction iteratively, by focusing on features guided by previous coarse pose predictions. As a result, our proposed model, Handformer2T, can achieve high performance while maintaining lightweight. Extensive experiments on public benchmarks demonstrate that our model can achieve state-of-the-art performance on interacting-hand pose estimation with higher throughput, less memory and faster speed.
Deying Kong
WACV2
2024 Medical image registration via neural fields
abstract
Image registration is an essential step in many medical image analysis tasks. Traditional methods for image registration are primarily optimization-driven, finding the optimal deformations that maximize the similarity between two images. Recent learning-based methods, trained to directly predict transformations between two images, run much faster, but suffer from performance deficiencies due to domain shift. Here we present a new neural network based image registration framework, called NIR (Neural Image Registration), which is based on optimization but utilizes deep neural networks to model deformations between image pairs. NIR represents the transformation between two images with a continuous function implemented via neural fields, receiving a 3D coordinate as input and outputting the corresponding deformation vector. NIR provides two ways of generating deformation field: directly output a displacement vector field for general deformable registration, or output a velocity vector field and integrate the velocity field to derive the deformation field for diffeomorphic image registration. The optimal registration is discovered by updating the parameters of the neural field via stochastic mini-batch gradient descent. We describe several design choices that facilitate model optimization, including coordinate encoding, sinusoidal activation, coordinate sampling, and intensity sampling. NIR is evaluated on two 3D MR brain scan datasets, demonstrating highly competitive performance in terms of both registration accuracy and regularity. Compared to traditional optimization-based methods, our approach achieves better results in shorter computation times. In addition, our methods exhibit performance on a cross-dataset registration task, compared to the pre-trained learning-based methods.
Shanlin Sun, Chenyu You, Hao Tang 0010, Deying Kong, Junayed Naushad, Xiangyi Yan, Pooya Khosravi, James S. Duncan, Xiaohui Xie
Medical Image Anal.5
2023 Diffeomorphic Image Registration with Neural Velocity Field
abstract
Diffeomorphic image registration, offering smooth transformation and topology preservation, is required in many medical image analysis tasks. Traditional methods impose certain modeling constraints on the space of admissible transformations and use optimization to find the optimal transformation between two images. Specifying the right space of admissible transformations is challenging: the registration quality can be poor if the space is too restrictive, while the optimization can be hard to solve if the space is too general. Recent learning-based methods, utilizing deep neural networks to learn the transformation directly, achieve fast inference, but face challenges in accuracy due to the difficulties in capturing the small local deformations and generalization ability. Here we propose a new optimization-based method named DNVF (Diffeomorphic Image Registration with Neural Velocity Field) which utilizes deep neural network to model the space of admissible transformations. A multilayer perceptron (MLP) with sinusoidal activation function is used to represent the continuous velocity field and assigns a velocity vector to every point in space, providing the flexibility of modeling complex deformations as well as the convenience of optimization. Moreover, we propose a cascaded image registration framework (Cas-DNVF) by combining the benefits of both optimization and learning based methods, where a fully convolutional neural network (FCN) is trained to predict the initial deformation, followed by DNVF for further refinement. Experiments on two large-scale 3D MR brain scan datasets demonstrate that our proposed methods significantly outperform the state-of-the-art registration methods.
Shanlin Sun, Xiangyi Yan, Chenyu You, Hao Tang 0010, Junayed Naushad, Deying Kong, Xiaohui Xie
WACV8
2023 Representation Recovering for Self-Supervised Pre-training on Medical Images
abstract
Advances in self-supervised learning have drawn attention to developing techniques to extract effective visual representations from unlabeled images. Contrastive learning (CL) trains a model to extract consistent features by generating different views. Recent success of Masked Autoencoders (MAE) highlights the benefit of generative modeling in self-supervised learning. The generative approaches encode the input into a compact embedding and empower the model’s ability of recovering the original input. However, in our experiments, we found vanilla MAE mainly recovers coarse high level semantic information and is inadequate in recovering detailed low level information. We show that in dense downstream prediction tasks like multi-organ segmentation, directly applying MAE is not ideal. Here, we propose RepRec, a hybrid visual representation learning framework for self-supervised pre-training on large-scale unlabelled medical datasets, which takes advantage of both contrastive and generative modeling. To solve the aforementioned dilemma that MAE encounters, a convolutional encoder is pre-trained to provide low-level feature information, in a contrastive way; and a transformer encoder is pre-trained to produce high level semantic dependency, in a generative way – by recovering masked representations from the convolutional encoder. Extensive experiments on three multi-organ segmentation datasets demonstrate that our method outperforms current state-of-the-art methods.
Xiangyi Yan, Junayed Naushad, Shanlin Sun, Hao Tang 0010, Deying Kong, Chenyu You, Xiaohui Xie
WACV6
2022 Topology-Preserving Shape Reconstruction and Registration via Neural Diffeomorphic Flow
abstract
Deep Implicit Functions (DIFs) represent 3D geometry with continuous signed distance functions learned through deep neural nets. Recently DIFs-based methods have been proposed to handle shape reconstruction and dense point correspondences simultaneously, capturing semantic relationships across shapes of the same class by learning a DIFs-modeled shape template. These methods provide great flexibility and accuracy in reconstructing 3D shapes and inferring correspondences. However, the point correspondences built from these methods do not intrinsically preserve the topology of the shapes, unlike mesh-based template matching methods. This limits their applications on 3D geometries where underlying topological structures exist and matter, such as anatomical structures in medical images. In this paper, we propose a new model called Neural Diffeomorphic Flow (NDF) to learn deep implicit shape templates, representing shapes as conditional diffeomorphic deformations of templates, intrinsically preserving shape topologies. The diffeomorphic deformation is realized by an autodecoder consisting of Neural Ordinary Differential Equation (NODE) blocks that progressively map shapes to implicit templates. We conduct extensive experiments on several medical image organ segmentation datasets to evaluate the effectiveness of NDF on reconstructing and aligning shapes. NDF achieves consistently state-of-the-art organ shape reconstruction and registration results in both accuracy and quality. The source code is publicly available at https://github.com/Siwensun/Neural_Diffeomorphic_Flow-NDF.
Shanlin Sun, Deying Kong, Hao Tang 0010, Xiangyi Yan, Xiaohui Xie
CVPR3
2022 Identity-Aware Hand Mesh Estimation and Personalization from RGB Images
Deying Kong, Linguang Zhang, Liangjian Chen, Xiangyi Yan, Shanlin Sun, Xingwei Liu, Xiaohui Xie
ECCV (5)1
2022 PPT: Token-Pruned Pose Transformer for Monocular and Multi-view Human Pose Estimation
Yifei Chen 0021, Deying Kong, Liangjian Chen, Xingwei Liu, Xiangyi Yan, Hao Tang 0010, Xiaohui Xie
ECCV (5)4
2022 AFTer-UNet: Axial Fusion Transformer UNet for Medical Image Segmentation
abstract
Recent advances in transformer-based models have drawn attention to exploring these techniques in medical image segmentation, especially in conjunction with the UNet model (or its variants), which has shown great success in medical image segmentation, under both 2D and 3D settings. Current 2D based methods either directly replace convolutional layers with pure transformers or consider a transformer as an additional intermediate encoder between the encoder and decoder of U-Net. However, these approaches only consider the attention encoding within one single slice and do not utilize the axial-axis information naturally provided by a 3D volume. In the 3D setting, convolution on volumetric data and transformers both consume large GPU memory. One has to either downsample the image or use cropped local patches to reduce GPU memory usage, which limits its performance. In this paper, we propose Axial Fusion Transformer UNet (AFTer-UNet), which takes both advantages of convolutional layers’ capability of extracting detailed features and transformers’ strength on long sequence modeling. It considers both intra-slice and inter-slice long-range cues to guide the segmentation. Meanwhile, it has fewer parameters and takes less GPU memory to train than the previous transformer-based models. Extensive experiments on three multi-organ segmentation datasets demonstrate that our method outperforms current state-of-the-art methods.
Xiangyi Yan, Hao Tang 0010, Shanlin Sun, Deying Kong, Xiaohui Xie
WACV5
2021 TransFusion: Cross-view Fusion with Transformer for 3D Human Pose Estimation
Liangjian Chen, Deying Kong, Xingwei Liu, Hao Tang 0010, Xiangyi Yan, Yusheng Xie, Shih-Yao Lin 0001, Xiaohui Xie
BMVC3
2020 SIA-GCN: A Spatial Information Aware Graph Neural Network with 2D Convolutions for Hand Pose Estimation
Deying Kong, Xiaohui Xie
BMVC1
2020 Nonparametric Structure Regularization Machine for 2D Hand Pose Estimation
abstract
Hand pose estimation is more challenging than body pose estimation due to severe articulation, self-occlusion and high dexterity of the hand. Current approaches often rely on a popular body pose algorithm, such as the Convolutional Pose Machine (CPM), to learn 2D keypoint features. These algorithms cannot adequately address the unique challenges of hand pose estimation, because they are trained solely based on keypoint positions without seeking to explicitly model structural relationship between them. We propose a novel Nonparametric Structure Regularization Machine (NSRM) for 2D hand pose estimation, adopting a cascade multi-task architecture to learn hand structure and keypoint representations jointly. The structure learning is guided by synthetic hand mask representations, which are directly computed from keypoint positions, and is further strengthened by a novel probabilistic representation of hand limbs and an anatomically inspired composition strategy of mask synthesis. We conduct extensive studies on two public datasets - OneHand 10k and CMU Panoptic Hand. Experimental results demonstrate that explicitly enforcing structure learning consistently improves pose estimation accuracy of CPM baseline models, by 1.17% on the first dataset and 4.01% on the second one. The implementation and experiment code is freely available online1. Our proposal of incorporating structural learning to hand pose estimation requires no additional training information, and can be a generic add-on module to other pose estimation models.
Yifei Chen 0021, Deying Kong, Xiangyi Yan, Jianbao Wu, Xiaohui Xie
WACV3
2020 Rotation-invariant Mixed Graphical Model Network for 2D Hand Pose Estimation
abstract
In this paper, we propose a new architecture named Rotation-invariant Mixed Graphical Model Network (R-MGMN) to solve the problem of 2D hand pose estimation from a monocular RGB image. By integrating a rotation net, the R-MGMN is invariant to rotations of the hand in the image. It also has a pool of graphical models, from which a combination of graphical models could be selected, conditioning on the input image. Belief propagation is performed on each graphical model separately, generating a set of marginal distributions, which are taken as the confidence maps of hand keypoint positions. Final confidence maps are obtained by aggregating these confidence maps together. We evaluate the R-MGMN on two public hand pose datasets. Experiment results show our model outperforms the state-of-the-art algorithm which is widely used in 2D hand pose estimation by a noticeable margin.
Deying Kong, Yifei Chen 0021, Xiaohui Xie
WACV1
2019 Adaptive Graphical Model Network for 2D Handpose Estimation
Deying Kong, Yifei Chen 0021, Xiangyi Yan, Xiaohui Xie
BMVC1
2016 Channel Estimation Under Staggered Frame Structure for Massive MIMO System
abstract
In this paper, a staggered frame structure is proposed for single-cell massive multiple-input multiple-output (MIMO) systems, and the key idea is that different users transmit training pilots at nonoverlapped time. As a result, users do not have to be synchronized strictly to send pilots and orthogonal pilots are not required. Moreover, we also propose two interference suppressed channel estimation methods, i.e., the linear minimum mean square error (LMMSE)-based and orthogonal projection based least squares (OPLS) methods for the staggered frame structure. Specifically, the LMMSE-based method minimizes the mean square error of the estimation, whereas the OPLS method only estimates the part of the user's channel response that is orthogonal to the other users. All conducted simulation results demonstrate that the massive MIMO system with the both proposed methods can achieve high average achievable rate. Furthermore, when the conjugate beamforming is employed, the massive MIMO system could obtain even higher average achievable data rate with the proposed OPLS method than that of the system with perfect channel state information. Moreover, the computational complexity of the proposed OPLS method is very low.
Deying Kong, Daiming Qu, Tao Jiang 0002
IEEE Trans. Wirel. Commun.1