Xiaoke Jiang

dblp:54/10608 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Computer networks · 7 · 4 first-author
YearPublicationVenuePosition
2025 HandOS: 3D Hand Reconstruction in One Stage
abstract
Existing approaches of hand reconstruction predominantly adhere to a multi-stage framework, encompassing detection, left-right classification, and pose estimation. This paradigm induces redundant computation and cumulative errors. In this work, we propose HandOS, an end-to-end framework for 3D hand reconstruction. Our central motivation lies in leveraging a frozen detector as the foundation while incorporating auxiliary modules for 2D and 3D keypoint estimation. In this manner, we integrate the pose estimation capacity into the detection framework, while at the same time obviating the necessity of using the left-right category as a prerequisite. Specifically, we propose an interactive 2D-3D decoder, where 2D joint semantics is derived from detection cues while 3D representation is lifted from those of 2D joints. Furthermore, hierarchical attention is designed to enable the concurrent modeling of 2D joints, 3D vertices, and camera translation. Consequently, we achieve an end-to-end integration of hand detection, 2D pose estimation, and 3D mesh reconstruction within a one-stage framework, so that the above multi-stage drawbacks are overcome. Meanwhile, the HandOS reaches state-of-the-art performances on public benchmarks, e.g., 5.0 PA-MPJPE on FreiHand and 64.6% [email protected] on HInt-Ego4D.
Xingyu Chen 0002, Zhuheng Song, Xiaoke Jiang, Yaoqing Hu, Junzhi Yu 0001, Lei Zhang 0001
CVPR3
2025 LeanGaussian: Breaking Pixel or Point Cloud Correspondence in Modeling 3D Gaussians
abstract
Rencently, Gaussian splatting has demonstrated significant success in novel view synthesis. Current methods often regress Gaussians with pixel or point cloud correspondence, linking each Gaussian with a pixel or a 3D point. This leads to the redundancy of Gaussians being used to overfit the correspondence rather than the objects represented by the 3D Gaussians themselves, consequently wasting resources and lacking accurate geometries or textures. In this paper, we introduce LeanGaussian, a novel approach that treats each query in deformable Transformer as one 3D Gaussian ellipsoid, breaking the pixel or point cloud correspondence constraints. We leverage deformable decoder to iteratively refine the Gaussians layer-by-layer with the image features as keys and values. Notably, the center of each 3D Gaussian is defined as 3D reference points, which are then projected onto the image for deformable attention in 2D space. On both the ShapeNet SRN dataset (category level) and the Google Scanned Objects dataset (open-category level, trained with the Objaverse dataset), our approach, outperforms prior methods by approximately 6.1%, achieving a PSNR of 25.44 and 22.36, respectively. Additionally, our method achieves a 3D reconstruction speed of 7.2 FPS and rendering speed 500 FPS. Codes are available at https://github.com/jwubz123/LeanGaussian.
Kenkun Liu, Xiaoke Jiang, Yuan Yao 0011, Lei Zhang 0001
CVPR4
2025 HumanMM: Global Human Motion Recovery from Multi-shot Videos
abstract
In this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions. Such long-sequence in-the-wild motions are highly valuable to applications such as motion generation and motion understanding, but are of great challenge to be recovered due to abrupt shot transitions, partial occlusions, and dynamic backgrounds presented in such videos. Existing methods primarily focus on single-shot videos, where continuity is maintained within a single camera view, or simplify multi-shot alignment in camera space only. In this work, we tackle the challenges by integrating an enhanced camera pose estimation with Human Motion Recovery (HMR) by incorporating a shot transition detector and a robust alignment module for accurate pose and orientation continuity across shots. By leveraging a custom motion integrator, we effectively mitigate the problem of foot sliding and ensure temporal consistency in human pose. Extensive evaluations on our created multi-shot dataset from public 3D human datasets demonstrate the robustness of our method in reconstructing realistic human motion in world coordinates.
Guanlin Wu, Zhuokai Zhao, Xiaoke Jiang, Zhuoheng Li, Hao (Frank) Yang, Haoqian Wang, Lei Zhang 0001
CVPR6
2025 UniGS: Modeling Unitary 3D Gaussians for Novel View Synthesis from Sparse-View Images
Kenkun Liu, Xiaoke Jiang, Yuan Yao 0011, Lei Zhang 0001
ICCV3
2025 SeqPose: An End-to-End Framework to Unify Single-frame and Video-based RGB Category-Level Pose Estimation
abstract
Category-level object pose estimation is a longstanding and fundamental task crucial for augmented reality and robotic manipulation applications. Existing RGB-based approaches struggle with multi-stage settings and heavily rely on off-the-shelf techniques, such as object detectors, depth estimators, non-differentiable NOCS shape alignment, etc. Extra dependencies lead to the accumulation of errors and complicate the whole pipeline, limiting the deployment of these approaches in practical applications. This paper streamlined an end-to-end framework unifying the single-frame and video-based category-level pose estimation. Specifically, instead of explicitly introducing extra dependencies, the DINOv2 encoder and depth decoder, as robust semantic and geometric prior extractors, are leveraged to produce intra-frame hierarchical semantic and geometric features. A spatial-temporal sparse query network is developed to model the implicit correspondence and inter-frame correlations between a set of implicit 3D query anchors and intra-frame features. Finally, a pose prediction head is employed using the bipartite matching algorithm. Experimental results demonstrate that our model achieves state-of-the-art performance compared with RGB-based categorical pose estimation methods on the REAL275 and CAMERA25 datasets. Our code is available at https://andrewchiyz.github.io/vision.3dv.seqpose/.
Yuzhu Ji, Mingshan Sun, Jianyang Shi, Xiaoke Jiang, Yiqun Zhang 0006, Haijun Zhang 0002
IJCAI4
2025 Geo6D: Geometric-Constraints-Guided Direct Object 6D Pose Estimation Network
abstract
Direct pose estimation networks aim to directly regress the 6D poses of target objects in the scene image using a neural network. These direct methods offer efficiency and an optimal optimization target, presenting significant potential for practical applications. However, due to the complex and implicit mappings between input features and target pose parameters, direct methods are challenging to train and prone to overfitting on mappings seen during training, resulting in limited effectiveness and generalization capability on unseen mappings. Existing methods focus primarily on improvements of the network architecture and training strategies, with less attention given to mappings. In this work, we propose a geometric constraints learning approach, which enables networks to explicitly capture and utilize the geometric mappings between inputs and optimization targets for pose estimation. Specifically, we introduce a residual pose transformation formula that preserves pose transformation constraints within both the 2D image plane and the 3D space while decoupling the absolute pose distribution, thereby addressing the pose distribution gap issue. We further design a Geo6D mechanism based on the formula, which enables the network to explicitly utilize geometric constraints for pose estimation by reconstructing the inputs and outputs. We select two different methods as our baseline and extensive experiments show that Geo6D enhances the performance and reduces the dependence on extensive training data, remaining effective even with only 10% of the typical data volume.
Jianqiu Chen, Mingshan Sun, Tianpeng Bao, Zhenyu He 0001, Donghai Li, Guoqiang Jin, Rui Zhao 0001, Xiaoke Jiang
IEEE Trans. Multim.10
2023 Uni6Dv2: Noise Elimination for 6D Pose Estimation
abstract
Uni6D is the first 6D pose estimation approach to employ a unified backbone network to extract features from both RGB and depth images. We discover that the principal reasons of Uni6D performance limitations are Instance-Outside and Instance-Inside noise. Uni6D’s simple pipeline design inherently introduces Instance-Outside noise from background pixels in the receptive field, while ignoring Instance-Inside noise in the input depth data. In this paper, we propose a two-step denoising approach for dealing with the aforementioned noise in Uni6D. To reduce noise from non-instance regions, an instance segmentation network is utilized in the first step to crop and mask the instance. A lightweight depth denoising module is proposed in the second step to calibrate the depth feature before feeding it into the pose regression network. Extensive experiments show that our Uni6Dv2 reliably and robustly eliminates noise, outperforming Uni6D without sacrificing too much inference efficiency. It also reduces the need for annotated real data that requires costly labeling.
Mingshan Sun, Tianpeng Bao, Jianqiu Chen, Guoqiang Jin, Rui Zhao 0001, Xiaoke Jiang
AISTATS8
2022 Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose Estimation
abstract
As RGB-D sensors become more affordable, using RGB- D images to obtain high-accuracy 6D pose estimation results becomes a better option. State-of-the-art approaches typically use different backbones to extract features for RGB and depth images. They use a 2D CNN for RGB images and a perpixel point cloud network for depth data, as well as a fusion network for feature fusion. We find that the essential reason for using two independent backbones is the “projection breakdown” problem. In the depth image plane, the projected 3D structure of the physical world is preserved by the 1D depth value and its built-in 2D pixel coordinate (UV). Any spatial transformation that modifies UV, such as resize, flip, crop, or pooling operations in the CNN pipeline, breaks the binding between the pixel value and UV coordinate. As a consequence, the 3D structure is no longer preserved by a modified depth image or feature. To address this issue, we propose a simple yet effective method denoted as Uni6D that explicitly takes the extra UV data along with RGB-D images as input. Our method has a Unified CNN framework for 6D pose estimation with a single CNN backbone. In particular, the architecture of our method is based on Mask R-CNN with two extra heads, one named RT head for directly predicting 6D pose and the other named abc head for guiding the network to map the visible points to their coordinates in the 3D model as an auxiliary module. This end-to-end approach balances simplicity and accuracy, achieving comparable accuracy with state of the arts and 7.2x faster inference speed on the YCB-Video dataset.
Xiaoke Jiang, Donghai Li, Hao Chen 0103, Rui Zhao 0001
CVPR1
2021 SSN3D: Self-Separated Network to Align Parts for 3D Convolution in Video Person Re-Identification
abstract
Temporal appearance misalignment is a crucial problem in video person re-identification. The same part of person (e.g. head or hand) appearing on different locations in video sequence weakens its discriminative ability, especially when we apply standard temporal aggregation such as 3D convolution or LSTM. To address this issue, we propose Self-Separated network (SSN) to seek out the same parts in different images. As the name implies, SSN, if trained in an unsupervised strategy, guarantees the selected parts distinct. With a few samples of labeled parts to guide SSN training, this semi-supervised trained SSN seeks out the parts that are human-understandable within a frame and stable across a video snippet. Given the distinct and stable person parts, rather than performing aggregation on features, we then apply 3D convolution across different frames for person re-identification. This SSN + 3D pipeline, dubbed SSN3D, is proved to be efficient through extensive experiments on both synthetic and real data.
Xiaoke Jiang, Qichen Li, Wanrong Zheng, Dapeng Chen
AAAI1
2017 NDNS: A DNS-Like Name Service for NDN
abstract
DNS provides a global-scale distributed lookup service to retrieve data of all types for a given name, be it IP addresses, service records, or cryptographic keys. This service has proven essential in today's operational Internet. Our experience with the design and development of Named Data Networking (NDN) suggests the need for a similar always-on lookup service. To fulfill this need we have designed the NDNS (NDN DNS) protocol, and learned several interesting lessons through the process. Although DNS's request-response operations seem closely resembling NDN's Interest-Data packet exchanges, they operate at different layers in the protocol stack. Comparing DNS's implementations over IP protocol stack with NDNS's implementation over NDN reveals several fundamental differences between applications designs for host-centric IP architecture and data-centric NDN architecture.
Alexander Afanasyev, Xiaoke Jiang, Yingdi Yu, Jiewen Tan, Yumin Xia, Allison Mankin, Lixia Zhang 0001
ICCCN2
2016 A multi-anchoring approach in mobile IP networks
abstract
Abstract Many recent mobility solutions, including derivatives of the well‐known Mobile IP as well as emerging protocols employed by future Internet architectures, propose to realize mobility management by distributing anchoring nodes (Home Agents or other indirection agents) over the Internet. One of their main goals is to address triangle routing by optimizing routes between mobile nodes and correspondent nodes. Thus, a key component of such proposals is the algorithm to select proper mobility anchoring nodes for mobile nodes. However, most current solutions adopt a single‐anchoring approach, which means each mobile node attaches to a sole mobility anchor at one time. In this paper, “we argue that the single‐anchoring approach has drawbacks when facing various mobility scenarios. Then, we offer a novel multi‐anchoring approach that allows each mobile node to select an independent mobility anchor for each correspondent node. We show that in most cases our proposal gains more performance benefits with an acceptable additional cost by evaluation based on real network topologies. For the cases that lead to potential high cost, we also provide a lightweight version of our solution which aims to preserve most performance benefits while keeping a lower cost. At last, we demonstrate how our proposal can be integrated into current Mobile IP networks. Copyright © 2017 John Wiley & Sons, Ltd.
You Wang 0004, Jun Bi, Xiaoke Jiang
Wirel. Commun. Mob. Comput.3
2014 What benefits does NDN have in supporting mobility
abstract
Inspired by forwarding hint and previous IP mobility solutions, we adopt forwarding hint, which intent to solve scalability problem, to support producer mobility. In this paper, we point out how those new elements, such as cache, content-oriented security, content-centric data transmission, benefit mobility. We implement a prototype to analyze the benefits that NDN has in supporting mobility. Our analysis and evaluation conclude that during mapping updating delay following mobility, only popular contents can get benefits from caching while unpopular contents gain little. What's more, only if the content is able to be accessed by partial consumers, caching would amplify the benefits and extend receivers to the rest of consumers, which has significant contribution during routing convergence or mapping updating delay after mobility happens.
Xiaoke Jiang, Jun Bi, You Wang 0004
ISCC1
2013 Interest set mechanism to improve the transport of named data networking
abstract
In this paper, we proposal an Interest Set mechanism which aggregate similar Interest packets from same flow to one packet to improve the efficient of transport of NDN. The trick here is to reset lifetime of corresponding PIT entry in the immediate routers every time when valid Data packet is passed by. This mechanism covers the time and space uncertainty of data generating, reduce the cost of maintaining the pipeline and improve the transport of NDN.
Xiaoke Jiang, Jun Bi
SIGCOMM1
2012 A content provider mobility solution of named data networking
abstract
In this paper, we proposal an content provider mobility solution of Named Data Networking (NDN) [1]. Here content provider means the host of NDN network which provide content originally. We add a Locator are to the Interest packet [2]. Mapping System also introduced into the network, which maps identifier to locator. The original name is used as identifier. Thus, matching lookup in Content Store (CS) and Pending Interest Table (PIT) employs identifier, while forwarding lookup in Forwarding Information Base (FIB) employs locator. In reality, Mapping System should be a DNS-like distributed system, so the record updating has time latency. It's hard to solve provider mobility when there is no explicit "where" information, that's why locator is imported in NDN, however, "where" still serves as the secondary information of the network.
Xiaoke Jiang, Jun Bi, You Wang 0004, Pingping Lin, Zhaogeng Li
ICNP1
2011 IPv6 evolution, stability and deployment
abstract
Our subject focuses on IPv6 network, which develops for more than 10 years. How IPv6 evolve in those years? Is IPv6 network mature enough to undertake the load produced by users? Can we find some principles to guide IPv6 deployment, which make the whole network more robust and efficiency? This paper tries to answer these questions with in-depth statistics. Good news is that network is growing at a speed of O(d2) (d is time) after 2006, moreover, network itself and its routing system become more and more stable. And we explore special properties of this preliminary network, We find that distribution of AS degree follows "Power-Law Distribution", but AS-level topology cannot be described as "Small-World Model" properly. We also propose a method to define the importance of AS and give a simple principle of IPv6 deployment. We even build "6Stats Project"[1] to provide data which help deploy IPv6.
Xiaoke Jiang, Jun Bi, Yangyang Wang 0001, Zhijie He, Wei Zhang 0040, Hongcheng Tian
ICNP1
2011 EasyTrace: An easily-deployable light-weight IP traceback on an AS-level overlay network
abstract
IP traceback can be used to find the origins and paths of attacking traffic. However, so far, no Internet-level IP traceback system has ever been deployed because of deployment difficulties. In this paper, we present an easily-deployable light-weight IP traceback based on flow (EasyTrace). In EasyTrace, it is not necessary to deploy any dedicated traceback software and hardware at routers, and an AS-level overlay network is built for incremental deployment. We theoretically analyze the quantitative relation among the probability that a flow is successfully traced back various AS-level hop number, independently sampling probability, and the number of packets that the flow comprises.
Hongcheng Tian, Jun Bi, Wei Zhang 0040, Xiaoke Jiang
ICNP4