Kejian Wu

dblp:19/8512 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Systems, architecture and hardware · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unveiling hidden vulnerabilities in digital human generation via adversarial attacks
Zhiying Li 0003, Yeying Jin, Michael Shen, Kejian Wu, Zhaoxin Fan
Pattern Recognit.10
2025 Modeling the interplay between regional heterogeneity and critical dynamics underlying brain functional networks
Jijin Zhang, Kejian Wu, Jianfeng Feng, Lianchun Yu
Neural Networks2
2025 $\sqrt{\mathbf {VINS}}$: Robust and Ultrafast Square-Root Filter-Based 3D Motion Tracking
Yuxiang Peng 0002, Chuchu Chen, Kejian Wu, Guoquan Huang 0001
IEEE Trans. Robotics3
2024 FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction Text
abstract
Rearrangement operations form the crux of interactions between humans and their environment. The ability to generate natural, fluid sequences of this operation is of essential value in AR/VR and CG. Bridging a gap in the field, our study introduces FAVOR: a novel dataset for Full-body AR-driven Virtual Object Rearrangement that uniquely employs motion capture systems and AR eyeglasses. Comprising 3k diverse motion rearrangement sequences and 7.17 million interaction data frames, this dataset breaks new ground in research data. We also present a pipeline FAVORITE for producing digital human rearrangement motion sequences guided by instructions. Experimental results, both qualitative and quantitative, suggest that this dataset and pipeline deliver high-quality motion sequences. Our dataset, code, and appendix are available at https://kailinli.github.io/FAVOR.
Kailin Li 0001, Lixin Yang 0001, Zenan Lin, Jian Xu 0027, Xinyu Zhan 0001, Yifei Zhao 0003, Pengxiang Zhu, Wenxiong Kang, Kejian Wu, Cewu Lu
AAAI9
2024 RPBG: Towards Robust Neural Point-Based Graphics in the Wild
Qingtian Zhu, Zizhuang Wei, Zhongtian Zheng, Yifan Zhan, Zhuyu Yao, Jiawang Zhang, Kejian Wu, Yinqiang Zheng
ECCV (15)7
2024 ACR-Pose: Adversarial Canonical Representation Reconstruction Network for Category Level 6D Object Pose Estimation
abstract
In the realm of category-level 6D object pose estimation, canonical 3D representation reconstruction is pivotal, yet current methods show limitations in reconstruction quality, a key step in current pose estimation pipeline. To address this, we introduce an innovative Adversarial Canonical Representation Reconstruction Network (ACR-Pose) in this paper. In particular, ACR-Pose comprises a Reconstructor, with novel sub-modules: a Pose-Irrelevant Module (PIM) for robustness to rotation and translation, and a Relational Reconstruction Module (RRM) for extracting relational information between input modalities. A Discriminator is incorporated to guide the generation of realistic canonical representations through adversarial optimization. Evaluated on the prevalent NOCS-CAMERA and NOCS-REAL datasets, our method significantly improves the performance of baseline models and achieves comparable performance with existing state-of-the-art methods, representing a promising advancement in the field of category-level 6D object pose estimation.
Zhaoxin Fan, Zhenbo Song, Zhicheng Wang 0007, Jian Xu 0027, Kejian Wu, Hongyan Liu 0002, Jun He 0008
ICMR5
2023 POEM: Reconstructing Hand in a Point Embedded Multi-view Stereo
abstract
Enable neural networks to capture 3D geometrical-aware features is essential in multi-view based vision tasks. Previous methods usually encode the 3D information of multi-view stereo into the 2D features. In contrast, we present a novel method, named POEM, that directly operates on the 3D POints Embedded in the Multi-view stereo for reconstructing hand mesh in it. Point is a natural form of 3D information and an ideal medium for fusing features across views, as it has different projections on different views. Our method is thus in light of a simple yet effective idea, that a complex 3D hand mesh can be represented by a set of 3D points that 1) are embedded in the multi-view stereo, 2) carry features from the multi-view images, and 3) encircle the hand. To leverage the power of points, we design two operations: point-based feature fusion and cross-set point attention mechanism. Evaluation on three challenging multi-view datasets shows that POEM outperforms the state-of-the-art in hand mesh reconstruction. Code and models are available for research at github.com/lixiny/POEM
Lixin Yang 0001, Jian Xu 0027, Licheng Zhong, Xinyu Zhan 0001, Zhicheng Wang 0007, Kejian Wu, Cewu Lu
CVPR6
2023 Chord: Category-level Hand-held Object Reconstruction via Shape Deformation
abstract
In daily life, humans utilize hands to manipulate objects. Modeling the shape of objects that are manipulated by the hand is essential for AI to comprehend daily tasks and to learn manipulation skills. However, previous approaches have encountered difficulties in reconstructing the precise shapes of hand-held objects, primarily owing to a deficiency in prior shape knowledge and inadequate data for training. As illustrated, given a particular type of tool, such as a mug, despite its infinite variations in shape and appearance, humans have a limited number of ‘effective’ modes and poses for its manipulation. This can be attributed to the fact that humans have mastered the shape prior of the ‘mug’ category, and can quickly establish the corresponding relations between different mug instances and the prior, such as where the rim and handle are located. In light of this, we propose a new method, Chord, for Category-level Hand-held Object Reconstruction via shape Deformation. Chord deforms a categorical shape prior for reconstructing the intra-class objects. To ensure accurate reconstruction, we empower Chord with three types of awareness: appearance, shape, and interacting pose. In addition, we have constructed a new dataset, Comic, of category-level hand-object interaction. Comic contains a rich array of object instances, materials, hand interactions, and viewing directions. Extensive evaluation shows that Chord outperforms state-of-the-art approaches in both quantitative and qualitative measures. Code, model, and datasets are available at https://kailinli.github.io/CHORD
Kailin Li 0001, Lixin Yang 0001, Haoyu Zhen, Zenan Lin, Xinyu Zhan 0001, Licheng Zhong, Jian Xu 0027, Kejian Wu, Cewu Lu
ICCV8
2023 Reconstruction-Aware Prior Distillation for Semi-supervised Point Cloud Completion
abstract
Real-world sensors often produce incomplete, irregular, and noisy point clouds, making point cloud completion increasingly important. However, most existing completion methods rely on large paired datasets for training, which is labor-intensive. This paper proposes RaPD, a novel semi-supervised point cloud completion method that reduces the need for paired datasets. RaPD utilizes a two-stage training scheme, where a deep semantic prior is learned in stage 1 from unpaired complete and incomplete point clouds, and a semi-supervised prior distillation process is introduced in stage 2 to train a completion network using only a small number of paired samples. Additionally, a self-supervised completion module is introduced to improve performance using unpaired incomplete point clouds. Experiments on multiple datasets show that RaPD outperforms previous methods in both homologous and heterologous scenarios.
Zhaoxin Fan, Yu-Lin He, Zhicheng Wang 0007, Kejian Wu, Hongyan Liu 0002, Jun He 0008
IJCAI4
2022 Object Level Depth Reconstruction for Category Level 6D Object Pose Estimation from Monocular RGB Image
Zhaoxin Fan, Zhenbo Song, Jian Xu 0027, Zhicheng Wang 0007, Kejian Wu, Hongyan Liu 0002, Jun He 0008
ECCV (2)5
2019 RISE-SLAM: A Resource-aware Inverse Schmidt Estimator for SLAM
abstract
In this paper, we present the RISE-SLAM algorithm for performing visual-inertial simultaneous localization and mapping (SLAM), while improving estimation consistency. Specifically, in order to achieve real-time operation, existing approaches often assume previously-estimated states to be perfectly known, which leads to inconsistent estimates. Instead, based on the idea of the Schmidt-Kalman filter, which has processing cost linear in the size of the state vector but quadratic memory requirements, we derive a new consistent approximate method in the information domain, which has linear memory requirements and adjustable (constant to linear) processing cost. In particular, this method, the resource-aware inverse Schmidt estimator (RISE), allows trading estimation accuracy for computational efficiency. Furthermore, and in order to better address the requirements of a SLAM system during an exploration vs. a relocalization phase, we employ different configurations of RISE (in terms of the number and order of states updated) to maximize accuracy while preserving efficiency. Lastly, we evaluate the proposed RISE-SLAM algorithm on publicly-available datasets and demonstrate its superiority, both in terms of accuracy and efficiency, as compared to alternative visual-inertial SLAM systems.
Tong Ke, Kejian Wu, Stergios I. Roumeliotis
IROS2
2017 A comparative analysis of tightly-coupled monocular, binocular, and stereo VINS
abstract
In this paper, a sliding-window two-camera vision-aided inertial navigation system (VINS) is presented in the square-root inverse domain. The performance of the system is assessed for the cases where feature matches across the two-camera images are processed with or without any stereo constraints (i.e., stereo vs. binocular). To support the comparison results, a theoretical analysis on the information gain when transitioning from binocular to stereo is also presented. Additionally, the advantage of using a two-camera (both stereo and binocular) system over a monocular VINS is assessed. Furthermore, the impact on the achieved accuracy of different image-processing frontends and estimator design choices is quantified. Finally, a thorough evaluation of the algorithm's processing requirements, which runs in real-time on a mobile processor, as well as its achieved accuracy as compared to alternative approaches is provided, for various scenes and motion profiles.
Mrinal K. Paul, Kejian Wu, Joel A. Hesch, Esha D. Nerurkar, Stergios I. Roumeliotis
ICRA2
2017 VINS on wheels
abstract
In this paper, we present a vision-aided inertial navigation system (VINS) for localizing wheeled robots. In particular, we prove that VINS has additional unobservable directions, such as the scale, when deployed on a ground vehicle that is constrained to move along straight lines or circular arcs. To address this limitation, we extend VINS to incorporate low-frequency wheel-encoder data, and show that the scale becomes observable. Furthermore, and in order to improve the localization accuracy, we introduce the manifold-(m)VINS that exploits the fact that the vehicle moves on an approximately planar surface. In our experiments, we first show the performance degradation of VINS due to special motions, and then demonstrate that by utilizing the additional sources of information, our system achieves significantly higher positioning accuracy, while operating in real-time on a commercial-grade mobile device.
Kejian Wu, Chao X. Guo, Georgios A. Georgiou, Stergios I. Roumeliotis
ICRA1
2014 C-KLAM: Constrained keyframe-based localization and mapping
abstract
In this paper, we present C-KLAM, a Maximum A Posteriori (MAP) estimator-based keyframe approach for SLAM. Instead of discarding information from non-keyframes for reducing the computational complexity, the proposed C-KLAM presents a novel, elegant, and computationally-efficient technique for incorporating most of this information in a consistent manner, resulting in improved estimation accuracy. To achieve this, C-KLAM projects both proprioceptive and exteroceptive information from the non-keyframes to the keyframes, using marginalization, while maintaining the sparse structure of the associated information matrix, resulting in fast and efficient solutions. The performance of C-KLAM has been tested in experiments, using visual and inertial measurements, to demonstrate that it achieves performance comparable to that of the computationally-intensive batch MAP-based 3D SLAM, that uses all available measurement information.
Esha D. Nerurkar, Kejian Wu, Stergios I. Roumeliotis
ICRA2
2013 Detecting and dealing with hovering maneuvers in vision-aided inertial navigation systems
abstract
In this paper, we study the problem of hovering (i.e., absence of translational motion) detection and compensation in Vision-aided Inertial Navigation Systems (VINS). We examine the system's unobservable directions for two common hovering conditions (with and without rotational motion) and propose a robust motion-classification algorithm, based on both visual and inertial measurements. By leveraging our observability analysis and the proposed motion classifier, we modify existing state-of-the-art filtering algorithms, so as to ensure that the number of the system's unobservable directions is minimized. Finally, we validate experimentally the proposed modified sliding window filter, by demonstrating its robustness on a quadrotor with rapid transitions between hovering and forward motions, within an indoor environment.
Dimitrios G. Kottas, Kejian Wu, Stergios I. Roumeliotis
IROS2
2010 Optimal Resource Allocation for Cognitive Radio Networks with Imperfect Spectrum Sensing
abstract
In this paper, an optimal resource allocation scheme is proposed for multi-user multi-channel cognitive radio networks under imperfect spectrum sensing. The channel dynamic model and the sensing errors are considered together to derive the metric of mean delay for each user-channel combination based on the vacation queueing model. Finally, the optimal resource allocation is determined according to the average system delay by bipartite graph matching. The simulation results indicate that the proposed mean delay metric can represent the transmission performance successfully.
Kejian Wu, Wei Wang 0021, Haiyan Luo, Guanding Yu, Zhaoyang Zhang 0001
VTC Spring1