VLDB 2026 Research / reviewers in the wild / expert
Kejie Li
dblp:44/3202
· DBLP profile ↗
38ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 13 since 2021Systems, architecture and hardware · 15Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Agentic Very Long Video UnderstandingabstractAniket Rege, Arka Sadhu, Yuliang Li, Kejie Li, Ramya Korlakai Vinayak, Yuning Chai, Yong Jae Lee, Hyo Jin Kim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Aniket Rege, Arka Sadhu, Kejie Li, Ramya Korlakai Vinayak, Yuning Chai, Yong Jae Lee |
ACL (1) | 4 |
| 2026 | RapidMV: Leveraging Spatio-Angular Latent Space for Efficient and Consistent Text-to-Multi-View SynthesisabstractGenerating consistent multi-view images given a text prompt is an essential bridge to generating synthetic 3D assets. In this work, we introduce RapidMV, a novel text-to-multi-view generative model that can produce 32 multi-view synthetic images in just around 5 seconds. In essence, we introduce a novel spatio-angular latent space, where we encode not only the spatial appearance of a single frame, but also the angular viewpoint deviations across multiple frames into a single latent for improved efficiency and multi-view consistency. We achieve effective training of RapidMV by strategically decomposing our training process into multiple steps. We demonstrate that RapidMV outperforms existing methods in terms of consistency and latency, with competitive quality and text-image alignment. Seungwook Kim 0005, Yichun Shi, Kejie Li, Minsu Cho, Peng Wang 0001 |
WACV | 3 |
| 2025 | CamFreeDiff: Camera-free Image to Panorama Generation with Diffusion ModelabstractThis paper introduces Camera-free Diffusion (CamFreeDiff) model for 360° image outpainting from a single camera-free image and text description. This method distinguishes itself from existing strategies, such as MVDiffusion, by eliminating the requirement for predefined camera poses. CamFreeDiff seamlessly incorporates a mechanism for predicting homography within the multi-view diffusion framework. The key component of our approach is to formulate camera estimation by directly predicting the homography transformation from the input view to the predefined canonical view. In contrast to the direct two-stage approach of image transformation and outpainting, CamFreeDiff utilizes predicted homography to establish point-level correspondences between the input view and the target panoramic view. This enables consistency through correspondence-aware attention, which is learned in a fully differentiable manner. Qualitative and quantitative experimental results demonstrate the strong robustness and performance of CamFreeDiff for 360° image outpainting in the challenging context of camera-free inputs. Xiaoding Yuan, Shitao Tang, Kejie Li |
CVPR | 3 |
| 2025 | DEHand: Deformable Encoding for Photo-Realistic Free-View and Free-Pose Hand RenderingabstractInput encoding has proven crucial in the success of methods based on neural radiance field. Compared to the literature on general static scene modeling, input encoding for dynamic hand modeling has been less explored. However, this aspect is critical to the modeling of deformation and rendering, as it maps a sampled point in space to the representation containing all the information associated with dynamic hand for inferring the geometry and appearance property of this point. The design of input encoding determines how well the neural network can learn for photo-realistic hand rendering. We offer an in-depth examination of this key component and introduceDEHand, a new representation utilizingDeformableEncoding for photo-realistic free-view and free-poseHandrendering. DEHand leverages deformable encoding with a latent code map to achieve high-quality, pose-controlled rendering. Deformable encoding is achieved by adapting static input encoding techniques for the view synthesis of dynamic hands, using parametric hand mesh model as a proxy to construct encodings that map sampled points into a space capable of integrating over different poses and providing rich information for hand modeling. Our findings demonstrate that with our deformable encoding, a single Multilayer Perceptron (MLP) can achieve high-quality dynamic hand rendering, learning solely from images. Extensive experiments on InterHand2.6M validate the superior rendering quality of our method and the effectiveness of each component in our design. Yunzhi Teng, Xiaoke Huang 0001, Kejie Li, Xiao-Ping Zhang 0002, Yansong Tang |
IEEE Trans. Multim. | 3 |
| 2024 | Consistent-1-to-3: Consistent Image to 3D View Synthesis via Geometry-aware Diffusion ModelsabstractZero-shot novel view synthesis (NVS) from a single image is an essential problem in 3D object understanding. While recent approaches that leverage pre-trained generative models can synthesize high-quality novel views from in-the-wild inputs, they still struggle to maintain 3D consistency across different views. In this paper, we present Consistent-1-to-3, which is a generative framework that significantly mitigates this issue. Specifically, we decompose the NVS task into two stages: (i) transforming observed regions to a novel view, and (ii) hallucinating unseen regions. We design a scene representation transformer and view-conditioned diffusion model for performing these two stages respectively. Inside the models, to enforce 3D consistency, we propose to employ epipolar-guided attention to incorporate geometry constraints, and multi-view attention to better aggregate multi-view information. Finally, we design a hierarchy generation paradigm to generate long sequences of consistent views, allowing a full 360° observation of the provided object image. Qualitative and quantitative evaluation over multiple datasets demonstrates the effectiveness of the proposed mechanisms against state-of-the-art approaches. Our project page is at https://jianglongye.com/consistent123/. Jianglong Ye, Kejie Li, Yichun Shi |
3DV | 3 |
| 2024 | Neural Refinement for Absolute Pose Regression with Feature SynthesisabstractAbsolute Pose Regression (APR) methods use deep neural networks to directly regress camera poses from RGB images. However, the predominant APR architectures only rely on 2D operations during inference, resulting in limited accuracy of pose estimation due to the lack of 3D geometry constraints or priors. In this work, we propose a test-time refinement pipeline that leverages implicit geometric constraints using a robust feature field to enhance the ability of APR methods to use 3D information during inference. We also introduce a novel Neural Feature Synthesizer (NeFeS) model, which encodes 3D geometric features during training and directly renders dense novel view features at test time to refine APR methods. To enhance the robustness of our model, we introduce a feature fusion module and a progressive training strategy. Our proposed method achieves state-of-the-art single-image APR accuracy on indoor and outdoor datasets. Code will be released at https://github.com/ActiveVisionLab/NeFeS. Yash Bhalgat, Xinghui Li, Jiawang Bian, Kejie Li, Victor Adrian Prisacariu |
CVPR | 5 |
| 2024 | Enhancing 3D Fidelity of Text-to-3D using Cross-View CorrespondencesabstractLeveraging multi-view diffusion models as priors for 3D optimization have alleviated the problem of 3D consistency, e.g., the Janus face problem or the content drift problem, in zero-shot text-to-3D models. However, the 3D geomet-ric fidelity of the output remains an unresolved issue; albeit the rendered 2D views are realistic, the underlying geom-etry may contain errors such as unreasonable concavities. In this work, we propose CorrespondentDream, an effective method to leverage annotation-free, cross-view corre- spondences yielded from the diffusion U-Net to provide additional 3D prior to the NeRF optimization process. We find that these correspondences are strongly consistent with hu-man perception, and by adopting it in our loss design, we are able to produce NeRF models with geometries that are more coherent with common sense, e.g., more smoothed ob-ject surface, yielding higher 3D fidelity. We demonstrate the efficacy of our approach through various comparative qualitative results and a solid user study. Kejie Li, Xueqing Deng, Yichun Shi, Minsu Cho |
CVPR | 2 |
| 2024 | MVDream: Multi-view Diffusion for 3D GenerationabstractWe introduce MVDream, a diffusion model that is able to generate consistent multi-view images from a given text prompt. Learning from both 2D and 3D data, a multi-view diffusion model can achieve the generalizability of 2D diffusion models and the consistency of 3D renderings. We demonstrate that such a multi-view diffusion model is implicitly a generalizable 3D prior agnostic to 3D representations. It can be applied to 3D generation via Score Distillation Sampling, significantly enhancing the consistency and stability of existing 2D-lifting methods. It can also learn new concepts from a few 2D examples, akin to DreamBooth, but for 3D generation. Yichun Shi, Jianglong Ye, Long Mai, Kejie Li |
ICLR | 5 |
| 2023 | NoPe-NeRF: Optimising Neural Radiance Field with No Pose PriorabstractTraining a Neural Radiance Field (NeRF) without precomputed camera poses is challenging. Recent advances in this direction demonstrate the possibility of jointly optimising a NeRF and camera poses in forward-facing scenes. However, these methods still face difficulties during dramatic camera movement. We tackle this challenging problem by incorporating undistorted monocular depth priors. These priors are generated by correcting scale and shift parameters during training, with which we are then able to constrain the relative poses between consecutive frames. This constraint is achieved using our proposed novel loss functions. Experiments on real-world indoor and outdoor scenes show that our method can handle challenging camera trajectories and outperforms existing methods in terms of novel view rendering quality and pose estimation accuracy. Our project page is https://nope-nerf.active.vision. Wenjing Bian, Kejie Li, Jiawang Bian |
CVPR | 3 |
| 2023 | MobileBrick: Building LEGO for 3D Reconstruction on Mobile DevicesabstractHigh-quality 3D ground-truth shapes are critical for 3D object reconstruction evaluation. However, it is difficult to create a replica of an object in reality, and even 3D reconstructions generated by 3D scanners have artefacts that cause biases in evaluation. To address this issue, we introduce a novel multi-view RGBD dataset captured using a mobile device, which includes highly precise 3D ground-truth annotations for 153 object models featuring a diverse set of 3D structures. We obtain precise 3D ground-truth shape without relying on high-end 3D scanners by utilising LEGO models with known geometry as the 3D structures for image capture. The distinct data modality offered by high-resolution RGB images and low-resolution depth maps captured on a mobile device, when combined with precise 3D geometry annotations, presents a unique opportunity for future research on high-fidelity 3D reconstruction. Furthermore, we evaluate a range of 3D reconstruction algorithms on the proposed dataset. Kejie Li, Jiawang Bian, Robert Castle, Philip Torr 0001, Victor Adrian Prisacariu |
CVPR | 1 |
| 2023 | ObjectSDF++: Improved Object-Compositional Neural Implicit SurfacesabstractIn recent years, neural implicit surface reconstruction has emerged as a popular paradigm for multi-view 3D reconstruction. Unlike traditional multi-view stereo approaches, the neural implicit surface-based methods leverage neural networks to represent 3D scenes as signed distance functions (SDFs). However, they tend to disregard the reconstruction of individual objects within the scene, which limits their performance and practical applications. To address this issue, previous work ObjectSDF introduced a nice framework of object-composition neural implicit surfaces, which utilizes 2D instance masks to supervise individual object SDFs. In this paper, we propose a new framework called ObjectSDF++ to overcome the limitations of ObjectSDF. First, in contrast to ObjectSDF whose performance is primarily restricted by its converted semantic field, the core component of our model is an occlusion-aware object opacity rendering formulation that directly volume-renders object opacity to be supervised with instance masks. Second, we design a novel regularization term for object distinction, which can effectively mitigate the issue that ObjectSDF may result in unexpected reconstruction in invisible regions due to the lack of constraint to prevent collisions. Our extensive experiments demonstrate that our novel framework not only produces superior object reconstruction results but also significantly improves the quality of scene reconstruction. Code and more resources can be found in https://qianyiwu.github.io/objectsdf++. Qianyi Wu, Kaisiyuan Wang, Kejie Li, Jianmin Zheng, Jianfei Cai 0001 |
ICCV | 3 |
| 2022 | BNV-Fusion: Dense 3D Reconstruction using Bi-level Neural Volume FusionabstractDense 3D reconstruction from a stream of depth images is the key to many mixed reality and robotic applications. Although methods based on Truncated Signed Distance Function (TSDF) Fusion have advanced the field over the years, the TSDF volume representation is confronted with striking a balance between the robustness to noisy measurements and maintaining the level of detail. We present Bi-level Neural Volume Fusion (BNV-Fusion), which leverages recent advances in neural implicit representations and neural rendering for dense 3D reconstruction. In order to incrementally integrate new depth maps into a global neural implicit representation, we propose a novel bi-level fusion strategy that considers both efficiency and reconstruction quality by design. We evaluate the proposed method on multiple datasets quantitatively and qualitatively, demonstrating a significant improvement over existing methods. Kejie Li, Yansong Tang, Victor Adrian Prisacariu, Philip Torr 0001 |
CVPR | 1 |
| 2022 | Object-Compositional Neural Implicit Surfaces
Qianyi Wu, Yuedong Chen, Kejie Li, Chuanxia Zheng, Jianfei Cai 0001, Jianmin Zheng |
ECCV (27) | 4 |
| 2022 | A Low Computational Cost Lightning Mapping Algorithm With a Nonuniform L-Shaped Array: Principle and VerificationabstractRecently, 2-D multiple signal classification (MUSIC) algorithm was applied to map lighting propagation with improved quality. However, the quality is improved at the cost of increasing the calculation load. To improve the efficiency, this article proposes an improved MUSIC-based lightning mapping algorithm, which transforms the 2-D direction of arrival (DOA) estimation problem for an L-shaped array into 1-D DOA estimation problem for the two arms of the array. Additionally, the lightning mapping array is arranged in a nonuniform way to reduce the risk of trivial ambiguity problem. The performance of the developed lightning mapping system is verified with numerical Monte Carlo simulation and further with real lightning observation results. The simulation results show that the nonuniform array outperforms the uniform array in terms of detection accuracy, spatial resolution, and antiambiguity. The observation results show that the MUSIC-based methods produce better lightning mapping quality compared with the interferometry method (15 777 sources mapped), due to better robustness for low signal-to-noise ratio (SNR) lightning very high-frequency (VHF) signals. Compared with the 2-D MUSIC method (34 620 sources mapped), the proposed low computational cost method (30 461 sources mapped) is proven to produce fairly similar path of lightning leader with significant lower computational cost. Specifically, the proposed method saves about 84.91% of calculation time compared with the 2-D MUSIC method in practice. Huaifei Chen, Weijiang Chen, Yu Wang 0150, Kai Bian, Nianwen Xiang, Kejie Li |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | Ray-ONet: Efficient 3D Reconstruction From A Single RGB Image
Wenjing Bian, Kejie Li, Victor Adrian Prisacariu |
BMVC | 3 |
| 2021 | ODAM: Object Detection, Association, and Mapping using Posed RGB VideoabstractLocalizing objects and estimating their extent in 3D is an important step towards high-level 3D scene understanding, which has many applications in Augmented Reality and Robotics. We present ODAM, a system for 3D Object Detection, Association, and Mapping using posed RGB videos. The proposed system relies on a deep learning front-end to detect 3D objects from a given RGB frame and associate them to a global object-based map using a graph neural network (GNN). Based on these frame-to-model associations, our back-end optimizes object bounding volumes, represented as super-quadrics, under multi-view geometry constraints and the object scale prior. We validate the proposed system on ScanNet where we show a significant improvement over existing RGB-only methods. Kejie Li, Daniel DeTone, Steven Chen, Minh Vo, Ian D. Reid 0001, Seyed Hamid Rezatofighi, Chris Sweeney, Julian Straub, Richard A. Newcombe |
ICCV | 1 |
| 2020 | FroDO: From Detections to 3D ObjectsabstractObject-oriented maps are important for scene understanding since they jointly capture geometry and semantics, allow individual instantiation and meaningful reasoning about objects. We introduce FroDO, a method for accurate 3D reconstruction of object instances from RGB video that infers their location, pose and shape in a coarse to fine manner. Key to FroDO is to embed object shapes in a novel learnt shape space that allows seamless switching between sparse point cloud and dense DeepSDF decoding. Given an input sequence of localized RGB frames, FroDO first aggregates 2D detections to instantiate a 3D bounding box per object. A shape code is regressed using an encoder network before optimizing shape and pose further under the learnt shape priors using sparse or dense shape representations. The optimization uses multi-view geometric, photometric and silhouette losses. We evaluate on real-world datasets, including Pix3D, Redwood-OS, and ScanNet, for single-view, multi-view, and multi-object reconstruction. Martin Rünz, Kejie Li, Meng Tang 0001, Lingni Ma, Chen Kong, Tanner Schmidt, Ian D. Reid 0001, Lourdes Agapito, Julian Straub, Steven Lovegrove, Richard A. Newcombe |
CVPR | 2 |
| 2019 | Single-view Object Shape Reconstruction Using Deep Shape Prior and Silhouette
Kejie Li, Ravi Garg, Ian D. Reid 0001 |
BMVC | 1 |
| 2019 | Real-Time Monocular Object-Model Aware Sparse SLAMabstractSimultaneous Localization And Mapping (SLAM) is a fundamental problem in mobile robotics. While sparse point-based SLAM methods provide accurate camera localization, the generated maps lack semantic information. On the other hand, state of the art object detection methods provide rich information about entities present in the scene from a single image. This work incorporates a real-time deep-learned object detector to the monocular SLAM framework for representing generic objects as quadrics that permit detections to be seamlessly integrated while allowing the real-time performance. Finer reconstruction of an object, learned by a CNN network, is also incorporated and provides a shape prior for the quadric leading further refinement. To capture the structure of the scene, additional planar landmarks are detected by a CNN-based plane detector and modelled as independent landmarks in the map. Extensive experiments support our proposed inclusion of semantic objects and planar structures directly in the bundle-adjustment of SLAM - Semantic SLAM- that enriches the reconstructed map semantically, while significantly improving the camera localization. Mehdi Hosseinzadeh 0003, Kejie Li, Yasir Latif, Ian D. Reid 0001 |
ICRA | 2 |
| 2018 | Unsupervised Learning of Monocular Depth Estimation and Visual Odometry With Deep Feature ReconstructionabstractDespite learning based methods showing promising results in single view depth estimation and visual odometry, most existing approaches treat the tasks in a supervised manner. Recent approaches to single view depth estimation explore the possibility of learning without full supervision via minimizing photometric error. In this paper, we explore the use of stereo sequences for learning depth and visual odometry. The use of stereo sequences enables the use of both spatial (between left-right pairs) and temporal (forward backward) photometric warp error, and constrains the scene depth and camera motion to be in a common, real-world scale. At test time our framework is able to estimate single view depth and two-view odometry from a monocular sequence. We also show how we can improve on a standard photometric warp loss by considering a warp of deep features. We show through extensive experiments that: (i) jointly training for single view depth and visual odometry improves depth prediction because of the additional constraint imposed on depths and achieves competitive results for visual odometry; (ii) deep feature-based warping loss improves upon simple photometric warp loss for both single view depth estimation and visual odometry. Our method outperforms existing learning based methods on the KITTI driving dataset in both tasks. The source code is available at https://github.com/Huangying-Zhan/Depth-VO-Feat. Huangying Zhan, Ravi Garg, Chamara Saroj Weerasekera, Kejie Li, Ian D. Reid 0001 |
CVPR | 4 |
| 2018 | Efficient Dense Point Cloud Object Reconstruction Using Deformation Vector Fields
Kejie Li, Trung Pham, Huangying Zhan, Ian D. Reid 0001 |
ECCV (12) | 1 |
| 2018 | Methodical Approach for Immunity Assessment of Electronic Devices Excited by High Power EMP
Vladimir Chepelev, Yury Parfenov, William Radasky, Boris Titov, Leonid Zdoukhov, Kejie Li, Xu Kong, Yan-Zhao Xie |
J. Electron. Test. | 6 |
| 2014 | Research on high resolution & high sensitivity panoramic surveillance systemabstractCatadioptric panoramic vision systems have been widely used in many fields because of their advantages such as a wide field of view, integrated imaging, and rotational symmetry. They also play a very important role in the monitoring of unmanned platforms. However, the resolution of catadioptric panoramic vision systems is not very high, usually less than 5 million pixels at present. Even if the resolution is greater than 5 million pixels, the unwrapping and rectification of these images is carried out off-line. Further, the systems are also deployed in a stationary or near-stationary state. This paper proposes an unwrapping and rectification method for a high-resolution catadioptric panoramic vision system that can be used while moving. It can segment dynamic circular mark regions from panoramic videos accurately, determine the center coordinates of circular images in real time, eliminate the processing of unwanted video regions, and shorten the image processing time. In addition, the center coordinates and radius of the circular mark regions are obtained simultaneously so that the image distortion caused by inaccurate center coordinates is reduced. The method uses radial and decentering distortion parameters and a correction factor for fitting a polynomial to rectify the panoramic video without distortion. Kejie Li, Bo Pu, Jiyang Li, Zijie Jiang, Mengying Liu |
AVSS | 2 |
| 2013 | Study on algorithm for panoramic image basing on high sensitivity and high resolution panoramic surveillance cameraabstractA single panoramic annular lens optical system and high-sensitivity, high-resolution CCD sensors form the basis of a 360° panoramic night vision image processing hardware platform. The unwrapped and correcting algorithm based on Coordinate Rotation Digital Computer (CORDIC) and bilinear interpolation algorithm was presented in this paper, with the purpose of processing dynamic panoramic annular image. The night vision image enhancement algorithm, based on adaptive piecewise linear gray transformation (APLGT) and Laplacian of Gaussian (LOG) edge detection, were given. An original annular panoramic image captured by panoramic annular lens (PAL) can be unwrapped and corrected to conventional rectangular image without distortion, which is much more coincident with people's vision. APLGT algorithm can be adaptively truncate the image histogram on both ends to obtain a smaller dynamic range so as to enhance the contrast of the night vision image. LOG algorithm can be propitious to find and detect dim small targets in night vision circumstance. The algorithm for panoramic image processing is modeled by VHDL and implemented in FPGA. The experimental results show that the proposed panoramic image algorithm for unwrapped and distortion correction has the lower computation complexity and the architecture for dynamic panoramic image processing has lower hardware cost and power consumption. Kejie Li |
AVSS | 2 |
| 2008 | Pattern Matching in RNA Structures
Kejie Li, Reazur Rahman, Aditi Gupta 0001, Prasad Siddavatam, Michael Gribskov |
ISBRA | 1 |
| 2006 | Real-time Object Tracking of a Robot Head Based on Multiple Visual Cues IntegrationabstractMost visual tracking systems use single visual cue and usually fail in a complex environment. In this paper, first, different visual cues are analyzed for a pan-tilt robot head tracking system. Then, a visual cues integration method combining disparity, color and shape is proposed. Two computers linked with Memolink communication module ensure the robot head to track a moving object rapidly. The high robustness and real time performance of the system are confirmed by the experiments Yunting Pang, Qiang Huang 0002, Zhangfeng Hu, Altaf Hussain Rajpar, Kejie Li |
IROS | 6 |
| 2006 | Walking Pattern Generation for Humanoid Robot Considering Upper Body MotionabstractWalking pattern generation is a main issue for humanoid robot. We have already proposed a method for planning stable walking pattern. Based on this method, this paper mainly discusses generating stable and harmonious walking pattern by considering upper body motion, and planning hip trajectories in both sagittal plane and lateral plane. To reduce the iterative computation cost, some constraints of the relationship between sagittal hip motion and lateral hip motion are formulated, and only the trajectories satisfy these constraints are worked out. Finally, we determine the trajectory with a large stability margin from these generated trajectories. The effectiveness of the proposed method is confirmed by simulations and experiments with our developed humanoid robot BHR-02 with 32 DOF Jie Yang 0018, Qiang Huang 0002, Jianxi Li, Kejie Li |
IROS | 5 |
| 2005 | Humanoid On-line Pattern Generation Based on Parameters of Off-line Typical Walk PatternsabstractThe complexity of nonlinear differential equations of dynamics makes it practically impossible to obtain the walk pattern on-line through computing the whole dynamics. This paper proposed a method of online trajectory generation based on key parameters of off-line typical walk patterns for a biped humanoid. The key parameters include hip parameters, step length, walking cycle and so on. The walking pattern can be obtained according to these parameters. In order to generate walking patterns online, first the key parameters of the on-line walking pattern are computed based on the parameters of off-line typical patterns, then stability optimization has been done on-line and on-line trajectories are derived. The effectiveness of the proposed method is confirmed by simulations and experiments with our developed humanoid robot with 33 DOF. Zhaoqin Peng, Qiang Huang 0002, Lige Zhang, Ali Raza Jafri, Kejie Li |
ICRA | 6 |
| 2005 | Design of humanoid complicated dynamic motion based on human motion captureabstractCaptured human data must be adapted for the humanoid because its kinematics and dynamics differ from those of the human actor. On the other hand, it is desirable that humanoid movements are highly similar to those of the human actor, since the human actor's motion is regarded as a teaching motion. This paper explores the design of a humanoid complicated dynamic motion based on human motion capture. First, the kinematic constraints, including ground contact conditions, are formulated. Next, the similarity evaluation and dynamic stability based on ZMP (zero moment point) of the humanoid motion are discussed, and the method to derive humanoid motion with a high similarity, and satisfying kinematic constraints and dynamic stability, is presented. Finally, the effectiveness of the proposed method is confirmed by simulations and experiments with the "sword" motion - a complicated and dynamic Chinese kungfu movement - using our developed humanoid robot with 32 degree of freedom. Qiang Huang 0002, Zhaoqin Peng, Lige Zhang, Kejie Li |
IROS | 5 |
| 2004 | Stability Criterion and Pattern Planning for Humanoid RunningabstractAlthough some researchers have studied the humanoid running based on simplified models, the issue of humanoid running based on the whole dynamics has not been sufficiently discussed before. The objective of this paper is to study the stability criterion and the dynamic pattern generation for humanoid running based on the whole dynamics. First, the cycle and the dynamics of running are analyzed. Next, the stability criterion of humanoid running is presented. Then, the method to plan a running pattern consisting of a foot trajectory and a hip trajectory is proposed. Finally, the effectiveness of the proposed method is illustrated by the dynamic simulation examples in DADS (Dynamic Analysis and Design System). Qiang Huang 0002, Kejie Li, Xingguang Duan |
ICRA | 3 |
| 2004 | Towards Automated Micromachining of PMMA Micro Channels using CO/Sub 2/ Laser and Sacrificial Mask ProcessabstractA novel system for 3D microchannel fabrication based on CO/sub 2/ laser-micromachining is presented. The system consists of a CO/sub 2/ laser focusing system and a 3D precision positioning platform. The CO/sub 2/ laser focusing system can regulate a laser beam which is of appropriate energy level and of micron dimension in beam diameter. The fabrication of 3D microchannel system in PMMA (polymethyl methacrylate) can be realized by controlling the CO/sub 2/ laser and a 3-axis platform with micron-resolution movement. A special 'sacrificial mask' process was used to produce translucent channels of micron dimensions with low surface roughness using the developed system. The effectiveness of our developed system is confirmed by experimental results. Potentially, our developed system can be automated to produce 3D micro channels in PMMA substrates without the requirement for the costly and time-consuming lithography and hot-embossing processes that are needed currently. Guangyi Shi, Qiang Huang 0002, Wen J. Li, Wenqian Huang, Gengchen Shi, Kejie Li |
ICRA | 6 |
| 2004 | Kinematics mapping and similarity evaluation of humanoid motion based on human motion captureabstractThe captured data must be adapted for the humanoid because its kinematics and dynamic differ from those of the human actor. The kinematics constraints such as ground contact conditions are crucial for humanoid locomotion. Furthermore, it is desirable that the humanoid motion have of high similarity with those of the human actor. In this paper, first the similarity function of the humanoid motion is proposed. Then, the kinematics constrains including ground contact conditions are formulated, and the algorithm to derive the humanoid motion with a high similarity and satisfying kinematics constraints is present. Finally, the effectiveness is confirmed by the experiment of Chinese Kungfu "Taiji" using our developed 33 DOF humanoid robot. Xiaojun Zhao, Qiang Huang 0002, Zhaoqin Peng, Kejie Li |
IROS | 4 |
| 2003 | Cooperation of dynamic patterns and sensory reflex for humanoid walkingabstractThis paper proposes a walk structure consisting of a dynamic pattern, a sensory reflex and a motion adjustment. The dynamic pattern is generated off-line based on the constraint of dynamic stability, assuming that the models of the humanoid and the environment are known. The sensory reflex is simple, but rapid motion programmed in respect to sensory information. The sensory reflex increases the humanoid adaptability to environmental uncertainties, but it may conflict with humanoid its constraints. To solve this problem, the walking constraints violated easily by the sensory reflex are formulated, and the method to coordinate the dynamic pattern and the sensory reflex through the motion adjustment is present. The effectiveness was confirmed by walk experiments of our developed 31 DOF humanoid. Gunag Wang, Qiang Huang 0002, Juhong Geng, Hongbin Deng, Kejie Li |
ICRA | 5 |
| 2002 | Uncalibrated Visual Servoing of Planar RobotsabstractThe calibration accuracy of the intrinsic and extrinsic parameters of the vision system greatly affects the performance of visual servoing. We address the problem of controlling a planar manipulator using a fixed single camera without calibrating its intrinsic parameters and the transformation matrix between the robot base frame and the camera frame, and without measuring manipulator's depth. Based on an important observation that the unknown parameters can be separated from the unknown composite image Jacobian matrix, we propose an adaptive algorithm to estimate the unknown and mixed parameters on-line. It is proved with a full consideration of dynamics of the system by Lyapunov approach that the feature points of planar manipulator approach asymptotically to the desired ones on image plane and the estimated parameters are bounded under the control of the proposed visual servo controller. The performance has been confirmed by simulations and experiments. Yantao Shen 0001, Guoliang Xiang, Yun-Hui Liu 0001, Kejie Li |
ICRA | 4 |
| 2001 | Asymptotic Motion Control of Robot Manipulators Using Uncalibrated Visual FeedbackabstractTo implement a visual feedback controller, it is necessary to calibrate the homogeneous transformation matrix between the robot base frame and the vision frame besides the intrinsic parameters of the vision system. The calibration accuracy greatly affects the control performance. In this paper, we address the problem of controlling a robot manipulator using visual feedback without calibrating the transformation matrix. We propose an adaptive algorithm to estimate the unknown matrix online. It is proved by the Lyapunov method that the robot motion approaches asymptotically to the desired one and the estimated matrix is bounded under the control of the proposed visual feedback controller. The performance was confirmed by simulations and experiments. Yantao Shen 0001, Yun-Hui Liu 0001, Kejie Li, Jianwei Zhang 0001, Alois C. Knoll |
ICRA | 3 |
| 2001 | Analysis of physical capability of a biped humanoid: walking speed and actuator specificationsabstractThe reliability of stable walk and the development of high performance components are two crucial issues to develop a humanoid with human-like physical capability. In order to for the humanoid to walk smoothly and to adapt to unknown environments, we first propose a balance control that combines a feedforward dynamic pattern and a feedback sensory reflection. Then, we present a method for clarifying the relationship between the physical capability and actuator's specifications. Using this method, it is possible to predict the walking speed based on known actuator specifications and to obtain the necessary specifications to accomplish a desired walking speed. Finally, experiments of an 26-DOF humanoid and simulation examples are provided to illustrate the effectiveness of the proposed method. Qiang Huang 0002, Kejie Li, Yoshihiko Nakamura, Kazuo Tanie |
IROS | 2 |
| 2001 | Adaptive visual feedback control of manipulators in uncalibrated environmentabstractTo implement a position-based visual feedback controller for a manipulator, it is necessary to calibrate the homogeneous transformation matrix between its base frame and the vision frame besides the intrinsic parameters of the vision system. In this paper, based on an important observation that the unknown transformation matrix can be separated from the visual Jacobian matrix, we design an adaptive controller for manipulators when the matrix is not calibrated. It is proved, with a full dynamics of the system, by the Lyapunov approach that the motion of the manipulator approaches asymptotically to the desired trajectory. Simulations and experimental results both demonstrate the performance of this new controller. Yantao Shen 0001, Yun-Hui Liu 0001, Kejie Li |
IROS | 3 |
| 2000 | Asymptotic position control of robot manipulators using uncalibrated visual feedbackabstractTo implement a visual feedback controller, it a's necessary to calibrate the homogeneous transformation matrix between the robot base frame and the vision frame besides the intrinsic parameters of the vision system. The calibration accuracy greatly affects the control performance. We address the problem of controlling a robot manipulator using visual feedback without calibrating the transformation matrix. It is assumed that the vision system can measure the 3D position and orientation of the robot in real-time. Based on the fact that the visual Jacobian matrix can be represented in a linear form of elements of the transformation matrix, we propose a simple adaptive algorithm to estimate the unknown matrix on-line. This visual feedback controller greatly simplifies the implementation process of a robot-vision workcell and is especially useful when a pre-calibration is not possible, such as when a robot works with an active vision system carried by a mobile robot. It is proved by the Lyapunov approach that the robot position approaches asymptotically to the desired one and the estimated matrix is bounded under the control of this visual feedback controller. The performance has been confirmed by simulations and experiment. Yantao Shen 0001, Yun-Hui Liu 0001, Kejie Li |
IROS | 3 |