Jiafeng Xu

dblp:175/5124 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Hybrid CNN-transformer framework with dynamic feature fusion for enhanced passport background texture classification
Maoqin Tian, Jiafeng Xu, Lingpei Zeng, Eryang Chen, Yuanlun Xie
Vis. Comput.3
2025 World Model-Based Perception for Visual Legged Locomotion
abstract
Legged locomotion over various terrains is challenging and requires precise perception of the robot and its surroundings from both proprioception and vision. However, learning directly from high-dimensional visual input is often data-inefficient and intricate. To address this issue, traditional methods attempt to learn a teacher policy with access to privileged information first and then learn a student policy to imitate the teacher's behavior with visual input. Despite some progress, this imitation framework prevents the student policy from achieving optimal performance due to the information gap between inputs. Furthermore, the learning process is unnatural since animals intuitively learn to traverse different terrains based on their understanding of the world without privileged knowledge. Inspired by this natural ability, we propose a simple yet effective method, World Model-based Perception (WMP), which builds a world model of the environment and learns a policy based on the world model. We illustrate that though completely trained in simulation, the world model can make accurate predictions of real-world trajectories, thus providing informative signals for the policy controller. Extensive simulated and real-world experiments demonstrate that WMP outperforms state-of-the-art baselines in traversability and robustness. Videos and Code are available at: https://wmp-loco.github.io/.
Hang Lai, Jiahang Cao, Jiafeng Xu, Yunfeng Lin, Tao Kong, Yong Yu 0001, Weinan Zhang 0001
ICRA3
2025 Flow-Based Policy for Online Reinforcement Learning
abstract
We present $\textbf{FlowRL}$, a novel framework for online reinforcement learning that integrates flow-based policy representation with Wasserstein-2-regularized optimization. We argue that in addition to training signals, enhancing the expressiveness of the policy class is crucial for the performance gains in RL. Flow-based generative models offer such potential, excelling at capturing complex, multimodal action distributions. However, their direct application in online RL is challenging due to a fundamental objective mismatch: standard flow training optimizes for static data imitation, while RL requires value-based policy optimization through a dynamic buffer, leading to difficult optimization landscapes. FlowRL first models policies via a state-dependent velocity field, generating actions through deterministic ODE integration from noise. We derive a constrained policy search objective that jointly maximizes Q through the flow polciy while bounding the Wasserstein-2 distance to a behavior-optimal policy implicitly derived from the replay buffer. This formulation effectively aligns the flow optimization with the RL objective, enabling efficient and value-aware policy learning despite the complexity of the policy class. Empirical evaluations on DMControl and Humanoidbench demonstrate that FlowRL achieves competitive performance in online reinforcement learning benchmarks.
Yu Luo 0021, Fuchun Sun 0001, Tao Kong, Jiafeng Xu
NeurIPS6
2025 A deep learning approach for non-invasive Alzheimer's monitoring using microwave radar data
abstract
Over 50 million people globally suffer from Alzheimer's disease (AD), emphasizing the need for efficient, early diagnostic tools. Traditional methods like Magnetic Resonance Imaging (MRI) and Computed Tomography (CT) scans are expensive, bulky, and slow. Microwave-based techniques offer a cost-effective, non-invasive, and portable solution, diverging from conventional neuroimaging practices. This article introduces a deep learning approach for monitoring AD , using realistic numerical brain phantoms to simulate scattered signals via the CST Studio Suite. The obtained data is preprocessed using normalization, standardization, and outlier removal to ensure data integrity. Furthermore, we propose a novel data augmentation technique to enrich the dataset across various AD stages. Our deep learning approach combines Recursive Feature Elimination (RFE) with Principal Component Analysis (PCA) and Autoencoders (AE) for optimal feature selection. Convolution Neural Network (CNN) is combined with Gated Recurrent Unit (GRU), Bidirectional Long Short Term Memory (Bidirectional-LSTM), and Long Short-Term Memory (LSTM) to improve classification performance. The integration of RFE-PCA-AE significantly elevates performance, with the CNN+GRU model achieving an 87% accuracy rate, thus outperforming existing studies.
Farhatullah, Xin Chen 0012, Deze Zeng, Rahmat Ullah, Rab Nawaz, Jiafeng Xu, Tughrul Arslan
Neural Networks6
2024 Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation
abstract
Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of this effectiveness by showing that visual robot manipulation can significantly benefit from large-scale video generative pre-training. We introduce GR-1, a GPT-style model designed for multi-task language-conditioned visual robot manipulation. GR-1 takes as inputs a language instruction, a sequence of observation images, and a sequence of robot states. It predicts robot actions as well as future images in an end-to-end manner. Thanks to a flexible design, GR-1 can be seamlessly finetuned on robot data after pre-trained on a large-scale video dataset. We perform extensive experiments on the challenging CALVIN benchmark and a real robot. On CALVIN benchmark, our method outperforms state-of-the-art baseline methods and improves the success rate from 88.9% to 94.9%. In the setting of zero-shot unseen scene generalization, GR-1 improves the success rate from 53.3% to 85.4%. In real robot experiments, GR-1 also outperforms baseline methods and shows strong potentials in generalization to unseen scenes and objects. We provide inaugural evidence that a unified GPT-style transformer, augmented with large-scale video generative pre-training, exhibits remarkable generalization to multi-task visual robot manipulation. Project page: https://GR1-Manipulation.github.io
Ya Jing, Chilam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Tao Kong
ICLR5
2024 Multi-objective trajectory planning in the multiple strata drilling process:A bi-directional constrained co-evolutionary optimizer with Pareto front learning
Jiafeng Xu, Xin Chen 0012, Min Wu 0002
Expert Syst. Appl.1
2023 MOMA-Force: Visual-Force Imitation for Real-World Mobile Manipulation
abstract
In this paper, we present a novel method for mobile manipulators to perform multiple contact-rich manipulation tasks. While learning-based methods have the potential to generate actions in an end-to-end manner, they often suffer from insufficient action accuracy and robustness against noise. On the other hand, classical control-based methods can enhance system robustness, but at the cost of extensive parameter tuning. To address these challenges, we present MOMA-Force, a visual-force imitation method that seamlessly combines representation learning for perception, imitation learning for complex motion generation, and admittance whole-body control for system robustness and controllability. MOMA-Force enables a mobile manipulator to learn multiple complex contact-rich tasks with high success rates and small contact forces. In a real household setting, our method outperforms baseline methods in terms of task success rates. Moreover, our method achieves smaller contact forces and smaller force variances compared to baseline methods without force imitation. Overall, we offer a promising approach for efficient and robust mobile manipulation in the real world. Videos and more details can be found on https://visual-force-imitation.github.io.
Taozheng Yang, Ya Jing, Jiafeng Xu, Kuankuan Sima, Guangzeng Chen, Qie Sima, Tao Kong
IROS4
2022 Real-time Inertial Parameter Identification of Floating-Base Robots Through Iterative Primitive Shape Division
abstract
Dynamic models play a key role in robot motion generation and control and the identification of inertial parameters is a critical component for obtaining an accurate dynamic model of a robot. This paper presents a novel iterative primitive shape division method for the inertia parameter identification of floating-base robots. Describing a robot by a set of primitive shapes with uniform mass distributions, the method iteratively divides the primitive shapes into smaller ones and refines their masses, which quickly converges to yielding the true inertia parameters of the robot. This method guarantees the physical consistency of the obtained parameters, possesses a high computational efficiency for online deployment, and works without contact force measurement. Furthermore, it can be used to estimate the position and magnitude of an external load applied to the robot. Simulations and experiments on a quadruped robot have been conducted to verify the effectiveness and efficiency of the proposed method.
Jiafeng Xu, Yu Zheng 0001, Xinyang Jiang, Lingzhu Xiang, Zhengyou Zhang
ICRA1
2021 Highest Wellbore Stability Obstacle Avoidance Drilling Trajectory Optimization in Complex Multiple Strata Geological Environment
abstract
Drilling trajectory optimization (DTO) is an important step in intelligent control of directional drilling process. DTO in complex multiple strata geological environment is a new nonlinear, strongly constrained, black-box multi-objective optimization problem with discrete objectives. In this paper, a multi-objective evolutionary algorithm (MOEA) optimization framework is present for an obstacle avoidance drilling trajectory design model. First, a new three-segment trajectory model is established by natural curve method to measure the whole trajectory positions. Meanwhile, a drilling point wellbore stability model is derived by Mohr-Coulomb failure criteria. Based on this calculation model, the trajectory data is mapping to get a multi-strata parameter field. Several representative MOEAs are adopted to design the highest wellbore stability trajectory in this multi-strata geological environment. A case study from actual drilling site shows that the result of adaptive non-dominated sorting genetic algorithm (ANSGAIII) is better than that of traditional trail and error drilling trajectory design, which has a good application prospect in drilling engineering.
Jiafeng Xu, Xin Chen 0012, Min Wu 0002
IECON1
2020 Stable Parking Control of a Robot Astronaut in a Space Station Based on Human Dynamics
abstract
Controlling a robot astronaut to move in the same way as a human astronaut to realize a wide range of motion in a space station is an important requirement for the robot astronauts that are meant to assist or replace human astronauts. However, a robot astronaut is a nonlinear and strongly coupled multibody dynamic system with multiple degrees of freedom, whose dynamic characteristics are complex. Therefore, implementing a robot astronaut with wide-ranging motion control in a space station is a tremendous challenge for robotic technology. This article presents a wide-ranging stable motion control method for robot astronauts in space stations based on human dynamics. Focusing on the astronauts' parking motion in a space station, a viscoelastic dynamic humanoid model of parking under microgravity environment was established using a mass-spring-damper system. The model was used as the expected model for stable parking control of a robot astronaut, and the complex dynamic characteristics were mapped into the robot astronaut system to control the stable parking of the robot astronaut in a manner similar to a human astronaut. This provides a critical basis for implementing robots that are capable of steady wide-ranging motion in space stations. The method was verified on a dynamic system of a robot astronaut that was constructed for this research. The experimental results showed that the method is feasible and effective and that it is a highly competitive solution for robot astronauts with human-like moving capabilities in space stations.
Zhihong Jiang, Jiafeng Xu, Hui Li 0047, Qiang Huang 0002
IEEE Trans. Robotics2