Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jiawei Mo

dblp:189/5967 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Systems, architecture and hardware · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Face, body and person analysis · 35% Robot navigation and mapping · 20% Representation and self-supervised learning · 18%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining
1.012026
Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification · AAAI 2026
Computer vision › Face, body and person analysis
person re-identification
1.012026
Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification · AAAI 2026
Computer vision › Video understanding and tracking
temporal modeling
1.012026
Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification · AAAI 2026
Computer vision › Face, body and person analysis › person re-identification
video-based person re-identification
1.012026
Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification · AAAI 2026
Robotics › Motion planning and robot control
constrained nonlinear optimization
0.612022
Continuous-Time Spline Visual-Inertial Odometry · ICRA 2022
Robotics › Robot navigation and mapping
state estimation
0.612022
Continuous-Time Spline Visual-Inertial Odometry · ICRA 2022
Robotics › Robot navigation and mapping › visual odometry
visual-inertial odometry
0.612022
Continuous-Time Spline Visual-Inertial Odometry · ICRA 2022

Methods — techniques the papers use, named apart from their topics

prototype fusion · 1.0contrastive learning · 1.0nonlinear optimization · 0.6cubic spline · 0.6
YearPublicationVenuePosition
2026 Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification
abstract
Multimodal pretraining has revolutionized visual understanding, but its impact on video-based person re-identification (ReID) remains underexplored. Existing approaches often rely on video-text pairs, yet suffer from two fundamental limitations: (1) lack of genuine multimodal pretraining, and (2) text poorly captures fine-grained temporal motion—an essential cue for distinguishing identities in video. In this work, we take a bold departure from text-based paradigms by introducing the first skeleton-driven pretraining framework for ReID. To achieve this, we propose Contrastive Skeleton-Image Pretraining for ReID (CSIP-ReID), a novel two-stage method that leverages skeleton sequences as a spatiotemporally informative modality aligned with video frames. In the first stage, we employ contrastive learning to align skeleton and visual features at sequence level. In the second stage, we introduce a dynamic Prototype Fusion Updater (PFU) to refine multimodal identity prototypes, fusing motion and appearance cues. Moreover, we propose a Skeleton Guided Temporal Modeling (SGTM) module that distills temporal cues from skeleton data and integrates them into visual features. Extensive experiments demonstrate that CSIP-ReID achieves new state-of-the-art results on standard video ReID benchmarks (MARS, LS-VID, iLIDS-VID). Moreover, it exhibits strong generalization to skeleton-only ReID tasks (BIWI, IAS), significantly outperforming previous methods. CSIP-ReID pioneers an annotation-free and motion-aware pretraining paradigm for ReID, opening a new frontier in multimodal representation learning.
Rifen Lin, Alex Jinpeng Wang, Jiawei Mo, Min Li 0007
AAAI3
2026 AxialUNet: A Lightweight Network for Medical Image Segmentation with Axial Operators
Meiyi Lyu, Jiawei Mo
MMM (2)2
2026 MoChat: Joints-Grouped Spatio-Temporal Grounding Multimodal Large Language Model for Multi-Turn Motion Comprehension and Description
abstract
Despite continuous advancements in deep learning for understanding human motion, existing models often struggle to accurately identify action timing and specific body parts, typically supporting only single-round interaction. This limitation is particularly pronounced in home exercise monitoring, neurological disorder assessment, and rehabilitation, where precise motion analysis is crucial for ensuring exercise efficacy, detecting early signs of neurological conditions, and guiding personalized recovery programs. In this paper, we propose MoChat, a multimodal large language model capable of spatio-temporal grounding of human motion and multi-turn dialogue understanding. To achieve this, we first group spatial features in skeleton frames according to human anatomical structures and process them through a Joints-Grouped Skeleton Encoder. The encoder's outputs are fused with large language model embeddings to generate spatio-aware representations. A cross-attention-based Regression Head module is then designed to align hidden-layer embeddings and skeletal sequence embeddings, enabling precise temporal grounding. Furthermore, we develop a pipeline for temporal grounding task to extract timestamps from skeleton-text pairs and construct a multi-turn instruction dialogues for spatial grounding task. Finally, various task instructions are generated for jointly training. Experimental results demonstrate that MoChat achieves state-of-the-art performance across multiple metrics in motion understanding tasks, making it as the first model capable of fine-grained spatio-temporal grounding of human motion.
Jiawei Mo, Yixuan Chen 0019, Rifen Lin, Yongkang Ni, Feng Liang 0004, Min Zeng 0004, Xiping Hu, Min Li 0007
IEEE J. Biomed. Health Informatics1
2025 VMS2-UNet: A Lightweight Non-Causal Vision Mamba2 Model with State Space Duality for Medical Image Segmentation
abstract
CNN-based and Transformer-based models are widely applied in medical image segmentation. However, CNN models exhibit limitations in capturing long-range dependencies, while Transformer models demand a substantially higher number of parameters due to their self-attention mechanism. To address these challenges, some methods adopt State Space Models (SSM), which excel at modeling long-range interactions with a compact parameter design. Recently, Mamba2 introduces State Space Duality (SSD), an improved variant of SSM that enhances model performance and efficiency. Nevertheless, the inherent causal property of SSM/SSD restricts their applicability in medical image segmentation. To overcome this limitation, we propose a novel method, Vision Mamba2 State-Space UNet (VMS2-UNet), a U-shaped lightweight Vision Mamba2 model with Non-Causal State Space Duality, specifically designed for medical image segmentation. VMS2-UNet integrates multiple NCState Blocks (NCS Blocks), which adopt the non-causal format of SSD to efficiently model long-range dependencies while preserving parameter efficiency. To mitigate spatial information loss caused by downsampling, we introduce FIMA (Feature Integration and Modulation Attention), which enhances feature integration in the skip connections, improving segmentation performance while adding only 0.04M parameters. We conduct extensive experiments on three challenging benchmarks, and the results show that VMS2-UNet achieves competitive performance in medical image segmentation, outperforming several state-of-the-art methods while maintaining a lightweight design.
Jiawei Mo
IJCNN1
2022 Continuous-Time Spline Visual-Inertial Odometry
abstract
We propose a continuous-time spline-based formulation for visual-inertial odometry (VIO). Specifically, we model the poses as a cubic spline, whose temporal derivatives are used to synthesize linear acceleration and angular velocity, which are compared to the measurements from the inertial measurement unit (IMU) for optimal state estimation. The spline boundary conditions create constraints between the camera and the IMU, with which we formulate VIO as a constrained nonlinear optimization problem. Continuous-time pose representation makes it possible to address many VIO challenges, e.g., rolling shutter distortion and sensors that may lack synchronization. We conduct experiments on two publicly available datasets that demonstrate the state-of-the-art accuracy and real-time computational efficiency of our method.
Jiawei Mo, Junaed Sattar
ICRA1
2020 Design and Experiments with LoCO AUV: A Low Cost Open-Source Autonomous Underwater Vehicle
abstract
In this paper we present the LoCO AUV, a Low-Cost, Open Autonomous Underwater Vehicle. LoCO is a general-purpose, single-person-deployable, vision-guided AUV, rated to a depth of 100 meters. We discuss the open and expandable design of this underwater robot, as well as the design of a simulator in Gazebo. Additionally, we explore the platform's preliminary local motion control and state estimation abilities, which enable it to perform maneuvers autonomously. In order to demonstrate its usefulness for a variety of tasks, we implement a variety of our previously presented human-robot interaction capabilities on LoCO, including gestural control, diver following, and robot communication via motion. Finally, we discuss the practical concerns of deployment and our experiences in using this robot in pools, lakes, and the ocean. All design details, instructions on assembly, and code will be released under a permissive, open-source license.
Chelsey Edge, Sadman Sakib Enan, Michael Fulton, Jungseok Hong, Jiawei Mo, Kimberly Barthelemy, Hunter Bashaw, Berik Kallevig, Corey Knutson, Kevin Orpen, Junaed Sattar
IROS5
2020 A Fast and Robust Place Recognition Approach for Stereo Visual Odometry Using LiDAR Descriptors
abstract
Place recognition is a core component of Simultaneous Localization and Mapping (SLAM) algorithms. Particularly in visual SLAM systems, previously-visited places are recognized by measuring the appearance similarity between images representing these locations. However, such approaches are sensitive to visual appearance change and also can be computationally expensive. In this paper, we propose an alternative approach adapting LiDAR descriptors for 3D points obtained from stereo-visual odometry for place recognition. 3D points are potentially more reliable than 2D visual cues (e.g., 2D features) against environmental changes (e.g., variable illumination) and this may benefit visual SLAM systems in long-term deployment scenarios. Stereo-visual odometry generates 3D points with an absolute scale, which enables us to use LiDAR descriptors for place recognition with high computational efficiency. Through extensive evaluations on standard benchmark datasets, we demonstrate the accuracy, efficiency, and robustness of using 3D points for place recognition over 2D methods.
Jiawei Mo, Junaed Sattar
IROS1
2019 Extending Monocular Visual Odometry to Stereo Camera Systems by Scale optimization
abstract
This paper proposes a novel approach for extending monocular visual odometry to a stereo camera system. The proposed method uses an additional camera to accurately estimate and optimize the scale of the monocular visual odometry, rather than triangulating 3D points from stereo matching. Specifically, the 3D points generated by the monocular visual odometry are projected onto the other camera of the stereo pair, and the scale is recovered and optimized by directly minimizing the photometric error. It is computationally efficient, adding minimal overhead to the stereo vision system compared to straightforward stereo matching, and is robust to repetitive texture. Additionally, direct scale optimization enables stereo visual odometry to be purely based on the direct method. Extensive evaluation on public datasets (e.g., KITTI), and outdoor environments (both terrestrial and underwater) demonstrates the accuracy and efficiency of a stereo visual odometry approach extended by scale optimization, and its robustness in environments with challenging textures.
Jiawei Mo, Junaed Sattar
IROS1
2018 A Secure AODV Protocol Improvement Scheme Based on Fuzzy Neural Network
Tongyi Xie, Jiawei Mo, Baohua Huang
SecureComm (2)2
2018 Improving Security and Stability of AODV with Fuzzy Neural Network in VANET
Baohua Huang, Jiawei Mo, Xiaolu Cheng
WASA2