Junyi Ma

dblp:163/2390 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1
YearPublicationVenuePosition
2026 MADiff: Motion-Aware Mamba Diffusion Models for Hand Trajectory Prediction on Egocentric Videos
abstract
Understanding human intentions and actions through egocentric videos is important on the path to embodied artificial intelligence. As a branch of egocentric vision techniques, hand trajectory prediction plays a vital role in comprehending human motion patterns, benefiting downstream tasks in extended reality and robot manipulation. However, capturing high-level human intentions consistent with reasonable temporal causality is challenging when only egocentric videos are available. This difficulty is exacerbated under camera egomotion interference and the absence of affordance labels to explicitly guide the optimization of hand waypoint distribution. In this work, we propose a novel hand trajectory prediction method dubbed MADiff, which forecasts future hand waypoints with diffusion models. The devised denoising operation in the latent space is achieved by our proposed motion-aware Mamba, where the camera wearer's egomotion is integrated to achieve motion-driven selective scan (MDSS). To discern the relationship between hands and scenarios without explicit affordance supervision, we leverage a foundation model that fuses visual and language features to capture high-level semantics from video clips. Comprehensive experiments conducted on five public datasets with the existing and our new evaluation metrics demonstrate that MADiff predicts comparably reasonable hand trajectories compared to the state-of-the-art baselines.
Junyi Ma, Xieyuanli Chen, Wentao Bao, Hesheng Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Spatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecasting
abstract
The task of occupancy forecasting (OCF) involves utilizing past and present perception data to predict future occupancy states of autonomous vehicle surrounding environments, which is critical for downstream tasks such as obstacle avoidance and path planning. Existing 3D OCF approaches struggle to predict plausible spatial details for movable objects and suffer from slow inference speeds due to neglecting the bias and uneven distribution of changing occupancy states in both space and time. In this paper, we propose a novel spatiotemporal decoupling vision-based paradigm to explicitly tackle the bias and achieve both effective and efficient 3D OCF. To tackle spatial bias in empty areas, we introduce a novel spatial representation that decouples the conventional dense 3D format into 2D bird’s-eye view (BEV) occupancy with corresponding height values, enabling 3D OCF derived only from 2D predictions thus enhancing efficiency. To reduce temporal bias on static voxels, we design temporal decoupling to improve end-to-end OCF by temporally associating instances via predicted flows. We develop an efficient multi-head network EfficientOCF to achieve 3D OCF with our devised spatiotemporally decoupled representation. A new metric, conditional IoU (C-IoU), is also introduced to provide a robust 3D OCF performance assessment, especially in datasets with missing or incomplete annotations. The experimental results demonstrate that EfficientOCF surpasses existing baseline methods on accuracy and efficiency, achieving state-of-the-art performance with a fast inference time of 82.33 ms with a single GPU. Our code is released at: https://github.com/BIT-XJY/EfficientOCF.
Xieyuanli Chen, Junyi Ma, Jintao Xu 0001, Yue Wang 0020, Ling Pei
CVPR3
2025 Novel Diffusion Models for Multimodal 3D Hand Trajectory Prediction
abstract
Predicting hand motion is critical for understanding human intentions and bridging the action space between human movements and robot manipulations. Existing hand trajectory prediction (HTP) methods forecast the future hand waypoints in 3D space conditioned on past egocentric observations. However, such models are only designed to accommodate 2D egocentric video inputs. There is a lack of awareness of multimodal environmental information from both 2D and 3D observations, hindering the further improvement of 3D HTP performance. In addition, these models overlook the synergy between hand movements and headset camera egomotion, either predicting hand trajectories in isolation or encoding egomotion only from past frames. To address these limitations, we propose novel diffusion models (MMTwin) for multimodal 3D hand trajectory prediction. MMTwin is designed to absorb multi-modal information as input encompassing 2D RGB images, 3D point clouds, past hand waypoints, and text prompt. Besides, two latent diffusion models, the egomotion diffusion and the HTP diffusion as twins, are integrated into MMTwin to predict camera egomotion and future hand trajectories concurrently. We propose a novel hybrid Mamba-Transformer module as the denoising model of the HTP diffusion to better fuse multimodal features. The experimental results on three publicly available datasets and our self-recorded data demonstrate that our proposed MMTwin can predict plausible future 3D hand trajectories compared to the state-of-the-art baselines, and generalizes well to unseen environments. The code and pretrained models will be released at https://github.com/IRMVLab/MMTwin.
Junyi Ma, Wentao Bao, Guanzhong Sun, Xieyuanli Chen, Hesheng Wang 0001
IROS1
2025 Diff-IP2D: Diffusion-Based Hand-Object Interaction Prediction on Egocentric Videos
abstract
Understanding how humans would behave during hand-object interaction (HOI) is vital for applications in service robot manipulation and extended reality. To achieve this, some recent works simultaneously forecast hand trajectories and object affordances on human egocentric videos. The joint prediction serves as a comprehensive representation of future HOI in 2D space, indicating potential human motion and motivation. However, the existing approaches mostly adopt the autoregressive paradigm, which lacks bidirectional constraints within the holistic future sequence, and accumulates errors along the time axis. Meanwhile, they overlook the effect of camera egomotion on first-person view predictions. To address these limitations, we propose a novel diffusion-based HOI prediction method, namely Diff-IP2D, to forecast future hand trajectories and object affordances with bidirectional constraints in an iterative non-autoregressive manner on egocentric videos. Motion features are further integrated into the conditional denoising process to enable Diff-IP2D aware of the camera wearer’s dynamics for more accurate interaction prediction. Extensive experiments demonstrate that Diff-IP2D significantly outperforms the state-of-the-art baselines on both the off-the-shelf and our newly proposed evaluation metrics. This highlights the efficacy of leveraging a generative paradigm for 2D HOI prediction. The code and the video have been released at https://github.com/IRMVLab/Diff-IP2D.
Junyi Ma, Xieyuanli Chen, Hesheng Wang 0001
IROS1
2025 Improved 2D Hand Trajectory Prediction with Multi-View Consistency
abstract
Forecasting how human hands would move around target objects on egocentric videos can provide prior knowledge to enhance the path planning capabilities of service robots and assistive wearable devices. During the hand-object interaction process, head movements always occur concurrently to provide observations for the interaction scene from different egocentric views. Although some prior works have successfully integrated head motion information into hand trajectory prediction (HTP), they basically overlook the multi-view consistency (MVC) inherent in headset camera egomotion. We argue that multi-view consistency reveals geometric and semantic relationships during hand-object interaction, and can be regarded as additional supervision signals for predicting more realistic hand trajectories. Therefore, in this work, we propose a novel learning scheme dubbed EER to improve diffusion-based 2D hand trajectory prediction methods, which involves exploiting the geometric consistency, enhancing the multi-canvas consistency, and reconstructing the semantic consistency inherent in MVC. The experimental results show that our proposed EER scheme significantly improves the prediction accuracy of existing diffusion-based 2D HTP methods on the publicly available datasets. We will release the code as open-source at https://github.com/IRMVLab/EER-HTP.
Junyi Ma, Erhang Zhang, Xieyuanli Chen, Hesheng Wang 0001
IROS1
2025 GSPR: Multimodal Place Recognition Using 3D Gaussian Splatting for Autonomous Driving
abstract
Place recognition is a crucial component that enables autonomous vehicles to obtain localization results in GPS-denied environments. In recent years, multimodal place recognition methods have gained increasing attention. They overcome the weaknesses of unimodal sensor systems by leveraging complementary information from different modalities. Most existing methods explore cross-modality correlations through feature-level or descriptor-level fusion. Conversely, the recently proposed 3D Gaussian Splatting provides a new perspective on multimodal spatio-temporal fusion by harmonizing temporally continuous multimodal data into an explicit scene representation. In this paper, we propose a 3D Gaussian Splatting-based multimodal place recognition network dubbed GSPR. It explicitly combines multi-view RGB images and LiDAR point clouds into a spatio-temporally unified scene representation with the proposed Multimodal Gaussian Splatting. A network composed of 3D graph convolution and transformer is designed to extract global descriptors from the Gaussian scenes for place recognition. Extensive evaluations on three datasets demonstrate that our method can effectively leverage complementary strengths of both multi-view cameras and LiDAR, achieving SOTA place recognition performance while maintaining solid generalization ability. Our open-source code will be released at https://github.com/QiZS-BIT/GSPR.
Zhangshuo Qi, Junyi Ma, Luqi Cheng, Guangming Xiong
IROS2
2025 Zero-Shot Temporal Interaction Localization for Egocentric Videos
abstract
Locating human-object interaction (HOI) actions within video serves as the foundation for multiple downstream tasks, such as human behavior analysis and human-robot skill transfer. Current temporal action localization methods typically rely on annotated action and object categories of interactions for optimization, which leads to domain bias and low deployment efficiency. Although some recent works have achieved zero-shot temporal action localization (ZS-TAL) with large vision-language models (VLMs), their coarse-grained estimations and open-loop pipelines hinder further performance improvements for temporal interaction localization (TIL). To address these issues, we propose a novel zero-shot TIL approach dubbed EgoLoc to locate the timings of grasp actions for human-object interaction in egocentric videos. EgoLoc introduces a self-adaptive sampling strategy to generate reasonable visual prompts for VLM reasoning. By absorbing both 2D and 3D observations, it directly samples high-quality initial guesses around the possible contact/separation timestamps of HOI according to 3D hand velocities, leading to high inference accuracy and efficiency. In addition, EgoLoc generates closed-loop feedback from visual and dynamic cues to further refine the localization results. Comprehensive experiments on the publicly available dataset and our newly proposed benchmark demonstrate that EgoLoc achieves better temporal interaction localization for egocentric videos compared to state-of-the-art baselines. We will release our code and relevant data as open-source at https://github.com/IRMVLab/EgoLoc.
Erhang Zhang, Junyi Ma, Yin-Dong Zheng, Yixuan Zhou 0003
IROS2
2025 A three-tiered semi supervised MTL mechanism and its application in dating apps
abstract
Abstract A thorough understanding of the purpose of dating applications is crucial for service providers in order to optimize the design and user experience of the application. Despite the fact that many APPs prompt users to provide their usage purpose, many do not reveal this attribute. In this study, a three-module framework with semi-supervised and multitask learning mechanisms is proposed (T-SSMTL). Using the T-SSMTL mechanism, the purpose of the dating APP usage can be automatically inferred from the publicly available heterogeneous data of the user. The heterogeneous feature extraction module employs a number of techniques to extract semantic representations, maximizing the use of heterogeneous dating APP data. The multi-task module extracts task-specific knowledge for learning and solves the classification problem involving multiple labels. To alleviate the problem of label insufficiency, the semi-supervised module utilizes a large quantity of unlabeled data generated by users who do not report their usage purpose. A large-scale dataset containing 34,364 active dating APP users with their self-reported usage purpose, portrait image, profile, and posts was collected to evaluate the T-SSMTL framework. In the context of this dataset, simulation experiments have confirmed the efficacy of all three modules of the T-SSMTL framework, demonstrating its substantial theoretical significance as well as its excellent application value.
Junyi Ma, Yasha Wang, Xuanliang Wang, Jiangtao Wang 0001, Junfeng Zhao 0001
Neural Comput. Appl.1
2025 GTCFN: A Graph-Based Transformer and Convolution Fusion Network for Hyperspectral Image Classification
abstract
Graph Neural Networks (GNN) are capable of modeling complex non-Euclidean structures through information transfer, and thus have been party widely used in the field of Hyperspectral Image (HSI) classification. However, conventional GNNs often have difficulty in handling regular grid data, which in turn loses positional information or spatial coherence, as well as in capturing long-range dependencies, which affects their performance in heterogeneous and limited-sample condition. To address these limitations, this paper proposes a novel Graph-based Transformer and Convolution Fusion Network (GTCFN) that integrates the local representation power of Convolutional Neural Networks (CNNs) with the global reasoning capability of graph-based Transformers. GTCFN consists of two synergistic branches: a Graph Transformer sub-network (GTsN) that models high-level semantic structures among superpixels via attention-based topology learning, and a Spectral–Spatial Convolutional sub-network (S2CsN) that extracts multi-scale fine-grained features using 5×5, 7×7, and 9×9 convolutional kernels. To enhance efficiency and generalization, GTCFN incorporates kernelized attention with random feature mapping, reducing the complexity fromO(M2) toO(M). At the same time, attention oversmoothing is avoided by introducing a Gumbel-based multi-head random aggregation mechanism. Experiments conducted on four benchmark datasets, namely Indian Pines, Pavia University, Salinas and WHU-Hi-HongHu, show that GTCFN achieves state-of-the-art performance with OA of 95.62%, 98.34%, 97.88% and 96.69%, which is significantly better than 12 other algorithms, such as CNNs, graph-based models and hybrid network models. The core code for GTCFN is posted on https://github.com/ Majunyi310321/GTCFN.
Junyi Ma, Yao Ding 0010, Jie Feng 0003
IEEE Trans. Geosci. Remote. Sens.2
2024 Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications
abstract
Understanding how the surrounding environment changes is crucial for performing downstream tasks safely and reliably in autonomous driving applications. Recent occupancy estimation techniques using only camera images as input can provide dense occupancy representations of large-scale scenes based on the current observation. However, they are mostly limited to representing the current 3D space and do not consider the future state of surrounding objects along the time axis. To extend camera-only occupancy estimation into spatiotemporal prediction, we propose Cam4DOcc, a new benchmark for camera-only 4D occupancy forecasting, evaluating the surrounding scene changes in a near future. We build our benchmark based on multiple publicly available datasets, including nuScenes, nuScenes-Occupancy, and Lyft-Level5, which provides sequential occupancy states of general movable and static objects, as well as their 3D backward centripetal flow. To establish this benchmark for future research with comprehensive comparisons, we introduce four baseline types from diverse camera-based perception and prediction implementations, including a static-world occupancy model, voxelization of point cloud prediction, 2D-3D instance-based prediction, and our proposed novel end-to-end 4D occupancy forecasting network. Furthermore, the standardized evaluation protocol for preset multiple tasks is also provided to compare the performance of all the proposed baselines on present and future occupancy estimation with respect to objects of interest in autonomous driving scenarios. The dataset and our implementation of all four baselines in the proposed Cam4DOcc benchmark are released as open source at https://github.com/haomo-ai/Cam4DOcc.
Junyi Ma, Xieyuanli Chen, Jintao Xu 0001, Weihao Gu, Rui Ai 0001, Hesheng Wang 0001
CVPR1
2024 Explicit Interaction for Fusion-Based Place Recognition
abstract
Fusion-based place recognition is an emerging technique jointly utilizing multi-modal perception data, to recognize previously visited places in GPS-denied scenarios for robots and autonomous vehicles. Recent fusion-based place recognition methods combine multi-modal features in implicit manners. While achieving remarkable results, they do not explicitly consider what the individual modality affords in the fusion system. Therefore, the benefit of multi-modal feature fusion may not be fully explored. In this paper, we propose a novel fusion-based network, dubbed EINet, to achieve explicit interaction of the two modalities. EINet uses LiDAR ranges to supervise more robust vision features for long time spans, and simultaneously uses camera RGB data to improve the discrimination of LiDAR point clouds. In addition, we develop a new benchmark for the place recognition task based on the nuScenes dataset. To establish this benchmark for future research with comprehensive comparisons, we introduce both supervised and self-supervised training schemes alongside evaluation protocols. We conduct extensive experiments on the proposed benchmark, and the experimental results show that our EINet exhibits better recognition performance as well as solid generalization ability compared to the state-of-the-art fusion-based place recognition approaches. Our open-source code and benchmark are released at: https://github.com/BIT-XJY/EINet.
Junyi Ma, Qi Wu 0007, Yue Wang 0020, Xieyuanli Chen, Wenxian Yu, Ling Pei
IROS2
2024 CVTNet: A Cross-View Transformer Network for LiDAR-Based Place Recognition in Autonomous Driving Environments
abstract
LiDAR-based place recognition (LPR) is one of the most crucial components of autonomous vehicles to identify previously visited places in GPS-denied environments. Most existing LPR methods use mundane representations of the input point cloud without considering different views, which may not fully exploit the information from LiDAR sensors. In this article, we propose across-viewtransformer-based network, dubbed CVTNet, to fuse the range image views and bird's eye views generated from the LiDAR data. It extracts correlations within the views using intratransformers and between the two different views using intertransformers. Based on that, our proposed CVTNet generates a yaw-angle-invariant global descriptor for each laser scan end-to-end online and retrieves previously seen places by descriptor matching between the current query scan and the prebuilt database. We evaluate our approach on three datasets collected with different sensor setups and environmental conditions. The experimental results show that our method outperforms the state-of-the-art LPR methods with strong robustness to viewpoint changes and long-time spans. Furthermore, our approach has better real-time performance that can run faster than the typical LiDAR frame rate does.
Junyi Ma, Guangming Xiong, Xieyuanli Chen
IEEE Trans. Ind. Informatics1
2023 An Adaptive Fusion Risk-Zone Detection Network and its Application
abstract
COVID-19 has caused a pandemic and adverse effects in many fields on a global scale. The city scale quarantine has demonstrated its effectiveness in controlling the epidemic. Conversely, it is costly and risky in inducing economic and social challenges. A compromised solution is to place quarantine measures at high-risk zones on a local scale. Therefore, it is important to investigate risk zones for conducting cost insensitive precautionary measures. The urban data depict the characteristics of different city zones, which offers an opportunity for detecting the high-risk zones. Yet, the high noise-to-signal ratio requires an efficient procedure to rule out irrelevant information in the informative raw urban data and adapt to the risk detection task. In this paper, we propose an Adaptive Fusion Risk-zone Detection Network (AFRDN), which fuses the static and dynamic multi-sourced urban data in an adaptive manner. Specifically, AFRDN first extracts diverse information-rich features from raw urban data with various encoders in the embedding learning module. Then, the AFRDN takes a hierarchical late fusion strategy by fusing the static embedding and the attentive hidden state of dynamic features in the deep latent space. To capture the most relevant information for risk-zone detection, the AFRDN adapts each dimension in the fused embedding with multi-head self-attention blocks. We have collected a real-world dataset including six Chinese cities and conducted extensive experiments to evaluate our framework. Simulation experiments and comparative analysis results show that the AFRDN is effective and feasible for early detection of infectious diseases high-risk zones.
Junyi Ma, Xuanliang Wang, Yasha Wang, Junfeng Zhao 0001
Int. J. Pattern Recognit. Artif. Intell.1
2022 A Comparative Study of Deep Reinforcement Learning-based Transferable Energy Management Strategies for Hybrid Electric Vehicles
abstract
The deep reinforcement learning-based energy management strategies (EMS) have become a promising solution for hybrid electric vehicles (HEVs). When driving cycles are changed, the neural network will be retrained, which is a time-consuming and laborious task. A more efficient way of choosing EMS is to combine deep reinforcement learning (DRL) with transfer learning, which can transfer knowledge of one domain to the other new domain, making the network of the new domain reach convergence values quickly. Different exploration methods of DRL, including adding action space noise and parameter space noise, are compared against each other in the transfer learning process in this work. Results indicate that the network added parameter space noise is more stable and faster convergent than the others. In conclusion, the best exploration method for transferable EMS is to add noise in the parameter space, while the combination of action space noise and parameter space noise generally performs poorly. Our code is available at https://github.com/BIT-XJY/RL-based-Transferable-EMS.git.
Junyi Ma, Qi Liu 0020, Yanan Zhao 0004
IV4
2020 Understanding and Predicting the Burst of Burnout via Social Media
abstract
Job burnout is a special type of work-related stress that is prevalent in our modern society, and constant burnout is extremely harmful for people's physical health and emotional wellbeing. Traditional studies for burnout mainly rely on surveys/questionnaires, which have revealed several interesting findings but are of high cost and very time consuming. With the prevalence of social networking applications, we aim to re-investigate the burnout phenomenon in a novel perspective. In this paper, we collected a dataset consisting of 1532 burnout Weibo users with their postings. Based on the previous literature, we propose a number of hypotheses about what might be the burst signal of the burnout from the perspective of language, time and interaction. Furthermore, extensive correlation analysis is conducted to investigate if these hypotheses are supported, which leads to a number of interesting findings. Finally, we develop machine learning models to predict the burst of burnout based on extracted features and achieve a relatively high accuracy, which reveals potential implications in early-stage intervention.
Jue Wu, Junyi Ma, Yasha Wang, Jiangtao Wang 0001
Proc. ACM Hum. Comput. Interact.2
2017 RT-Fall: A Real-Time and Contactless Fall Detection System with Commodity WiFi Devices
abstract
This paper presents the design and implementation of RT-Fall, a real-time, contactless, low-cost yet accurate indoor fall detection system using the commodity WiFi devices. RT-Fall exploits the phase and amplitude of the fine-grained Channel State Information (CSI) accessible in commodity WiFi devices, and for the first time fulfills the goal of segmenting and detecting the falls automatically in real-time, which allows users to perform daily activities naturally and continuously without wearing any devices on the body. This work makes two key technical contributions. First, we find that the CSI phase difference over two antennas is a more sensitive base signal than amplitude for activity recognition, which can enable very reliable segmentation of fall and fall-like activities. Second, we discover the sharp power profile decline pattern of the fall in the time-frequency domain and further exploit the insight for new feature extraction and accurate fall segmentation/detection. Experimental results in four indoor scenarios demonstrate that RT-fall consistently outperforms the state-of-the-art approach WiFall with 14 percent higher sensitivity and 10 percent higher specificity on average.
Hao Wang 0035, Daqing Zhang 0001, Yasha Wang, Junyi Ma, Shengjie Li 0001
IEEE Trans. Mob. Comput.4
2016 Human respiration detection with commodity wifi devices: do user location and body orientation matter?
abstract
Recent research has demonstrated the feasibility of detecting human respiration rate non-intrusively leveraging commodity WiFi devices. However, is it always possible to sense human respiration no matter where the subject stays and faces? What affects human respiration sensing and what's the theory behind? In this paper, we first introduce the Fresnel model in free space, then verify the Fresnel model for WiFi radio propagation in indoor environment. Leveraging the Fresnel model and WiFi radio propagation properties derived, we investigate the impact of human respiration on the receiving RF signals and develop the theory to relate one's breathing depth, location and orientation to the detectability of respiration. With the developed theory, not only when and why human respiration is detectable using WiFi devices become clear, it also sheds lights on understanding the physical limit and foundation of WiFi-based sensing systems. Intensive evaluations validate the developed theory and case studies demonstrate how to apply the theory to the respiration monitoring system design.
Hao Wang 0035, Daqing Zhang 0001, Junyi Ma, Yasha Wang, Dan Wu 0007, Tao Gu 0001
UbiComp3
2015 Anti-fall: A Non-intrusive and Real-Time Fall Detector Leveraging CSI from Commodity WiFi Devices
Daqing Zhang 0001, Hao Wang 0035, Yasha Wang, Junyi Ma
ICOST4