Jiawei Fu 0001

dblp:194/2277-1 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0003-3964-5992ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CNDM: Customized Noise Diffusion Model for Trajectory Prediction
abstract
Precise and reliable multi-agent trajectory prediction is fundamental to enabling safe autonomous navigation. While denoising diffusion models have demonstrated remarkable capability in capturing the multimodal uncertainty inherent in this task, their reliance on a generic, isotropic Gaussian noise prior poses a critical limitation. This uniform prior is fundamentally misaligned with the highly structured, heterogeneous uncertainty of agent motion—shaped by dynamics, type, and environmental constraints—forcing the denoising network to implicitly relearn complex motion priors from scratch. To bridge this gap, we propose the Customized Noise Diffusion Model (CNDM), a novel framework that introduces a learned, agent-specific noise prior. At the core of CNDM is a Prior-Guidance Network (PGN) that distills an agent’s history, type, and scene context into a parametric, anisotropic Gaussian distribution. This customized prior provides a physically-grounded starting point for the diffusion process. To enable efficient training on such anisotropic noise, we leverage a Mahalanobis whitening transformation to standardize the denoising task. Extensive experiments on the Waymo Open Motion and Argoverse 2 datasets show that CNDM achieves competitive performance against state-of-the-art methods, excelling particularly in capturing diverse motion patterns and improving probabilistic calibration. Ablation studies confirm that the performance gains are directly attributable to our customized noise design, underscoring the importance of integrating structured domain knowledge into the generative foundation of diffusion models.
Entao Chang, Jiawei Fu 0001, Wenjie Gao 0001, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.2
2025 Hybrid Reciprocal Transformer with Triplet Feature Alignment for Scene Graph Generation
abstract
Scene graph generation is a pivotal task in computer vision, focusing on comprehensive identification of visual relation tuples embedded within images. The advancement of methods involving triplets has sought to enhance task performance by integrating triplets as contextual features for more precise predicate identification from component level. However, challenges remain due to interference from multi-role objects in overlapping tuples within complex environments, which impairs the model’s ability to distinguish and align specific triplet features for reasoning diverse semantics of multi-role objects. To address these issues, we introduce a novel framework that incorporates a triplet alignment model into a hybrid reciprocal transformer architecture, starting from using triplet mask features to guide the learning of component-level relation graphs. To effectively distinguish multi-role objects characterized by overlapping visual relation tuples, we introduce a triplet alignment loss, which provides multi-role objects with aligned features from triplet and helps customize them. Additionally, we explore the inherent connectivity between hybrid aligned triplet and component features through a bidirectional refinement module, which enhances feature interaction and reciprocal reinforcement. Experimental results demonstrate that our model achieves state-of-the-art performance on the Visual Genome and Action Genome datasets, underscoring its effectiveness and adaptability. Project page: hq-sg.github.io.
Jiawei Fu 0001, Kai Chen 0028, Qi Dou 0001
CVPR1
2025 SAMap: Semantic Alignment for HD Map Detection Domain Generalization Under Varying Weather and Lighting
abstract
High-definition (HD) maps are crucial for autonomous driving systems. Despite recent advances in learning-based HD map prediction methods, these approaches experience significant performance degradation when encountering unseen weather or lighting conditions due to feature distribution discrepancies (domain gaps) of input images. To address this issue, we propose SAMap, a novel map learning framework that enhances domain generalization capabilities of existing models by reducing domain discrepancies in input images. SAMap innovatively introduces a Semantic Aligner, an image-to-image transformation module that aligns images from different domains into a unified domain space while preserving semantic consistency. To train this aligner, we leverage Vision-Language Models (VLMs) that have acquired image-text alignment capabilities. Specifically, we first train a Prompt Learner that combines handcrafted and learnable prompts to capture domain-invariant semantic information. We then train Semantic Aligner through dual supervision mechanisms: a content preservation loss that maintains feature consistency across transformations and a semantic alignment loss that leverages VLM’s encoders to align transformed images with domain-invariant textual representations. Adequate experiments on the NuScenes dataset demonstrate that when integrated with three existing HD map prediction methods, SAMap achieves a performance improvement of up to 11.6% on unseen domains (rain or night conditions), effectively validating its generalization capabilities across domains.
Wenjie Gao 0001, Haodong Jing, Jiawei Fu 0001, Shi-tao Chen, Nanning Zheng 0001
IROS3
2024 Multi-objective Cross-task Learning via Goal-conditioned GPT-based Decision Transformers for Surgical Robot Task Automation
abstract
Surgical robot task automation has been a promising research topic for improving surgical efficiency and quality. Learning-based methods have been recognized as an interesting paradigm and been increasingly investigated. However, existing approaches encounter difficulties in long-horizon goal-conditioned tasks due to the intricate compositional structure, which requires decision-making for a sequence of sub-steps and understanding of inherent dynamics of goal-reaching tasks. In this paper, we propose a new learning-based framework by leveraging the strong reasoning capability of the GPT-based architecture to automate surgical robotic tasks. The key to our approach is developing a goal-conditioned decision transformer to achieve sequential representations with goal-aware future indicators in order to enhance temporal reasoning. Moreover, considering to exploit a general understanding of dynamics inherent in manipulations, thus making the model’s reasoning ability to be task-agnostic, we also design a cross-task pretraining paradigm that uses multiple training objectives associated with data from diverse tasks. We have conducted extensive experiments on 10 tasks using the surgical robot learning simulator SurRoL [1]. The results show that our new approach achieves promising performance and task versatility compared to existing methods. The learned trajectories can be deployed on the da Vinci Research Kit (dVRK) for validating its practicality in real surgical robot settings. Our project website is at: https://med-air.github.io/SurRoL.
Jiawei Fu 0001, Yonghao Long 0001, Kai Chen 0028, Qi Dou 0001
ICRA1
2024 Complementing Onboard Sensors with Satellite Maps: A New Perspective for HD Map Construction
abstract
High-definition (HD) maps play a crucial role in autonomous driving systems. Recent methods have attempted to construct HD maps in real-time using vehicle onboard sensors. Due to the inherent limitations of onboard sensors, which include sensitivity to detection range and susceptibility to occlusion by nearby vehicles, the performance of these methods significantly declines in complex scenarios and long-range detection tasks. In this paper, we explore a new perspective that boosts HD map construction through the use of satellite maps to complement onboard sensors. We initially generate the satellite map tiles for each sample in nuScenes and release a complementary dataset for further research. To enable better integration of satellite maps with existing methods, we propose a hierarchical fusion module, which includes feature-level fusion and BEV-level fusion. The feature-level fusion, composed of a mask generator and a masked cross-attention mechanism, is used to refine the features from onboard sensors. The BEV-level fusion mitigates the coordinate differences between features obtained from onboard sensors and satellite maps through an alignment module. The experimental results on the augmented nuScenes showcase the seamless integration of our module into three existing HD map construction methods. The satellite maps and our proposed module notably enhance their performance in both HD map semantic segmentation and instance detection tasks. Our code will be available at https://github.com/xjtu-csgao/SatforHDMap.
Wenjie Gao 0001, Jiawei Fu 0001, Yanqing Shen, Haodong Jing, Shi-tao Chen, Nanning Zheng 0001
ICRA2
2023 InteractionNet: Joint Planning and Prediction for Autonomous Driving with Transformers
abstract
Planning and prediction are two important modules of autonomous driving and have experienced tremendous advancement recently. Nevertheless, most existing methods regard planning and prediction as independent and ignore the correlation between them, leading to the lack of consideration for interaction and dynamic changes of traffic scenarios. To address this challenge, we propose InteractionNet, which leverages transformer to share global contextual reasoning among all traffic participants to capture interaction and interconnect planning and prediction to achieve joint. Besides, InteractionNet deploys another transformer to help the model pay extra attention to the perceived region containing critical or unseen vehicles. InteractionNet outperforms other baselines in several benchmarks, especially in terms of safety, which benefits from the joint consideration of planning and forecasting. The code will be available at https://github.com/fujiawei0724/InteractionNet.
Jiawei Fu 0001, Yanqing Shen, Zhiqiang Jian, Shi-tao Chen, Jingmin Xin, Nanning Zheng 0001
IROS1