Jiajia Fu

dblp:189/1656 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0005-9923-9904ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
3D vision · 48% Video understanding and tracking · 32% Autonomous driving · 21%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › Video understanding and tracking › object tracking
3d object tracking
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Robotics › Autonomous driving › perception
3d perception
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › Video understanding and tracking
multi-object tracking
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › 3D vision
spatio-temporal alignment
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › 3D vision › 3d object detection
temporal 3d detection
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026

Methods — techniques the papers use, named apart from their topics

multi-hypothesis decoding · 1.0motion model · 1.0attention mechanism · 1.0
YearPublicationVenuePosition
2026 Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception
abstract
Spatio-temporal alignment is crucial for temporal modeling of end-to-end (E2E) perception in autonomous driving (AD), providing valuable structural and textural prior information. Existing methods typically rely on the attention mechanism to align objects across frames, simplifying the motion model with a unified explicit physical model (constant velocity, etc.). These approaches prefer semantic features for implicit alignment, challenging the importance of explicit motion modeling in the traditional perception paradigm. However, variations in motion states and object features across categories and frames render this alignment suboptimal. To address this, we propose HAT, a spatio-temporal alignment module that allows each object to adaptively decode the optimal alignment proposal from multiple hypotheses without direct supervision. Specifically, HAT first utilizes multiple explicit motion models to generate spatial anchors and motion-aware feature proposals for historical instances. It then performs multi-hypothesis decoding by incorporating semantic and motion cues embedded in cached object queries, ultimately providing the optimal alignment proposal for the target frame. On nuScenes, HAT consistently improves 3D temporal detectors and trackers across diverse baselines. It achieves state-of-the-art tracking results with 46.0% AMOTA on the test set when paired with the DETR3D detector. In an object-centric E2E AD method, HAT enhances perception accuracy (+1.3% mAP, +3.1% AMOTA) and reduces the collision rate by 32%. When semantics are corrupted (nuScenes-C), the enhancement of motion modeling by HAT enables more robust perception and planning in the E2E AD.
Peidong Li, Dedong Liu, Jiajia Fu, Dixiao Cui, Lijun Zhao 0003, Lining Sun
AAAI7
2026 Spatially Aware Adaptive Diffusion: Unifying Low-Resolution Image Fusion and Super-Resolution
abstract
Low-resolution visible-infrared image fusion and super-resolution (LRVIF) are critical for enhancing image quality in low-resolution scenarios, yet limited information in the input images often constrains performance. To address these challenges, we propose SaDiff, a spatially-aware adaptive diffusion model that introduces diffusion processes into LRVIF for the first time, representing a major breakthrough in the field. Leveraging the generative capabilities of diffusion models, our approach unifies and enhances image fusion and super-resolution within a cohesive framework. A key component of SaDiff is the Spatial Residual Adaptation Block, which extends the diffusion process by dynamically adapting feature representations to spatial variations in the local regions of the input images. This module maximally preserves crucial information from the input images, such as texture details and contrast, while effectively suppressing noise, ensuring robust and context-aware feature refinement. Then we further propose Direct Diffusion Synthesis, a novel mechanism that utilizes noise predictions during diffusion to generate fused images, enabling joint training of the fusion and super-resolution networks. Additionally, a Cross-Feature Fusion Module integrates texture and contrast details, producing super-resolution fused images with improved clarity and structural integrity. Extensive experiments show that SaDiff achieves state-of-the-art performance, offering a robust and unified solution to infrared-visible image fusion and super-resolution. The code for the proposed method will be made available at https://github.com/guobaoxiao/SaDiff.
Jiajia Fu, Zhenni Yu, Haosheng Chen 0001, Songlin Du, Changcai Yang, Lianghua He, Guobao Xiao
IEEE Trans. Circuits Syst. Video Technol.1
2021 Service Fault Diagnosis Algorithm of Noise Network under 5G Network Slice
abstract
In order to solve the problems of low fault diagnosis accuracy and high false alarm rate in noisy environment, this paper proposes a service fault diagnosis algorithm of noise network under 5G network slice. First, the relationship between the service and the underlying network resources based on the mapping relationship, and fault diagnosis model are established. Secondly, in order to reduce the negative impact of network noise on the fault diagnosis algorithm, the fault propagation model is optimized based on the number of simultaneous faults and the value of the failure rate of each link. Finally, the fault set with the largest interpretation ability is selected as the fault set. In the experimental part, it is verified that the algorithm in this paper improves the accuracy of fault diagnosis and reduces the false alarm rate.
Jiajia Fu, Zanhong Wu, Song Kang
IWCMC1
2021 Service Fault Location Algorithm based on Network Characteristics under 5G Network Slicing
abstract
In order to solve the problem of long diagnosis time for fault diagnosis algorithms in large-scale environments, this paper proposes a service fault location algorithm based on network characteristics. First, based on the virtual network mapping data, the service is associated with the underlying resources to build a Bayesian fault location model. Second, by analyzing the network topology and the running status of each service, the probability of network node failure is calculated based on the multi-attribute characteristics. Finally, the fault location is achieved by calculating the set of suspected faults with the strongest ability to explain abnormal symptoms. The experimental part compares the algorithm of this paper with the classical algorithm, and verifies that the algorithm of this paper improves the accuracy of fault diagnosis and reduces the time cost of fault diagnosis.
Jiangang Lu, Jiajia Fu, Linna Ruan
IWCMC2