Jiaxu Kang

dblp:357/5908 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0006-2135-9532ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Towards Ultrasound-based Reliable Disease Diagnosis Using Causal Inference
abstract
Aligning the decision-making process of deep learning models with that of experienced sonographers is essential for ultrasound-based reliable disease diagnosis. Although existing methods have made significant progress in this aspect, their alignments are primarily associational rather than causal, leading to pseudo-correlations between features and diagnostic results. Such a biased diagnosis blindly models the sonographer's diagnostic skills and attention to specific patterns, which we argue hardly produces an AI diagnoser that is comparable to human experts. To address this issue, we propose a causality-based diagnostic framework to align the model's diagnostic behaviors with those of experts. Specifically, by delving into both conspicuous and inconspicuous confounders within the ultrasound images, the back-door and front-door adjustment causal learning modules are proposed to promote unbiased learning by mitigating potential pseudo-correlations. In addition, we integrate causal inference into a well-designed dual-branch model with feature interaction bridges for compatibility with multimodal ultrasound inputs. To fully evaluate our method, we conduct comparative studies on different diseases and ultrasound modalities. In particular, we publish a carefully constructed multimodal ultrasound dataset for breast lesion diagnosis and segmentation. Sufficient comparative and ablation studies on this dataset emphasize that our method outperforms state-of-the-art methods.
Bolei Chen, Jiaxu Kang, Haonan Yang 0001, Ping Zhong 0002, Yixiong Liang, Rui Fan 0001, Jianxin Wang 0001
AAAI2
2026 Perspective from a Broader Context: Can Room Style Knowledge Help Visual Floorplan Localization?
abstract
Since a building's floorplan remains consistent over time and is inherently robust to changes in visual appearance, visual Floorplan Localization (FLoc) has received increasing attention from researchers. However, as a compact and minimalist representation of the building's layout, floorplans contain many repetitive structures (e.g., hallways and corners), thus easily result in ambiguous localization. Existing methods either pin their hopes on matching 2D structural cues in floorplans or rely on 3D geometry-constrained visual pre-trainings, ignoring the richer contextual information provided by visual images. In this paper, we suggest using broader visual scene context to empower FLoc algorithms with scene layout priors to eliminate localization uncertainty. In particular, we propose an unsupervised learning technique with clustering constraints to pre-train a room discriminator on self-collected unlabeled room images. Such a discriminator can empirically extract the hidden room type of the observed image and distinguish it from other room types. By injecting the scene context information summarized by the discriminator into an FLoc algorithm, the room style knowledge is effectively exploited to guide definite visual FLoc. We conducted sufficient comparative studies on two standard visual Floc benchmarks. Our experiments show that our approach outperforms state-of-the-art methods and achieves significant improvements in robustness and accuracy.
Bolei Chen, Shengsheng Yan, Yongzheng Cui, Jiaxu Kang, Ping Zhong 0002, Jianxin Wang 0001
AAAI4
2026 Treasure Hunting: Embodied Contrastive Learning-Enhanced Coarse-to-Fine Object Seeking With Explorer and Discriminator Cooperation
abstract
Object navigation (ObjcetNav), which enables an agent to seek any instance of an object category, has shown great advances. However, current agents are built upon occlusion-prone visual observations or compressed 2-D maps, which hinder their embodied perception of 3-D scene geometry. Furthermore, existing methods usually decouple ObjectNav into the exploration and exploitation subtasks, easily leading to ambiguous object localization and blind exploration. To address these issues, we first propose an embodied contrastive learning (ECL) method with geometric consistency (GC) and behavioral awareness (BA), which motivates agents to encode 3-D scene layouts and semantic cues actively. The BA is modeled by predicting navigational actions based on multiframe visual images, as behaviors causing differences between adjacent visual sensations are crucial for learning correlations among continuous visions. The GC is modeled by aligning the behavior-aware visual stimulus with 3-D semantic shapes through unsupervised contrastive learning. Then, based on the above ECL pretraining, a coarse-to-fine ObjectNav policy with explorer and discriminator cooperation is proposed, inspired by the treasure-hunting mindset. Concretely, the explorer is designed to adaptively switch the action spaces, thereby switching the global and local exploration thoughts according to the accumulated scene priors. The discriminator is designed to discriminate the target's authenticity using behavior-aware visual features and geometric invariance priors, which permits mimicking the human behavior of "approaching to confirm" when distinguishing objects from a distance. As expected, our ECL method performs well on object detection (ObjDet) and instance segmentation (InstSeg) tasks. Our ECL-enhanced ObjectNav strategy outperforms state-of-the-art (SOTA) methods on Matterport3D (MP3D), Gibson, and HM3D datasets.
Bolei Chen, Jiaxu Kang, Haonan Yang 0001, Ping Zhong 0002, Rui Fan 0001, Jianxin Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Vehicle-Arm Coordination-Based Active 3D Reconstruction using Deep Reinforcement Learning
abstract
Active 3D reconstruction has a wide range of applications, including augmented/virtual reality and various robotics tasks such as navigation, object recognition and manipulation, and scene perception. Most existing approaches rely on offline images or online streams captured by human-operated cameras, requiring significant manual effort and depending on the operator’s expertise. Although some methods have explored using robots to assist data collection by planning the Next-Best-View and then moving the camera to cover these viewpoints, they are often constrained by fixed camera perspectives and suboptimal robot control strategies. These limitations result in inaccurate camera positioning and inaccessible viewpoints, leading to ineffective data collection. In this paper, we explore the feasibility of using a vehicle-arm collaborative robot to tackle the challenges of active 3D reconstruction. Specifically, we decompose this task into two mutually iterative subtasks: the object-centric planning task, which focuses solely on the objects to generate candidate viewpoints, and the agent-centric interaction task, where the robot moves the camera to cover these viewpoints. To achieve this, we propose a collaborative control framework that integrates the motions of the base and arm, enabling efficient recovery of the 3D shape of objects. Sufficient comparative and ablation studies validate that our approach achieves finer surface reconstruction with reduced reconstruction time.
Liangbai Liu, Ping Zhong 0002, Bolei Chen, Jiaxu Kang, Haonan Yang 0001, Yifei Wang 0006
IJCNN4
2025 Perspective from a Higher Dimension: Can 3D Geometric Priors Help Visual Floorplan Localization?
abstract
Since a building's floorplans are easily accessible, consistent over time, and inherently robust to changes in visual appearance, self-localization within the floorplan has attracted researchers' interest. However, since floorplans are minimalist representations of a building's structure, modal and geometric differences between visual perceptions and floorplans pose challenges to this task. While existing methods cleverly utilize 2D geometric features and pose filters to achieve promising performance, they fail to address the localization errors caused by frequent visual changes and view occlusions due to variously shaped 3D objects. To tackle these issues, this paper views the 2D Floorplan Localization (FLoc) problem from a higher dimension by injecting 3D geometric priors into the visual FLoc algorithm. For the 3D geometric prior modeling, we first model geometrically aware view invariance using multi-view constraints, i.e., leveraging imaging geometric principles to provide matching constraints between multiple images that see the same points. Then, we further model the view-scene aligned geometric priors, enhancing the cross-modal geometry-color correspondences by associating the scene's surface reconstruction with the RGB frames of the sequence. Both 3D priors are modeled through self-supervised contrastive learning, thus no additional geometric or semantic annotations are required. These 3D priors summarized in extensive realistic scenes bridge the modal gap while improving localization success without increasing the computational burden on the FLoc algorithm. Sufficient comparative studies demonstrate that our method significantly outperforms state-of-the-art methods and substantially boosts the FLoc accuracy.
Bolei Chen, Jiaxu Kang, Haonan Yang 0001, Ping Zhong 0002, Jianxin Wang 0001
ACM Multimedia2
2025 Unbiased Embodied Visual Representation Learning with Causal Inference and Cross-Modality Alignment
abstract
Object Goal Navigation (ObjectNav) in novel environments relies on comprehensive scene understanding, including precise visual perception and accurate modeling of spatial-semantic regularities. However, excessive attention to the hand-crafted scene representation in prevailing approaches leads to the neglect of the negative influence of the perception bias hidden in the visual observations. The hand-crafted semantic distribution in domestic environments causes the spurious association bias, while the semantic conflict bias arises due to the dynamic perspective changes. Biased visual perception significantly limits the generalization of the navigation strategy. In this article, we propose the U nbiased E mbodied V isual R epresentation ( UEVR ), which overcomes the perception biases using causal inference and cross-modality alignment. Specifically, we establish reasonable assumptions about confounders for multi-object features through our proposed Unbiased Causal R-CNN framework and eliminate the spurious associations bias through B ackdoor I ntervention C ausal A djustment ( BICA ) module during navigation. To overcome the dynamic-view bias hidden in 2D image features, we propose to employ the cross-modality alignment mechanism with the Geometric Consistency ( GeoCon ) to encode 3D geometry prior into the 2D representations. Finally, we design a modular ObjectNav framework integrated with UEVR named Causal-ObjectNav , which consists of the corner-based scene exploration module and target object discrimination module. Extensive experiments on the MP3D and HM3D datasets demonstrate the superiority of the unbiased navigation model over existing ObjectNav methods.
Jiaxu Kang, Bolei Chen, Ping Zhong 0002, Yifei Wang 0006, Haonan Yang 0001, Yu Sheng
ACM Trans. Multim. Comput. Commun. Appl.1
2024 HSPNav: Hierarchical Scene Prior Learning for Visual Semantic Navigation Towards Real Settings
abstract
Visual Semantic Navigation (VSN) aims at navigating a robot to a given target object in a previously unseen scene. To tackle this task, the robot must learn a nimble navigation policy by utilizing spatial patterns and semantic co-occurrence relations among objects in the scene. Prevailing approaches extract scene priors from the instant visual observations and solidify them in neural episodic memory to achieve flexible navigation. However, due to the oblivion and underuse of the scene priors, these methods are plagued by repeated exploration, effective-knowledge sparsity, and wrong decisions. To alleviate these issues, we propose a novel VSN policy, HSPNav, based on Hierarchical Scene Priors (HSP) and Deep Reinforcement Learning (DRL). The HSP contains two components, i.e., the egocentric semantic map-based Local Scene Priors (LSP) and the commonsense relational graph-based Global Scene Priors (GSP). Then, efficient semantic navigation is achieved by employing an immediate LSP to retrieve conducive contextual memories from the GSP. By utilizing the MP3D dataset, the experimental results in the Habitat simulator demonstrate that our HSP brings a significant boost over the baselines. Furthermore, we take an essential step from simulation to reality by bridging the gap from Habitat to ROS. The migration evaluations show that HSPNav can generalize to realistic settings well and achieve promising performance.
Jiaxu Kang, Bolei Chen, Ping Zhong 0002, Haonan Yang 0001, Yu Sheng, Jianxin Wang 0001
ICRA1
2024 Embodied Contrastive Learning with Geometric Consistency and Behavioral Awareness for Object Navigation
abstract
Object Navigation (ObjcetNav), which enables an agent to seek any instance of an object category specified by a semantic label, has shown great advances. However, current agents are built upon occlusion-prone visual observations or compressed 2D semantic maps, which hinder their embodied perception of 3D scene geometry and easily lead to ambiguous object localization and blind exploration. To address these limitations, we present an Embodied Contrastive Learning (ECL) method with Geometric Consistency (GC) and Behavioral Awareness (BA), which motivates agents to actively encode 3D scene layouts and semantic cues. Driven by our embodied exploration strategy, BA is modeled by predicting navigational actions based on multi-frame visual images, as behaviors that cause differences between adjacent visual sensations are crucial for learning correlations among continuous visions. The GC is modeled as the alignment of behavior-aware visual stimulus with 3D semantic shapes by employing unsupervised contrastive learning. The aligned behavior-aware visual features and geometric invariance priors are injected into a modular ObjectNav framework to enhance object recognition and exploration capabilities. As expected, our ECL method performs well on object detection and instance segmentation tasks. Our ObjectNav strategy outperforms state-of-the-art methods on MP3D and Gibson datasets, showing the potential of our ECL in embodied navigation.
Bolei Chen, Jiaxu Kang, Ping Zhong 0002, Yixiong Liang, Yu Sheng, Jianxin Wang 0001
ACM Multimedia2
2024 Think Holistically, Act Down-to-Earth: A Semantic Navigation Strategy With Continuous Environmental Representation and Multi-Step Forward Planning
abstract
The Object goal Navigation (ObjectNav) task requires an agent to navigate through a previously unknown domestic scenario using spatial and semantic contextual information, where the goal is specified by a semantic label (e.g., find a TV). Such a task is especially challenging as it requires formulating and understanding the complex co-occurrence relations among objects in diverse settings, which is critical for long-sequence navigational decision-making. Existing methods learn to either explicitly represent co-occurrence relationships as discrete semantic priors, or implicitly encode them from raw observations, thus can not benefit from the rich environmental semantics. In this work, we propose a novel Deep Reinforcement Learning (DRL) based ObjectNav strategy by actively imagining spatial and semantic clues outside the agent’s Field of View (FoV) and further mining Continuous Environmental Representations (CER) using self-supervised learning. Additionally, the illusion of spatial and semantic patterns allows the agent to perform Multi-Step Forward-Looking Planning (MSFLP) by considering the temporal evolution of egocentric local observations. Our approach is thoroughly evaluated and ablated in the visually realistic environments of the Matterport3D (MP3D) dataset. The experimental results reflect that our method combining CER and imagination-based MSFLP facilitates learning complicated semantic priors and navigation skills, thus achieving state-of-the-art performance on the ObjectNav task. In addition, adequate quantitative and qualitative analyses validate the excellent generalization ability and superiority of our method.
Bolei Chen, Jiaxu Kang, Ping Zhong 0002, Yongzheng Cui, Siyi Lu, Yixiong Liang, Jianxin Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2