Haojun Jiang

dblp:267/5754 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 UltraSeP: Sequence-aware pre-training for echocardiography probe movement guidance
Haojun Jiang, Zhenguo Sun, Yulin Wang 0002, Yu Sun 0020, Meng Li 0087, Shaqi Luo, Shiji Song, Gao Huang 0001
Pattern Recognit.1
2025 EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
abstract
Echocardiography is crucial for cardiovascular disease detection but relies heavily on experienced sonographers. Echocardiography probe guidance systems, which provide real-time movement instructions for acquiring standard plane images, offer a promising solution for AI-assisted or fully autonomous scanning. However, developing effective machine learning models for this task remains challenging, as they must grasp heart anatomy and the intricate interplay between probe motion and visual signals. To address this, we present EchoWorld, a motion-aware world modeling framework for probe guidance that encodes anatomical knowledge and motion-induced visual dynamics, while effectively leveraging past visual-motion sequences to enhance guidance precision. EchoWorld employs a pre-training strategy inspired by world modeling principles, where the model predicts masked anatomical regions and simulates the visual outcomes of probe adjustments. Built upon this pre-trained model, we introduce a motion-aware attention mechanism in the fine-tuning stage that effectively integrates historical visual-motion data, enabling precise and adaptive probe guidance. Trained on more than one million ultrasound images from over 200 routine scans, EchoWorld effectively captures key echocar-diographic knowledge, as validated by qualitative analysis. Moreover, our method significantly reduces guidance errors compared to existing visual backbones and guidance frameworks, excelling in both single-frame and sequential evaluation protocols. Code is available at https://github.com/LeapLabTHU/EchoWorld.
Yulin Wang 0002, Haojun Jiang, Pan Liu 0004, Shiji Song, Gao Huang 0001
CVPR3
2025 An Operator-Centric Framework for Risk-Aware Low-Altitude Urban Security: a UAV-as-a-Service Approach
abstract
The proliferation of Unmanned Aerial Vehicles (UAVs) for urban public safety is critically hindered by operational risks inherent in complex environments. To address these challenges, this paper introduces a “UAV-as-a-Service” (UaaS) paradigm, an application paradigm innovation that leverages the core infrastructure of telecom operators. Our primary contribution is a closed-loop intelligent method, representing an algorithmic innovation, centered on a 3D Dynamic Risk Map (DRM). The DRM is generated in real-time by a Dynamic Bayesian Network (DBN) and explicitly informs both a risk-aware Multi-Agent Reinforcement Learning (MARL) dispatcher and a hybrid Rapidly-exploring Random Tree Star (RRT*)-Model Predictive Control (MPC) path planner. This tight coupling ensures that tactical decisions are grounded in a holistic understanding of risk. Simulation results demonstrate that the proposed UaaS model substantially enhances operational outcomes, reducing the time to achieve critical situational awareness by over 75 % and decreasing firefighter risk exposure by$\mathbf{7 8 \%}$. These technical advancements validate a novel and viable Business-to-Government (B2G) service model, demonstrating significant technologycommercial synergy.
Enwan Zhang, Jianxun Jason Ding, Xingbin Zhan, Sheng Nie, Yutong Xing, Chaolun Wang, Ning Yin, Xiandong Zhang, Haojun Jiang
HPCC11
2025 Cross-modal adapter for vision-language retrieval
Haojun Jiang, Jianke Zhang, Rui Huang 0012, Chunjiang Ge, Zanlin Ni, Shiji Song, Gao Huang 0001
Pattern Recognit.1
2024 Cardiac Copilot: Automatic Probe Guidance for Echocardiography with World Model
Haojun Jiang, Zhenguo Sun, Meng Li 0087, Yu Sun 0020, Shaqi Luo, Shiji Song, Gao Huang 0001
MICCAI (1)1
2024 Joint representation learning for text and 3D point cloud
Rui Huang 0012, Xuran Pan, Henry Zheng, Haojun Jiang, Cheng Wu 0002, Shiji Song, Gao Huang 0001
Pattern Recognit.4
2023 Deep Incubation: Training Large Models by Divide-and-Conquering
abstract
Recent years have witnessed a remarkable success of large deep learning models. However, training these models is challenging due to high computational costs, painfully slow convergence, and overfitting issues. In this paper, we present Deep Incubation, a novel approach that enables the efficient and effective training of large models by dividing them into smaller sub-modules which can be trained separately and assembled seamlessly. A key challenge for implementing this idea is to ensure the compatibility of the independently trained sub-modules. To address this issue, we first introduce a global, shared meta model, which is leveraged to implicitly link all the modules together, and can be designed as an extremely small network with negligible computational overhead. Then we propose a module incubation algorithm, which trains each sub-module to replace the corresponding component of the meta model and accomplish a given learning task. Despite the simplicity, our approach effectively encourages each sub-module to be aware of its role in the target large model, such that the finally-learned sub-modules can collaborate with each other smoothly after being assembled. Empirically, our method can outperform end-to-end (E2E) training in well-established training setting and shows transferable performance gain for downstream tasks (e.g., object detection and image segmentation on COCO and ADE20K). Our code is available at https://github.com/LeapLabTHU/Deep-Incubation.
Zanlin Ni, Yulin Wang 0002, Jiangwei Yu, Haojun Jiang, Yue Cao 0001, Gao Huang 0001
ICCV4
2023 Glance and Focus Networks for Dynamic Visual Recognition
abstract
Spatial redundancy widely exists in visual recognition tasks, i.e., discriminative features in an image or video frame usually correspond to only a subset of pixels, while the remaining regions are irrelevant to the task at hand. Therefore, static models which process all the pixels with an equal amount of computation result in considerable redundancy in terms of time and space consumption. In this paper, we formulate the image recognition problem as a sequential coarse-to-fine feature learning process, mimicking the human visual system. Specifically, the proposed Glance and Focus Network (GFNet) first extracts a quick global representation of the input image at a low resolution scale, and then strategically attends to a series of salient (small) regions to learn finer features. The sequential process naturally facilitates adaptive inference at test time, as it can be terminated once the model is sufficiently confident about its prediction, avoiding further redundant computation. It is worth noting that the problem of locating discriminant regions in our model is formulated as a reinforcement learning task, thus requiring no additional manual annotations other than classification labels. GFNet is general and flexible as it is compatible with any off-the-shelf backbone models (such as MobileNets, EfficientNets and TSM), which can be conveniently deployed as the feature extractor. Extensive experiments on a variety of image classification and video recognition tasks and with various backbone models demonstrate the remarkable efficiency of our method. For example, it reduces the average latency of the highly efficient MobileNet-V3 on an iPhone XS Max by 1.3x without sacrificing accuracy. Code and pre-trained models are available at https://github.com/blackfeather-wang/GFNet-Pytorch.
Gao Huang 0001, Yulin Wang 0002, Kangchen Lv, Haojun Jiang, Shiji Song
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Pseudo-Q: Generating Pseudo Language Queries for Visual Grounding
abstract
Visual grounding, i.e., localizing objects in images ac-cording to natural language queries, is an important topic in visual language understanding. The most effective approaches for this task are based on deep learning, which generally require expensive manually labeled image-query or patch-query pairs. To eliminate the heavy depen-dence on human annotations, we present a novel method, named Pseudo-Q, to automatically generate pseudo language queries for supervised training. Our method lever-ages an off-the-shelf object detector to identify visual ob-jects from unlabeled images, and then language queries for these objects are obtained in an unsupervised fashion with a pseudo-query generation module. Then, we design a task-related query prompt module to specifically tailor generated pseudo language queries for visual grounding tasks. Further, in order to fully capture the contextual re-lationships between images and language queries, we de-velop a visual-language model equipped with multi-level cross-modality attention mechanism. Extensive experimen-tal results demonstrate that our method has two notable benefits: (1) it can reduce human annotation costs signifi-cantly, e.g., 31% on Ref Coco [65] without degrading orig-inal model's performance under the fully supervised set-ting, and (2) without bells and whistles, it achieves supe-rior or comparable performance compared to state-of-the-art weakly-supervised visual grounding methods on all the five datasets we have experimented. Code is available at https://github.com/LeapLabTHU/Pseudo-Q.
Haojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song, Gao Huang 0001
CVPR1
2022 AdaFocus V2: End-to-End Training of Spatial Dynamic Networks for Video Recognition
abstract
Recent works have shown that the computational efficiency of video recognition can be significantly improved by reducing the spatial redundancy. As a representative work, the adaptive focus method (AdaFocus) has achieved a favorable trade-off between accuracy and inference speed by dynamically identifying and attending to the informative regions in each video frame. However, AdaFocus requires a complicated three-stage training pipeline (involving reinforcement learning), leading to slow convergence and is unfriendly to practitioners. This work reformulates the training of AdaFocus as a simple one-stage algorithm by introducing a differentiable interpolation-based patch selection operation, enabling efficient end-to-end optimization. We further present an improved training scheme to address the issues introduced by the one-stage formulation, including the lack of supervision, input diversity and training stability. Moreover, a conditional-exit technique is proposed to perform temporal adaptive computation on top of AdaFocus without additional training. Extensive experiments on six benchmark datasets (i.e., ActivityNet, FCVID, Mini-Kinetics, Something-Something V1&V2, and Jester) demonstrate that our model significantly outperforms the original AdaFocus and other competitive baselines, while being considerably more simple and efficient to train. Code is available at https://github.com/LeapLabTHU/AdaFocusV2.
Yulin Wang 0002, Yuanze Lin, Haojun Jiang, Zihang Lai, Victor Kulikov, Nikita Orlov, Humphrey Shi, Gao Huang 0001
CVPR4
2021 CondenseNet V2: Sparse Feature Reactivation for Deep Networks
abstract
Reusing features in deep networks through dense connectivity is an effective way to achieve high computational efficiency. The recent proposed CondenseNet [14] has shown that this mechanism can be further improved if redundant features are removed. In this paper, we propose an alternative approach named sparse feature reactivation (SFR), aiming at actively increasing the utility of features for reusing. In the proposed network, named CondenseNetV2, each layer can simultaneously learn to 1) selectively reuse a set of most important features from preceding layers; and 2) actively update a set of preceding features to increase their utility for later layers. Our experiments show that the proposed models achieve promising performance on image classification (ImageNet and CIFAR) and object detection (MS COCO) in terms of both theoretical efficiency and practical speed.
Le Yang 0007, Haojun Jiang, Ruojin Cai, Yulin Wang 0002, Shiji Song, Gao Huang 0001, Qi Tian 0001
CVPR2
2021 Adaptive Focus for Efficient Video Recognition
abstract
In this paper, we explore the spatial redundancy in video recognition with the aim to improve the computational efficiency. It is observed that the most informative region in each frame of a video is usually a small image patch, which shifts smoothly across frames. Therefore, we model the patch localization problem as a sequential decision task, and propose a reinforcement learning based approach for efficient spatially adaptive video recognition (AdaFocus). In specific, a light-weighted ConvNet is first adopted to quickly process the full video sequence, whose features are used by a recurrent policy network to localize the most task-relevant regions. Then the selected patches are inferred by a high-capacity network for the final prediction. During offline inference, once the informative patch sequence has been generated, the bulk of computation can be done in parallel, and is efficient on modern GPU devices. In addition, we demonstrate that the proposed method can be easily extended by further considering the temporal redundancy, e.g., dynamically skipping less valuable frames. Extensive experiments on five benchmark datasets, i.e., ActivityNet, FCVID, MiniKinetics, Something-Something V1&V2, demonstrate that our method is significantly more efficient than the competitive baselines. Code is available at https://github.com/blackfeather-wang/AdaFocus.
Yulin Wang 0002, Zhaoxi Chen 0007, Haojun Jiang, Shiji Song, Yizeng Han, Gao Huang 0001
ICCV3
2021 Spatially Adaptive Feature Refinement for Efficient Inference
abstract
Spatial redundancy commonly exists in the learned representations of convolutional neural networks (CNNs), leading to unnecessary computation on high-resolution features. In this paper, we propose a novel Spatially Adaptive feature Refinement (SAR) approach to reduce such superfluous computation. It performs efficient inference by adaptively fusing information from two branches: one conducts standard convolution on input features at a lower spatial resolution, and the other one selectively refines a set of regions at the original resolution. The two branches complement each other in feature learning, and both of them evoke much less computation than standard convolution. SAR is a flexible method that can be conveniently plugged into existing CNNs to establish models with reduced spatial redundancy. Experiments on CIFAR and ImageNet classification, COCO object detection and PASCAL VOC semantic segmentation tasks validate that the proposed SAR can consistently improve the network performance and efficiency. Notably, our results show that SAR only refines less than 40% of the regions in the feature representations of a ResNet for 97% of the samples in the validation set of ImageNet to achieve comparable accuracy with the original model, revealing the high computational redundancy in the spatial dimension of CNNs.
Yizeng Han, Gao Huang 0001, Shiji Song, Le Yang 0007, Haojun Jiang
IEEE Trans. Image Process.6