Hao Li 0075

dblp:17/5705-75 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-2419-2700ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CoSurfGS: 3D Surface Gaussian Splatting with Collaborative Distributed Learning for Large-scale Scene Reconstruction
Yalun Dai, Hao Li 0075, Weicai Ye, Danpeng Chen, Dingwen Zhang, Tong He 0001, Guofeng Zhang 0001, Junwei Han 0001
Int. J. Comput. Vis.3
2026 Patch-of-Interest ViT Inference Acceleration System for Edge-Assisted Video Analytics
abstract
The advent of edge computing has made real-time intelligent video analytics feasible. Previous works, based on traditional model architecture (e.g., CNN, RNN, etc.), employ various strategies to filter out non-region-of-interest content to minimize bandwidth and computation consumption but show inferior performance in adverse environments. Recently, visual foundation models based on transformers have shown great performance in adverse environments due to their amazing generalization capability. However, they require a large amount of computation power, which limits their applications in realtime intelligent video analytics. In this paper, we find visual foundation models like Vision Transformer (ViT) also have a dedicated acceleration mechanism for video analytics. To this end, we introduce Arena, an end-to-end edge-assisted video inference acceleration system based on ViT. We leverage the capability of ViT that can be accelerated through token pruning by only offloading and feeding Patches-of-Interest to the downstream models. Additionally, we design an adaptive keyframe inference switching algorithm tailored to different videos, capable of adapting to the current video content to jointly optimize accuracy and bandwidth. Through extensive experiments, our findings reveal that Arena can boost inference speeds by up to 1.58×, 1.82× and 1.98× on average while consuming only 47%, 31% and 27% of the bandwidth, respectively, all with high inference accuracy.
Haosong Peng, Hao Li 0075, Yufeng Zhan, Ren Jin, Yuanqing Xia
IEEE Trans. Computers3
2026 ChatStitch: Visualizing Through Structures via Surround-View Unsupervised Deep Image Stitching With Collaborative LLM-Agents
Hao Liang 0016, Hao Li 0075, Jiyuan Guo, Yufeng Yue, Mengyin Fu, Yi Yang 0009
IEEE Trans. Circuits Syst. Video Technol.4
2026 Radiant: Efficient Timely Large-Scale Scene Analytics Based on Hierarchical Framework
abstract
With the advancement of computer vision, the recently emerged 3D Gaussian Splatting (3DGS) has increasingly become a popular scene analytics algorithm due to its outstanding performance. Existing cloud-based 3DGS architectures overlook the challenges in real-world environments when handling large-scale scene analysis. This exposes issues such as inefficiency, low security, lack of privacy, and limited scalability. In this paper, we propose Radiant, a hierarchical framework for large scene analytics in a heterogeneous cloud-edge-device system, which jointly considers high efficiency, privacy and security, and scalability. Via extensive empirical study, we find that it is crucial to partition the regions for each edge appropriately and allocate varying camera positions to each device for image collection and training. The core of Radiant is partitioning regions based on heterogeneous environment information and allocating workloads to each device accordingly. Furthermore, we provide a 3DGS model aggregation algorithm that enhances the quality and ensures the continuity of models' boundaries. Finally, we develop a testbed, and experiments demonstrate that Radiant improved reconstruction quality by up to 25.7% and reduced up to 79.6% end-to-end latency.
Haosong Peng, Tianyu Qi, Yufeng Zhan, Ren Jin, Hao Li 0075, Yalun Dai, Yuanqing Xia
IEEE Trans. Serv. Comput.5
2025 XLD: A Cross-Lane Dataset for Benchmarking Novel Driving View Synthesis
abstract
Comprehensive testing of autonomous systems through simulation is essential to ensure the safety of autonomous driving vehicles. This requires the generation of safety-critical scenarios that extend beyond the limitations of real-world data collection, as many of these scenarios are rare or rarely encountered on public roads. However, evaluating most existing novel view synthesis (NVS) methods relies on sporadic sampling of image frames from the training data, comparing the rendered images with ground-truth images. Unfortunately, this evaluation protocol falls short of meeting the actual requirements in closed-loop simulations. Specifically, the true application demands the capability to render novel views that extend beyond the original trajectory (such as cross-lane views), which are challenging to capture in the real world. To address this, this paper presents a synthetic dataset for novel driving view synthesis evaluation, which is specifically designed for autonomous driving simulations. This unique dataset includes testing images captured by deviating from the training trajectory by 1–4 meters. It comprises six sequences that cover various times and weather conditions. Each sequence contains 450 training images, 120 testing images, and their corresponding camera poses and intrinsic parameters. Leveraging this novel dataset, we establish the first realistic benchmark for evaluating existing NVS approaches under frontonly and multicamera settings. The experimental findings underscore the significant gap in current approaches, revealing their inadequate ability to fulfill the demanding prerequisites of cross-lane or closed-loop simulation. Our dataset and code are released publicly on the project page: https://3d-aigc.github.io/XLD.
Hao Li 0075, Chenming Wu, Chen Zhao 0011, Chunyu Song, Haocheng Feng, Errui Ding, Dingwen Zhang, Jingdong Wang 0001
3DV1
2025 DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast Scenes
abstract
Novel-view synthesis approaches play a critical role in vast scene reconstruction. However, these methods rely heavily on dense image inputs and prolonged training times, making them unsuitable where computational resources are limited. Additionally, few-shot methods often struggle with poor reconstruction quality in vast environments. This paper presents DGTR, a novel distributed framework for efficient Gaussian reconstruction for sparse-view vast scenes. Our approach divides the scene into regions, processed independently by drones with sparse image inputs. Using a feed-forward Gaussian model, we predict high-quality Gaussian primitives, followed by a global alignment algorithm to ensure geometric consistency. Depth priors is incorporated to further enhance training, while a distillation-based model aggregation mechanism enables efficient reconstruction. Our method achieves high-quality large-scale scene reconstruction and novel-view synthesis in significantly reduced training times, outperforming existing approaches in both speed and scalability. We demonstrate the effectiveness of our framework on vast aerial scenes, achieving high-quality results within minutes. Code will released on our project page https://3d-aigc.github.io/DGTR.
Hao Li 0075, Haosong Peng, Chenming Wu, Weicai Ye, Yufeng Zhan, Chen Zhao 0011, Dingwen Zhang, Jingdong Wang 0001, Junwei Han 0001
ICRA1
2025 UDSH: An Unsupervised Deep Image Stitching and De-Occlusion Method for Heavy Occlusion Scene
abstract
Image stitching in heavy occlusion scenarios faces the dual challenges of accurate alignment and occlusion removal. On one hand, occlusion causes the loss of key texture and structural information in the image. On the other hand, it affects the image’s integrity. Existing stitching methods perform well in cases with small occlusion coverage, but they often fail in heavy occlusion. This failure is mainly due to three reasons: 1) they cannot identify occluded regions, 2) they cannot suppress interference from the occluded regions, 3) they cannot remove the occluded regions. To address these issues, we propose an unsupervised deep image stitching and de-occlusion method. First, to solve the issue of occluded region identification, we design an Occlusion-Aware Feature Weighted module (OAFW) that explicitly distinguishes between occluded and non-occluded regions by learning the occlusion masks of the images. Second, to address the issue of interference from occlusion, we use the learned occlusion masks to filter out features from the occluded regions. To further suppress the impact of occlusion-induced errors, we design a Mask-Guided Dual-Granularity Alignment loss function (MGDGA) that only calculates alignment errors for non-occluded regions, effectively reducing occlusion error interference during network training. Finally, to resolve the content gap in the occluded regions, we replace the pixels in the occluded areas with those from the aligned overlapping regions and incorporate a Progressive Content Inpainting module (PCI) to recover the missing content in the non-overlapping regions caused by occlusion, ultimately achieving a complete and natural de-occlusion stitched image. Experimental results show that our method improves the mean squared error metric by 17.45% compared to the state-of-the-art stitching method.
Hao Li 0075, Rundong Sun, Yi Yang 0009, Mengyin Fu
IROS2
2025 STRIDER: Navigation via Instruction-Aligned Structural Decision Space Optimization
abstract
The Zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) task requires agents to navigate previously unseen 3D environments using natural language instructions, without any scene-specific training. A critical challenge in this setting lies in ensuring agents’ actions align with both spatial structure and task intent over long-horizon execution. Existing methods often fail to achieve robust navigation due to a lack of structured decision-making and insufficient integration of feedback from previous actions. To address these challenges, we propose STRIDER (Instruction-Aligned Structural Decision Space Optimization), a novel framework that systematically optimizes the agent’s decision space by integrating spatial layout priors and dynamic task feedback. Our approach introduces two key innovations: 1) a Structured Waypoint Generator that constrains the action space through spatial structure, and 2) a Task-Alignment Regulator that adjusts behavior based on task progress, ensuring semantic alignment throughout navigation. Extensive experiments on the R2R-CE and RxR-CE benchmarks demonstrate that STRIDER significantly outperforms strong SOTA across key metrics; in particular, it improves Success Rate (SR) from 29\% to 35\%, a relative gain of 20.7\%. Such results highlight the importance of spatially constrained decision-making and feedback-guided execution in improving navigation fidelity for zero-shot VLN-CE.
Diqi He, Xuehao Gao, Hao Li 0075, Junwei Han 0001, Dingwen Zhang
NeurIPS3
2025 Unsupervised Pre-Training With Language-Vision Prompts for Low-Data Instance Segmentation
abstract
In recent times, following the paradigm of DETR (DEtection TRansformer), query-based end-to-end instance segmentation (QEIS) methods have exhibited superior performance compared to CNN-based models, particularly when trained on large-scale datasets. Nevertheless, the effectiveness of these QEIS methods diminishes significantly when confronted with limited training data. This limitation arises from their reliance on substantial data volumes to effectively train the pivotal queries/kernels that are essential for acquiring localization and shape priors. To address this problem, we propose a novel method for unsupervised pre-training in low-data regimes. Inspired by the recently successful prompting technique, we introduce a new method, Unsupervised Pre-training with Language-Vision Prompts (UPLVP), which improves QEIS models' instance segmentation by bringing language-vision prompts to queries/kernels. Our method consists of three parts: (1) Masks Proposal: Utilizes language-vision models to generate pseudo masks based on unlabeled images. (2) Prompt-Kernel Matching: Converts pseudo masks into prompts and injects the best-matched localization and shape features to their corresponding kernels. (3) Kernel Supervision: Formulates supervision for pre-training at the kernel level to ensure robust learning. With the help of our pre-training method, QEIS models can converge faster and perform better than CNN-based models in low-data regimes. Experimental evaluations conducted on MS COCO, Cityscapes, and CTW1500 datasets indicate that the QEIS models' performance can be significantly improved when pre-trained with our method.
Dingwen Zhang, Hao Li 0075, Diqi He, Nian Liu 0002, Lechao Cheng, Jingdong Wang 0001, Junwei Han 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Focus-TransUnet3D: High-Precision Model for 3D Segmentation of Medical Point Targets
abstract
Deep learning has been extensively applied in medical image segmentation, providing significant support for disease diagnosis. However, traditional encoder-decoder networks struggle with segmenting scale-sensitive point target lesions. To address this challenge, this paper proposes an innovative incremental fusion architecture that can integrate different models and achieve significant performance improvements through complementary fusion. Based on this architecture, we developed Focus-TransUnet3D by combining the Trans-FusionNet3D model and the 3D Unet model. This model adopts a global-to-local segmentation strategy, effectively addressing the challenges of medical point target segmentation, thereby expanding the application of deep learning in the field of medical image processing. Furthermore, we design a deep fusion strategy suitable for the transformer model to adapt to multi-scale feature learning. The integration of the transformer model with convolutional neural networks brings improvements in local and global feature extraction capabilities, enhancing the applicability of our model. We evaluate our model on three clinical datasets with different target scales: the Intracranial Artery dataset, the Intracranial Aneurysm dataset, and the LiTS17 dataset. The results indicate that in the external test for intracranial aneurysm auxiliary diagnosis, the model trained with only 47 annotated samples achieved the state-of-the-art performance, attaining a Dice coefficient of 84.14% and a sensitivity of 100%. This effectively addresses the challenges of annotation scarcity and tiny targets. Our code will be released athttps://github.com/caijilia/FTUnet3D.
Dihua Zhai, Hao Li 0075, Ke Tian, Yi Yang 0009, Zhenyao Chang, Shuo Wang 0001, Yuanqing Xia
IEEE Trans. Circuits Syst. Video Technol.2
2025 Weakly Supervised Semantic Segmentation via Alternate Self-Dual Teaching
abstract
Weakly supervised semantic segmentation (WSSS) is a challenging yet important research field in vision community. In WSSS, the key problem is to generate high-quality pseudo segmentation masks (PSMs). Existing approaches mainly depend on the discriminative object part to generate PSMs, which would inevitably miss object parts or involve surrounding image background, as the learning process is unaware of the full object structure. In fact, both the discriminative object part and the full object structure are critical for deriving of high-quality PSMs. To fully explore these two information cues, we build a novel end-to-end learning framework, alternate self-dual teaching (ASDT), based on a dual-teacher single-student network architecture. The information interaction among different network branches is formulated in the form of knowledge distillation (KD). Unlike the conventional KD, the knowledge of the two teacher models would inevitably be noisy under weak supervision. Inspired by the Pulse Width (PW) modulation, we introduce a PW wave-like selection signal to alleviate the influence of the imperfect knowledge from either teacher model on the KD process. Comprehensive experiments on the PASCAL VOC 2012 and COCO-Stuff 10K demonstrate the effectiveness of the proposed ASDT framework, and new state-of-the-art results are achieved.
Dingwen Zhang, Hao Li 0075, Wenyuan Zeng, Chaowei Fang, Lechao Cheng, Ming-Ming Cheng, Junwei Han 0001
IEEE Trans. Image Process.2
2024 GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene Understanding
abstract
Applying Neural Radiance Fields (NeRF) to downstream perception tasks for scene understanding and representation is becoming increasingly popular. Most existing methods treat semantic prediction as an additional rendering task, i.e., the “label rendering” task, to build semantic NeRFs. However, by rendering semantic/instance labels per pixel without considering the contextual information of the rendered image, these methods usually suffer from unclear boundary segmentation and abnormal segmentation of pixels within an object. To solve this problem, we propose Generalized Perception NeRF (GP-NeRF), a novel pipeline that makes the widely used segmentation model and NeRF work compatibly under a unified framework, for facilitating context-aware 3D scene perception. To accomplish this goal, we introduce transformers to aggregate radiance as well as semantic embedding fields jointly for novel views and facilitate the joint volumetric rendering of both fields. In addition, we propose two self-distillation mechanisms, i.e., the Semantic Distill Loss and the Depth-Guided Semantic Distill Loss, to enhance the discrimination and quality of the semantic field and the maintenance of geometric consistency. In evaluation, as shown in Fig. 1 we conduct experimental comparisons under two perception tasks (i.e. semantic and instance segmentation) using both synthetic and real-world datasets. Notably, our method outperforms SOTA approaches by 6.94%,11.76%, and 8.47% on generalized semantic segmentation, finetuning semantic segmentation, and instance segmentation, respectively. Project.
Hao Li 0075, Dingwen Zhang, Yalun Dai, Nian Liu 0002, Lechao Cheng, Jingfeng Li, Jingdong Wang 0001, Junwei Han 0001
CVPR1
2024 LTGC: Long-Tail Recognition via Leveraging LLMs-Driven Generated Content
abstract
Long-tail recognition is challenging because it requires the model to learn good representations from tail categories and address imbalances across all categories. In this paper, we propose a novel generative and fine-tuning framework, LTGC, to handle long-tail recognition via leveraging generated content. Firstly, inspired by the rich implicit knowledge in large-scale models (e.g., large language models, LLMs), LTGC leverages the power of these models to parse and reason over the original tail data to produce diverse tail-class content. We then propose several novel designs for LTGC to ensure the quality of the generated data and to efficiently fine-tune the model using both the generated and original data. The visualization demonstrates the effectiveness of the generation module in LTGC, which produces accurate and diverse tail data. Additionally, the experimental results demonstrate that our LTGC outperforms existing state-of-the-art methods on popular long-tailed benchmarks.
Qihao Zhao, Yalun Dai, Hao Li 0075, Wei Hu 0004, Fan Zhang 0007, Jun Liu 0036
CVPR3
2024 GGRt: Towards Pose-Free Generalizable 3D Gaussian Splatting in Real-Time
Hao Li 0075, Chenming Wu, Dingwen Zhang, Yalun Dai, Chen Zhao 0011, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Junwei Han 0001
ECCV (71)1
2024 ERDUnet: An Efficient Residual Double-Coding Unet for Medical Image Segmentation
abstract
Medical image segmentation is widely used in clinical diagnosis, and methods based on convolutional neural networks have been able to achieve high accuracy. However, it is still difficult to extract global context features, and the parameters are too large to be clinically applied. In this regard, we propose a novel network structure to improve the traditional encoder-decoder network model, which saves parameters while maintaining segmentation accuracy. We improve the feature extraction efficiency by constructing an encoder module that can simultaneously extract local features and global continuity information. A novel attention module is designed to optimize segmentation boundary regions while improving training efficiency. The feature transfer structure of the decoding part is also improved, which fully integrates the features of different levels to restore the spatial resolution more finely. We evaluate our model on seven different medical segmentation datasets, the 2018 Data Science Bowl Challenge (DSBC2018), the 2018 Lesion Boundary Segmentation Challenge (ISIC2018), the Gland Segmentation in Colon Histology Images Challenge (GlaS), Kvasir-SEG, CVC-ClinicDB, Kvasir-Instrument and Polypgen. Extensive experimental results show that our model can achieve good segmentation performance while maintaining a small number of parameters and computational load, which can further facilitate the generalization of the theoretical approach to clinical practice. Our code will be released athttps://github.com/caijilia/ERDUnet.
Hao Li 0075, Dihua Zhai, Yuanqing Xia
IEEE Trans. Circuits Syst. Video Technol.1
2023 Boosting Low-Data Instance Segmentation by Unsupervised Pre-training with Saliency Prompt
abstract
Inspired by DETR variants, query-based end-to-end instance segmentation (QEIS) methods have recently outperformed CNN-based models on large-scale datasets. Yet they would lose efficacy when only a small amount of training data is available since it's hard for the crucial queries/kernels to learn localization and shape priors. To this end, this work offers a novel unsupervised pre-training solution for low-data regimes. Inspired by the recent success of the Prompting technique, we introduce a new pre-training method that boosts QEIS models by giving Saliency Prompt for queries/kernels. Our method contains three parts: 1) Saliency Masks Proposal is responsible for generating pseudo masks from unlabeled images based on the saliency mechanism. 2) Prompt-Kernel Matching transfers pseudo masks into prompts and injects the corresponding localization and shape priors to the best-matched kernels. 3) Kernel Supervision is applied to supply supervision at the kernel level for robust learning. From a practical perspective, our pre-training method helps QEIS models achieve a similar convergence speed and comparable performance with CNN-based models in low-data regimes. Experimental results show that our method significantly boosts several QEIS models on three datasets.11Code: https://github.com/lifuguan/saliency.prompt
Hao Li 0075, Dingwen Zhang, Nian Liu 0002, Lechao Cheng, Yalun Dai, Xinggang Wang, Junwei Han 0001
CVPR1
2017 Intersection scan model and probability inference for vision based small-scale urban intersection detection
abstract
Large-scale intersections stamped on maps have diverse visual features for detection, while small-scale urban intersections are hard to be identified especially when GPS signals are missing. In this paper, we propose a Hidden Markov Model (HMM) based small-scale intersection detection method utilizing monocular vision. We extract visual cues of road transformations and dynamic vehicles' tracks, and then design an Intersection Scan Model to obtain the potential traversable direction of the current road, which is the primary criterion of the intersection estimation. For better performances, we take the detections of consecutive frames into consideration and finally integrate them into HMM to estimate the probabilities of intersections. Results from KITTI datasets and real-world experiments have shown the functionality of the presented approach.
Yi Yang 0009, Hao Li 0075, Hao Zhu 0002, Songtian Shang, Ningyi Lyu, Wenjie Song 0001
Intelligent Vehicles Symposium2