Yutao Hu 0002

dblp:37/5215-2 · DBLP profile ↗
← Back
22ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0002-1287-6515ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking the detail-preserved completion of complex tubular structures based on point cloud: A dataset and a benchmark
Yaolei Qi, Yikai Yang, Wenbo Peng, Shumei Miao, Yutao Hu 0002, Guanyu Yang 0001
Medical Image Anal.5
2026 Toward Semantic-Aware Aerial Video Anomaly Detection by Exploiting Multimodal Large Language Model
abstract
Drones have become increasingly widely applied in surveillance systems due to their mobility, making aerial video anomaly detection methods more crucial. Anomalies in aerial videos often present as semantic conflicts, such as the presence of unexpected objects or unusual behaviors that do not align with the context. Previous approaches often relied on manually crafted knowledge graphs to detect such conflicts, which suffer from poor scalability. Recently, owing to their sufficient alignment training, multimodal large language models (MLLMs) have emerged as a generalized solution for semantic understanding. However, the direct application of MLLMs does not yield satisfactory anomaly detection performance in aerial videos. First, aerial videos often manifest platform-induced pseudo-motion, which obscures the true motion of objects and exacerbates detection errors. Second, without sufficient labeled data for fine-tuning, generic MLLMs often lack scene-level semantic guidance to reliably distinguish abnormal events that include contextually inappropriate behaviors. To address these challenges, we propose SemAero, an MLLM-based framework to address these challenges by: 1) designing an ego-motion reduction module to enhance model perception on object movement, 2) generating scene-specific prompts adaptively with step-by-step guidance for reasonable output, and 3) refining scores with dual-stream consistent feature for better domain-specific anomaly detection. Evaluated across 8 diverse aerial scenes and 73 sub-datasets, SemAero achieves a 3.09% improvement in AUC-ROC over the second-best model, demonstrating its ability in aerial video anomaly detection.
Ruoheng Li, Xuhui Liu, Yutao Hu 0002, Xianbin Cao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 AEGIS: Using Conditional Multi-View Diffusion Model to Achieve Angiographic Enhancement in Non-Contrast CT
abstract
Angiographic enhancement of non-contrast CT (NCCT) using AI techniques is essential for diagnosing patients unable to use contrast agents. However, AI angiography remains a challenging task because of the feature fragility, structural complexity, and spatial continuity. In this paper, we propose an angiographic framework based on a conditional multi-view diffusion model called AEGIS with three innovations: multi-view hybrid learning (MHL), conditional angiographic diffusion estimation (CADE), and multi-view map fusion (MMF). 1) MHL targets Contrast Map (CM), the difference between NCCT and CT angiography, from multiple views to perceive 3D features in 2D space, enhancing the stability of feature representation. 2) CADE is a conditional diffusion model using NCCT as spatial guidance, providing crucial information for CM generation. 3) MMF adopts a lightweight AutoEncoder for filtering and fusing multi-view CMs, maintaining coherence between adjacent slices while modifying slight bias in low-dimensional representations, thus optimizing data quality and accuracy. Experiments demonstrate our superior performance, which achieve state-of-the-art image quality (PSNR+6.69, SSIM+3.17, MSE-46.38), segmentation evaluation (CADIR×10.49, HSDIR×5.57) and feature distance (FID-64.27). Visualizations and positive evaluation scores from clinicians further reveals that AEGIS has significant potential in clinical applications.
Jiahao Xia 0005, Xiaolei Zhang 0005, Yuting He 0001, Yaolei Qi, Yutao Hu 0002, Pascal Haigron, Chunxiang Tang, Longjiang Zhang, Guanyu Yang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space
abstract
Feature matching plays a fundamental role in many computer vision tasks, yet existing methods heavily rely on scarce and clean multi-view image collections, which constrains their generalization to diverse and challenging scenarios. Moreover, conventional feature encoders are typically trained on single-view 2D images, limiting their capacity to capture 3D-aware correspondences. In this paper, we propose a novel two-stage framework that lifts 2D images to 3D space, named as \textbf{Lift to Match (L2M)}, taking full advantage of large-scale and diverse single-view images. To be specific, in the first stage, we learn a 3D-aware feature encoder using a combination of multi-view image synthesis and 3D feature Gaussian representation, which injects 3D geometry knowledge into the encoder. In the second stage, a novel-view rendering strategy, combined with large-scale synthetic data generation from single-view images, is employed to learn a feature decoder for robust feature matching, thus achieving generalization across diverse domains. Extensive experiments demonstrate that our method achieves superior generalization across zero-shot evaluation benchmarks, highlighting the effectiveness of the proposed framework for robust feature matching.
Yingping Liang, Yutao Hu 0002, Wenqi Shao, Ying Fu 0001
ICCV2
2025 RobuSTereo: Robust Zero-Shot Stereo Matching under Adverse Weather
abstract
Learning-based stereo matching models struggle in adverse weather conditions due to the scarcity of corresponding training data and the challenges in extracting discriminative features from degraded images. These limitations significantly hinder zero-shot generalization to out-of-distribution weather conditions. In this paper, we propose \textbf{RobuSTereo}, a novel framework that enhances the zero-shot generalization of stereo matching models under adverse weather by addressing both data scarcity and feature extraction challenges. First, we introduce a diffusion-based simulation pipeline with a stereo consistency module, which generates high-quality stereo data tailored for adverse conditions. By training stereo matching models on our synthetic datasets, we reduce the domain gap between clean and degraded images, significantly improving the models' robustness to unseen weather conditions. The stereo consistency module ensures structural alignment across synthesized image pairs, preserving geometric integrity and enhancing depth estimation accuracy. Second, we design a robust feature encoder that combines a specialized ConvNet with a denoising transformer to extract stable and reliable features from degraded images. The ConvNet captures fine-grained local structures, while the denoising transformer refines global representations, effectively mitigating the impact of noise, low visibility, and weather-induced distortions. This enables more accurate disparity estimation even under challenging visual conditions. Extensive experiments demonstrate that \textbf{RobuSTereo} significantly improves the robustness and generalization of stereo matching models across diverse adverse weather scenarios.
Yingping Liang, Yutao Hu 0002, Ying Fu 0001
ICCV3
2025 Flow-Anything: Learning Real-World Optical Flow Estimation From Large-Scale Single-View Images
abstract
Optical flow estimation is a crucial subfield of computer vision, serving as a foundation for video tasks. However, the real-world robustness is limited by animated synthetic datasets for training. This introduces domain gaps when applied to real-world applications and limits the benefits of scaling up datasets. To address these challenges, we propose Flow-Anything, a large-scale data generation framework designed to learn optical flow estimation from any single-view images in the real world. We employ two effective steps to make data scaling-up promising. First, we convert a single-view image into a 3D representation using advanced monocular depth estimation networks. This allows us to render optical flow and novel view images under a virtual camera. Second, we develop an Object-Independent Volume Rendering module and a Depth-Aware Inpainting module to model the dynamic objects in the 3D representation. These two steps allow us to generate realistic datasets for training from large-scale single-view images, namely FA-Flow Dataset. For the first time, we demonstrate the benefits of generating optical flow training data from large-scale real-world images, outperforming the most advanced unsupervised methods and supervised methods on synthetic datasets. Moreover, our models serve as a foundation model and enhance the performance of various downstream video tasks.
Yingping Liang, Ying Fu 0001, Yutao Hu 0002, Wenqi Shao, Debing Zhang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) have made significant strides in various multimodal tasks. Notably, GPT4V, Claude, Gemini, and others showcase exceptional multimodal capabilities, marked by profound comprehension and reasoning skills. This study introduces a comprehensive and efficient evaluation framework, TinyLVLM-eHub, to assess LVLMs’ performance, including proprietary models. TinyLVLM-eHub covers six key multimodal capabilities, such as visual perception, knowledge acquisition, reasoning, commonsense understanding, object hallucination, and embodied intelligence. The benchmark, utilizing 2.1K image-text pairs, provides a user-friendly and accessible platform for LVLM evaluation. The evaluation employs the ChatGPT Ensemble Evaluation (CEE) method, which improves alignment with human evaluation compared to word-matching approaches. Results reveal that closed-source API models like GPT4V and GeminiPro-V excel in most capabilities compared to previous open-source LVLMs, though they show some vulnerability in object hallucination. This evaluation underscores areas for LVLM improvement in real-world applications and serves as a foundational assessment for future multimodal advancements.
Wenqi Shao, Meng Lei, Yutao Hu 0002, Peng Gao 0007, Peng Xu 0035, Kaipeng Zhang, Fanqing Meng, Siyuan Huang 0004, Hongsheng Li 0001, Yu Qiao 0001, Ping Luo 0002
IEEE Trans. Big Data3
2025 Diffusion Self-Distillation for Remote Sensing Scene Classification
abstract
Remote sensing scene classification, a fundamental task in remote image analysis, has obtained rapid progress due to the powerful capabilities of Convolutional Neural Networks (CNNs). Achieving precise classification performance heavily relies on the feature extraction capacity of the network. However, due to the large variation and severe distortion within the images, extracting robust feature representations is necessary but challenging. Self-distillation could enhance the shallow layers by providing stronger gradients and more accurate supervision from deeper layers, thereby promoting the extraction of spatially detailed features. Nonetheless, due to the limited capacity of shallow layers to learn truly valuable knowledge, shallow layer features can be viewed as the noisy version of deep layer features and contain more disruptive factors, which significantly impedes the effectiveness of self-distillation. To address this issue, in this paper, we establish the Diffusion Self-Distillation Network (DSDNet), which incorporates the conditional diffusion denoising model into the self-distillation framework. Specifically, DSDNet filters noise from shallow features through the diffusion denoising process, enabling more precise and accurate distillation between the refined student features and the teacher features. Extensive experiments on four challenging remote sensing datasets emonstrate that the proposed DSDNet achieves significant performance improvements over various backbone networks with negligible increases in parameters, delivering state-of-the-art classification performance. Our code and dataset are available on https://github.com/toggle1995/DSDNet.
Yutao Hu 0002, Lei Zhang 0001, Xiaoyan Luo, Xianbin Cao 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Hierarchical Self-Distilled Feature Learning for Fine-Grained Visual Categorization
abstract
Fine-grained visual categorization (FGVC) relies on hierarchical features extracted by deep convolutional neural networks (CNNs) to recognize closely alike objects. Particularly, shallow layer features containing rich spatial details are vital for specifying subtle differences between objects but are usually inadequately optimized due to gradient vanishing during backpropagation. In this article, hierarchical self-distillation (HSD) is introduced to generate well-optimized CNNs features for accurate fine-grained categorization. HSD inherits from the widely applied deep supervision and implements multiple intermediate losses for reinforced gradients. Besides that, we observe that the hard (one-hot) labels adopted for intermediate supervision hurt the performance of FGVC by enforcing overstrict supervision. As a solution, HSD seeks self-distillation where soft predictions generated by deeper layers of the network are hierarchically exploited to supervise shallow parts. Moreover, self-information entropy loss (SIELoss) is designed in HSD to adaptively soften intermediate predictions and facilitate better convergence. In addition, the gradient detached fusion (GDF) module is incorporated to produce an ensemble result with multiscale features via effective feature fusion. Extensive experiments on four challenging fine-grained datasets show that, with neglectable parameter increase, the proposed HSD framework and the GDF module both bring significant performance gains over different backbones, which also achieves state-of-the-art classification performance.
Yutao Hu 0002, Xuhui Liu, Xiaoyan Luo, Yao Hu 0002, Xianbin Cao 0001, Baochang Zhang 0001, Jun Zhang 0007
IEEE Trans. Neural Networks Learn. Syst.1
2024 OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM
abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in various multimodal tasks. However, their potential in the medical domain re-mains largely unexplored. A significant challenge arises from the scarcity of diverse medical images spanning various modalities and anatomical regions, which is essential in real-world medical applications. To solve this problem, in this paper, we introduce OmniMedVQA, a novel comprehensive medical Visual Question Answering (VQA) benchmark. This benchmark is collected from 73 different medical datasets, including 12 different modalities and covering more than 20 distinct anatomical regions. Importantly, all images in this benchmark are sourced from authentic medical scenarios, ensuring alignment with the requirements of the medical field and suitability for evaluating LVLMs. Through our extensive experiments, we have found that existing LVLMs struggle to address these medical VQA problems effectively. Moreover, what surprises us is that medical-specialized LVLMs even exhibit inferior performance to those general-domain models, calling for a more versatile and robust LVLM in the biomedical field. The evaluation results not only reveal the current limitations of LVLM in understanding real medical images but also highlight our dataset's significance. Our code with dataset are available at https://github.com/OpenGVLab/ Multi Modality-Arena.
Yutao Hu 0002, Tianbin Li, Quanfeng Lu, Wenqi Shao, Junjun He, Yu Qiao 0001, Ping Luo 0002
CVPR1
2024 Learning Foreground Information Bottleneck for few-shot semantic segmentation
Yutao Hu 0002, Xiaoyan Luo, Jungong Han, Xianbin Cao 0001, Jun Zhang 0007
Pattern Recognit.1
2023 Beyond One-to-One: Rethinking the Referring Image Segmentation
abstract
Referring image segmentation aims to segment the target object referred by a natural language expression. However, previous methods rely on the strong assumption that one sentence must describe one target in the image, which is often not the case in real-world applications. As a result, such methods fail when the expressions refer to either no objects or multiple objects. In this paper, we address this issue from two perspectives. First, we propose a Dual Multi-Modal Interaction (DMMI) Network, which contains two decoder branches and enables information flow in two directions. In the text-to-image decoder, text embedding is utilized to query the visual feature and localize the corresponding target. Meanwhile, the image-to-text decoder is implemented to reconstruct the erased entity-phrase conditioned on the visual feature. In this way, visual features are encouraged to contain the critical semantic information about target entity, which supports the accurate segmentation in the text-to-image decoder in turn. Secondly, we collect a new challenging but realistic dataset called Ref-ZOM, which includes image-text pairs under different settings. Extensive experiments demonstrate our method achieves state-of-the-art performance on different datasets, and the Ref-ZOM-trained model performs well on various types of text inputs. Codes and datasets are available at https://github.com/toggle1995/RIS-DMMI.
Yutao Hu 0002, Qixiong Wang, Wenqi Shao, Enze Xie, Zhenguo Li, Jungong Han, Ping Luo 0002
ICCV1
2023 Boosting Variational Inference With Margin Learning for Few-Shot Scene-Adaptive Anomaly Detection
abstract
Anomaly detection in surveillance videos aims to identify frames where abnormal events happen. Existing approaches assume that the training and testing videos are from the same scene, exhibiting poor generalization performance when encountering an unseen scene. In this paper, we propose a Variational Anomaly Detection Network (VADNet), which is characterized by its high scene-adaptation - it can identify abnormal events in a new scene only via referring to a few normal samples without fine-tuning. Our model embodies two major innovations. First, a novel Variational Normal Inference (VNI) module is proposed to formulate image reconstruction in a conditional variational auto-encoder (CVAE) framework, which learns a probabilistic decision model instead of a traditional deterministic one. Secondly, a Margin Learning Embedding (MLE) module is leveraged to boost the variational inference and aid in distinguishing normal events. We theoretically demonstrate that minimizing the triplet loss in MLE module facilitates maximizing the evidence lower bound (ELBO) of CVAE, which promotes the convergence of VNI. By incorporating variational inference with margin learning, VADNet becomes much more generative that is able to handle the uncertainty caused by the changed scene and limited reference data. Extensive experiments on several datasets demonstrate that the proposed VADNet can adapt to a new scene effectively without fine-tuning and achieve remarkable performance, which outperforms other methods significantly and establishes new state-of-the-art in the case of few-shot scene-adaptive anomaly detection. We believe our method is closer to real-world application due to its strong generalization ability. All codes are released inhttps://github.com/huangxx156/VADNet.
Yutao Hu 0002, Xiaoyan Luo, Jungong Han, Baochang Zhang 0001, Xianbin Cao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Attentive encoder-decoder networks for crowd counting
Xuhui Liu, Yutao Hu 0002, Baochang Zhang 0001, Xiantong Zhen, Xiaoyan Luo, Xianbin Cao 0001
Neurocomputing2
2022 Variational Self-Distillation for Remote Sensing Scene Classification
abstract
Supported by deep learning techniques, remote sensing scene classification, a fundamental task in remote image analysis, has recently obtained remarkable progress. However, due to the severe uncertainty and perturbation within an image, it is still a challenging task and remains many unsolved problems. In this paper, we note that regular one-hot labels cannot precisely describe remote sensing images, and they fail to provide enough information for supervision and limiting the discriminative feature learning of the network. To solve this problem, we propose a Variational Self-Distillation Network (VSDNet), in which the class entanglement information from the prediction vector acts as the supplement to the category information. Then, the exploited information is hierarchically distilled from the deep layers into the shallow parts via a Variational Knowledge Transfer (VKT) module. Notably, the VKT module performs knowledge distillation in a probabilistic way through variational estimation, which enables end-to-end optimization for mutual information and promotes robustness to uncertainty within the image. Extensive experiments on four challenging remote sensing datasets demonstrate that, with a negligible parameter increase, the proposed VSDNet brings a significant performance improvement over different backbone networks and delivers state-of-the-art results.
Yutao Hu 0002, Xiaoyan Luo, Jungong Han, Xianbin Cao 0001, Jun Zhang 0007
IEEE Trans. Geosci. Remote. Sens.1
2021 Adaptive Anomaly Detection Network for Unseen Scene Without Fine-Tuning
Yutao Hu 0002, Xiaoyan Luo
PRCV (2)1
2021 Long-range Attention Network for Multi-View Stereo
abstract
Learning-based multi-view stereo (MVS) has recently gained great popularity, which can efficiently infer depth map and reconstruct fine-grained scene geometry. Previous methods calculate the variance of the corresponding pixel pairs to determine whether they are matched mostly based on the pixel-wise measure, which fails to consider the interdependence among pixels and is ineffective on the matching of texture-less or occluded regions. These false matching problems challenge MVS and result in its most failure cases. To address the issues, we introduce a Long-range Attention Network (LANet) to selectively aggregate reference features to each position to capture the long-range interdependence across the entire space. As a result, similar features relate to each other regardless of their distance, propagating more guiding information for the effective match. Furthermore, we introduce a new loss to supervise the intermediate probability volume by constraining its distribution reasonably centered at the true depth. Extensive experiments on large-scale DTU dataset demonstrate that the proposed LANet achieves the new state-of-the-art performance, outperforming previous methods by a large margin. Our method is generic and also achieves comparable results on outdoor Tanks and Temples dataset without any fine-tuning, which validates our method's generalization ability.
Yutao Hu 0002, Xianbin Cao 0001, Baochang Zhang 0001
WACV2
2021 Attentional Kernel Encoding Networks for Fine-Grained Visual Categorization
abstract
Fine-grained visual categorization aims to recognize objects from different sub-ordinate categories, which is a challenging task due to subtle visual differences between images. It is highly desired to identify discriminative regions while achieving highly non-linear compact representation for fine-grained visual categorization. However, existing methods either rely on manually defined part-based annotations to indicate the distinctive regions or operate on longitudinal vectors to capture the non-linear information, which may lose important spatial layout information. In this paper, we propose the Attentional Kernel Encoding Networks (AKEN) for fine-grained visual categorization. Specifically, the AKEN aggregates feature maps from the last convolutional layer of ConvNets to obtain a holistic feature representation. By Fourier embedding, it encodes features from both the longitudinal and transverse directions, which largely retains the spatial layout information. Moreover, we incorporate a Cascaded Attention (Cas-Attention) module to highlight local regions that distinguish among subordinate categories, enabling the AKEN to extract the most discriminative features. Working in conjunction with the attention mechanism, the proposed AKEN combines the strengths of ConvNets and kernels for non-linear feature learning, which can establish discriminative and descriptive feature representations for fine-grained image categorization. Experiments on three benchmark datasets show that the proposed AKEN delivers highly competitive performance, surpassing most existed methods and achieving state-of-the-art results.
Yutao Hu 0002, Yandan Yang, Jun Zhang 0007, Xianbin Cao 0001, Xiantong Zhen
IEEE Trans. Circuits Syst. Video Technol.1
2021 Alignment Enhancement Network for Fine-grained Visual Categorization
abstract
Fine-grained visual categorization (FGVC) aims to automatically recognize objects from different sub-ordinate categories. Despite attracting considerable attention from both academia and industry, it remains a challenging task due to subtle visual differences among different classes. Cross-layer feature aggregation and cross-image pairwise learning become prevailing in improving the performance of FGVC by extracting discriminative class-specific features. However, they are still inefficient to fully use the cross-layer information based on the simple aggregation strategy, while existing pairwise learning methods also fail to explore long-range interactions between different images. To address these problems, we propose a novel Alignment Enhancement Network (AENet), including two-level alignments, Cross-layer Alignment (CLA) and Cross-image Alignment (CIA). The CLA module exploits the cross-layer relationship between low-level spatial information and high-level semantic information, which contributes to cross-layer feature aggregation to improve the capacity of feature representation for input images. The new CIA module is further introduced to produce the aligned feature map, which can enhance the relevant information as well as suppress the irrelevant information across the whole spatial region. Our method is based on an underlying assumption that the aligned feature map should be closer to the inputs of CIA when they belong to the same category. Accordingly, we establish Semantic Affinity Loss to supervise the feature alignment within each CIA block. Experimental results on four challenging datasets show that the proposed AENet achieves the state-of-the-art results over prior arts.
Yutao Hu 0002, Xuhui Liu, Baochang Zhang 0001, Jungong Han, Xianbin Cao 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2020 NAS-Count: Counting-by-Density with Neural Architecture Search
Yutao Hu 0002, Xuhui Liu, Baochang Zhang 0001, Jungong Han, Xianbin Cao 0001, David S. Doermann
ECCV (22)1
2020 Few-Shot Semantic Segmentation with Democratic Attention Networks
Yutao Hu 0002, Yandan Yang, Xianbin Cao 0001, Xiantong Zhen
ECCV (13)3
2020 Improving Backbones Performance by Complex Architectures
Jinxin Shao, Yutao Hu 0002, Teli Ma, Baochang Zhang 0001
PRCV (2)2