VLDB 2026 Research / reviewers in the wild / expert
Haoyang He
dblp:272/5008
· DBLP profile ↗
18ranked-venue papers
7as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-City Traffic Prediction with Semantic-Topological Decoupling and Spatial Attention Enhancement
Yuwei Gu, Weilong Ding 0002, Haoyang He |
DASFAA (3) | 3 |
| 2026 | MambaGesture2: Co-Speech Gesture Generation via Hierarchical Fusion and Spatiotemporal AggregationabstractCo-speech gesture generation plays a vital role in producing synchronized and natural human gestures, thereby enhancing the realism of avatars in virtual environments. Although diffusion models have shown strong generative capabilities, their combination with transformer-based architectures often incurs high computational costs due to the quadratic complexity of self-attention. Moreover, as a temporal sequence modeling task, existing methods frequently struggle to effectively capture multi-scale temporal dynamics inherent in speech and gesture signals. To address these challenges, we propose MambaGesture2, a novel framework that integrates a Mamba-based denoising network, Hierarchical U-Net Gesture Mamba (HUG-Mamba), with a multimodal feature fusion module, SEAD. HUG-Mamba combines the efficient state-space modeling of Mamba blocks with the hierarchical sampling of the U-Net architecture, significantly improving temporal coherence and computational efficiency. We further introduce the Temporal-Stratified Fusion (TSF) module to capture diverse temporal scales via multi-scale learning, and the Spatial-Temporal Cascaded Aggregation (STCA) module to enhance spatial-temporal feature aggregation. Extensive experiments on the multi-modal BEAT2 and SHOW datasets demonstrate that our approach achieves state-of-the-art performance across multiple quantitative metrics, while substantially reducing model complexity and inference time. The results validate the effectiveness of our architectural innovations in generating diverse, realistic, and temporally consistent co-speech gestures. Project page:https://fcchit.github.io/mambagesture2. Chencan Fu, Yabiao Wang, Haoyang He, Chengjie Wang 0001, Ying Tai, Yong Liu 0007, Jiangning Zhang |
IEEE Trans. Multim. | 3 |
| 2026 | Learning Multi-View Anomaly Detection With Efficient Adaptive SelectionabstractThis study explores the recently proposed and challenging multi-view Anomaly Detection (AD) task. Single-view tasks will encounter blind spots from other perspectives, resulting in inaccuracies in sample-level prediction. Therefore, we introduce theMulti-ViewAnomalyDetection (MVAD) approach, which learns and integrates features from multi-views. Specifically, we propose aMulti-ViewAdaptiveSelection (MVAS) algorithm for feature learning and fusion across multiple views. The feature maps are divided into neighbourhood attention windows to calculate a semantic correlation matrix between single-view windows and all other views, which is an attention mechanism conducted for each single-view window and the top-$k$most correlated multi-view windows. Adjusting the window sizes and top-$k$can minimise the complexity to$O((hw)^\frac{4}{3})$. Extensive experiments on the Real-IAD dataset under the multi-class setting validate the effectiveness of our approach, achieving state-of-the-art performance with an average improvement of+2.5$\uparrow$across10 metricsat the sample/image/pixel levels, using only18Mparameters and requiring fewer FLOPs and training time. The codes are available athttps://github.com/lewandofskee/MVAD. Haoyang He, Jiangning Zhang, Guanzhong Tian, Chengjie Wang 0001, Lei Xie 0007 |
IEEE Trans. Multim. | 1 |
| 2025 | PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud LearningabstractTransformers have revolutionized the point cloud learning task, but the quadratic complexity hinders its extension to long sequence and makes a burden on limited computational resources. The recent advent of RWKV, a fresh breed of deep sequence models, has shown immense potential for sequence modeling in NLP tasks. In this paper, we present PointRWKV, a model of linear complexity derived from the RWKV model in the NLP field with necessary modifications for point cloud learning tasks. Specifically, taking the embedded point patches as input, we first propose to explore the global processing capabilities within PointRWKV blocks using modified multi-headed matrix-valued states and a dynamic attention recurrence mechanism. To extract local geometric features simultaneously, we design a parallel branch to encode the point cloud efficiently in a fixed radius near-neighbors graph with a graph stabilizer. Furthermore, we design PointRWKV as a multi-scale framework for hierarchical feature learning of 3D point clouds, facilitating various downstream tasks. Extensive experiments on different point cloud learning tasks show our proposed PointRWKV outperforms the transformer- and mamba-based counterparts, while significantly saving about 42\% FLOPs, demonstrating the potential option for constructing foundational 3D models. Qingdong He, Jiangning Zhang, Jinlong Peng, Haoyang He, Xiangtai Li, Yabiao Wang, Chengjie Wang 0001 |
AAAI | 4 |
| 2025 | FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and ChallengingabstractZichen Tang, Haihong E, Ziyan Ma, Haoyang He, Jiacheng Liu, Zhongjun Yang, Zihua Rong, Rongjin Li, Kun Ji, Qing Huang, Xinyang Hu, Yang Liu, Qianhe Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zichen Tang, Haihong E, Ziyan Ma, Haoyang He, Zhongjun Yang, Zihua Rong, Rongjin Li, Kun Ji, Xinyang Hu, Qianhe Zheng |
ACL (1) | 4 |
| 2025 | MobileMamba: Lightweight Multi-Receptive Visual Mamba NetworkabstractPrevious research on lightweight models has primarily focused on CNNs and Transformer-based designs. CNNs, with their local receptive fields, struggle to capture long-range dependencies, while Transformers, despite their global modeling capabilities, are limited by quadratic computational complexity in high-resolution scenarios. Recently, state-space models have gained popularity in the visual domain due to their linear computational complexity. Despite their low FLOPs, current lightweight Mamba-based models exhibit suboptimal throughput. In this work, we propose the MobileMamba framework, which balances efficiency and performance. We design a three-stage network to enhance inference speed significantly. At a fine-grained level, we introduce the Multi-Receptive Field Feature Interaction (MRFFI) module, comprising the Long-Range Wavelet Transform-Enhanced Mamba (WTE-Mamba), Efficient Multi-Kernel Depthwise Convolution (MK-DeConv), and Eliminate Redundant Identity components. This module integrates multi-receptive field information and enhances high-frequency detail extraction. Additionally, we employ training and testing strategies to further improve performance and efficiency. MobileMamba achieves up to 83.6% on Top-1, surpassing existing state-of-the-art methods which is maximum ×21↑ faster than LocalVim on GPU. Extensive experiments on high-resolution downstream tasks demonstrate that MobileMamba surpasses current efficient models, achieving an optimal balance between speed and accuracy. Haoyang He, Jiangning Zhang, Xiaobin Hu, Zhenye Gan, Yabiao Wang, Chengjie Wang 0001, Yunsheng Wu, Lei Xie 0007 |
CVPR | 1 |
| 2025 | LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsabstractThe success of Large Language Models (LLMs) has inspired the development of Multimodal Large Language Models (MLLMs) for unified understanding of vision and language. However, the increasing model size and computational complexity of large-scale MLLMs (l-MLLMs) limit their use in resource-constrained scenarios. Although small-scale MLLMs (s-MLLMs) are designed to reduce computational costs, they typically suffer from performance degradation. To mitigate this limitation, we propose a novel LLaVA-KD framework to transfer knowledge from l-MLLMs to s-MLLMs. Specifically, we introduce Multimodal Distillation (MDist) to transfer teacher model's robust representations across both visual and linguistic modalities, and Relation Distillation (RDist) to transfer teacher model's ability to capture visual token relationships. Additionally, we propose a three-stage training scheme to fully exploit the potential of the proposed distillation strategy: 1) Distilled Pre-Training to strengthen the alignment between visual-linguistic representations in s-MLLMs, 2) Supervised Fine-Tuning to equip the s-MLLMs with multimodal understanding capacity, and 3) Distilled Fine-Tuning to refine s-MLLM's knowledge. Our approach significantly improves s-MLLMs performance without altering the model architecture. Extensive experiments and ablation studies validate the effectiveness of each proposed component. Code will be available at https://github.com/Fantasyele/LLaVA-KD. Jiangning Zhang, Haoyang He, Xinwei He 0001, Ao Tong, Zhenye Gan, Chengjie Wang 0001, Zhucun Xue, Yong Liu 0007, Xiang Bai |
ICCV | 3 |
| 2025 | $\mathcal{F}_{M}$ FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
Zichen Tang, Haihong E, Zhongjun Yang, Rongjin Li, Zihua Rong, Haoyang He, Zhuodi Hao, Xinyang Hu, Kun Ji, Ziyan Ma, Mengyuan Ji, Chenghao Ma, Qianhe Zheng, Zijian Xie, Shiyao Peng |
ICCV | 7 |
| 2025 | RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and ExplorationabstractOpen-set semantic mapping is crucial for openworld robots. Current mapping approaches either are limited by the depth range or only map beyond-range entities in constrained settings, where overall they fail to combine within-range and beyond-range observations. Furthermore, these methods make a trade-off between fine-grained semantics and efficiency. We introduce RayFronts, a unified representation that enables both dense and beyond-range efficient semantic mapping. RayFronts encodes task-agnostic openset semantics to both in-range voxels and beyond-range rays encoded at map boundaries, empowering the robot to reduce search volumes significantly and make informed decisions both within & beyond sensory range, while running at 8.84 Hz on an Orin AGX. Benchmarking the within-range semantics shows that RayFronts’s fine-grained image encoding provides 1.34× zero-shot 3D semantic segmentation performance while improving throughput by 16.5×. Traditionally, online mapping performance is entangled with other system components, complicating evaluation. We propose a planner-agnostic evaluation framework that captures the utility for online beyond-range search and exploration, and show RayFronts reduces search volume 2.2× more efficiently than the closest online baselines. Omar Alama, Avigyan Bhattacharya, Haoyang He, Seungchan Kim, Yuheng Qiu, Cherie Ho, Nikhil Varma Keetha, Sebastian A. Scherer |
IROS | 3 |
| 2025 | UltraVideo: High-Quality UHD Video Dataset with Comprehensive CaptionsabstractThe quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. %The growing demand for video applications sets higher requirements for high-quality video generation models. %For example, the generation of movie-level Ultra-High Definition (UHD) videos and the creation of 4K short video content. %However, the existing public datasets cannot support related research and applications. %In this paper, we first propose a high-quality open-sourced UHD-4K (22.4\% of which are 8K) text-to-video dataset named UltraVideo, which contains a wide range of topics (more than 100 kinds), and each video has 9 structured captions with one summarized caption (average of 824 words). %Specifically, we carefully design a highly automated curation process with four stages to obtain the final high-quality dataset: i) collection of diverse and high-quality video clips. ii statistical data filtering. iii) model-based data purification. iv) generation of comprehensive, structured captions. %In addition, we expand Wan to UltraWan-1K/-4K, which can natively generate high-quality 1K/4K videos with more consistent text controllability, demonstrating the effectiveness of our data curation.%We believe that this work can make a significant contribution to future research on UHD video generation. UltraVideo dataset and UltraWan models are available at https://xzc-zju.github.io/projects/UltraVideo. Zhucun Xue, Jiangning Zhang, Haoyang He, Yabiao Wang, Chengjie Wang 0001, Yong Li 0008, Xiangtai Li, Dacheng Tao |
NeurIPS | 4 |
| 2025 | Job Scheduling in Hybrid Clouds With Privacy Constraints: A Deep Reinforcement Learning ApproachabstractABSTRACT With the proliferation of cloud computing and the escalating demand for extensive data processing capabilities, an increasing number of enterprises are embracing hybrid cloud solutions. However, as more businesses move toward hybrid clouds, the need for effective solutions to privacy and security concerns becomes increasingly important. Although current scheduling approaches for cloud computing have addressed privacy protection to some extent, few have adequately considered the unique challenges posed by hybrid clouds. To address this gap, we propose a novel approach for scheduling jobs in hybrid clouds that prioritizes privacy protection. Our approach, called PH‐DRL, leverages Deep Reinforcement Learning (DRL) to intelligently allocate jobs to virtual machines, optimizing both privacy and Quality of Service (QoS), while minimizing response time. We present the detailed implementation of our approach and our experimental results demonstrate the superior performance of PH‐DRL in terms of privacy protection compared to existing methods. Haoyang He, Qingzhi Liu, Hao Wu 0017, Long Cheng 0003 |
Concurr. Comput. Pract. Exp. | 1 |
| 2025 | Real-time workflow scheduling in hybrid clouds with privacy and security constraints: A deep reinforcement learning approach
Haoyang He, Yang Hu 0009, Fang Fang 0007, Xin Ning 0001, Long Cheng 0003 |
Expert Syst. Appl. | 1 |
| 2025 | EMOv2: Pushing 5M Vision Model FrontierabstractThis work focuses on developing parameter-efficient and lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Our goal is to set up the new frontier of the 5 M magnitude lightweight model on various downstream tasks. Inverted Residual Block (IRB) serves as the infrastructure for lightweight CNNs, but no counterparts have been recognized by attention-based design. Our work rethinks the lightweight infrastructure of efficient IRB and practical components in Transformer from a unified perspective, extending CNN-based IRB to attention-based models and abstracting a one-residual Meta Mobile Block (MMBlock) for lightweight model design. Following neat but effective design criterion, we deduce a modern Improved Inverted Residual Mobile Block (i$^{2}$2RMB) and improve a hierarchical Efficient MOdel (EMOv2) with no elaborate complex structures. Considering the imperceptible latency for mobile users when downloading models under 4 G/5 G bandwidth and ensuring model performance, we investigate the performance upper limit of lightweight models with a magnitude of 5 M. Extensive experiments on various vision recognition, dense prediction, and image generation tasks demonstrate the superiority of our EMOv2 over state-of-the-art methods, e.g., EMOv2-1 M/2M/5 M achieve 72.3, 75.8, and 79.4 Top-1 that surpass equal-order CNN-/Attention-based models significantly. At the same time, EMOv2-5 M equipped RetinaNet achieves 41.5 mAP for object detection tasks that surpasses the previous EMO-5 M by +2.6$\uparrow$↑ . When employing the more robust training recipe, our EMOv2-5M eventually achieves 82.9 Top-1 accuracy, which elevates the performance of 5M magnitude models to a new level. Jiangning Zhang, Haoyang He, Zhucun Xue, Yabiao Wang, Chengjie Wang 0001, Yong Liu 0007, Xiangtai Li, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | A Diffusion-Based Framework for Multi-Class Anomaly DetectionabstractReconstruction-based approaches have achieved remarkable outcomes in anomaly detection. The exceptional image reconstruction capabilities of recently popular diffusion models have sparked research efforts to utilize them for enhanced reconstruction of anomalous images. Nonetheless, these methods might face challenges related to the preservation of image categories and pixel-wise structural integrity in the more practical multi-class setting. To solve the above problems, we propose a Difusion-based Anomaly Detection (DiAD) framework for multi-class anomaly detection, which consists of a pixel-space autoencoder, a latent-space Semantic-Guided (SG) network with a connection to the stable diffusion’s denoising network, and a feature-space pre-trained feature extractor. Firstly, The SG network is proposed for reconstructing anomalous regions while preserving the original image’s semantic information. Secondly, we introduce Spatial-aware Feature Fusion (SFF) block to maximize reconstruction accuracy when dealing with extensively reconstructed areas. Thirdly, the input and reconstructed images are processed by a pre-trained feature extractor to generate anomaly maps based on features extracted at different scales. Experiments on MVTec-AD and VisA datasets demonstrate the effectiveness of our approach which surpasses the state-of-the-art methods, e.g., achieving 96.8/52.6 and 97.2/99.0 (AUROC/AP) for localization and detection respectively on multi-class MVTec-AD dataset. Code will be available at https://lewandofskee.github.io/projects/diad. Haoyang He, Jiangning Zhang, Xuhai Chen, Zhishan Li, Xu Chen 0024, Yabiao Wang, Chengjie Wang 0001, Lei Xie 0007 |
AAAI | 1 |
| 2024 | MARS: Multi-Agent Deep Reinforcement Learning for Real-Time Workflow Scheduling in Hybrid Clouds with Privacy ProtectionabstractScheduling workflows in hybrid cloud environments presents significant challenges due to the inherent complexity of workflows and the dynamic nature of cloud resources. This complexity is further increased when attempting to balance workflow performance with privacy protection. Recent efforts have leveraged deep reinforcement learning (DRL) to address these challenges. However, most of these approaches rely on single-agent models, which can lead to security issues and scalability problems due to their centralized processing. Specifically, the properties of workflows are transferred to the single agent, which risks leaking privacy information. Our paper addresses these issues by introducing MARS, a real-time workflow scheduling method that prioritizes privacy protection in hybrid clouds. MARS leverages multi-agent deep reinforcement learning (MADRL) to optimize the workflow scheduling of cloud virtual machines (VMs). The benefit of our solution is that it relies on the collaborative learning of multi-agents on multiple VMs, which could assign user data to specific cloud servers for privacy protection while sharing training experiences between agents. In our implementation, MARS aims to reduce workflow completion time and operational costs while complying with strict privacy protection guidelines. The experimental results demonstrate that MARS can significantly surpass existing methods, reducing makespan by an average of $53.18 \%$ and costs by $61.98 \%$ compared to basic techniques, and achieving $20.26 \%$ and $25.71 \%$ improvements over the latest advanced methods, respectively. Long Cheng 0003, Haoyang He, Qingzhi Liu, Zhiming Zhao, Fang Fang 0007 |
ICPADS | 2 |
| 2024 | MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly DetectionabstractRecent advancements in anomaly detection have seen the efficacy of CNN- and transformer-based approaches. However, CNNs struggle with long-range dependencies, while transformers are burdened by quadratic computational complexity. Mamba-based models, with their superior long-range modeling and linear efficiency, have garnered substantial attention. This study pioneers the application of Mamba to multi-class unsupervised anomaly detection, presenting MambaAD, which consists of a pre-trained encoder and a Mamba decoder featuring (Locality-Enhanced State Space) LSS modules at multi-scales. The proposed LSS module, integrating parallel cascaded (Hybrid State Space) HSS blocks and multi-kernel convolutions operations, effectively captures both long-range and local information. The HSS block, utilizing (Hybrid Scanning) HS encoders, encodes feature maps into five scanning methods and eight directions, thereby strengthening global connections through the (State Space Model) SSM. The use of Hilbert scanning and eight directions significantly improves feature sequence modeling. Comprehensive experiments on six diverse anomaly detection datasets and seven metrics demonstrate state-of-the-art performance, substantiating the method's effectiveness. The code and models are available at https://lewandofskee.github.io/projects/MambaAD. Haoyang He, Yuhu Bai, Jiangning Zhang, Qingdong He, Zhenye Gan, Chengjie Wang 0001, Xiangtai Li, Guanzhong Tian, Lei Xie 0007 |
NeurIPS | 1 |
| 2023 | Infrared Moving Small Target Detection Based on Consistency of Sparse TrajectoryabstractInfrared search and track (IRST) systems require reliable detection of small targets in complex backgrounds. Outlier based methods are prone to high false positive rates due to the resemblance of point-like background features to small targets. The difference image-based method is an effective approach for suppressing point-like background interference; however, it has limitations in detecting slow-moving targets. In this letter, a novel sparse trajectory is proposed for moving target detection in IR videos. With a trajectory growing strategy, two kinds of trajectories from difference images, namely short sparse trajectories and long sparse trajectories, are correlated to avoid the slow-moving targets being dismissed. The strategy matches the trajectories based on the sparse trajectory intensity composed of similarity measures and optical flow consistency. Finally, real targets are extracted from candidate trajectories using trajectory filtering. Experimental results show that, in the scene with point-like background features, our method achieves the best detection rate and lowest false alarm compared to state-of-the-art methods. Mo Wu, Xiubin Yang, Zongqiang Fu, Haoyang He, Jiamin Du, Ziming Tu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Towards accurate dense pedestrian detection via occlusion-prediction aware label assignment and hierarchical-NMS
Haoyang He, Zhishan Li, Guanzhong Tian, Lei Xie 0007, Shan Lu 0009 |
Pattern Recognit. Lett. | 1 |