Haocheng Xu

dblp:251/3400 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Characterizing State Space Model and Hybrid Language Model Performance with Long Context
abstract
Emerging applications such as AR are driving demands for machine intelligence capable of processing continuous and/or long-context inputs on local devices. However, currently dominant models based on Transformer architecture suffers from the quadratic computational and memory overhead, which hinders applications required to process long contexts. This has spurred a paradigm shift towards new architectures like State Space Models (SSMs) and SSM-Transformer hybrid models, which provide near-linear scaling. The near-linear scaling enabled efficient handling of millions of tokens while delivering high performance in recent studies. Although such works present promising results, their workload characteristics in terms of computational performance and hardware resource requirements are not yet thoroughly explored, which limits our understanding of their implications to the system level optimizations. To address this gap, we present a comprehensive, comparative benchmarking of carefully selected Transformers, SSMs, and hybrid models specifically for long-context inference on consumer and embedded GPUs. Our analysis shows that SSMs are well-suited for on-device AI on consumer and embedded GPUs for long context inferences. While Transformers are up to $1.9 \times$ faster at short sequences ($ \lt 8 \mathrm{~K}$ tokens), SSMs demonstrate a dramatic performance inversion, becoming up to $4 \times$ faster at very long contexts ($\sim 57 \mathrm{~K}$ tokens), thanks to their linear computational complexity and $\boldsymbol{\sim} \mathbf{6 4 \%}$ reduced memory footprint. Our operator-level analysis reveals that custom SSM kernels like selective scan despite being hardware-aware to minimize memory IO, dominate the inference runtime on edge platforms, accounting for over $55 \%$ of latency due to their sequential, element-wise nature. To foster further research, we are sharing CPU/GPU profiling traces and have made our characterization framework SSM-Scope open-sourced at https://github.com/sapmitra/ssm-scope
Saptarshi Mitra, Rachid Karami, Haocheng Xu, Sitao Huang, Hyoukjun Kwon
ISPASS3
2026 Dual-network proportional variable speed limit control for mixed traffic flows with multi-level automated vehicles
Ci Liang, Yusheng Ci, Haocheng Xu
Expert Syst. Appl.4
2026 Improving transferability of adversarial examples with mixed representation attack
Zilin Tian, Haocheng Xu, Huosheng Xu
Vis. Comput.4
2025 RSEND: Retinex-based Squeeze and Excitation Network with Dark Region Detection for Efficient Low Light Image Enhancement
abstract
Images captured under low-light scenarios often suffer from low quality. Efficient low-light image enhancement with mobile computing has become an urgent need. Previous CNN-based low-light image enhancement methods often involve using Retinex theory. Nevertheless, most of them do not perform well in complicated datasets like LOL-v2 while using too much computational resources. Besides, some of these methods require sophisticated training at different stages, making the procedure even more time-consuming and tedious. In this paper, we propose an accurate, concise, and one-stage Retinex theory-based framework with a novel dark region detection module and Squeeze and Excitation blocks for enhanced detail retention, RSEND, for efficient low-light image enhancement. RSEND first divides the low-light image into the illumination map and reflectance map, then detects the different dark regions in the illumination map and performs light enhancement. After this step, it refines the enhanced gray-scale image and does element-wise matrix multiplication with the reflectance map. By denoising the output it has from the previous step, it obtains the final result. In all the steps, RSEND utilizes Squeeze and Excitation network to better capture the details. Comprehensive quantitative and qualitative experiments show that our efficient Retinex model significantly outperforms other CNN-based state-of-the-art models, achieving a PSNR improvement ranging from 1.69 dB to 3.63 dB in different datasets. Compared to Transformer-based models, RSEND achieves higher PSNR values ranging from 1.22 dB to 2.44 dB in the LOL-v2-real dataset. Importantly, RSEND achieves these performance improvements with remarkable efficiency, utilizing only 0.41 million parameters, which represents a substantial reduction (3.93–9.78×) in computational resources compared to existing state-of-the-art methods. The code can be found at https://github.com/jeffconqueror/RSEND/tree/main.
Jingcheng Li, Ye Qiao, Haocheng Xu, Sitao Huang
IJCNN3
2025 MONAS: Efficient Zero-Shot Neural Architecture Search for MCUs
abstract
Neural Architecture Search (NAS) has proven effective in discovering new Convolutional Neural Network (CNN) architectures, particularly for scenarios with well-defined ac-curacy optimization goals. However, previous approaches often involve time-consuming training on super networks or intensive architecture sampling and evaluations. Although various zero-cost proxies correlated with CNN model accuracy have been proposed for efficient architecture search without training, their lack of hardware consideration makes it challenging to target highly resource-constrained edge devices such as microcontroller units (MCUs). To address these challenges, we introduce MONAS, a novel hardware-aware zero-shot NAS framework specifically designed for MCUs in edge computing. MONAS incorporates hardware optimality considerations into the search process through our proposed MCU hardware latency estimation model. By combining this with specialized performance indicators (proxies), MONAS identifies optimal neural architectures without incurring heavy training and evaluation costs, optimizing for both hardware latency and accuracy under resource constraints. MONAS achieves up to a 1104× improvement in search efficiency over previous work targeting MCUs and can discover CNN models with over 3.23× faster inference on MCUs while maintaining similar accuracy compared to more general NAS approaches.
Ye Qiao, Haocheng Xu, Sitao Huang
IJCNN2
2025 Unified Cross-Structural Motion Retargeting for Humanoid Characters
abstract
Motion retargeting for animation characters has potential applications in fields such as animation production and virtual reality. However, current methods either assume that the source and target characters have the same skeletal structure, or require designing and training specific model architectures for each structure. In this article, we aim to address the challenge of motion retargeting across previously unseen skeletal structures with a unified dynamic graph network. The proposed approach utilizes a dynamic graph transformation module to dynamically transfer latent motion features to different structures. We also take into consideration for intricate hand movements and model both torso and hand joints as graphs in a unified manner for whole-body motion retargeting. Our model allows the use of motion data from different structures to train a unified model and learns cross-structural motion retargeting in an unsupervised manner with unpaired data. Experimental results demonstrate the superiority of the proposed method in terms of data efficiency and performance on both seen and unseen structures.
Zhike Chen, Haocheng Xu, Songcen Xu, Rong Xiong, Yue Wang 0020
IEEE Trans. Vis. Comput. Graph.3
2024 Semantics-Aware Motion Retargeting with Vision-Language Models
abstract
Capturing and preserving motion semantics is essential to motion retargeting between animation characters. However, most of the previous works neglect the semantic information or rely on human-designed joint-level representations. Here, we present a novel Semantics-aware Motion reTargeting (SMT) method with the advantage of vision-language models to extract and maintain meaningful motion semantics. We utilize a differentiable module to ren-der 3D motions. Then the high-level motion semantics are incorporated into the motion retargeting process by feeding the vision-language model with the rendered images and aligning the extracted semantic embeddings. To en-sure the preservation of fine-grained motion details and high-level semantics, we adopt a two-stage pipeline consisting of skeleton-aware pretraining and fine-tuning with semantics and geometry constraints. Experimental results show the effectiveness of the proposed method in producing high-quality motion retargeting results while accurately preserving motion semantics. Project page can be found at https://sites.google.com/view/smtnet.
Zhike Chen, Haocheng Xu, Songcen Xu, Zhensong Zhang, Yue Wang 0020, Rong Xiong
CVPR3
2024 MicroNAS: Zero-Shot Neural Architecture Search for MCUs
abstract
Neural architecture search (NAS) effectively discovers new convolutional neural network (CNN) architectures, particularly for accuracy optimization. However, prior approaches often require resource-intensive training on super networks or extensive architecture evaluations, limiting practical applications. To address these challenges, we propose MicroNAS, a hardware-aware zero-shot NAS framework designed for microcontroller units (MCVs) in edge computing. MicroNAS considers target hardware optimality during the search, utilizing specialized performance indicators to identify optimal neural architectures without heavy computational costs. Compared to previous works, MicroNAS achieves up to$1104\times$improvement in search efficiency and discovers models with over$3.23\times$faster MCU inference while maintaining similar accuracy.
Ye Qiao, Haocheng Xu, Sitao Huang
DATE2
2024 HyperDetect: A Real-Time Hyperdimensional Solution for Intrusion Detection in IoT Networks
abstract
Network-based security has emerged as an increasingly critical challenge in the domain of the Internet of Things (IoT). A number of network intrusion detection systems (NIDS), typically relying on sophisticated machine learning (ML) algorithms, have been proposed to monitor network traffic and detect malicious activity. However, these NIDS designs require extensive memory and computational power, exceeding the capability of today’s IoT devices, and often fail to provide timely detection of network attacks. To tackle this issue, we propose HyperDetect, the first attempt at NIDS modeling that leverages the highly efficient and parallel operations of brain-inspired hyperdimensional computing (HDC). Our innovative model updating method effectively mitigates model saturation and significantly reduces the number of retraining iterations needed to reach convergence. Additionally, we employ a novel dynamic encoding technique to regenerate insignificant dimensions, considerably lowering the dimensionalities required to achieve high-quality performance and further accelerating the learning process. HyperDetect delivers on average 5.02× faster training and 31.83× faster inference compared to state-of-the-art (SOTA) learning approaches on a wide range of network intrusion classification tasks. We also extensively evaluate HyperDetect on embedded hardware to demonstrate its low-latency and resource-efficient characteristics.
Junyao Wang 0001, Haocheng Xu, Yonatan Gizachew Achamyeleh, Sitao Huang, Mohammad Abdullah Al Faruque
IEEE Internet Things J.2
2019 Time-aware Session Embedding for Click-Through-Rate Prediction
abstract
TV series correlation computing is one of the most important tasks of personalized online streaming services. With the relevance of TV series and viewer feedback, we can calculate the TV series correlation table based on the viewer's implicit feedback which does not perform well for the newly added "cold start" TV series. In this paper, we aim to improve correlation computing within the cold-start phase. We propose a framework named Time-aware Session Embedding (TSE), with Item Embedding in Session and Time Decay Factor for a multimodal recommendation. We apply an lower- dimensional vector as item embedding and calculate their factor considering the time decay. The framework performed well in the Content-based Video Relevance Prediction Challenge and we get the first place in this competition.
Qidi Xu, Haocheng Xu, Weilong Chen, Chaojun Han, Haoyang Li 0002, Wenxin Tan, Fumin Shen, Heng Tao Shen
ACM Multimedia2