VLDB 2026 Research / reviewers in the wild / expert
Haoxin Wang 0003
dblp:203/0135-3
· DBLP profile ↗
20ranked-venue papers
9as first author
14since 2021 · last 2025
0000-0002-8732-6200ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 7 first-author · 5 since 2021Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MamBEV: Enabling State Space Models to Learn Birds-Eye-View Representationsabstract3D visual perception tasks, such as 3D detection from multi-camera images, are essential components of autonomous driving and assistance systems. However, designing computationally efficient methods remains a significant challenge. In this paper, we propose a Mamba-based framework called MamBEV, which learns unified Bird's Eye View (BEV) representations using linear spatio-temporal SSM-based attention. This approach supports multiple 3D perception tasks with significantly improved computational and memory efficiency. Furthermore, we introduce SSM based cross-attention, analogous to standard cross attention, where BEV query representations can interact with relevant image features. Extensive experiments demonstrate MamBEV's promising performance across diverse visual perception metrics, highlighting its advantages in input scaling efficiency compared to existing benchmark models. Hongyu Ke, Jack Morris, Kentaro Oguchi 0001, Xiaofei Cao, Yongkang Liu 0005, Haoxin Wang 0003, Yi Ding 0010 |
ICLR | 6 |
| 2025 | lm-Meter: Unveiling Runtime Inference Latency for On-Device Language ModelsabstractLarge Language Models (LLMs) are increasingly integrated into everyday applications, but their prevalent cloud-based deployment raises growing concerns around data privacy and long-term sustainability. Running LLMs locally on mobile and edge devices (on-device LLMs) offers the promise of enhanced privacy, reliability, and reduced communication costs. However, realizing this vision remains challenging due to substantial memory and compute demands, as well as limited visibility into performance-efficiency trade-offs on resource-constrained hardware. We propose lm-Meter, the first lightweight, online latency profiler tailored for on-device LLM inference. lm-Meter captures fine-grained, real-time latency at both phase (e.g., embedding, prefill, decode, softmax, sampling) and kernel levels without auxiliary devices. We implement lm-Meter on commercial mobile platforms and demonstrate its high profiling accuracy with minimal system overhead, e.g., only 2.58% throughput reduction in prefill and 0.99% in decode under the most constrained Powersave governor. Leveraging lm-Meter, we conduct comprehensive empirical studies revealing phase- and kernel-level bottlenecks in on-device LLM inference, quantifying accuracy-efficiency trade-offs, and identifying systematic optimization opportunities. lm-Meter provides unprecedented visibility into the runtime behavior of LLMs on constrained platforms, laying the foundation for informed optimization and accelerating the democratization of on-device LLM systems. Code and tutorials are available at github.com/amai-gsu/LM-Meter. Haoxin Wang 0003, Xiaolong Tu, Hongyu Ke, Huirong Chai, Kyungtae Han |
SEC | 1 |
| 2025 | TinyBEV: Compact Temporal Fusion for Multi-View 3D PerceptionabstractMulti-view camera-based 3D object detection through unified Bird's Eye View (BEV) representation has become popular for autonomous driving due to its low cost, but efficiently inferring precise spatial and temporal information from cameras alone remains a significant challenge. Transformer-based approaches have shown substantial performance improvements but have the drawback of quadratic memory complexity — making these architectures ill-suited for edge deployment. Recently, State Space Models (SSMs) offer a more favorable balance of computational efficiency and performance in 2D vision, suggesting that they could help here as well. We present TinyBEV, an efficient BEV framework for multi-view 3D perception. For spatial modeling, we replace cross attention with SSMs that fusing BEV and camera images with linear complexity. For temporal modeling, we adopt a lightweight, linear-complexity history-fusion scheme that uses explicit time conditioning and channel-level aggregation instead of cross-frame attention. Both fusion strategies follow small constant scaling with respect to history length and enabling edge-friendly deployment. Experiments on NuScenes datasets demonstrate that TinyBEV is comparable with other state-of-the-art methods across diverse visual perception metrics with advantages in computational efficiency. Hongyu Ke, Jack Morris, Yongkang Liu 0005, Satoshi Kitai, Kentaro Oguchi 0001, Yi Ding 0041, Haoxin Wang 0003 |
SEC | 7 |
| 2025 | PlatformX: An End-to-End Transferable Platform for Energy-Efficient Neural Architecture SearchabstractHardware-Aware Neural Architecture Search (HW-NAS) has emerged as a powerful tool for designing efficient deep neural networks (DNNs) tailored to edge devices. However, existing methods remain largely impractical for real-world deployment due to their high time cost, extensive manual profiling, and poor scalability across diverse hardware platforms with complex, device-specific energy behavior. Xiaolong Tu, Kyungtae Han, Onur Altintas, Haoxin Wang 0003 |
SEC | 5 |
| 2023 | High Definition Map Data Optimization for Autonomous Driving in Vehicular Named Data NetworksabstractHigh-definition (HD) map is an essential building block in the autonomous driving era, which enables fine-grained environmental awareness, exact localization, and route planning. However, because HD maps include rich, multidimensional information, the volume of HD map data is enormous, making it expensive and time-consuming to transmit on vehicular networks. Therefore, in this paper, we propose a data optimization scheme for effective HD map updates in vehicular named data networking (NDN) scenarios. We formulate the HD map data optimization problem as a convex optimization problem and solve it with modified convolutional neural networks (CNNs) from YOLOX's real-time object detection system. Specifically, we modify the YOLOX object detection algorithm to detect and compress redundant pixels in local map data before transmission to the MEC server. To deploy our proposed scheme, we construct a vehicular NDN environment for data collection, processing, and transmission using the CARLA simulator and robot operating system 2 (ROS2). Extensive simulations show that our proposed scheme can significantly reduce the transmission data size and time by 48.25% - 65.78% and 46.85% - 78.84% compared with state-of-the-art HD map update techniques like RLSS, Pro-RTT, and Loss-based systems. Daniel Mawunyo Doe, Kyungtae Han, Haoxin Wang 0003, Jiang (Linda) Xie, Zhu Han 0001 |
ICC | 4 |
| 2023 | EPAM: A Predictive Energy Model for Mobile AIabstractArtificial intelligence (AI) has enabled a new paradigm of smart applications - changing our way of living entirely. Many of these AI-enabled applications have very stringent latency requirements, especially for applications on mobile devices (e.g., smartphones, wearable devices, and vehicles). Hence, smaller and quantized deep neural network (DNN) models are developed for mobile devices, which provide faster and more energy-efficient computation for mobile AI applications. However, how AI models consume energy in a mobile device is still unexplored. Predicting the energy consumption of these models, along with their different applications, such as vision and non-vision, requires a thorough investigation of their behavior using various processing sources. In this paper, we introduce a comprehensive study of mobile AI applications considering different DNN models and processing sources, focusing on computational resource utilization, delay, and energy consumption. We measure the latency, energy consumption, and memory usage of all the models using four processing sources through extensive experiments. We explain the challenges in such investigations and how we propose to overcome them. Our study highlights important insights, such as how mobile AI behaves in different applications (vision and non-vision) using CPU, GPU, and NNAPI. Finally, we propose a novel Gaussian process regression-based general predictive energy model based on DNN structures, computation resources, and processors, which can predict the energy for each complete application cycle irrespective of device configuration and application. This study provides crucial facts and an energy prediction mechanism to the AI research community to help bring energy efficiency to mobile AI applications. Anik Mallik, Haoxin Wang 0003, Jiang (Linda) Xie, Kyungtae Han |
ICC | 2 |
| 2023 | Poster: Real-Time Object Substitution for Mobile Diminished Reality with Edge ComputingabstractDiminished Reality (DR) is considered as the conceptual counterpart to Augmented Reality (AR), and has recently gained increasing attention from both industry and academia. Unlike AR which adds virtual objects to the real world, DR allows users to remove physical content from the real world. When combined with object replacement technology, it presents an further exciting avenue for exploration within the metaverse. Although a few researches have been conducted on the intersection of object substitution and DR, there is no real-time object substitution for mobile diminished reality architecture with high quality. In this paper, we propose an end-to-end architecture to facilitate immersive and real-time scene construction for mobile devices with edge computing. Hongyu Ke, Haoxin Wang 0003 |
SEC | 2 |
| 2023 | Unveiling Energy Efficiency in Deep Learning: Measurement, Prediction, and Scoring Across Edge DevicesabstractToday, deep learning optimization is primarily driven by research focused on achieving high inference accuracy and reducing latency. However, the energy efficiency aspect is often overlooked, possibly due to a lack of sustainability mindset in the field and the absence of a holistic energy dataset. In this paper, we conduct a threefold study, including energy measurement, prediction, and efficiency scoring, with an objective to foster transparency in power and energy consumption within deep learning across various edge devices. Firstly, we present a detailed, first-of-its-kind measurement study that uncovers the energy consumption characteristics of on-device deep learning. This study results in the creation of three extensive energy datasets for edge devices, covering a wide range of kernels, state-of-the-art DNN models, and popular AI applications. Secondly, we design and implement the first kernel-level energy predictors for edge devices based on our kernel-level energy dataset. Evaluation results demonstrate the ability of our predictors to provide consistent and accurate energy estimations on unseen DNN models. Lastly, we introduce two scoring metrics, PCS and IECS, developed to convert complex power and energy consumption data of an edge device into an easily understandable manner for edge device end-users. We hope our work can help shift the mindset of both end-users and the research community towards sustainability in edge computing, a principle that drives our research. Find data, code, and more up-to-date information at https://amai-gsu.github.io/DeepEn2023. Xiaolong Tu, Anik Mallik, Kyungtae Han, Onur Altintas, Haoxin Wang 0003, Jiang (Linda) Xie |
SEC | 6 |
| 2023 | DSORL: Data Source Optimization With Reinforcement Learning Scheme for Vehicular Named Data NetworksabstractHighly-dynamic (HD) map is an indispensable building block in the future of autonomous driving, allowing for fine-grained environmental awareness, precise localization, and route planning. However, since HD maps include rich, multidimensional information, the volume of HD map data is substantial and cannot be transmitted frequently by several vehicles over vehicular networks in real-time. Therefore, in this paper, we propose a data source selection scheme for effective HD map transmissions in vehicular named data networking (NDN) scenarios. To achieve our goal, we created a vehicular NDN environment for data collection, processing, and transmission using the CARLA simulator and robot operating system 2 (ROS2). Next, due to our vehicular NDN’s dynamic and complex nature, we formulate the data source selection problem as a Markov decision process (MDP) and solve it using a reinforcement learning approach. For simplicity, we termed our proposed scheme data source optimization with reinforcement learning (DSORL), which selects suitable vehicles for HD map data transmission to MEC servers. The experiment results indicate that our suggested method outperformed existing baseline schemes, such as RLSS, Pro-RTT, and HDM-RTT, across all performance criteria in the evaluation. For instance, the system throughput increases by$65\%-72.68\%$compared to other baseline systems. Similarly, the proposed approach can minimize packet loss rate, data size, and transmission time by up to 60.6%, 77.5%, and 54.1%, respectively. Daniel Mawunyo Doe, Kyungtae Han, Haoxin Wang 0003, Jiang (Linda) Xie, Zhu Han 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | LEAF + AIO: Edge-Assisted Energy-Aware Object Detection for Mobile Augmented RealityabstractToday very few deep learning-based mobile augmented reality (MAR) applications are applied in mobile devices because they are significantly energy-guzzling. In this paper, we design an edge-based energy-aware MAR system that enables MAR devices to dynamically change their configurations, such as CPU frequency, computation model size, and image offloading frequency based on user preferences, camera sampling rates, and available radio resources. Our proposed dynamic MAR configuration adaptations can minimize the per frame energy consumption of multiple MAR clients without degrading their preferred MAR performance metrics, such as latency and detection accuracy. To thoroughly analyze the interactions among MAR configurations, user preferences, camera sampling rate, and energy consumption, we propose, to the best of our knowledge, the first comprehensive analytical energy model for MAR devices. Based on the proposed analytical model, we design a LEAF optimization algorithm to guide the MAR configuration adaptation and server radio resource allocation. An image offloading frequency orchestrator, coordinating with the LEAF, is developed to adaptively regulate the edge-based object detection invocations and to further improve the energy efficiency of MAR devices. Extensive evaluations are conducted to validate the performance of the proposed analytical model and algorithms. Haoxin Wang 0003, BaekGyu Kim, Jiang (Linda) Xie, Zhu Han 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | EdgeMap: CrowdSourcing High Definition Map in Automotive Edge ComputingabstractHigh definition (HD) map needs to be updated frequently to capture road changes, which is constrained by limited specialized collection vehicles. To maintain an up-to-date map, we explore crowdsourcing data from connected vehicles. Updating the map collaboratively is, however, challenging under constrained transmission and computation resources in dynamic networks. In this paper, we propose EdgeMap, a crowdsourcing HD map to minimize the usage of network resources while maintaining the latency requirements. We design a DATE algorithm to adaptively offload vehicular data on a small time scale and reserve network resources on a large time scale, by leveraging the multi-agent deep reinforcement learning and Gaussian process regression. We evaluate the performance of EdgeMap with extensive network simulations in a time-driven end-to-end simulator. The results show that EdgeMap reduces more than 30% resource usage as compared to state-of-the-art solutions. Qiang Liu 0013, Haoxin Wang 0003 |
ICC | 3 |
| 2022 | Poster: Enabling High-Fidelity and Real-Time Mobility Digital Twin with Edge ComputingabstractA Mobility Digital Twin is an emerging implementation of Digital Twin in the transportation domain, and has been attracting extensive attention from both industry and academia. Although a few research have been conducted on the mobility digital twin, there is no systematic work with an end-to-end digital twin model construction framework. In this paper, we propose an end-to-end system framework, including sensory data collection, offloading, and processing, that aims to facilitate a high-fidelity and real-time digital twin model construction for connected and automated vehicles. Additionally, preliminary experiments are conducted to demonstrate our research motivation and to guide the future system framework design. Haoxin Wang 0003, Zhipeng Cai 0001, Kyungtae Han |
SEC | 2 |
| 2022 | Mobility Digital Twin: Concept, Architecture, Case Study, and Future ChallengesabstractA Digital Twin is a digital replica of a living or nonliving physical entity, and this emerging technology attracted extensive attention from different industries during the past decade. Although a few Digital Twin studies have been conducted in the transportation domain very recently, there is no systematic research with a holistic framework connecting various mobility entities together. In this study, a mobility digital twin (MDT) framework is developed, which is defined as an artificial intelligence (AI)-based data-driven cloud–edge–device framework for mobility services. This MDT consists of three building blocks in the physical space (namely,Human,Vehicle, andTraffic), and their associated Digital Twins in the digital space. An example cloud–edge architecture is built with Amazon Web Services (AWS) to accommodate the proposed MDT framework and to fulfill its digital functionalities of storage, modeling, learning, simulation, and prediction. A case study of the personalized adaptive cruise control (P-ACC) system is conducted, which integrates the key microservices of all three digital building blocks of the MDT framework: 1) theHuman Digital Twinwith user management and driver type classification; 2) theVehicle Digital Twinwith cloud-based advanced driver-assistance systems (ADAS); and 3) theTraffic Digital Twinwith traffic flow monitoring and variable speed limit. Future challenges of the proposed MDT framework are discussed toward the end of the article, including standardization, AI for computing, public or private cloud service, and network heterogeneity. Ziran Wang, Kyungtae Han, Haoxin Wang 0003, Akila Ganlath, Nejib Ammar, Prashant Tiwari |
IEEE Internet Things J. | 4 |
| 2021 | You Can Enjoy Augmented Reality While Running Around: An Edge-based Mobile AR System
Haoxin Wang 0003, Jiang (Linda) Xie |
SEC | 1 |
| 2020 | User Preference Based Energy-Aware Mobile AR System with Edge ComputingabstractThe advancement in deep learning and edge computing has enabled intelligent mobile augmented reality (MAR) on resource limited mobile devices. However, today very few deep learning based MAR applications are applied in mobile devices because they are significantly energy-guzzling. In this paper, we design a user preference based energy-aware edge-based MAR system that enables MAR clients to dynamically change their configuration parameters, such as CPU frequency and computation model size, based on their user preferences, camera sampling rates, and available radio resources at the edge server. Our proposed dynamic MAR configuration adaptations can minimize the per frame energy consumption of multiple MAR clients without degrading their preferred MAR performance metrics, such as service latency and detection accuracy. To thoroughly analyze the interactions among MAR configuration parameters, user preferences, camera sampling rate, and per frame energy consumption, we propose, to the best of our knowledge, the first comprehensive analytical energy model for MAR clients. Based on the proposed analytical model, we develop a LEAF optimization algorithm to guide the MAR configuration adaptation and server radio resource allocation. Extensive evaluations are conducted to validate the performance of the proposed analytical model and LEAF algorithm. Haoxin Wang 0003, Jiang (Linda) Xie |
INFOCOM | 1 |
| 2019 | How Is Energy Consumed in Smartphone Deep Learning Apps? Executing Locally vs. RemotelyabstractApplying deep learning to object detection provides the capability to accurately detect and classify complex objects in the real world. However, currently, few mobile applications use deep learning because such technology is computation- and energy-intensive. This paper, to the best of our knowledge, presents the first detailed experimental study of the smartphone's energy consumption and the detection latency of executing deep Convolutional Neural Networks (CNN) optimized object detec- tion, either locally on the smartphone or remotely on an edge server. We experiment with a variety of smartphones, obtaining different levels of computation capacities, in order to ensure that we are not profiling a specific device. Our detailed measurements refine the energy analysis of smartphones and reveal some interesting perspectives regarding the energy consumption of executing the deep CNN optimized object detection. We believe that these findings will guide the design of energy efficient processing pipeline of the CNN optimized object detection. Haoxin Wang 0003, BaekGyu Kim, Jiang (Linda) Xie, Zhu Han 0001 |
GLOBECOM | 1 |
| 2019 | E-Auto: A Communication Scheme for Connected Vehicles with Edge-Assisted Autonomous DrivingabstractWith the rapid advancement of automobile industry, autonomous driving in connected vehicles are expected to be the key technology to satisfy the expansion of human demands on more comfortable and safer driving experience. However, only on-board computation resources are insufficient to satisfy tough computation requirements of achieving full or even high automation. Therefore, autonomous driving with cloud/edge participation is desirable. In this paper, we propose E-Auto, a novel communication scheme to enable fast, stable, and accurate edge-assisted autonomous driving service for connected vehicles within any road types (e.g., driving on highway with very high speed or local roads with slow speed due to traffic congestion). In addition, as two key components of the proposed E-Auto scheme, a service period allocation algorithm and a frame resolution selection algorithm are designed to guarantee a sufficient frame rate for connected vehicles acquiring either uplink application (offload camera captured frames to the edge server) or downlink application (download entertainment videos). Through network simulations, we evaluate the performance of the proposed E-Auto scheme. Simulation results demonstrate that E-Auto can provide a high frame rate and low energy consumption autonomous driving service for connected vehicles. Haoxin Wang 0003, BaekGyu Kim, Jiang (Linda) Xie, Zhu Han 0001 |
ICC | 1 |
| 2018 | A Smart Service Rebuilding Scheme across Cloudlets via Mobile AR Frame Feature MappingabstractMobile edge computing platforms, such as cloudlets, bring computation resources closer to mobile users, as compared to the cloud, which decreases the end-to-end network latency. This benefit enables a myriad of real-time mobile applications, especially augmented reality (AR), that require low latency and high computation power. However, when mobile users move away from the attached cloudlet, the offloaded services have to be migrated or rebuilt on a new nearby cloudlet. However, this service rebuilding process takes a lot of time and may deteriorate user experience. In this paper, we propose a smart service rebuilding scheme which seamlessly restores the offloading services on the target cloudlet while the mobile user is moving. The service rebuilding process includes the radio handoff stage and service handoff stage. A seamless service rebuilding process is achieved via predicting user's target cloudlet before being triggered a radio handoff, by leveraging extracted features from the captured frames of the mobile user's camera. Furthermore, based on the proposed service rebuilding scheme, we design a feature mapping algorithm to achieve a high prediction precision and a short prediction latency. We implement our scheme on a testbed and conduct experiments using real world AR applications. The experimental results show that our proposed scheme decreases the service rebuilding latency by around 65.8%, as compared to the conventional rebuilding process. In addition, we conduct extensive simulations to evaluate the performance of our proposed feature mapping algorithm. Simulation confirms that our algorithm is robust and can predict users' target cloudlet with high precision and low latency. Haoxin Wang 0003, Jiang (Linda) Xie, Tao Han 0002 |
ICC | 1 |
| 2018 | Rethinking Mobile Devices' Energy Efficiency in WLAN Management ServicesabstractWith the rapid popularization of large data stream mobile applications, wireless local area networks (WLANs) have been a top choice for mobile users (MUs), because of the high data rate and low monetary cost. However, the battery life of mobile devices, which is the most concerned feature of MUs, may suffer from WLAN management services, such as mobility management and load balancing services. Unfortunately, few existing WLAN systems take into account both the energy efficiency of mobile devices and the performance of management services. Even worse, to improve the performance of WLAN management services, various existing management mechanisms sacrifice mobile devices' energy. In this paper, we propose BELL, a novel WLAN system that provides two energy-efficient management services for its associated MUs by reproducing and scheduling the beacons broadcast from access points (APs). We name them BELL- handoff and BELL-2M services. We have implemented the proposed BELL-handoff using commercial Wi-Fi adapters. The experimental results reveal that BELL-handoff significantly decreases both mobile devices' energy consumption and latency during handoffs, compared with the commercial WLAN mobility management service. Furthermore, we conduct extensive simulations to evaluate APs' load and mobile devices' battery life within a large-scale deployment of BELL. Simulation results demonstrate that BELL not only balances the load among APs, but also prolongs the battery life of mobile devices. Haoxin Wang 0003, Jiang (Linda) Xie, Xingya Liu |
SECON | 1 |
| 2017 | V-handoff: A practical energy efficient handoff for 802.11 infrastructure networksabstractWireless local area networks (WLANs) are currently among the most important technologies for wireless access. Because of its higher data rate and lower monetary cost compared with cellular networks, mobile users are likely to choose WiFi when they are using mobile applications. However, keeping continuous connectivity with access points (APs) may require frequent handoffs, which may consume much energy in the handoff process. Unfortunately, most of the existing work only focused on reducing the handoff delay of IEEE 802.11-based handoffs and many handoff approaches may even increase the energy consumption of mobile nodes (MNs) in order to reduce the handoff latency. In this paper, we introduce virtual handoff (V-handoff), an energy efficiency-based handoff protocol via generating virtual access points (VAPs) in the corresponding physical access points (PAPs). The main idea of our proposed V-handoff protocol is to create an evenly spaced periodic schedule of beacon periods for all the VAPs in one virtual AP grid. To the best of our knowledge, this is the first paper that investigates the application of the wireless virtualization technique in MN's handoff energy efficiency. Simulation results show that our proposed V-handoff protocol can significantly reduce the MN's handoff energy consumption and the average handoff delay compared with IEEE 802.11-based full scanning and selective scanning handoff protocol. Haoxin Wang 0003, Jiang (Linda) Xie, Tao Han 0002 |
ICC | 1 |