Hao (Frank) Yang

dblp:324/7893 · also Hao Yang 0018 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0001-6431-8956ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Toward Optimal Mixture of Experts System for 3D Object Detection: A Game of Accuracy, Efficiency and Adaptivity
abstract
Autonomous vehicles, open-world robots, and other automated systems rely on accurate, efficient perception modules for real-time object detection. Although high-precision models improve reliability, their processing time and computational overhead can hinder real-time performance and raise safety concerns. This paper introduces an Edge-based Mixture-of-Experts Optimal Sensing (EMOS) System that addresses the challenge of co-achieving accuracy, latency and scene adaptivity, further demonstrated in the open-world autonomous driving scenarios. Algorithmically, EMOS fuses multimodal sensor streams via an Adaptive Multimodal Data Bridge and uses a scenario-aware MoE switch to activate only a complementary set of specialized experts as needed. The proposed hierarchical backpropagation and a multiscale pooling layer let model capacity scale with real-world demand complexity. System-wise, an edge-optimized runtime with accelerator-aware scheduling (e.g., ONNX/TensorRT), zero-copy buffering, and overlapped I/O-compute enforces explicit latency/accuracy budgets across diverse driving conditions. Experimental results establish EMOS as the new state of the art: on KITTI, it increases average AP by 3.17% while running $2.6\times$2.6× faster on Nvidia Jetson. On nuScenes, it improves accuracy by 0.2% mAP and 0.5% NDS, with 34% fewer parameters and a $15.35\times$15.35× Nvidia Jetson speedup. Leveraging multimodal data and intelligent experts cooperation, EMOS delivers accurate, efficient and edge-adaptive perception system for autonomous vehicles, thereby ensuring robust, timely responses in real-world scenarios.
Linshen Liu, Guanlin Wu, Junyue Jiang, Hao (Frank) Yang
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 HumanMM: Global Human Motion Recovery from Multi-shot Videos
abstract
In this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions. Such long-sequence in-the-wild motions are highly valuable to applications such as motion generation and motion understanding, but are of great challenge to be recovered due to abrupt shot transitions, partial occlusions, and dynamic backgrounds presented in such videos. Existing methods primarily focus on single-shot videos, where continuity is maintained within a single camera view, or simplify multi-shot alignment in camera space only. In this work, we tackle the challenges by integrating an enhanced camera pose estimation with Human Motion Recovery (HMR) by incorporating a shot transition detector and a robust alignment module for accurate pose and orientation continuity across shots. By leveraging a custom motion integrator, we effectively mitigate the problem of foot sliding and ensure temporal consistency in human pose. Extensive evaluations on our created multi-shot dataset from public 3D human datasets demonstrate the robustness of our method in reconstructing realistic human motion in world coordinates.
Guanlin Wu, Zhuokai Zhao, Xiaoke Jiang, Zhuoheng Li, Hao (Frank) Yang, Haoqian Wang, Lei Zhang 0001
CVPR9
2025 Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
abstract
Spiking Neural Networks (SNNs) are highly efficient due to their spike-based activation, which inherently produces bit-sparse computation patterns. Existing hardware implementations of SNNs leverage this sparsity pattern to avoid wasteful zero-value computations, yet this approach fails to fully capitalize on the potential efficiency of SNNs. This study introduces a novel sparsity paradigm called Product Sparsity, which leverages combinatorial similarities within matrix multiplication operations to reuse the inner product result and reduce redundant computations. Product Sparsity significantly enhances sparsity in SNNs without compromising the original computation results compared to traditional bit sparsity methods. For instance, in the SpikeBERT SNN model, Product Sparsity achieves a density of only 1.23% and reduces computation by $11 \times$, compared to bit sparsity, which has a density of 13.19%. To efficiently implement Product Sparsity, we propose Prosperity, an architecture that addresses the challenges of identifying and eliminating redundant computations in real-time. Compared to prior SNN accelerator PTB and the A100 GPU, Prosperity achieves an average speedup of $7.4 \times$ and $1.8 \times$, respectively, along with energy efficiency improvements of $8.0 \times$ and $193 \times$, respectively. The code for Prosperity is available at https://github.com/dubcyfor3/Prosperity.
Chiyue Wei, Cong Guo 0003, Shiyu Li 0001, Hao (Frank) Yang, Hai Li 0001, Yiran Chen 0001
HPCA5
2025 Towards Accurate and Efficient 3D Object Detection for Autonomous Driving: A Mixture of Experts Computing System on Edge
abstract
This paper presents Edge-based Mixture of Experts (MoE) Collaborative Computing (EMC2), an optimal computing system designed for autonomous vehicles (AVs) that simultaneously achieves low-latency and high-accuracy 3D object detection. Unlike conventional approaches, EMC2 incorporates a scenario-aware MoE architecture specifically optimized for edge platforms. By effectively fusing LiDAR and camera data, the system leverages the complementary strengths of sparse 3D point clouds and dense 2D images to generate robust multimodal representations. To enable this, EMC2 employs an adaptive multimodal data bridge that performs multi-scale preprocessing on sensor inputs, followed by a scenario-aware routing mechanism that dynamically dispatches features to dedicated expert models based on object visibility and distance. In addition, EMC2 integrates joint hardware-software optimizations, including hardware resource utilization optimization and computational graph simplification, to ensure efficient and real-time inference on resource-constrained edge devices. Experiments on open-source benchmarks clearly show the EMC2 advancements as an end-to-end system. On the KITTI dataset, it achieves an average accuracy improvement of 3.58% and a 159.06% inference speedup compared to 15 baseline methods on Jetson platforms, with similar performance gains on the nuScenes dataset, highlighting its capability to advance reliable, real-time 3D object detection tasks for AVs. The official implementation is available at https://github.com/LinshenLiu622/EMC2.
Linshen Liu, Boyan Su, Junyue Jiang, Guanlin Wu, Cong Guo 0003, Ceyu Xu, Hao (Frank) Yang
ICCV7
2025 Min-K%++: Improved Baseline for Pre-Training Data Detection from Large Language Models
abstract
The problem of pre-training data detection for large language models (LLMs) has received growing attention due to its implications in critical issues like copyright violation and test data contamination. Despite improved performance, existing methods (including the state-of-the-art, Min-K%) are mostly developed upon simple heuristics and lack solid, reasonable foundations. In this work, we propose a novel and theoretically motivated methodology for pre-training data detection, named Min-K%++. Specifically, we present a key insight that training samples tend to be local maxima of the modeled distribution along each input dimension through maximum likelihood training, which in turn allow us to insightfully translate the problem into identification of local maxima. Then, we design our method accordingly that works under the discrete distribution modeled by LLMs, whose core idea is to determine whether the input forms a mode or has relatively high probability under the conditional categorical distribution. Empirically, the proposed method achieves new SOTA performance across multiple settings (evaluated with 5 families of 10 models and 2 benchmarks). On the WikiMIA benchmark, Min-K%++ outperforms the runner-up by 6.2% to 10.5% in detection AUROC averaged over five models. On the more challenging MIMIR benchmark, it consistently improves upon reference-free methods while performing on par with reference-based method that requires an extra reference model.
Jingyang Zhang, Jingwei Sun 0002, Eric C. Yeats, Yang Ouyang, Martin Kuo, Hao (Frank) Yang, Hai Li 0001
ICLR7
2025 Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
abstract
How to integrate and verify spatial intelligence in foundation models remains an open challenge. Current practice often proxies Visual-Spatial Intelligence (VSI) with purely textual prompts and VQA-style scoring, which obscures geometry, invites linguistic shortcuts, and weakens attribution to genuinely spatial skills. We introduce Spatial Intelligence Grid (SIG): a structured, grid-based schema that explicitly encodes object layouts, inter-object relations, and physically grounded priors. As a complementary channel to text, SIG provides a faithful, compositional representation of scene structure for foundation-model reasoning. Building on SIG, we derive SIG-informed evaluation metrics that quantify a model’s intrinsic VSI, which separates spatial capability from language priors. In few-shot in-context learning with state-of-the-art multimodal LLMs (e.g. GPT- and Gemini-family models), SIG yields consistently larger, more stable, and more comprehensive gains across all VSI metrics compared to VQA-only representations, indicating its promise as a data-labeling and training schema for learning VSI. We also release SIGBench, a benchmark of 1.4K driving frames annotated with ground-truth SIG labels and human gaze traces, supporting both grid-based machine VSI tasks and attention-driven, human-like VSI tasks in autonomous-driving scenarios.
Guanlin Wu, Boyan Su, Yang Zhao 0013, Yichen Lin, Hao (Frank) Yang
NeurIPS6
2025 How to Auto-optimize Prompts for Domain Tasks? Adaptive Prompting and Reasoning through Evolutionary Domain Knowledge Adaptation
abstract
Designing optimal prompts and reasoning processes for large language models (LLMs) on domain-specific tasks is both necessary and challenging in real-world applications. Determining how to integrate domain knowledge, enhance reasoning efficiency, and even provide domain experts with refined knowledge integration hints are particularly crucial yet unresolved tasks. In this research, we propose Evolutionary Graph Optimization for Prompting (EGO-Prompt), an automated framework to designing better prompts, efficient reasoning processes and providing enhanced causal-informed process. EGO-Prompt begins with a general prompt and fault-tolerant initial Semantic Causal Graph (SCG) descriptions, constructed by human experts, which is then automatically refined and optimized to guide LLM reasoning. Recognizing that expert-defined SCGs may be partial or imperfect and that their optimal integration varies across LLMs, EGO-Prompt integrates a novel causal-guided textual gradient process in two steps: first, generating nearly deterministic reasoning guidance from the SCG for each instance, and second, adapting the LLM to effectively utilize the guidance alongside the original input. The iterative optimization algorithm further refines both the SCG and the reasoning mechanism using textual gradients with ground-truth. We tested the framework on real-world public health, transportation and human behavior tasks. EGO-Prompt achieves 7.32\%–12.61\% higher F1 than cutting-edge methods, and allows small models to reach the performence of larger models at under 20\% of the original cost. It also outputs a refined, domain-specific SCG that improves interpretability.
Yang Zhao 0013, Hao (Frank) Yang
NeurIPS3
2025 Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
abstract
Decision-making models for individuals, particularly in high-stakes scenarios like vaccine uptake, often diverge from population optimal predictions. This gap arises from the uniqueness of the individual decision-making process, shaped by numerical attributes (e.g., cost, time) and linguistic influences (e.g., personal preferences and constraints). Developing upon Utility Theory and leveraging the textual-reasoning capabilities of Large Language Models (LLMs), this paper proposes an Adaptive Textual-symbolic Human-centric Reasoning framework (ATHENA) to address the optimal information integration. ATHENA uniquely integrates two stages: First, it discovers robust, group-level symbolic utility functions via LLM-augmented symbolic discovery; Second, it implements individual-level semantic adaptation, creating personalized semantic templates guided by the optimal utility to model personalized choices. Validated on real-world travel mode and vaccine choice tasks, ATHENA consistently outperforms utility-based, machine learning, and other LLM-based models, lifting F1 score by at least 6.5\% over the strongest cutting-edge models. Further, ablation studies confirm that both stages of ATHENA are critical and complementary, as removing either clearly degrades overall predictive performance. By organically integrating symbolic utility modeling and semantic adaptation, ATHENA provides a new scheme for modeling human-centric decisions. The project page can be found at https://yibozh.github.io/Athena.
Yang Zhao 0013, Hongru Du, Hao (Frank) Yang
NeurIPS4
2025 MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning
abstract
Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In light of the recent advances in multimodal large language models (MLLMs), such as GPT-4v and LLaVA, which demonstrate their exceptional proficiency in multimodal tasks, such as image captioning and multimodal question answering. We introduce a novel federated learning framework, named Multimodal Large Language Model Assisted Federated Learning (MLLM-LLaVA-FL), which employs powerful MLLMs at the server end to address the heterogeneous and long-tailed challenges. Owing to the advanced cross-modality representation capabilities and the extensive open-vocabulary prior knowledge of MLLMs, our framework is adept at harnessing the extensive, yet previously underexploited, open-source data accessible from websites and powerful server-side computational resources. Hence, the MLLM-LLaVA-FL not only enhances the performance but also avoids increasing the risk of privacy leakage and the computational burden on local devices, distinguishing it from prior methodologies. Our framework has three key stages. Initially, we conduct global visual-text pretraining of the model. This pretraining is facilitated by utilizing the extensive open-source data available online, with the assistance of MLLMs. Subsequently, the pretrained model is distributed among various clients for local training. Finally, once the locally trained models are transmitted back to the server, a global alignment is carried out under the supervision of MLLMs to further enhance the performance. Experimental evaluations on established benchmarks, show that our framework delivers promising performance in the typical scenarios with data heterogeneity and long-tail distribution across different clients in FL.
Hao (Frank) Yang, Ang Li 0005, Xin Guo 0008, Haiming Wang 0002, Yiran Chen 0001, Hai Li 0001
WACV2
2024 OSR-ViT: A Simple and Modular Framework for Open-Set Object Detection and Discovery
abstract
An object detector’s ability to detect and flag novel objects during open-world deployments is critical for many real-world applications. Unfortunately, much of the work in open object detection today is disjointed and fails to adequately address applications that prioritize unknown object recall in addition to known-class accuracy. To close this gap, we present a new task called Open-Set Object Detection and Discovery (OSODD) and as a solution propose the Open-Set Regions with ViT features (OSR-ViT) detection framework. OSR-ViT combines a class-agnostic proposal network with a powerful ViT-based classifier. Its modular design simplifies optimization and allows users to easily swap proposal solutions and feature extractors to best suit their application. Using our multifaceted evaluation protocol, we show that OSR-ViT obtains performance levels that far exceed state-of-the-art supervised methods. Our method also excels in low-data settings, outperforming supervised baselines using a fraction of the training data.
Matthew Inkawhich, Nathan Inkawhich, Hao (Frank) Yang, Jingyang Zhang, Randolph Linderman, Yiran Chen 0001
IEEE Big Data3
2024 Cost-effective Vehicle Recognition System in Challenging Environment Empowered by Micro-Pulse LiDAR and Edge AI
abstract
Vehicle recognition and classification are critical for a number of traffic applications, e.g., traffic signal control, traffic flow modeling, tolling, and logistics optimization. Commonly used sensing systems are mainly counted on in-pavement loops or surveillance video cameras, while both of them have their inherent limitations. Leveraging micro high-speed pulse LiDAR mounted overhead of travel lanes, this study proposes Compact LiDAR Empowered Vehicle Enhancing-minority Recognition (CLEVER) system, a real-time cost-effective vehicle detection and classification framework that is empowered by edge Artificial Intelligence (AI). Based on the customized minority-enhancing vehicle classification deep neural network, the CLEVER system outperforms cutting-edge LiDAR-based vehicle classification methods up to 15.98% true-positive rate in classifying ten types of vehicles. Furthermore, by highly integrating the hardware, the pre-processing algorithm and the classification neural network into an edge computing node, the CLEVER system only consumes 10% of the cost in LiDAR systems and works perfectly in a plug-and-play mode with a negligible sub-second inference time (212ms to 459ms). The proposed CLEVER system offers an affordable end-to-end solution that can benefit traffic operators by collecting more accurate and reliable vehicle classification data streams and that can lead to a more efficient and flexible ITS.
Junyue Jiang, Meixin Zhu, Yiran Chen 0001, Hao (Frank) Yang
IV6
2024 Mitigating Bias of Deep Neural Networks for Trustworthy Traffic Perception in Autonomous Systems
abstract
With the rapid advancement of deep learning technology, feature extraction backbones that are effectively trained have found increasing use in various traffic perception tasks, such as vehicle recognition and roadway user detection and classification. However, given the naturally imbalanced distribution of objects in the real world, deep learning networks can inadvertently act as bias amplifiers, leading to unfair detection and classification outcomes. Addressing and quantifying this bias in traffic applications has thus become a pressing challenge. In response, this research introduces the first comprehensive traffic imbalance object recognition dataset tailored for autonomous vehicles, called the Autonomous-vehicle Long-tail Image Dataset (ALIDA). This dataset reflects real-world sample distribution and includes four categories—motorized users, non-motorized users, roadway facilities, and traffic signs—spanning 87 classes and totaling 37,558 images. Our experimental results confirm that these backbones may struggle to accurately recognize less common objects with limited training data, such as children and wheelchair users. To mitigate such biases and improve traffic perception equality, we introduce a DEbiased Traffic Object Recognition (DETOR) scheme. This scheme leverages both few-shot and representation learning techniques. Employing DETOR, the residual neural network achieved a 290% increase in accuracy for recognizing minority classes, such as children, motorcyclists, deer, and bears. This not only enhances the effectiveness but also significantly improves the fairness and scalability of traffic perception using deep neural networks.
Hao (Frank) Yang, Yang Zhao 0013, Jiarui Cai, Meixin Zhu, Jenq-Neng Hwang, Yiran Chen 0001
IV1
2024 Real-Time Multi-Task Environmental Perception System for Traffic Safety Empowered by Edge Artificial Intelligence
abstract
Traffic safety, reliability, and resilience are significantly influenced by environmental factors such as visibility, road surface, and weather conditions. Yet, current monitoring methods, including weather stations and onboard environmental sensors, often fall short due to their high costs, significant latency, and limited dissemination. This paper presents the Edge-based Multi-task Safety-oriented Environmental (Edge-MuSE) sensing system, designed to address these traffic safety challenges associated with environmental factors. Edge-MuSE departs from traditional single-task sensing methods by performing multidimensional traffic environment perception tasks. It estimates key safety-related environmental factors exclusively through camera inputs and incorporates four innovative sensing tasks: visibility estimation, image dehazing, road segmentation, and road surface condition classification. The system is tailored to edge devices to transition computational loads from central servers to distributed nodes, thereby enhancing privacy and reducing latency. Additionally, Edge-MuSE integrates communication functions based on TCP/IP and Wi-Fi protocols, enabling rapid dissemination of sensing results and warning messages to local road users. System structures and data streaming have been optimized to accommodate the constraints of edge devices, ensuring high-efficiency edge computing. Field testing of Edge-MuSE in multiple testbeds in Bellevue (WA, US) and Oslo (Norway) has demonstrated its reliable and precise performance in perception tasks (92.15% accuracy in visibility estimation and 92.25% in road surface condition classification) as well as an impressive processing speed of 21.3 FPS. As such, Edge-MuSE presents a promising solution for enhancing roadway safety, efficiency, and resilience.
Hao (Frank) Yang, Meixin Zhu, Torgeir Vaa, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.2
2022 Toward a real-time Smart Parking Data Management and Prediction (SPDMP) system by attributes representation learning
abstract
Managing and estimating the availability information of parking lots is of great importance to travelers and managers. However, the task is very challenging since the occupancy rate is affected by various factors, including spatial-temporal features, parking lot attributes features, and environmental changes. Previous studies mostly focus on the short-term prediction by capturing the historical sequential dependencies among inputs and outputs, which leads to low estimation accuracy for long-term prediction and limited scalability for real deployment in parking lots. To address the challenges, a comprehensive framework for real-time Smart Parking Data Management and Prediction (SPDMP) system is proposed. Three types of data sources, including historical sequential data, real-time sequential data, and attributes category data, are sufficiently integrated into a customized Parking Availability Prediction (PAP) neural network by representation learning and heterogeneous feature embedding. Specifically, instead of using parking lot property information and environmental data directly, the authors design an Attributes Tensor Embedding Component (ATEC) to integrate the intra-class affection and interclass representations and correlations by two steps of customized embedding process. To balance the impact of the indiscriminate features for various prediction targets, this study proposes a multi-factor attention mechanism to learn the weights and help the PAP achieve a stable performance for time-various tasks. From extensive experiments on two real-world large-scale data sets collected in China and USA (including 19 urban parking and 2 truck parking lots), PAP achieves superior prediction results on both urban parking (6.69% average MAPE) and truck parking prediction (7.63% average MAPE) for five prediction time slots (5 min, 10 min, 30 min, 1 h, and 2 h). Furthermore, even with limited training data, PAP still shows better transferability and adaptability for various types of lots. The proposed SPDMP is selected and used as a pilot test bed by the Washington State Department of Transportation (WSDOT) for truck parking information promotion.
Hao (Frank) Yang, Ruimin Ke, Zhiyong Cui, Yinhai Wang, Karthik Murthy
Int. J. Intell. Syst.1
2022 Toward a Dynamic Reversible Lane Management Strategy by Empowering Learning-Based Predictive Assignment Scheme
abstract
Traffic congestion is a long-lasting worldwide problem and even becoming more severe in well-developed regions. Reversible lanes have been used worldwide on various road types to mitigate the effects of congestion and optimize mobility since the 1930s. However, with the limitation of traditional control and management methods, existing solutions can not meet the increasing travel demands. Therefore, the paper introduces a Predictive Empowered Assignment scheme for Reversible Lane (PEARL). By integrating the advanced traffic flow prediction module and bi-level optimization model, PEARL can be a more flexible dynamic lane assignment strategy with foresight compared to traditional lane management methods. In the prediction module, taking advantage of the development of machine learning technologies, the input of PEARL covers not only historical sequence and real-time data but also the environmental conditions in the region. The advanced Bi-directional Long Short-term Memory (Abi-LSTM) model is employed for short-term traffic flow prediction. Then, the study introduces a bi-level optimization method to maximize the total throughput in both directions and minimize the total user costs which determine the lane deployment. The iterations on the prediction module and optimization module can help PEARL coordinate the future lane control plan and make the best decision. Finally, the paper builds up an experiment platform to simulate PEARL in a real-world scenario with heavy input flow for its performance evaluation.
Hao (Frank) Yang, Ruimin Ke, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.2
2022 Traffic-Informed Multi-Camera Sensing (TIMS) System Based on Vehicle Re-Identification
abstract
Surveillance cameras are widely deployed traffic sensors, due to their affordable prices and being able to capture rich information. However, current surveillance systems have not been fully exploited: these cameras are isolated and can only extract information from their own fixed views. To enable a collaborative sensing system, we propose a novel framework called Traffic-Informed Multi-camera Sensing (TIMS) system for network-level traffic information estimation. By pushing multi-camera Re-IDentification (ReID) workflow towards network-wide traffic information extraction, TIMS system integrates a customized metric-learning vision-based vehicle ReID method (TIM-ReID) and establishing a traffic-informed workflow. To integrate the traffic network connection information along with visual and vehicle attributes features, the road network is extracted as a weighted graph through the Spatial-temporal Camera Graph Inference Model (StCGIM) and serves for matching and re-ranking ReID candidates. Moreover, an Accuracy Model (AAM) is designed to provide accurate, reliable and comprehensive traffic information estimation, including both the values and distribution of parameters under a high penetration rate. In experiments based on real-world multi-camera datasets captured in the city of Seattle, the customized TIM-ReID outperforms existing state-of-the-art methods, and delivers accurate cross-camera information estimation, whose value error is less than 8% and the Kullback-Leibler (KL) distance between the estimated and real distribution is less than 3.42 among all the evaluated camera pairs. TIMS system empowers cameras to work collaboratively through an interactive brain, and provides users with valuable and comprehensive traffic information.
Hao (Frank) Yang, Jiarui Cai, Meixin Zhu, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.1
2022 How Fast You Will Drive? Predicting Speed of Customized Paths By Deep Neural Network
abstract
Customized path-based speed prediction is an eventful tool for congestion avoidance, route optimization and travel time prediction for navigation apps, cab-hailing companies and autonomous vehicles. Traditionally, the speed prediction algorithms are based on road segments and can only support several main roads. Path-based speed prediction is very challenging since the speed is always changing in different path locations and is jointly affected by lots of complicated factors. This article presents a novel deep learning framework for customized path-based speed prediction. A Path-based Speed Prediction Neural Network (PSPNN) is designed to achieve speed predictions for a given path and attributes information. A hierarchical Convolutional Neural Network (CNN) and deep Bidirectional Long Short-Term Memory (Bi-LSTM) structure for different kinds of feature extraction are applied for multiple levels: the path cell, sub-path and the whole path. The method narrows down the prediction unit from road segments to customized path cells (mean length: 59.52m) and achieves a mean absolute error (MAE) of 1.94 m/s and Mean Absolute Percentage Error (MAPE) of 18.14%, showing the potential of serving rigorous data-driven applications. So far, PSPNN is the first made-to-order path-based speed prediction algorithm and can help both travelers and managers to obtain large-scale bespoke paths speed information in advance.
Hao (Frank) Yang, Meixin Zhu, Xuegang Ban, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.1
2022 Truck Parking Pattern Aggregation and Availability Prediction by Deep Learning
abstract
With the significant increase of e-commerce, freight transportation demand has surged significantly over the past decade. Most of the demand has been served by trucks in the United States. One major problem commonly identified across the country is the worsening truck parking availability because the increase of truck parking facilities has lagged behind the growth of trucking activities. The lack of parking spaces and real-time parking availability information greatly exacerbate the uncertainty of trips, and often results in illegal and potentially dangerous parking or overtime driving. This paper elaborates on pilot research on improving truck parking facilities cooperated with the Washington State Department of Transportation (WSDOT), building and testing the advanced Truck Parking Information and Management System (TPIMS) with the real-time user visualization and prediction function empowered by artificial intelligence. Furthermore, by analyzing the activities of truck drivers, the researchers aggregated the regularity of truck parking patterns by a customized sequential similarity methodology. A Truck Parking Occupancy Prediction (TPOP) neural network for time-variant occupancy prediction by deep learning and attributes embedding is proposed and integrated into the TPIMS. The TPOP achieves 5.82%, 5.07%, 4.84%, and 4.19% mean average percentage error (MAPE) for 16, 8, 4, and 2 minutes ahead of occupancy prediction respectively, significantly outperforms other state-of-the-art methods. Clearly, the proposed solutions can benefit both the truck drivers and government agencies by a more efficient and smart TPIMS.
Hao (Frank) Yang, Yifan Zhuang, Karthik Murthy, Ziyuan Pu, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.1
2021 Delayed Propagation Transformer: A Universal Computation Engine towards Practical Control in Cyber-Physical Systems
abstract
Multi-agent control is a central theme in the Cyber-Physical Systems (CPS). However, current control methods either receive non-Markovian states due to insufficient sensing and decentralized design, or suffer from poor convergence. This paper presents the Delayed Propagation Transformer (DePT), a new transformer-based model that specializes in the global modeling of CPS while taking into account the immutable constraints from the physical world. DePT induces a cone-shaped spatial-temporal attention prior, which injects the information propagation and aggregation principles and enables a global view. With physical constraint inductive bias baked into its design, our DePT is ready to plug and play for a broad class of multi-agent systems. The experimental results on one of the most challenging CPS -- network-scale traffic signal control system in the open world -- show that our model outperformed the state-of-the-art expert methods on synthetic and real-world datasets. Our codes are released at: https://github.com/VITA-Group/DePT.
Wenqing Zheng, Qiangqiang Guo, Hao (Frank) Yang, Peihao Wang, Zhangyang Wang
NeurIPS3
2017 WLAN interference self-optimization using som neural networks
abstract
Summary In order to suppress the interference in local area networks, this paper presents a Wireless Local Area Networks (WLAN) interference self‐optimization method based on a Self‐Organizing Feature Map (SOM) neural network model. This method trains the model by using original data sets as the initial vector set and using the whole Signal to Interference plus Noise Ratio (SINR) vector generated by the change of one Wireless Access Point (AP) channel as the basic feature. After the training, the SOM neural network can quickly locate the fault AP and optimize the network according to the changes of the network environment. Simulation results reveal that the proposed scheme can efficiently locate the AP where interference happens and optimize the interference with an improved user experience. Copyright © 2016 John Wiley & Sons, Ltd.
Haipeng Yao, Hao (Frank) Yang, Chao Fang 0001, Yiru Guo
Concurr. Comput. Pract. Exp.2