VLDB 2026 Research / reviewers in the wild / expert
Tao Huang 0008
dblp:34/808-8
· DBLP profile ↗
40ranked-venue papers
4as first author
31since 2021 · last 2026
0000-0002-8098-8906ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AerialMind: Towards Referring Multi-Object Tracking in UAV ScenariosabstractReferring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intelligent robotic systems. However, current RMOT research remains mostly confined to ground-level scenarios, which constrains their ability to capture broad-scale scene contexts and perform comprehensive tracking and path planning. In contrast, Unmanned Aerial Vehicles (UAVs) leverage their expansive aerial perspectives and superior maneuverability to enable wide-area surveillance. Moreover, UAVs have emerged as critical platforms for Embodied Intelligence, which has given rise to an unprecedented demand for intelligent aerial systems capable of natural language interaction. To this end, we introduce AerialMind, the first large-scale RMOT benchmark in UAV scenarios, which aims to bridge this research gap. To facilitate its construction, we develop an innovative semi-automated collaborative agent-based labeling assistant (COALA) framework that significantly reduces labor costs while maintaining annotation quality. Furthermore, we propose HawkEyeTrack (HETrack), a novel method that collaboratively enhances vision-language representation learning and improves the perception of UAV scenarios. Comprehensive experiments validated the challenging nature of our dataset and the effectiveness of our method. Chenglizhao Chen, Shaofeng Liang, Runwei Guan, Xiaolou Sun, Haocheng Zhao, Haiyun Jiang, Tao Huang 0008, Henghui Ding, Qing-Long Han |
AAAI | 7 |
| 2026 | LEGO: Learning and Graph-Optimized Modular Tracker for Online Multi-Object Tracking With Point CloudsabstractOnline Multi-Object Tracking (MOT) plays a pivotal role in autonomous systems. The state-of-the-art approaches usually employ a tracking-by-detection method, and data association plays a critical role. This paper proposes a learning and graph-optimized (LEGO) modular tracker to improve data association performance in the existing literature. The proposed LEGO tracker integrates graph optimization, which efficiently formulates the association score map, facilitating the accurate and efficient matching of objects across time frames. To further enhance the state update process, the Kalman filter is added to ensure consistent tracking by incorporating temporal coherence in the object states to further enhance the state update process. Our proposed method, utilising LiDAR alone, has shown exceptional performance compared to other online tracking approaches, including LiDAR-based and LiDAR-camera fusion-based methods. LEGO ranked 3rdamong all trackers (both online and offline) and 2ndamong all online trackers in the KITTI MOT benchmark for cars1, at the time of submitting results to KITTI object tracking evaluation ranking board. Moreover, our method also achieves competitive performance on the Waymo open dataset benchmark. Yuxuan Xia, Tao Huang 0008, Qing-Long Han, Hongbin Liu 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | ScopeDrive: Text-Anchored Cross-Modal Calibration and Density-Aware Modulation for Autonomous DrivingabstractVision–language models are emerging as a unified paradigm for perception, prediction, and planning in autonomous driving. However, most existing approaches still rely on shallow fusion between visual and textual features, leading to weak cross-modal interaction, feature drift, and poor adaptability in complex scenes. This work present ScopeDrive, an end-to-end VLM framework that enables deep semantic alignment and adaptive reasoning through two novel components. The text-anchored calibrator transforms textual semantics into multiscale anchors that progressively calibrate visual features across layers, strengthening cross-modal correspondence. The density-aware agent modulator estimates scene complexity and dynamically adjusts attention distribution, allowing the model to focus on dense or dynamic regions when needed. Evaluated on the DriveLM, DriveBench, and NuScenes-QA benchmarks, ScopeDrive surpasses both lightweight and large-scale baselines, achieving a BLEU-4 of 53.27 and METEOR of 38.75 on DriveLM while maintaining only 328 M parameters. It also delivers state-of-the-art performance on perception and planning tasks under both clean and corrupted conditions. These results demonstrate that ScopeDrive effectively breaks the shallow-fusion barrier, offering a lightweight yet semantically aligned foundation for interpretable autonomous-driving intelligence. Minghui Hou, Runwei Guan, Tao Huang 0008, Qing-Long Han |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Adversarial Training and Cross-modal Feature Fusion in Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis recognizes emotions through text, audio, and visual modalities, but data incompleteness is a major challenge. Existing methods often focus on specific types of deficiencies and perform poorly when multiple types of noise are present simultaneously. To address this issue, we propose a noise-prompted adversarial training framework with a multimodal interaction model to enhance the model’s robustness to missing modalities. The model first extracts common and unique features from each modality using a BERT text encoder and a shared-private encoder. Correlation measurements are then used to calculate the similarity between modalities, and a weighting mechanism is applied to the shared features. These features are deeply fused using a Transformer, and adversarial training combined with semantic reconstruction supervision helps the model learn a unified representation of noisy and clean data. Experimental results show that this method significantly improves the performance of multimodal sentiment analysis. Junhuai Li, Huaijun Wang, Yuxing Zhi, Tao Huang 0008 |
ICASSP | 6 |
| 2025 | Priority-Driven Instant Delivery System with Drone ResupplyabstractThe unmanned aerial vehicle (UAV) resupply mode offers a viable alternative to parcel delivery by replenishing ground vehicles at intermediate open-space supply points. Usually, orders are treated as equally important with no guarantees of on-time delivery. In this paper, we propose a priority-based delivery framework to efficiently manage real-time delivery demands with varying priorities. Specifically, this framework integrates a fleet of UAVs for resupply and trucks for final delivery. By considering the inherent priority of orders and their elapsed waiting time, a dynamic priority discipline is introduced and customized to different order contexts. This approach ensures that urgent orders are prioritized to ensure timely fulfillment within their deadlines, while low-prioritized orders can still be completed without excessive waiting. Simulations using a real-world city map of Helsinki, along with mobility features of both trucks and UAVs, show significant performance gains over baseline schemes in terms of on-time delivery rate, average delivery time, and waiting fairness among all orders. Xu Zhang 0016, Xiaokang Zhou, Yi Ren 0001, Tao Huang 0008 |
SMC | 4 |
| 2025 | Resource Allocation and Trajectory Optimization in Multi-UAV Collaborative Vehicular Networks: An Extended Multiagent DRL ApproachabstractIn vehicular networks enhanced by uncrewed aerial vehicles (UAVs), vehicle state information is efficiently collected, and traffic safety is assured. UAVs, serving as aerial base stations, enable vehicle network access and provide edge computing services in the absence of roadside units (RSUs). This study explores a multi-UAV-assisted vehicular network, where multiple UAVs collaboratively offer services to vehicles. The goal is to minimize task completion time by optimizing trajectory planning, spectrum resource allocation, and dynamic data offloading. An enhanced multiagent deep deterministic policy gradient (MADDPG) algorithm is introduced to address the optimization challenge in cooperative multi-UAV scenarios. Within this framework, each UAV, acting as an agent, devises strategies for movement, data offloading, and resource allocation based on the current states of vehicles and fellow UAVs. The simulation results reveal that the proposed algorithm improves task completion efficiency and ensures vehicle Quality of Service (QoS) over existing benchmarks. Wenqian Zhang 0003, Tao Huang 0008, Xiaowen Huang 0002, Mengting Huang, Guanglin Zhang |
IEEE Internet Things J. | 3 |
| 2025 | Trustworthy Blockchain-Assisted Federated Learning: Decentralized Reputation Management and Performance OptimizationabstractBlockchain-assisted federated learning (BFL) can achieve decentralized storage and management of model data without relying on a central server. However, security issues caused by deliberate attacks in distributed systems and efficiency issues induced by heterogeneous computing consumption in resource-limited systems need to be urgently addressed in BFL. To address these issues, we propose a decentralized reputation management (DRM) mechanism for a trustworthy BFL (T-BFL) network, that explores, stores, and utilizes the endogenous reputation of distributed nodes to promote system security and efficiency. The proposed DRM includes three core modules, i.e., decentralized reputation evaluation, reputation-based model aggregation, and reputation-based blockchain consensus. Specifically, in the off-chain phase of T-BFL, the reputation value of each node is evaluated based on model quality, which other peer nodes can verify. This reputation value further determines the weight of global aggregation at each node. In the on-chain phase, the reputation of each node serves as the stake to dynamically adjust its consensus difficulty. Furthermore, we investigate the convergence rate of the T-BFL network under the poisoning attack, and dynamically optimize the energy allocation of local training, consensus, and communications by minimizing the upper bound of the global loss function. Extensive experiments are conducted to evaluate the performance of T-BFL on MNIST, Fashion-MNIST, and Cifar-10 datasets. The experimental results demonstrate that, compared with traditional BFL, T-BFL can achieve up to 56.12% accuracy improvement and$8.6\times $acceleration for reaching the target learning accuracy under the poisoning attack. Weihao Zhu, Long Shi 0001, Jun Li 0004, Bin Cao 0002, Kang Wei 0004, Zhe Wang 0005, Tao Huang 0008 |
IEEE Internet Things J. | 7 |
| 2025 | Deep Learning-Enabled RIS Massive MIMO Systems for Industrial IoT: A Joint Communication and Computation ApproachabstractAccurate estimation and detection, along with phase shift optimization, are vital for implementing reconfigurable intelligent surface (RIS)-enabled multi-antenna systems in highly disruptive industrial IoT environments. Motivated by the remarkable capabilities of deep learning (DL) techniques, this paper introduces a pioneering approach to address challenges in channel estimation, channel correlation prediction, and symbol detection for industrial IoT. We develop an optimization framework for large-scale IoT deployments to maximize the signal-to-interference-plus-noise ratio (SINR) while minimizing transmit power. We also propose a transformer-based channel correlation predictor for IoT devices, which enables adaptive pilot retransmissions and reduces training overhead through a co-design approach that integrates communication, computation, and control. Extensive simulations under realistic, time-varying industrial IoT channel conditions demonstrate the superiority of our DL-driven approach, achieving significant improvements in detection accuracy and SINR. Wei Xiang 0001, Muhammad Umer Zia, Jameel Ahmad, Peng Cheng 0002, Kan Yu 0002, Tao Huang 0008 |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | Vehicle-to-Everything Cooperative Perception for Autonomous DrivingabstractAchieving fully autonomous driving with enhanced safety and efficiency relies on vehicle-to-everything (V2X) cooperative perception (CP), which enables vehicles to share perception data, thereby enhancing situational awareness and overcoming the limitations of the sensing ability of individual vehicles. V2X CP plays a crucial role in extending the perception range, increasing detection accuracy, and supporting more robust decision-making and control in complex environments. This article provides a comprehensive survey of recent developments in V2X CP, introducing mathematical models that characterize the perception process under different collaboration strategies. Key techniques for enabling reliable perception sharing, such as agent selection, data alignment, and feature fusion, are examined in detail. In addition, major challenges are discussed, including differences in agents and models, uncertainty in perception outputs, and the impact of communication constraints such as transmission delay and data loss. This article concludes by outlining promising research directions, including privacy-preserving artificial intelligence methods, collaborative intelligence, and integrated sensing frameworks to support future advancements in V2X CP. Tao Huang 0008, Xi Zhou 0006, Dinh C. Nguyen, Mostafa Rahimi Azghadi, Yuxuan Xia, Qing-Long Han, Sumei Sun |
Proc. IEEE | 1 |
| 2025 | Joint Optimization of Task Partial Offloading and Resource Allocation in a Dual-Blockchain-Enabled MEC System With Parallelism ConstraintsabstractIntegrating data security with resource management enhances security, efficiency, and reliability of blockchain-enabled mobile edge computing (MEC) systems. However, challenges such as secure data storage, timely task execution, and limited parallelism introduce complexities in task offloading decisions and resource allocation strategies. To address these challenges, the task latency minimization problem in blockchain-enabled MEC networks is formulated as an NP-hard optimization problem. The model incorporates constraints on parallelism, partial task offloading, bandwidth and computation resource allocation among mobile users (MUs) and edge servers (ESs). To enhance the reliability and transparency of data storage, a dual-blockchain framework is proposed, consisting of multiple MU blockchains and a dedicated ES blockchain. To tackle the NP-hard problem, the original optimization problem is decomposed into multiple sub-problems, facilitating parameter decoupling. An alternating optimization algorithm is employed to refine task offloading decisions and resource allocation of MUs and ESs with limited parallelism. The ESs update their strategies iteratively based on feedback mechanisms. Additionally, a task prioritization formulation is developed to enhance scalability, considering sub-level task importance, urgency, and first-level task classification. Extensive simulation experiments demonstrate that the proposed algorithm achieves lower task latency compared to existing methods across varying network sizes, offloading schemes, and parallelism constraints. By optimizing the parallel processing of tasks, the waiting latency of this algorithm is reduced on average by 35. 35%, 57. 16% and 35. 35% compared to other methods, respectively. Xiaowen Huang 0002, Tao Huang 0008, Shuguang Zhao, Wei Xiang 0001, Wenqian Zhang 0003, Guanglin Zhang |
IEEE Trans. Commun. | 2 |
| 2025 | Robust Outage-Constrained Secrecy Rate of Hybrid Power Line and Wireless Communication With Artificial Noise-Aided Beamforming for Smart GridabstractPower line communication is a critical component of smart grids, which are vulnerable to eavesdropping. To address this challenge, we investigate a cooperative relay hybrid power line and wireless communication system where multiple eavesdroppers are considered. We propose an elaborate artificial noise (AN)-aided beamforming (BF) scheme to improve physical layer security. Our scheme maximizes the outage-constrained secrecy rate (OCSR) of the legitimate link while restricting the capacity of the eavesdroppers to a reasonable region. However, due to the imperfect channel state information of the wiretap channel and secrecy outage probability constraint, the robust OCSR problem becomes intractable because of the non-concave secrecy objective function and the non-convex constraints. To solve this issue, we utilize semidefinite programming and Bernstein-type inequality to transform the robust OCSR nonconvex problem into two convex sub-problems, which a block-coordinated descent algorithm can solve. Simulation results showcase the effectiveness of our robust AN-aided secure BF scheme and show that the proposed scheme outperforms the benchmark scheme in a security performance gain under various channel conditions, even in the worst case. Zhengmin Kong, Li Gan, Tao Huang 0008, Weijun Yin, Shihao Yan, Jinhong Yuan |
IEEE Trans. Commun. | 4 |
| 2025 | RaLiBEV: Radar and LiDAR BEV Fusion Learning for Anchor Box Free Object Detection SystemsabstractIn autonomous driving, LiDAR and radar are crucial for environmental perception. LiDAR offers precise 3D spatial sensing information but struggles in adverse weather like fog. Conversely, radar signals can penetrate rain or mist due to their specific wavelength but are prone to noise disturbances. Recent state-of-the-art works reveal that the fusion of radar and LiDAR can lead to robust detection in adverse weather. Current approaches typically fuse features from various data sources using basic convolutional/transformer network architectures and employ straightforward label assignment strategies for object detection. However, these methods have two main limitations: they fail to adequately capture feature interactions and lack consistent regression constraints. In this paper, we propose a bird’s-eye view fusion learning-based anchor box-free object detection system. Our approach introduces a novel interactive transformer module for enhanced feature fusion and an advanced label assignment strategy for more consistent regression, addressing key limitations in existing methods. Specifically, experiments show that, our approach’s average precision ranks$1^{st}$and significantly outperforms the state-of-the-art method by 13.1% and 19.0% at Intersection of Union (IoU) of 0.8 under “Clear+Foggy” training conditions for “Clear” and “Foggy” testing, respectively. Our code repository is available at:https://github.com/yyxr75/RaLiBEV. Yanlong Yang, Tao Huang 0008, Qing-Long Han, Bing Zhu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Explicit Abnormality Extraction for Unsupervised Motion Artifact Reduction in Magnetic Resonance ImagingabstractMotion artifacts compromise the quality of magnetic resonance imaging (MRI) and pose challenges to achieving diagnostic outcomes and image-guided therapies. In recent years, supervised deep learning approaches have emerged as successful solutions for motion artifact reduction (MAR). One disadvantage of these methods is their dependency on acquiring paired sets of motion artifact-corrupted (MA-corrupted) and motion artifact-free (MA-free) MR images for training purposes. Obtaining such image pairs is difficult and therefore limits the application of supervised training. In this paper, we propose a novel UNsupervised Abnormality Extraction Network (UNAEN) to alleviate this problem. Our network is capable of working with unpaired MA-corrupted and MA-free images. It converts the MA-corrupted images to MA-reduced images by extracting abnormalities from the MA-corrupted images using a proposed artifact extractor, which intercepts the residual artifact maps from the MA-corrupted MR images explicitly, and a reconstructor to restore the original input from the MA-reduced images. The performance of UNAEN was assessed by experimenting with various publicly available MRI datasets and comparing them with state-of-the-art methods. The quantitative evaluation demonstrates the superiority of UNAEN over alternative MAR methods and visually exhibits fewer residual artifacts. Our results substantiate the potential of UNAEN as a promising solution applicable in real-world clinical environments, with the capability to enhance diagnostic accuracy and facilitate image-guided therapies. Hao Li 0034, Zhengmin Kong, Tao Huang 0008, Euijoon Ahn, Zhihan Lyu, Jinman Kim, David Dagan Feng |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | OptiPMB: Enhancing 3D Multi-Object Tracking With Optimized Poisson Multi-Bernoulli FilteringabstractAccurate 3D multi-object tracking (MOT) is crucial for autonomous driving, as it enables robust perception, navigation, and planning in complex environments. While deep learning-based solutions have demonstrated impressive 3D MOT performance, model-based approaches remain appealing for their simplicity, interpretability, and data efficiency. Conventional model-based trackers typically rely on random vector-based Bayesian filters within the tracking-by-detection (TBD) framework but face limitations due to heuristic data association and track management schemes. In contrast, random finite set (RFS)-based Bayesian filtering handles object birth, survival, and death in a theoretically sound manner, facilitating interpretability and parameter tuning. In this paper, we present OptiPMB, a novel RFS-based 3D MOT method that employs an optimized Poisson multi-Bernoulli (PMB) filter while incorporating several key innovative designs within the TBD framework. Specifically, we propose a measurement-driven hybrid adaptive birth model for improved track initialization, employ adaptive detection probability parameters to effectively maintain tracks for occluded objects, and optimize density pruning and track extraction modules to further enhance overall tracking performance. Extensive evaluations on nuScenes and KITTI datasets show that OptiPMB achieves superior tracking accuracy compared with state-of-the-art methods, thereby establishing a new benchmark for model-based 3D MOT and offering valuable insights for future research on RFS-based trackers in autonomous driving. Guanhua Ding, Yuxuan Xia, Runwei Guan, Qinchen Wu, Tao Huang 0008, Weiping Ding 0001, Jinping Sun, Guoqiang Mao |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Adaptive Spatio-Temporal Voxel-Based Trajectory Planning and Optimization for Close-Quarters Ships Collision AvoidanceabstractClose-quarters encounter scenarios, characterized by their potentially severe consequences, represent one of the most formidable challenges in ship collision avoidance decision-making and trajectory planning. An efficient collision avoidance framework must simultaneously satisfy critical requirements for real-time decision-making, accurate trajectory planning, and effective vessel maneuvering, all while maintaining rigorous compliance with the International Regulations for Preventing Collisions at Sea (COLREGs). This framework serves as the critical safeguard for the navigation of maritime autonomous surface ships. To design an efficient framework, this study proposes an innovative method for real-time partitioning of collision-free navigable space, termed “adaptive spatio-temporal voxels”, which accounts for the ship’s maneuverability by incorporating the velocity obstacle concept. Furthermore, a spatio-temporal graph, incorporating COLREGs compliance, yields an optimal collision avoidance decision represented as a sequence of voxels. These voxel sequences subsequently serve as constraints within a model predictive control (MPC) framework to optimize and plan a safe, navigable trajectory. This framework has undergone rigorous testing in a variety of close-quarters encounter scenarios. Its computational performance, with both decision-making and trajectory optimization completed within 1s, effectively meets the demands of real-time maritime collision avoidance. The results demonstrate its ability to generate efficient collision avoidance trajectories in complex environments, significantly enhancing overall collision avoidance performance. Zhepeng Han, Da Wu, Jinfen Zhang, Tao Huang 0008, Qing-Long Han, Jonas W. Ringsberg |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Optimizing Task Migration for Public and Private Services in Vehicular Edge Networks: A Dual- Layer Graph Neural Network ApproachabstractIn the vehicular edge networks (VEN), task migration is complicated by issues like vehicle movement, diverse resource allocation, and integrating sensing with communication technologies. This paper presents a task migration strategy to optimize task flow under limited resources in PMN-assisted VEN. Vehicles can send public and private tasks to roadside units (RSUs), constrained by bandwidth, computational power, and storage space. Public tasks aim at data collection for road transportation management, while private tasks cover a spectrum of services from work to entertainment. To address the limitations imposed by resource scarcity and meet the demands of task migration, we have developed a dual-layer graph neural network (GNN) that leverages vehicle mobility patterns. In particular, the first layer of GNN acquires vehicle information and the latest surrounding information, and sends it to the nearby RSU. Considering the variety of tasks and multi-dimensional resource constraints, the second GNN layer forecasts RSU resource availability and vehicular trajectories. Subsequently, a task-based maximum flow algorithm (T-MFA) is proposed to refine task migration paths and resource allocation strategies to maximize task flow. Simulation experiments validate the efficacy of the proposed algorithm, demonstrating its capability to achieve optimal task migration by accommodating differences in tasks, resources, and capacities. Xiaowen Huang 0002, Tao Huang 0008, Peng Cheng 0002, Jinhong Yuan, Shuguang Zhao, Guanglin Zhang |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Study on the methods of hyperspectral image saliency detection based on MBCNN
Jiexi Chen, Jinming Guo, Xiaoxue Xing, Tao Huang 0008 |
Vis. Comput. | 7 |
| 2024 | LiDAR Point Cloud-Based Multiple Vehicle Tracking with Probabilistic Measurement-Region AssociationabstractMultiple extended target tracking (ETT) has gained increasing attention due to the development of high-precision LiDAR and radar sensors in automotive applications. For LiDAR point cloud-based vehicle tracking, this paper presents a probabilistic measurement-region association (PMRA) ETT model, which can describe the complex measurement distribution by partitioning the target extent into different regions. The PMRA model overcomes the drawbacks of previous data-region association (DRA) models by eliminating the approximation error of constrained estimation and using continuous integrals to more reliably calculate the association probabilities. Furthermore, the PMRA model is integrated with the Poisson multi-Bernoulli mixture (PMBM) filter for tracking multiple vehicles. Simulation results illustrate the superior estimation accuracy of the proposed PMRA-PMBM filter in terms of both the positions and extents of vehicles compared with PMBM filters using the gamma Gaussian inverse Wishart and DRA implementations. Guanhua Ding, Yuxuan Xia, Tao Huang 0008, Bing Zhu 0004, Jinping Sun |
FUSION | 4 |
| 2024 | Which Framework is Suitable for Online 3D Multi-Object Tracking for Autonomous Driving with Automotive 4D Imaging Radar?abstractOnline 3D multi-object tracking (MOT) has recently received significant research interests due to the expanding demand of 3D perception in advanced driver assistance systems (ADAS) and autonomous driving (AD). Among the existing 3D MOT frameworks for ADAS and AD, conventional point object tracking (POT) framework using the tracking-by-detection (TBD) strategy has been well studied and accepted for LiDAR and 4D imaging radar point clouds. In contrast, extended object tracking (EOT), another important framework which accepts the joint-detection-and-tracking (JDT) strategy, has rarely been explored for online 3D MOT applications. This paper provides the first systematical investigation of the EOT framework for online 3D MOT in real-world ADAS and AD scenarios. Specifically, the widely accepted TBD-POT framework, the recently investigated JDT-EOT framework, and our proposed TBD-EOT framework are compared via extensive evaluations on two open source 4D imaging radar datasets: View-of-Delft and TJ4DRadSet. Experiment results demonstrate that the conventional TBD-POT framework remains preferable for online 3D MOT with high tracking performance and low computational complexity, while the proposed TBD-EOT framework has the potential to outperform it in certain situations. However, the results also show that the JDT-EOT framework encounters multiple problems and performs inadequately in evaluation scenarios. After analyzing the causes of these phenomena based on various evaluation metrics and visualizations, we provide possible guidelines to improve the performance of these MOT frameworks on real-world data. These provide the first benchmark and important insights for the future development of 4D imaging radar-based online 3D MOT algorithms. Guanhua Ding, Yuxuan Xia, Jinping Sun, Tao Huang 0008, Lihua Xie 0001, Bing Zhu 0004 |
IV | 5 |
| 2024 | SMURF: Spatial Multi-Representation Fusion for 3D Object Detection with 4D Imaging RadarabstractConventional automotive radar has been extensively utilized in advanced driver assistance systems and autonomous driving, with potential applications in future cooperative perception systems. However, compared to LiDAR-based perception, conventional radar-based perception technologies often encounter limitations such as the absence of elevation information and low resolution. These limitations impede their ability to detect and localize objects in the surrounding environment accurately. In recent years, the development of 4D imaging radar emerged as a promising solution to overcome these limitations. 4D imaging radar can measure the pitch angle, enhancing the understanding of the environment and improving object detection and localization accuracy. Due to the cost-effectiveness and operability in adverse weather conditions of 4D radar, its emergence has attracted attention from both the academic and industrial communities. However, the measurements obtained from 4D radar are subject to noise, primarily stemming from the multi-path propagation of radar signals. Additionally, 4D radar captures less geometry and semantic information than the more dense LiDAR point cloud. As a result, existing 3D object detection algorithms specifically developed for dense LiDAR point cloud may yield suboptimal performance when directly applied to sparse 4D radar point cloud data. Qiuchi Zhao, Weiyi Xiong, Tao Huang 0008, Qing-Long Han, Bing Zhu 0004 |
IV | 4 |
| 2024 | LXL: LiDAR Excluded Lean 3D Object Detection with 4D Imaging Radar and Camera FusionabstractAs an emerging technology and a relatively affordable device, the 4D imaging radar has already been confirmed effective in performing 3D object detection in autonomous driving [1] . Nevertheless, the sparsity and noisiness of 4D radar point clouds hinder further performance improvement, and in-depth studies about its fusion with other modalities are lacking. On the other hand, as a new image view transformation strategy, sampling has been applied in a few image-based detectors and shown to outperform the widely applied depth-based splatting proposed in Lift-Splat-Shoot (LSS) [2] , even without image depth prediction [3] . However, the potential of sampling is not fully unleashed. As a result, this paper investigates the sampling strategy on the camera and 4D imaging radar fusion-based 3D object detection. In the proposed LiDAR Excluded Lean (LXL) model, predicted image depth distribution maps and radar 3D occupancy grids are generated from image perspective view (PV) features and radar bird’s eye view (BEV) features, respectively. They are sent to the core of LXL, called radar occupancy-assisted depth-based sampling , to aid image view transformation. Weiyi Xiong, Tao Huang 0008, Qing-Long Han, Yuxuan Xia, Bing Zhu 0004 |
IV | 3 |
| 2024 | An ADMM-LSTM framework for short-term load forecasting
Zhengmin Kong, Tao Huang 0008, Yang Du 0005, Wei Xiang 0001 |
Neural Networks | 3 |
| 2024 | Distributed Robust Artificial-Noise-Aided Secure Precoding for Wiretap MIMO Interference ChannelsabstractWe propose a distributed artificial noise-assisted precoding scheme for secure communications over wiretap multi-input multi-output (MIMO) interference channels, where K legitimate transmitter-receiver pairs communicate in the presence of a sophisticated eavesdropper having more receive-antennas than the legitimate user. Realistic constraints are considered by imposing statistical error bounds for the channel state information of both the eavesdropping and interference channels. Based on the asynchronous distributed pricing model, the proposed scheme maximizes the total utility of all the users, where each user’s utility function is defined as the secrecy rate minus the interference cost imposed on other users. Using the weighted minimum mean square error, Schur complement and sign-definiteness techniques, the original non-concave optimization problem is approximated with high accuracy as a quasi-concave problem, which can be solved by the alternating convex search method. Simulation results consolidate our theoretical analysis and show that the proposed scheme outperforms the artificial noise-assisted interference alignment and minimum total mean-square error-based schemes. Zhengmin Kong, Shaoshi Yang, Li Gan, Weizhi Meng 0001, Tao Huang 0008, Sheng Chen 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Pricing Optimization in MEC Systems: Maximizing Resource Utilization Through Joint Server Configuration and Dynamic OperationabstractThe resource allocation problem in Multi-access Edge Computing (MEC) has been widely studied to maximize its operation efficiency under limited resource constraint. However, the existing literatures overlooked the setup cost and the associated dynamic operations. In this work, we consider server configuration and overload in the multi-server scenario where servers are switched on/off depending on the network environment. A novel pricing mechanism maximizing the utility of base station (BS) monitoring multiple servers is proposed, which jointly optimizes the setup cost and server load. We aim to maximize the BS utility under one-day task requests, and divide the time into off-peak and peak periods based on task requests. In the off-peak period, we flexibly switch on/off servers for BS to reduce setup costs. In the peak period, to avoid overloading, we introduce crowdsourcing where servers as agents purchase idle resources from private users (PUs) for mobile users (MUs) and minimize MUs’ cost by a contract-based knapsack algorithm. Lastly, a pricing mechanism is proposed to solve the BS utility maximization problem with an exploratory Upper Confidence Bound (UCB)-based algorithm adjusting server prices dynamically. Simulation results show that the proposed algorithm is superior to others in minimizing MUs cost and maximizing BS utility. Xiaowen Huang 0002, Tao Huang 0008, Wenjie Zhang 0003, Chai Kiat Yeo, Shuguang Zhao, Guanglin Zhang |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | A novel neural network for improved in-hospital mortality prediction with irregular and incomplete multivariate data
Xi Zhou 0006, Wei Xiang 0001, Tao Huang 0008 |
Neural Networks | 3 |
| 2023 | Deep Unfolding Scheme for Grant-Free Massive-Access Vehicular NetworksabstractGrant-free random access is an effective solution to enable massive access for future Internet of Vehicles (IoV) scenarios based on massive machine-type communication (mMTC). Considering the uplink transmission of grant-free based vehicular networks, vehicular devices sporadically access the base station, the joint active device detection (ADD) and channel estimation (CE) problem can be addressed by compressive sensing (CS) recovery algorithms due to the sparsity of transmitted signals. However, traditional CS-based algorithms present high complexity and low recovery accuracy. In this manuscript, we propose a novel alternating direction method of multipliers (ADMM) algorithm with low complexity to solve this problem by minimizing the$\ell _{2,1}$norm. Furthermore, we design a deep unfolded network with learnable parameters based on the proposed ADMM, which can simultaneously improve convergence rate and recovery accuracy. The experimental results demonstrate that the proposed unfolded network performs better performance than other traditional algorithms in terms of ADD and CE. Xiaobing Dang, Wei Xiang 0001, Lei Yuan 0004, Yuan Yang 0006, Eric Wang 0001, Tao Huang 0008 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Energy-Aware Coded Caching Strategy Design With Resource Optimization for Satellite-UAV-Vehicle-Integrated NetworksabstractThe Internet of Vehicles (IoV) can offer safe and comfortable driving experience, by the enhanced advantages of space–air–ground-integrated networks (SAGINs), i.e., global seamless access, wide-area coverage, and flexible traffic scheduling. However, due to the huge popular traffic volume, limited cache/power resources, and the heterogeneous network infrastructures, the burden of backhaul link will be seriously enlarged, degrading the energy efficiency of IoV in SAGIN. In this article, to implement the popular content severing multiple vehicle users (VUs), we consider a cache-enabled satellite-UAV-vehicle-integrated network (CSUVIN), where the geosynchronous Earth orbit (GEO) satellite is regard as a cloud server, and unmanned aerial vehicles are deployed as edge caching servers. Then, we propose an energy-aware coded caching strategy employed in our system model to provide more multicast opportunities, and to reduce the backhaul transmission volume, considering the effects of file popularity, cache size, request frequency, and mobility in different road sections (RSs). Furthermore, we derive the closed-form expressions of total energy consumption both in single-RS and multi-RSs scenarios with asynchronous and synchronous services schemes, respectively. An optimization problem is formulated to minimize the total energy consumption, and the optimal content placement matrix, power allocation vector, and coverage deployment vector are obtained by well-designed algorithms. We finally show, numerically, our coded caching strategy can greatly improve energy efficient performance in CSUVINs, compared with other benchmarked caching schemes under the heterogeneous network conditions. Shushi Gu, Xinyi Sun, Zhihua Yang, Tao Huang 0008, Wei Xiang 0001, Keping Yu |
IEEE Internet Things J. | 4 |
| 2022 | A Novel Occlusion-Aware Vote Cost for Light Field Depth EstimationabstractCapturing the directions of light by light field cameras powers next-generation immersive multimedia applications. A critical problem in taking advantage of the rich visual information in light field images is depth estimation. Conventional light field depth estimation methods build a cost volume that measures the photo-consistency of pixels refocused to a range of depths, and the highest consistency indicates the correct depth. This strategy works well in most regions but usually generates blurry edges in the estimated depth map due to occlusions. Recent work shows that integrating occlusion models to light field depth estimation can largely reduce blurry edges. However, existing occlusion handling methods rely on complex edge-aided processing and post-refinement, and this reliance limits the resultant depth accuracy and impacts on the computational performance. In this paper, we propose a novel occlusion-aware vote cost (OAVC) which is able to accurately preserve edges in the depth map. Instead of using photo-consistency as an indicator of the correct depth, we construct a novel cost from a new perspective that counts the number of refocused pixels whose deviations from the central-view pixel are less than a small threshold, and utilizes that number to select the correct depth. The pixels from occluders are thus excluded in determining the correct depth. Without the use of any explicit occlusion handling methods, the proposed method can inherently preserve edges and produces high-quality depth estimates. Experimental results show that the proposed OAVC outperforms state-of-the-art light field depth estimation methods in terms of depth estimation accuracy and computational complexity. Kang Han, Wei Xiang 0001, Eric Wang 0001, Tao Huang 0008 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Deep Learning-Aided TR-UWB MIMO SystemabstractThis paper presents a novel deep learning-aided scheme dubbed$PR\rho $-net for improving the bit error rate (BER) of the Time Reversal (TR) Ultra-Wideband (UWB) Multiple Input Multiple Output (MIMO) system with imperfect Channel State Information (CSI). The designed system employs Frequency Division Duplexing (FDD) with explicit feedback in a scenario where the CSI is subject to estimation and quantization errors. Imperfect CSI causes a drastic increase in BER of the FDD-based TR-UWB MIMO system, and we tackle this problem by proposing a novel neural network-aided design for the conventional precoder at the transmitter and equalizer at the receiver. A closed-form expression for the initial estimation of the channel correlation is derived by utilizing transmitted data in time-varying channel conditions modeled as a Markov process. Subsequently, a neural network-aided design is proposed to improve the initial estimate of channel correlation. An adaptive pilot transmission strategy for a more efficient data transmission is proposed that uses channel correlation information. The theoretical analysis of the model under the Gaussian assumptions is presented, and the results agree with the Monte-Carlo simulations. The simulation results indicate high performance gains when the suggested neural networks are used to combat the effect of channel imperfections. Muhammad Umer Zia, Wei Xiang 0001, Tao Huang 0008, Ijaz Haider Naqvi |
IEEE Trans. Commun. | 3 |
| 2022 | Competitive Relationship Prediction for Points of Interest: A Neural Graphlet Based ApproachabstractCompetition between Points of Interest (POIs) refers to the situation in which two POIs directly or indirectly provide similar services to secure businesses. A large portion of prior studies on competition analysis focuses on mining textual data, e.g., news articles and social comments. However, the increasing availability of human mobility and mobile query data enables a new paradigm for analyzing the competitive relationships among POIs, which remains largely unexplored. To this end, in this paper, we attempt to mine large-scale online map search query data for better understanding POI competitive relationships. Based on a co-query POI graph built from the map search query data, we develop a novel neural graphlet-based prediction framework to predict the competitive relationships among POIs. A unique perspective of our model is to infer latent POI competitive relationships by integrating multiple distinct factors, e.g., graphlet structure, geographical distance, and regional features, reflected in map search query data and POI data. Finally, we conduct extensive experiments on real-world datasets to demonstrate the effectiveness of the proposed framework, and show that our framework outperforms all baselines with a significant margin in all evaluation metrics. Jingbo Zhou 0003, Tao Huang 0008, Shuangli Li, Renjun Hu, Yanchi Liu, Yanjie Fu, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Global repair bandwidth cost optimization of generalized regenerating codes in clustered distributed storage systemsabstractAbstract In clustered distributed storage systems (CDSSs), one of the main design goals is minimizing the transmission cost during the failed storage nodes repairing. Generalized regenerating codes (GRCs) are proposed to balance the intra‐cluster repair bandwidth and the inter‐cluster repair bandwidth for guaranteeing data availability. The trade‐off performance of GRCs illustrates that, it can reduce storage overhead and inter‐cluster repair bandwidths simultaneously. However, in practical big data storage scenarios, GRCs cannot give an effective solution to handle the heterogeneity of bandwidth costs among different clusters for node failures recovery. This paper proposes an asymmetric bandwidth allocation strategy (ABAS) of GRCs for the inter‐cluster repair in heterogeneous CDSSs. Furthermore, an upper bound of the achievable capacity of ABAS is derived based on the information flow graph (IFG), and the constraints of storage capacity and intra‐cluster repair bandwidth are also elaborated. Then, a metric termed global repair bandwidth cost (GRBC), which can be minimized regarding of the inter‐cluster repair bandwidths by solving a linear programming problem, is defined. The numerical results demonstrate that, maintaining the same data availability and storage overhead, the proposed ABAS of GRCs can effectively reduce the GRBC compared to the traditional symmetric bandwidth allocation schemes. Shushi Gu, Fugang Wang, Qinyu Zhang 0001, Tao Huang 0008, Wei Xiang 0001 |
IET Commun. | 4 |
| 2019 | Machine learning based optimization for vehicle-to-infrastructure communications
Wei Xiang 0001, Tao Huang 0008 |
Future Gener. Comput. Syst. | 2 |
| 2014 | Degrees of freedom of half-duplex MIMO multi-way relay channel with full data exchangeabstractIn this paper, we investigate the degrees of freedom (DoF) of the half-duplex multiple-input multiple-output (MIMO) multi-way relay channel (MWRC) with full data exchange, where each user wants to learn all the messages from the other users in the channel. We observe that, unlike the case of pairwise data exchange, the uplink and downlink traffic loads are asymmetric when full data exchange is considered. This asymmetry implies that unequal uplink/downlink time allocation, which is allowed in a half-duplex system, can improve the DoF of the MIMO MWRC. Based on that, we derive the DoF capacity of the half-duplex MIMO MWRC with full data exchange. We show that, as compared to the equal uplink/downlink time allocation, the optimized uplink/downlink time allocation achieves a significant DoF gain. Tao Huang 0008, Xiaojun Yuan 0002, Jinhong Yuan |
GLOBECOM | 1 |
| 2013 | Opportunistic pair-wise compute-and-forward in multi-way relay channelsabstractIn this paper, we propose a novel opportunistic pair-wise transmission scheme in a multi-way relay channel (MWRC), in which multiple users exchange information via a common relay. We investigate pair-wise compute-and-forward for MWRCs by exploiting the multi-user fading channels. Conventionally, a pair-wise physical-layer network coding scheme with binary phase shift keying modulation was studied for a MWRC. In this paper, the proposed opportunistic pair-wise compute-and-forward employs high level modulation with nested lattice codes to improve the sum-rate of multi-user transmission. We demonstrate that this novel opportunistic pair-wise transmission has a 2 bits/s/Hz improvement in the sum-rate performance at signal-to-noise ratio of 30 dB for a 4-user MWRC. For the same MWRC, up to 4.5 dB gain or 2.5 dB gain can be achieved for an uncoded or a channel-coded system, respectively, at the frame error probability of 10-2. Tao Huang 0008, Jinhong Yuan, Qifu Tyler Sun |
ICC | 1 |
| 2013 | Novel nested convolutional lattice codes for multi-way relaying systems over fading channelsabstractIn this paper, we focus on the realization of multiple interpretations (MI) in multi-way relay channels (MWRC) with fading, where multiple sources communicate with each other with the help of a relay. We first propose a novel nested convolutional lattice codes (NCLC) over the finite field, which can achieve the MI for each source in two time slots. Then we derive a theoretical upper bound for the codeword error rate (WER) of the NCLC. We further optimize our NCLC by developing a code design criterion which minimizes the derived WER. In simulations, we construct a specific NCLC based on our code design criterion. Simulation results show that our code can realize MI for each source in two time slots, and validate the derived upper bound in the high normalized signal-to-effective-noise ratio (SENRnorm) region. Yuanye Ma, Tao Huang 0008, Jun Li 0004, Jinhong Yuan, Zihuai Lin, Branka Vucetic |
WCNC | 2 |
| 2013 | Design of Irregular Repeat-Accumulate Coded Physical-Layer Network Coding for Gaussian Two-Way Relay ChannelsabstractThis paper addresses the design of irregular repeat accumulate (IRA) codes for coded physical-layer network coding (PNC) for the binary-input Gaussian two-way relay channel, assuming perfect synchronization and equal received power at the relay. The design is based on a nontrivial extension of EXIT-chart based design. Specifically, we analyze the components of the IRA-PNC scheme and propose an approach to model the soft information exchanged between these components. Then, we develop upper and lower bounds on the extrinsic information transfer functions to characterize the iterative process of computing the network-coded information. Based on that, we construct optimized IRA codes to minimize the computation error at the relay. The optimized IRA-PNC has considerable performance improvement over the existing regular RA coded PNC. For a rate 3/4 code, as an example, we observed improvements of 2.6 dB, and the optimized IRA-PNC scheme is only about 1.7 dB away from the capacity upper bound of the Gaussian two-way relay channel. Tao Huang 0008, Tao Yang 0004, Jinhong Yuan, Ingmar Land |
IEEE Trans. Commun. | 1 |
| 2013 | Lattice Network Codes Based on Eisenstein IntegersabstractIn this paper, we investigate lattice network codes (LNCs) constructed from Eisenstein integer based lattices. Quantization and encoding algorithms over Eisenstein integers are first introduced. Then, a union bound estimation (UBE) of the decoding error probability is derived when the shaping region of the LNC is a product of regular hexagons. Next, the Gaussian reduction algorithm is generalized to be applicable to complex lattices over Eisenstein integers such that an optimal coefficient vector can be found in the two-transmitter single-relay system. Based on the UBE, design criteria for optimal LNCs with minimum decoding error probability are formulated and applied to construct both Gaussian integer and Eisenstein integer based good LNCs from rate-1/2 feed-forward convolutional codes by Complex Construction A. The constructed codes provide up to 7.65 dB nominal coding gains over Rayleigh fading channels. Furthermore, we introduce the construction of LNCs from linear codes by Complex Construction B. The nominal coding gains and error performance of the LNCs thus constructed are explicitly analyzed. Examples show that the LNCs constructed by Complex Construction B provide a better tradeoff between code rate and nominal coding gain. Qifu Tyler Sun, Jinhong Yuan, Tao Huang 0008, Kenneth W. Shum |
IEEE Trans. Commun. | 3 |
| 2012 | Distance Spectrum and Performance of Channel-Coded Physical-Layer Network Coding for Binary-Input Gaussian Two-Way Relay ChannelsabstractWe investigate a channel-coded physical-layer network coding (CPNC) scheme for binary-input Gaussian two-way relay channels. In this scheme, the codewords of the two users are transmitted simultaneously. The relay computes and forwards a network-coded (NC) codeword without complete decoding of the two users' individual messages. We propose a new punctured codebook method to explicitly find the distance spectrum of the CPNC scheme. Based on that, we derive an asymptotically tight performance bound for the error probability. Our analysis shows that, compared to the single-user scenario, the CPNC scheme exhibits the same minimum Euclidean distance but an increased multiplicity of error events with minimum distance. At a high SNR, this leads to an SNR penalty of at most ln2 (in linear scale), for long channel codes of various rates. Our analytical results match well with the simulated performance. Tao Yang 0004, Ingmar Land, Tao Huang 0008, Jinhong Yuan, Zhuo Chen 0001 |
IEEE Trans. Commun. | 3 |
| 2011 | Distance properties and performance of physical layer network coding with binary linear codes for Gaussian two-way relay channelsabstractWe investigate joint channel and physical layer network coding (CPNC) for Gaussian two-way relay channels. The two users' messages are encoded using the same binary linear code and are transmitted simultaneously with equal power. At the relay node, the network-coded message is recovered directly from the received signal sequence, and is then broadcast to the users. We propose a new methodology to explicitly find the distance spectrum of the coding scheme. Based on that, we analyze the error probability at the relay and derive an asymptotically tight performance bound (for high SNRs). We show that, with a general binary linear code, the CPNC scheme is subject to an asymptotic SNR loss of approximately ln 2 relative to the single-user case, regardless of the coding rate. Numerical results show that our analysis matches very well with the performance of the CPNC scheme. Tao Yang 0004, Ingmar Land, Tao Huang 0008, Jinhong Yuan, Zhuo Chen 0001 |
ISIT | 3 |
| 2011 | Outage performance of analog network coding in generalized two-way multi-hop networksabstractWe investigate the performance of analog network coding (ANC) for multi-hop networks in this paper. With the amplify-and-forward (AF) protocol, relays broadcast the sum of two colliding signals to neighboring nodes, while the source node can subtract its own signal from the colliding signal to obtain the received information. We first give the transmission scheme expressions for the n-node m-frame two-way multi-hop network. For this scheme, we derive the end-to-end signal-to-noise ratio (SNR) expression. Without loss of generality, the closed-form expression of the outage probability for the generalized multi-hop network is evaluated. Numerical results demonstrate that the outage performance with ANC is better than the traditional scheme without ANC. Eric Wang 0001, Wei Xiang 0001, Jinhong Yuan, Tao Huang 0008 |
WCNC | 4 |