Abolfazl Razi

dblp:83/8819 · DBLP profile ↗
← Back
42ranked-venue papers
5as first author
29since 2021 · last 2026
0000-0002-3330-6132ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 15 since 2021Computer networks · 13 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Comparative Analysis of Patch Attack on VLM-Based Autonomous Driving Architectures
David Fernandez, Pedram MohajerAnsari, Amir Salarpour, Long Cheng 0005, Abolfazl Razi, Mert D. Pesé
IV5
2026 RobustFormer: Noise-Robust Pre-training for Images and Videos
abstract
While deep learning-based models like transformers, have revolutionized time-series and vision tasks, they remain highly susceptible to noise and often overfit on noisy patterns rather than robust features. This issue is exacerbated in vision transformers, which rely on pixel-level details that can easily be corrupt. To address this, we leverage the discrete wavelet transform (DWT) for its ability to decompose into multi-resolution layers, isolating noise primarily in the high frequency domain while preserving essential low-frequency information for resilient feature learning. Conventional DWT-based methods, however, struggle with computational inefficiencies due to the requirement for a subsequent inverse discrete wavelet transform (IDWT) step. In this work, we introduce RobustFormer, a novel framework that enables noise-robust masked autoencoder (MAE) pre-training for both images and videos by using DWT for efficient downsampling, eliminating the need for expensive IDWT reconstruction and simplifying the attention mechanism to focus on noise-resilient multi-scale representations. To our knowledge, RobustFormer is the first DWT-based method fully compatible with video inputs and MAE-style pre-training. Extensive experiments on noisy image and video datasets demonstrate that our approach achieves up to 8% increase in Top-1 classification accuracy under severe noise conditions in Imagenet-C and up to 2.7% in Imagenet-P standard benchmarks compared to the baseline and up to 13% higher Top-1 accuracy on UCF-101 under severe custom noise perturbations while maintaining similar accuracy scores for clean datasets. We also observe the reduction of computation complexity by up to 4.4% through IDWT removal compared to VideoMAE baseline without any performance drop.
Ashish Bastola, Nishant Luitel, Hao Wang 0176, Danda Pani Paudel, Roshni Poudel, Abolfazl Razi
WACV6
2026 Prompt-OT: An Optimal Transport Regularization Paradigm for Knowledge Preservation in Vision-Language Model Adaptation
abstract
Vision-language models (VLMs) such as CLIP demonstrate strong performance but struggle when adapted to downstream tasks. Prompt learning has emerged as an efficient and effective strategy to adapt VLMs while preserving their pre-trained knowledge. However, existing methods still lead to overfitting and degrade zero-shot generalization. To address this challenge, we propose an optimal transport (OT)-guided prompt learning framework that mitigates forgetting by preserving the structural consistency of feature distributions between pre-trained and fine-tuned models. Unlike conventional point-wise constraints, OT naturally captures cross-instance relationships and expands the feasible parameter space for prompt tuning, allowing a better tradeoff between adaptation and generalization. Our approach enforces joint constraints on both vision and text representations, ensuring a holistic feature alignment. Extensive experiments on benchmark datasets demonstrate that our simple yet effective method outperforms existing prompt learning strategies in base-to-novel generalization, cross-dataset evaluation, and domain generalization, without requiring additional augmentation or ensemble techniques.
Xiwen Chen, Peijie Qiu, Hao Wang 0176, Haiyu Wu, Aristeidis Sotiras, Yalin Wang 0001, Abolfazl Razi
WACV9
2026 Factorized Transport Alignment for Multimodal and Multiview E-commerce Representation Learning
abstract
The rapid growth of e-commerce requires robust multimodal representations that capture diverse signals from user-generated listings. Existing vision–language models (VLMs) typically align titles with primary images, i.e., single-view, but overlook non-primary images and auxiliary textual views that provide critical semantics in open marketplaces such as Etsy or Poshmark. To this end, we propose a framework that unifies multimodal and multi-view learning through Factorized Transport, a lightweight approximation of optimal transport, designed for scalability and deployment efficiency. During training, the method emphasizes primary views while stochastically sampling auxiliary ones, reducing training cost from quadratic in the number of views to constant per item. At inference, all views are fused into a single cached embedding, preserving the efficiency of two-tower retrieval with no additional online overhead. On an industrial dataset of 1M product listings and 0.3M interactions, our approach delivers consistent improvements in cross-view and query-to-item retrieval, achieving up to +7.9% Recall@500 over strong multimodal baselines. Overall, our framework bridges scalability with optimal transport–based learning, making multi-view pretraining practical for large-scale e-commerce search.
Xiwen Chen, Yen-Chieh Lien, Xueqing Liu 0001, María Castaños, Abolfazl Razi, Xiaoting Zhao, Congzhe Su
WSDM5
2026 Turbo-IRL: Enhancing multi-agent systems using turbo decoding-inspired deep maximum entropy inverse reinforcement learning
abstract
Inverse Reinforcement Learning (IRL) involves analyzing expert agents’ behavior to uncover the rationality behind their actions, with applications in robotics, autonomous systems, gaming, and the animation industry. Conventional Multi-Agent IRL (MA-IRL) approaches typically assume either fully cooperative or fully independent agent objectives. However, real-world scenarios often involve partially aligned goals comprising both shared and agent-specific components. To address this complexity, we propose Turbo-IRL, a novel, parallelizable MA-IRL framework that extends Deep Maximum Entropy IRL with an iterative information exchange mechanism inspired by Turbo Decoding (TD) in coding theory. Turbo-IRL consists of multiple agent-specific IRL modules, each leveraging deep neural networks to recover complex reward structures. Through iterative communication, agents refine their estimates of shared rewards while concurrently discovering individualized reward components. This information exchange mechanism allows them to benefit from others’ trajectories, which convey partial information about their common goals. This is particularly advantageous for under-trained agents with a small number of trajectories in unbalanced scenarios. Our simulations across multiple scenarios with different levels of overlap and similarity between agents’ reward functions show a considerable gain for the proposed Turbo-IRL framework compared to the benchmark Deep IRL method. Specifically, we achieve about 50% and 44% improvement in terms of MSE in recovering the true reward function for similarity levels 75% and 25%, utilizing the same number of total trajectories. Conversely, our method achieves equivalent performance utilizing much fewer (about 44% for a 3-agent system). Furthermore, on the standard MPE simple_spread benchmark, Turbo-IRL achieves the lowest MAE of 1.04 in episode return, outperforming several state-of-the-art baselines, including MA-GAIL, MA-AIRL, Behavior Cloning, and MIFQ (38% to 87% improvement) – underscoring its efficacy, scalability, and generalization capability in complex multi-agent environments with partially shared objectives.
Niloufar Mehrabi, Sayed Pedram Haeri Boroujeni, Abolfazl Razi
Expert Syst. Appl.3
2026 All you need for object detection: From pixels, points, and prompts to Next-Gen fusion and multimodal LLMs/VLMs in autonomous vehicles
abstract
Autonomous Vehicles (AVs) are transforming the future of transportation through advances in intelligent perception, decision-making, and control systems. However, their success is tied to one core capability, reliable object detection in complex and multimodal environments. While recent breakthroughs in Computer Vision (CV) and Artificial Intelligence (AI) have driven remarkable progress, the field still faces a critical challenge as knowledge remains fragmented across multimodal perception, contextual reasoning, and cooperative intelligence. This survey bridges that gap by delivering a forward-looking analysis of object detection in AVs, emphasizing emerging paradigms such as Vision-Language Models (VLMs), Large Language Models (LLMs), and Generative AI rather than re-examining outdated techniques. We begin by systematically reviewing the fundamental spectrum of AV sensors (camera, ultrasonic, LiDAR, and Radar) and their fusion strategies, highlighting not only their capabilities and limitations in dynamic driving environments but also their potential to integrate with recent advances in LLM/VLM-driven perception frameworks. We also review autonomous vehicle simulators as a critical layer for safe development, scalable testing, and reproducible benchmarking of perception and detection pipelines before real-world deployment. Next, we introduce a structured categorization of AV datasets that moves beyond simple collections, positioning ego-vehicle, infrastructure-based, and cooperative datasets (e.g., V2V, V2I, V2X, I2I), followed by a cross-analysis of data structures and characteristics. Ultimately, we analyze cutting-edge detection methodologies, ranging from 2D and 3D pipelines to hybrid sensor fusion, with particular attention to emerging transformer-driven approaches powered by Vision Transformers (ViTs), Large and Small Language Models (SLMs), and VLMs. By synthesizing these perspectives, our survey delivers a clear roadmap of current capabilities, open challenges, and future opportunities, highlighting underexplored avenues such as multimodal reasoning, cooperative perception, and foundation-model integration. We aim to establish this work as a definitive reference for researchers, practitioners, and developers, fostering accelerated innovation toward safer and more intelligent autonomous driving systems. • Comprehensive review of state-of-the-art object detection in autonomous vehicles. • Analysis of latest AV sensors, fusion strategies, and multimodal perception systems. • Novel categorization and comparison of ego-vehicle, roadside, and CP datasets. • In-depth evaluation of 2D, 3D, fusion, and emerging LLM/VLM-based detection methods. • Highlights open challenges and potential advancements in AV perception research.
Sayed Pedram Haeri Boroujeni, Niloufar Mehrabi, Hazim Alzorgan, Mahlagha Fazeli, Abolfazl Razi
Image Vis. Comput.5
2025 Sequence Complementor: Complementing Transformers for Time Series Forecasting with Learnable Sequences
abstract
Since its introduction, the transformer has shifted the development trajectory away from traditional models (e.g., RNN, MLP) in time series forecasting, which is attributed to its ability to capture global dependencies within temporal tokens. Follow-up studies have largely involved altering the tokenization and self-attention modules to better adapt Transformers for addressing special challenges like non-stationarity, channel-wise dependency, and variable correlation in time series. However, we found that the expressive capability of sequence representation is a key factor influencing Transformer performance in time forecasting after investigating several representative methods, where there is an almost linear relationship between sequence representation entropy and mean square error, with more diverse representations performing better. In this paper, we propose a novel attention mechanism with Sequence Complementors and prove feasible from an information theory perspective, where these learnable sequences are able to provide complementary information beyond current input to feed attention. We further enhance the Sequence Complementors via a diversification loss that is theoretically covered. The empirical evaluation of both long-term and short-term forecasting has confirmed its superiority over the recent state-of-the-art methods.
Xiwen Chen, Peijie Qiu, Hao Wang 0176, Aristeidis Sotiras, Yalin Wang 0001, Abolfazl Razi
AAAI8
2025 Multimodal Variational Autoencoder: A Barycentric View
abstract
Multiple signal modalities, such as vision and sounds, are naturally present in real-world phenomena. Recently, there has been growing interest in learning generative models, in particular variational autoencoder (VAE), for multimodal representation learning especially in the case of missing modalities. The primary goal of these models is to learn a modality-invariant and modality-specific representation that characterizes information across multiple modalities. Previous attempts at multimodal VAEs approach this mainly through the lens of experts, aggregating unimodal inference distributions with a product of experts (PoE), a mixture of experts (MoE), or a combination of both. In this paper, we provide an alternative generic and theoretical formulation of multimodal VAE through the lens of barycenter. We first show that PoE and MoE are specific instances of barycenters, derived by minimizing the asymmetric weighted KL divergence to unimodal inference distributions. Our novel formulation extends these two barycenters to a more flexible choice by considering different types of divergences. In particular, we explore the Wasserstein barycenter defined by the 2-Wasserstein distance, which better preserves the geometry of unimodal distributions by capturing both modality-specific and modality-invariant representations compared to KL divergence. Empirical studies on three multimodal benchmarks demonstrated the effectiveness of the proposed method.
Peijie Qiu, Sayantan Kumar, Xiwen Chen, Abolfazl Razi, Yalin Wang 0001, Aristeidis Sotiras
AAAI7
2025 Adaptive Data Transport Mechanism for UAV Surveillance Missions in Lossy Environments
abstract
Unmanned Aerial Vehicles (UAVs) play an increasingly critical role in Intelligence, Surveillance, and Reconnaissance (ISR) missions such as border patrolling and criminal detection due to their ability to access remote areas and transmit real-time imagery to servers. However, UAVs face limitations in payload, power, and communication bandwidth, necessitating selective data transmission strategies. While traditional methods strive to preserve maximal information in transferred video frames, missing the fact that only certain parts of images/video frames are relevant for Object Detection and Tracking (OD/OT) in ISR missions. This paper adopts a different perspective and offers an alternative AI-driven scheduling policy that prioritizes selecting regions of the image that significantly contribute to the mission objective. The key idea is tiling the image into small patches and developing a Deep Reinforcement Learning (DRL) framework that assigns higher transmission probabilities to patches that present higher overlaps with the detected object of interest while penalizing sharp transitions over consecutive frames to promote smooth scheduling shifts. Although we used YOLOv8 object detection and UDP transmission protocols as a benchmark testing scenario, the idea is general and applicable to different transmission protocols and OD/OT methods. To further boost the system's performance and avoid OD errors for cluttered image patches, we integrate it with inter-frame interpolations. With this method, we achieved about 45% improvement in terms of OD accuracy for the proposed method (F1 score:98%) compared to random selection (F1 score: 53%) when the transmission budget is 50% (we afford sending half of the image patches). Under an extremely constrained transmission budget (5%), this gain can be as high as 87%. The only cost for such improvement is a feedback channel from the ground server to drones.
Niloufar Mehrabi, Sayed Pedram Haeri Boroujeni, Jenna Hofseth, Abolfazl Razi, Long Cheng 0005, Manveen Kaur, Jim Martin 0001, Rahul Amin
CCNC4
2025 Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image Analysis
abstract
While multiple instance learning (MIL) has shown to be a promising approach for histopathological whole slide image (WSI) analysis, its reliance on permutation invariance significantly limits its capacity to effectively uncover semantic correlations between instances within WSIs. Based on our empirical and theoretical investigations, we argue that approaches that are not permutation-invariant but better capture spatial correlations between instances can offer more effective solutions. In light of these findings, we propose a novel alternative to existing MIL for WSI analysis by learning to restore the order of instances from their randomly shuffled arrangement. We term this task as cracking an instance jigsaw puzzle problem, where semantic correlations between instances are uncovered. To tackle the instance jigsaw puzzles, we propose a novel Siamese network solution, which is theoretically justified by optimal transport theory. We validate the proposed method on WSI classification and survival prediction tasks, where the proposed method outperforms the recent state-of-the-art MIL competitors. The code is available at https://github.com/xiwenc1/MIL-JigsawPuzzles.
Xiwen Chen, Peijie Qiu, Hao Wang 0176, Xuanzhao Dong, Yalin Wang 0001, Abolfazl Razi, Aristeidis Sotiras
ICCV10
2025 FIC-TSC: Learning Time Series Classification with Fisher Information Constraint
abstract
Analyzing time series data is crucial to a wide spectrum of applications, including economics, online marketplaces, and human healthcare. In particular, time series classification plays an indispensable role in segmenting different phases in stock markets, predicting customer behavior, and classifying worker actions and engagement levels. These aspects contribute significantly to the advancement of automated decision-making and system optimization in real-world applications. However, there is a large consensus that time series data often suffers from domain shifts between training and test sets, which dramatically degrades the classification performance. Despite the success of (reversible) instance normalization in handling the domain shifts for time series regression tasks, its performance in classification is unsatisfactory. In this paper, we propose $\textit{FIC-TSC}$, a training framework for time series classification that leverages Fisher information as the constraint. We theoretically and empirically show this is an efficient and effective solution to guide the model converges toward flatter minima, which enhances its generalizability to distribution shifts. We rigorously evaluate our method on 30 UEA multivariate and 85 UCR univariate datasets. Our empirical results demonstrate the superiority of the proposed method over 14 recent state-of-the-art methods.
Xiwen Chen, Peijie Qiu, Hao Wang 0176, Yalin Wang 0001, Aristeidis Sotiras, Abolfazl Razi
ICML9
2025 How Effective Can Dropout Be in Multiple Instance Learning ?
abstract
Multiple Instance Learning (MIL) is a popular weakly-supervised method for various applications, with a particular interest in histological whole slide image (WSI) classification. Due to the gigapixel resolution of WSI, applications of MIL in WSI typically necessitate a two-stage training scheme: first, extract features from the pre-trained backbone and then perform MIL aggregation. However, it is well-known that this suboptimal training scheme suffers from "noisy" feature embeddings from the backbone and inherent weak supervision, hindering MIL from learning rich and generalizable features. However, the most commonly used technique (i.e., dropout) for mitigating this issue has yet to be explored in MIL. In this paper, we empirically explore how effective the dropout can be in MIL. Interestingly, we observe that dropping the top-k most important instances within a bag leads to better performance and generalization even under noise attack. Based on this key observation, we propose a novel MIL-specific dropout method, termed MIL-Dropout, which systematically determines which instances to drop. Experiments on five MIL benchmark datasets and two WSI datasets demonstrate that MIL-Dropout boosts the performance of current MIL methods with a negligible computational cost. The code is available at https://github.com/ChongQingNoSubway/MILDropout.
Peijie Qiu, Xiwen Chen, Zhangsihao Yang, Aristeidis Sotiras, Abolfazl Razi, Yalin Wang 0001
ICML6
2025 RD-DPP: Rate-Distortion Theory Meets Determinantal Point Process to Diversify Learning Data Samples
abstract
Selecting representative samples plays an indispensable role in many machine learning and computer vision applications under limited resources (e.g., limited communication bandwidth and computational power). Determinantal Point Process (DPP) is a widely used method for selecting the most diverse representative samples that can summarize a dataset. However, its adaptability to different tasks remains an open challenge, as it is challenging for DPP to perform task-specific tuning. In contrast, Rate-Distortion (RD) theory provides a way to measure task-specific diversity. However, optimizing RD for a data selection problem remains challenging because the quantity that needs to be optimized is the index set of the selected samples. To tackle these challenges, we first draw an inherent relationship between DPP and RD theory. Our theoretical derivation paves the way to take advantage of both RD and DPP for a task-specific data selection. To this end, we propose a novel method for task-specific data selection for multi-level classification tasks, named RD-DPP. Empirical studies on seven different datasets using five benchmark models demonstrate the effectiveness of the proposed RD-DPP method. Our method also outperforms recent strong competing methods, while exhibiting high generalizability to a variety of learning tasks. The source code is available on https://github.com/xiwencl/RD-DPP1.
Xiwen Chen, Peijie Qiu, Rahul Amin, Abolfazl Razi
WACV6
2025 DISCOVER: A Cyberinfrastructure Testbed for Distributed Computing and Networking in Rural and Remote Environments
abstract
The Distributed Sensing and Computing Over Sparse Environments (DISCOVER) testbed is a pioneering cyberinfrastructure initiative designed to advance research in distributed computing and networking tailored to rural, remote, and sparsely populated regions. Supported by the National Science Foundation (NSF), DISCOVER integrates a network of configurable Internet-of-Things (IoT) nodes—including stationary sensors, drones, and terrestrial rovers—across three key sites: Northern Arizona University (NAU), Clemson University, and Navajo Technical University (NTU). This collaboration offers a unique platform to explore innovative algorithms and methodologies addressing the technical challenges of under-served areas, with an emphasis on environmental and civil disaster response. The testbed enables a wide range of experiments, such as regional-scale data collection, heterogeneous networked services, distributed artificial intelligence (AI), distributed multi-robot control, and communication-aware software for resource-constrained networks. An online portal enhances accessibility, allowing researchers to request resources, upload experimental code, and retrieve data, with pre-integrated deep learning models for applications like human posture detection, object detection, and wildfire detection. This paper outlines the testbed’s architecture, operational sites, supported experiment types, ongoing research efforts, and its educational and outreach impacts, highlighting its role in fostering scientific innovation.
Alireza Ebrahimi, Connor Gouin, Sayed Pedram Haeri Boroujeni, Juan Carlos Tique Rangel, Tolunay Seyfi, Truong Nghiem, Abolfazl Razi, Morgan Vigil-Hayes, Paul L. Heinrich, Fatemeh Afghah
WoWMoM7
2024 Geographical Information Alignment Boosts Traffic Analysis via Transpose Cross-attention
abstract
Traffic accident prediction is crucial for enhancing road safety and mitigating congestion, and recent Graph Neural Networks (GNNs) have shown promise in modeling the inherent graph-based traffic data. However, existing GNN-based approaches often overlook or do not explicitly exploit geographic position information, which often plays a critical role in understanding spatial dependencies. This is also aligned with our observation, where accident locations are often highly relevant. To address this issue, we propose a plug-in-and-play module for common GNN frameworks, termed Geographic Information Alignment (GIA). This module can efficiently fuse the node feature and geographic position information through a novel Transpose Cross-attention mechanism. Due to the large number of nodes for traffic data, the conventional cross-attention mechanism performing the node-wise alignment may be infeasible in computation-limited resources. Instead, we take the transpose operation for Query, Key, and Value in the Cross-attention mechanism, which substantially reduces the computation cost while maintaining sufficient information. Experimental results for both traffic occurrence prediction and severity prediction (severity levels based on the interval of recorded crash counts) on large-scale city-wise datasets confirm the effectiveness of our proposed method. For example, our method can obtain gains ranging from 1.3% to 10.9% in F1 score and 0.3% to 4.8% in AUC1.
Xiangyu Jiang, Xiwen Chen, Abolfazl Razi
IEEE Big Data4
2024 FLAME Diffuser: Wildfire Image Synthesis using Mask Guided Diffusion
abstract
Wildfires are a significant threat to ecosystems and human infrastructure, leading to widespread destruction and environmental degradation. Recent advancements in deep learning and generative models have enabled new methods for wildfire detection and monitoring. However, the scarcity of annotated wildfire images limits the development of robust models for these tasks. In this work, we present the FLAME Diffuser, a training-free, diffusion-based framework designed to generate realistic wildfire images with paired ground truth. Our framework uses augmented masks, sampled from real wildfire data, and applies Perlin noise to guide the generation of realistic flames. By controlling the placement of these elements within the image, we ensure precise integration while maintaining the original image’s style. We evaluate the generated images using normalized Fréchet Inception Distance (nFID), CLIP Score, and a custom CLIP Confidence metric, demonstrating the high quality and realism of the synthesized wildfire images. Specifically, the fusion of Perlin noise in this work significantly improved the quality of synthesized images. The proposed method is particularly valuable for enhancing datasets used in downstream tasks such as wildfire detection and monitoring.The resources for accessing the code and datasets are provided at the link: https://arazi2.github.io/aisends.github.io/project/flame
Sayed Pedram Haeri Boroujeni, Xiwen Chen, Ashish Bastola, Abolfazl Razi
IEEE Big Data7
2024 Motor Focus: Fast Ego-Motion Prediction for Assistive Visual Navigation
abstract
Assistive visual navigation systems for visually impaired individuals have become increasingly popular thanks to the rise of mobile computing. Most of these devices work by translating visual information into voice commands. In complex scenarios where multiple objects are present, it is imperative to prioritize object detection and provide immediate notifications for key entities in specific directions. This brings the need for identifying the observer's motion direction (ego-motion) by merely processing visual information, which is the key contribution of this paper. Specifically, we introduce Motor Focus, a lightweight image-based framework that predicts the ego-motion - the humans' (and humanoid machines') movement intentions based on their visual feeds, while filtering out camera motion without any camera calibration. To this end, we implement an optical flow-based pixel-wise temporal analysis method to compensate for the camera motion with a Gaussian aggregation to smooth out the movement prediction area. Subsequently, to evaluate the performance, we collect a dataset including 50 clips of pedestrian scenes in 5 different scenarios. We tested this framework with classical feature detectors such as SIFT and ORB to show the comparison. Our framework demonstrates its superiority in speed$(> 40\text{FPS})$, accuracy (MAE$=60\text{pixels})$, and robustness (SNR$=23\text{dB})$, confirming its potential to enhance the usability of vision-based assistive navigation tools in complex environments. The code is publicly available at https://arazi2.github.io/aisends.github.io/project/VisionGPT.
Jiayou Qin, Xiwen Chen, Ashish Bastola, John Suchanek, Zihao Gong, Abolfazl Razi
BSN7
2024 DGR-MIL: Exploring Diverse Global Representation in Multiple Instance Learning for Whole Slide Image Classification
Xiwen Chen, Peijie Qiu, Aristeidis Sotiras, Abolfazl Razi, Yalin Wang 0001
ECCV (38)5
2024 Enhancing Graph Neural Networks in Large-scale Traffic Incident Analysis with Concurrency Hypothesis
abstract
Despite recent progress in reducing road fatalities, the persistently high rate of traffic-related deaths highlights the necessity for improved safety interventions. Leveraging large-scale graph-based nationwide road network data across 49 states in the USA, our study first posits the Concurrency Hypothesis from intuitive observations, suggesting a significant likelihood of incidents occurring at neighboring nodes within the road network. To quantify this phenomenon, we introduce two novel metrics, Average Neighbor Crash Density (ANCD) and Average Neighbor Crash Continuity (ANCC), and subsequently employ them in statistical tests to validate the hypothesis rigorously. Building upon this foundation, we propose the Concurrency Prior (CP) method, a powerful approach designed to enhance the predictive capabilities of general Graph Neural Network (GNN) models in semi-supervised traffic incident prediction tasks. Our method allows GNNs to incorporate concurrent incident information, as mentioned in the hypothesis, via tokenization with negligible extra parameters. The extensive experiments, utilizing real-world data across states and cities in the USA, demonstrate that integrating CP into 12 state-of-the-art GNN architectures leads to significant improvements, with gains ranging from 3% to 13% in F1 score and 1.3% to 9% in AUC metrics. The code is publicly available at https://github.com/xiwenc1/Incident-GNN-CP1.
Xiwen Chen, Sayed Pedram Haeri Boroujeni, Xin Shu 0006, Abolfazl Razi
SIGSPATIAL/GIS5
2024 TimeMIL: Advancing Multivariate Time Series Classification via a Time-aware Multiple Instance Learning
abstract
Deep neural networks, including transformers and convolutional neural networks (CNNs), have significantly improved multivariate time series classification (MTSC). However, these methods often rely on supervised learning, which does not fully account for the sparsity and locality of patterns in time series data (e.g., quantification of diseases-related anomalous points in ECG and abnormal detection in signal). To address this challenge, we formally discuss and reformulate MTSC as a weakly supervised problem, introducing a novel multiple-instance learning (MIL) framework for better localization of patterns of interest and modeling time dependencies within time series. Our novel approach, TimeMIL, formulates the temporal correlation and ordering within a time-aware MIL pooling, leveraging a tokenized transformer with a specialized learnable wavelet positional token. The proposed method surpassed 26 recent state-of-the-art MTSC methods, underscoring the effectiveness of the weakly supervised TimeMIL in MTSC. The code is available https://github.com/xiwenc1/TimeMIL.
Xiwen Chen, Peijie Qiu, Hao Wang 0176, Aristeidis Sotiras, Yalin Wang 0001, Abolfazl Razi
ICML8
2024 SelfReg-UNet: Self-Regularized UNet for Medical Image Segmentation
Xiwen Chen, Peijie Qiu, Mohammad Farazi, Aristeidis Sotiras, Abolfazl Razi, Yalin Wang 0001
MICCAI (8)6
2024 IC-GAN: An Improved Conditional Generative Adversarial Network for RGB-to-IR image translation with applications to forest fire monitoring
Sayed Pedram Haeri Boroujeni, Abolfazl Razi
Expert Syst. Appl.2
2023 Design and Evaluation of an Application-Oriented Data-Centric Communication Framework for Emerging Cyber-Physical Systems
abstract
Emergent Cyber-Physical Systems (CPSs) like VANETs and UAV swarms are expected to fulfill essential roles in critical infrastructure domains. This increasing utility and the present nurturing economic conditions that enable their cost-effective deployment herald a period of significant growth and adoption. In addition, these systems are increasingly required to support complex data-intensive and QoS-sensitive applications in challenging operating conditions. However, the growth of these systems is limited by current Internet protocols that do not comprehensively meet the communication requirements of these systems. In this work, we present the design and evaluation of Software-Defined NAmed-data enabled Publish-subscribe (SNAP) communication framework that can effectively meet the communication requirements of demanding applications in emergent CPSs.
Manveen Kaur, Abolfazl Razi, Long Cheng 0005, Rahul Amin, Jim Martin 0001
CCNC2
2023 Diversity Maximized Scheduling in RoadSide Units for Traffic Monitoring Applications
abstract
This paper develops an optimal data aggregation policy for learning-based traffic control systems based on imagery collected from Road Side Units (RSUs) under imperfect communications. Our focus is optimizing semantic information flow from RSUs to a nearby edge server or cloud-based processing units by maximizing data diversity based on the target machine learning application while taking into account heterogeneous channel conditions and constrained total transmission rate. To this end, we enforce fairness among class labels to increase data diversity for classification problems. Furthermore, we propose a greedy interval-by-interval scheduling policy powered by coalition game theory to reduce the computation complexity. Once, RSUs are selected, we employ a maximum uncertainty method to handpick data samples that contribute the most to the learning performance. Our method yields higher learning accuracy compared to random selection, uniform selection, and network-based optimization methods (e.g., FedCS)1.
Ahmad Sarlak, Abolfazl Razi, Xiwen Chen, Rahul Amin
LCN2
2023 Invited Paper: Actuator Trajectory Planning for UAVs with Overhead Manipulator using Reinforcement Learning
abstract
In this paper, we investigate the operation of an aerial manipulator system, namely an Unmanned Aerial Vehicle (UAV) equipped with a controllable arm with two degrees of freedom to carry out actuation tasks on the fly. Our solution is based on employing a Q-learning method to control the trajectory of the tip of the arm, also called end-effector. More specifically, we develop a motion planning model based on Time To Collision (TTC), which enables a quadrotor UAV to navigate around obstacles while ensuring the manipulator’s reachability. Additionally, we utilize a model-based Q-learning model to independently track and control the desired trajectory of the manipulator’s end-effector, given an arbitrary baseline trajectory for the UAV platform. Such a combination enables a variety of actuation tasks such as high-altitude welding, structural monitoring and repair, battery replacement, gutter cleaning, sky scrapper cleaning, and power line maintenance in hard-to-reach and risky environments while retaining compatibility with flight control firmware. Our RL-based control mechanism results in a robust control strategy that can handle uncertainties in the motion of the UAV, offering promising performance. Specifically, our method achieves 92% accuracy in terms of average displacement error (i.e. the mean distance between the target and obtained trajectory points) using Q-learning with 15,000 episodes1.
Hazim Alzorgan, Abolfazl Razi, Ata Jahangir Moshayedi
PIMRC2
2022 A review of AI-enabled routing protocols for UAV networks: Trends, challenges, and future outlook
abstract
Unmanned Aerial Vehicles (UAVs), as a recently emerging technology, enabled a new breed of unprecedented applications in different domains. This technology's ongoing trend is departing from large remotely-controlled drones to networks of small autonomous drones to collectively complete intricate tasks time and cost-effectively. An important challenge is developing efficient sensing, communication, and control algorithms that can accommodate the requirements of highly dynamic UAV networks with heterogeneous mobility levels. Recently, the use of Artificial Intelligence (AI) in learning-based networking has gained momentum to harness the learning power of cognizant nodes to make more intelligent networking decisions by integrating computational intelligence into UAV networks. An important example of this trend is developing learning-powered routing protocols, where machine learning methods are used to model and predict topology evolution, channel status, traffic mobility, and environmental factors for enhanced routing. This paper reviews AI-enabled routing protocols designed primarily for aerial networks, including topology-predictive and self-adaptive learning-based routing algorithms, with an emphasis on accommodating highly-dynamic network topology. To this end, we justify the importance and adaptation of AI into UAV network communications. We also address, with an AI emphasis, the closely related topics of mobility and networking models for UAV networks, simulation tools and public datasets, and relations to UAV swarming, which serve to choose the right algorithm for each scenario. We conclude by presenting future trends, and the remaining challenges in AI-based UAV networking, for different aspects of routing, connectivity, topology control, security and privacy, energy efficiency, and spectrum sharing.1
Arnau Rovira-Sugranes, Abolfazl Razi, Fatemeh Afghah, Jacob Chakareski
Ad Hoc Networks2
2021 Boosting Belief Propagation for LDPC Codes with Deep Convolutional Neural Network Predictors
abstract
Conventional channel codes are designed to recover channel errors by adding controlled redundancy to transmit bits; however, the main underlying assumption is that information bits are independent and identically distributed (i.i.d.). Short term and linear temporal correlations are assumed to be exploited by the preceding source encoders. This assumption is flawed in some scenarios since many types of data (e.g, audio samples, video frames, and sensor measurements) exhibit long-term relations and intricate dependencies that are not exploitable by conventional source encoders. Furthermore, sending plain information is still commonplace in wireless networks. Therefore, it is essential to design channel encoders that accommodate these conditions. It is well-known that the underlying hidden patterns can be captured by deep learning methods. This important capability is not yet fully utilized in channel encoder design. This work is a primary step towards developing a predictive channel decoder that learns the intricate dependencies within and between data frames using an embedded learning module at the receiver to enhance the bit decoding performance, especially in high-noise regimes. The learning module is integrated with the belief propagation algorithm over bipartite graphs appropriate for low-density parity-check (LDPC) codes. The proposed method is universal since no specific correlation model is adopted and the learning-based prediction is performed at the bit level. The proposed method is fully implemented at the receiver side, making it compatible with generic LDPC encoders. Our simulations demonstrate the superior performance of the proposed method compared to standard LDPC decoders. For instance, about 1.7 dB gain at the 10-4BER level is achieved when recovering noisy audio files.
Xiwen Chen, Junsuo Qu, Abolfazl Razi
CCNC4
2021 Towards Boosting Channel Attention for Real Image Denoising: Sub-band Pyramid Attention
Haiyu Wu, Xiwen Chen, Hao Wang 0176, Abolfazl Razi
ICIG (3)5
2021 Aerial imagery pile burn detection using deep learning: The FLAME dataset
abstract
Wildfires are one of the costliest and deadliest natural disasters in the US, causing damage to millions of hectares of forest resources and threatening the lives of people and animals. Of particular importance are risks to firefighters and operational forces, which highlights the need for leveraging technology to minimize danger to people and property. FLAME (Fire Luminosity Airborne-based Machine learning Evaluation) offers a dataset of aerial images of fires along with methods for fire detection and segmentation which can help firefighters and researchers to develop optimal fire management strategies. This paper provides a fire image dataset collected by drones during a prescribed burning piled detritus in an Arizona pine forest. The dataset includes video recordings and thermal heatmaps captured by infrared cameras. The captured videos and images are annotated, and labeled frame-wise to help researchers easily apply their fire detection and modeling algorithms. The paper also highlights solutions to two machine learning problems: (1) Binary classification of video frames based on the presence [and absence] of fire flames. An Artificial Neural Network (ANN) method is developed that achieved a 76% classification accuracy. (2) Fire detection using segmentation methods to precisely determine fire borders. A deep learning method is designed based on the U-Net up-sampling and down-sampling approach to extract a fire mask from the video frames. Our FLAME method approached a precision of 92%, and recall of 84%. Future research will expand the technique for free burning broadcast fire using thermal images.
Alireza Shamsoshoara, Fatemeh Afghah, Abolfazl Razi, Peter Fule, Erik Blasch
Comput. Networks3
2020 Wildfire Spread Modeling with Aerial Image Processing
abstract
Currently, wildfire spread modeling has drawn a lot of attention from the research community since many countries are suffering from severe socioeconomic impacts of wildfires, every year. Fire spread modeling is a key requirement for effective fire management to deploy fire control equipment and forces at the right time and locations, and plan timely evacuations of residential areas. This paper proposes a new data-driven model for fire expansion which uses reference-based image segmentation for vegetation density estimation and incorporates it into the fire heat conduction modeling. Compared with the conventional parameter collection methods at fire scenes, our method relies on topview images taken by unmanned aerial vehicles, which provides significant advantages of flexibility, safety, low cost, and convenience. Our low-complexity and probabilistic model incorporates the terrain slope, vegetation density, and wind factors with adjustable model parameters which can be easily learned from experiments. The proposed model is flexible and applicable to forests with mixed vegetation and different geographical and climate conditions. We evaluate the fire propagation model by comparing the results with the propagation data available for California Rim fire in 2013.
Qiyuan Huang, Abolfazl Razi, Fatemeh Afghah, Peter Fule
WoWMoM2
2019 Optimized Compression Policy for Flying Ad hoc Networks
abstract
Managing energy consumption for computation and communication is a key requirement for flying ad hoc networks (FANET) to prolong the network lifetime. In many applications, the main role of drones is to collect imagery information and relay them to a ground station for further processing and decision making. In this paper, we present a predictive compression policy to maximize the end-to-end image quality penalized by the communication and computation costs. The idea is to predict the number of remaining links to the destination for a given routing algorithm and use it to re-compress image frames at intermediate nodes such that the overall energy consumption is minimized. Numerical results confirm that the performance of this method is within 4% of the global optima and higher than the current fixed-rate policies with a significant margin.
Arnau Rovira-Sugranes, Fatemeh Afghah, Abolfazl Razi
CCNC3
2019 Distributed Cooperative Spectrum Sharing in UAV Networks Using Multi-Agent Reinforcement Learning
abstract
In this paper, we develop a distributed mechanism for spectrum sharing among a network of unmanned aerial vehicles (UAV) and licensed terrestrial networks. This method can provide a practical solution for situations where the UAV network may need external spectrum when dealing with congested spectrum or need to change its operational frequency due to security threats. Here we study a scenario where the UAV network performs a remote sensing mission. In this model, the UAVs are categorized to two clusters of relaying and sensing UAVs. The relay UAVs provide a relaying service for a licensed network to obtain spectrum access for the rest of UAVs that perform the sensing task. We develop a distributed mechanism in which the UAVs locally decide whether they need to participate in relaying or sensing considering the fact that communications among UAVs may not be feasible or reliable. The UAVs learn the optimal task allocation using a distributed reinforcement learning algorithm. Convergence of the algorithm is discussed and simulation results are presented for different scenarios to verify the convergence.
Alireza Shamsoshoara, Mehrdad Khaledi, Fatemeh Afghah, Abolfazl Razi, Jonathan D. Ashdown
CCNC4
2019 A Solution for Dynamic Spectrum Management in Mission-Critical UAV Networks
abstract
In this paper, we study the problem of spectrum scarcity in a network of unmanned aerial vehicles (UAVs) during mission-critical applications such as disaster monitoring and public safety missions, where the pre-allocated spectrum is not sufficient to offer a high data transmission rate for real-time video-streaming. In such scenarios, the UAV network can lease part of the spectrum of a terrestrial licensed network in exchange for providing relaying service. In order to optimize the performance of the UAV network and prolong its lifetime, some of the UAVs will function as a relay for the primary network while the rest of the UAVs carry out their sensing tasks. Here, we propose a team reinforcement learning algorithm performed by the UAV's controller unit to determine the optimum allocation of sensing and relaying tasks among the UAVs as well as their relocation strategy at each time. We analyze the convergence of our algorithm and present simulation results to evaluate the system throughput in different scenarios.
Alireza Shamsoshoara, Mehrdad Khaledi, Fatemeh Afghah, Abolfazl Razi, Jonathan D. Ashdown, Kurt A. Turck
SECON4
2018 Energy Efficiency Analysis of UAV-Assisted mmWave HetNets
abstract
We study downlink transmission in a multi-band heterogeneous network comprising unmanned aerial vehicle (UAV) small base stations and ground-based dual mode mmWave small cells within the coverage area of a microwave (μW) macro base station. We formulate a two-layer optimization framework to simultaneously find efficient coverage radius for the UAVs and energy efficient radio resource management for the network, subject to minimum quality-of-service (QoS) and maximum transmission power constraints. The outer layer derives an optimal coverage radius/height for each UAV as a function of the maximum allowed path loss. The inner layer formulates an optimization problem to maximize the system energy efficiency (EE), defined as the ratio between the aggregate user data rate delivered by the system and its aggregate energy consumption (downlink transmission and circuit power). We demonstrate that at certain values of the target SINR τ introducing the UAV base stations doubles the EE. We also show that an increase in τ beyond an optimal EE point decreases the EE.
Syed Naqvi, Jacob Chakareski, Nicholas Mastronarde, Jie Xu 0001, Fatemeh Afghah, Abolfazl Razi
ICC6
2017 Delay Minimization by Adaptive Framing Policy in Cognitive Sensor Networks
abstract
In this paper, a delay-minimal joint framing and scheduling policy is proposed for cognitive sensor networks, where a secondary sensor node collects measurement samples, combines them together into packets and transmits the packets to a central data fusion center after proper scheduling, when a shared channel is released by primary nodes. The main objective of this study is to minimize the end-to-end delivery time for secondary sensor nodes based on the current input traffic rate, channel availability process and channel bit error probability. The proposed method outperforms conventional constant-length framing policies for any choice of packet length by minimizing the delay and preventing potential queue instability under dynamic channel conditions. This method can be utilized by secondary nodes in a wide variety of wireless sensing applications in order to collect time-sensitive data with minimal delays.
Abolfazl Razi, Ali Valehi, Elizabeth S. Bentley
WCNC1
2016 Optimizing low density parity check code for two parallel erasure links
Hassan Tavakoli, Abolfazl Razi
ISITA2
2016 Channel-Adaptive Packetization Policy for Minimal Latency and Maximal Energy Efficiency
abstract
This paper considers the problem of delay-optimal bundling of the input symbols into transmit packets in the entry point of a wireless sensor network such that the link delay is minimized under an arbitrary arrival rate and a given channel error rate. The proposed policy exploits the variable packet length feature of contemporary communications protocols in order to minimize the link delay via packet length regularization. This is performed through concrete characterization of the end-to-end link delay for zero-error tolerance system with first come first serve (FCFS) queuing discipline and automatic repeat request (ARQ) re-transmission mechanism. The derivations are provided for an uncoded system as well as a coded system with a given bit error rate. The proposed packetization policy provides an optimal packetization interval that minimizes the end-to-end delay for a given channel with certain bit error probability. This algorithm can also be used for near-optimal bundling of input symbols for dynamic channel conditions provided that the channel condition varies slowly over time with respect to symbol arrival rate. This algorithm complements the current network-based delay-optimal routing and scheduling algorithms in order to further reduce the end-to-end delivery time. Moreover, the proposed method is employed to solve the problem of energy efficiency maximization under an average delay constraint by recasting it as a convex optimization problem.
Abolfazl Razi, Fatemeh Afghah, Ali Abedi 0001
IEEE Trans. Wirel. Commun.1
2014 Nonlinear Information-Theoretic Compressive Measurement Design
abstract
We investigate design of general nonlinear functions for mapping high-dimensional data into a lower-dimensional (compressive) space. The nonlinear measurements are assumed contaminated by additive Gaussian noise. Depending on the application, we are either interested in recovering the high-dimensional data from the nonlinear compressive measurements, or performing classification directly based on these measurements. The latter case corresponds to classification based on nonlinearly constituted and noisy features. The nonlinear measurement functions are designed based on constrained mutual-information optimization. New analytic results are developed for the gradient of mutual information in this setting, for arbitrary input-signal statistics. We make connections to kernel-based methods, such as the support vector machine. Encouraging results are presented on multiple datasets, for both signal recovery and classification. The nonlinear approach is shown to be particularly valuable in high-noise scenarios.
Liming Wang 0004, Abolfazl Razi, Miguel R. D. Rodrigues, A. Robert Calderbank, Lawrence Carin
ICML2
2014 Convergence Analysis of Iterative Decoding for Binary CEO Problem
abstract
Estimation of a binary source using multiple observers, a variant of the so called Chief Executive Officer (CEO) problem, is considered. A low-complexity Distributed Joint Source Channel Coding (D-JSCC) based on the Parallel Concatenated Convolutional Codes (PCCC) is implemented in a cluster of sensors in a distributed fashion. Convergence of the iterative decoder is analyzed by utilizing EXtrinsic Information Transfer (EXIT) chart technique to determine the convergence region in terms of the sensors observation accuracy and channel SNR, where the iterative decoder outperforms the non-iterative one. This leads to design of a bi-modal decoder that adaptively switches between the iterative and non-iterative modes in order to avoid inefficient iterative information exchange without compromising the resulting Bit Error Rate (BER). This adaptive decoding algorithm saves the computational power and decoding time by a factor of about 10 by avoiding unnecessary iterations.
Abolfazl Razi, Ali Abedi 0001
IEEE Trans. Wirel. Commun.1
2013 Power Optimized DSTBC Assisted DMF Relaying in Wireless Sensor Networks with Redundant Super Nodes
abstract
In conventional two-tiered Wireless Sensor Networks (WSN), sensors in each cluster transmit observed data to a fusion center via an intermediate supernode. This structure is vulnerable to supernode failure. A double supernode system model with a new coding scheme is proposed to monitor a binary data source. A Distributed Joint Source Channel Code (D-JSCC) is proposed for sensors inside a cluster that provides two advantages of low complexity transmitters and scalability to a large number of sensors. In order to setup a robust communication channel from sensors to the data fusion center, Distributed Space-Time Block Coding (D-STBC) is employed at two supernodes prior to relaying that results in additional diversity gain. DeModulate and Forward (DMF) relaying mode is chosen to enable packet reformatting at the supernodes, which is not possible in widely used Amplify and Forward (AF) mode. The optimum power allocation for the two-hop multiple DMF relaying is calculated to minimize the system Bit Error Rate (BER). An upper bound is derived for the system end-to-end BER by analyzing a basic decoder operation over the system model. The simulation results validate this upper bound and also demonstrate considerable improvement in the system BER for the proposed coding scheme.
Abolfazl Razi, Fatemeh Afghah, Ali Abedi 0001
IEEE Trans. Wirel. Commun.1
2011 Throughput optimization in relay networks using Markovian game theory
abstract
In this paper, problem of throughput optimization in relay networks where all users transmit their packets on a multiple-access channel is studied by introducing a new Markovian game theoretical solution. Despite the previously reported works, in the proposed model, simultaneously transmitted packets on a multiple-access channel are not always discarded. In this article, the possibility of capturing one of these packets is considered. Both cases of cooperative and non-cooperative stochastic game solutions are investigated and compared. The main objective is to maximize the system throughput with minimum transmission delay and power consumption cost. Effect of different packet error rates due to possible collision occurrence is considered in the game definition that improves system performance further. Performance of the proposed non-cooperative game model approaches the cooperative case, however the non-cooperative game model adds less signaling load to the system, therefore it is more likely to be used in practical applications.
Fatemeh Afghah, Abolfazl Razi, Ali Abedi 0001
WCNC2
2011 On minimum number of wireless sensors required for reliable binary source estimation
abstract
The CEO problem of estimating a single binary data source using multiple observations is considered. A closed form equation for maximum achievable information capacity of the system is derived for different system parameters including number of sensors, observation accuracy and channel quality. A new criterion for sensor clustering is provided based on the minimum number of sensors required to achieve arbitrarily low estimation error using different coding rates. A distributed joint source-channel coding (D-JSCC) scheme is proposed as an implementation example. The proposed coding scheme is based on the parallel concatenated convolutional codes (PCCC) that achieves high BER performance utilizing the number of sensors determined by the derived criterion.
Abolfazl Razi, Keyvan Yasami, Ali Abedi 0001
WCNC1