Mohammad Loni

dblp:225/6712 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-9704-7117ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 DIANA: Drone Imagery for Archipelago Navigation and Analysis for Maritime Object Detection
abstract
The integration of unmanned aerial vehicles (UAVs) into maritime tasks such as debris detection, vessel monitoring, and tracking of marine life is crucial to advancing autonomous navigation. However, the effectiveness of UAVs depends on robust object detection algorithms, which are currently hindered by the limited scale, diversity, and relevance of existing maritime datasets. To address these limitations, (i) we introduce DIANA (Drone Imagery for Archipelago Navigation and Analysis), the first UAV dataset featuring 17,758 high-resolution images annotated with more than 353,141 objects. This dataset spans diverse maritime environments and contains a wide range of object sizes, from small buoys to large vessels. (ii) Our study assesses the impact of data augmentation techniques and a sliding window function on improving accuracy for small objects, as shown by an increase in AP in DiffusionDet (34.4% to 37.8%) with GAN-based augmentation. (iii) We also present a benchmark of state-of-the-art object detection models, including CNN-based and transformer-based architectures, evaluating performance on small, medium, and large objects. Key findings show that transformer-based models, such as DINO and DiffusionDet, excel in small object detection. On the other hand, one-stage CNN models, such as YOLOv8, are effective for medium-size objects. The source code and dataset are available at https://github.com/TurkuAISLAB/DIANA.
Mehdi Asadi, Amin Majd, Mohammad Loni, Juha Kalliovaara
IJCNN3
2025 Robust Few-Shot Semantic Segmentation for Blurred and Occluded Objects in Construction Environments
abstract
The increasing demand for autonomous machines in construction environments necessitates the development of robust object detection algorithms that can perform effectively across various weather and environmental conditions. However, challenging conditions at construction sites, such as mud splashes and vibrations, can degrade object detection performance by causing sensor occlusions and image blurriness. Traditional adversarial training methods, which enhance model robustness by using perturbed data, are limited in construction environments due to the scarcity of diverse real-world adversarial data and the dynamic nature of construction environments. To overcome these challenges, this paper explores utilizing few-shot learning (FSL) to improve the generalization performance and robustness of object detection models. FSL enables models to adapt quickly using minimal data, reducing the need for large datasets. In addition, we identify an often-overlooked issue: the hyperparameters used in FSL training are typically not optimized for this unique paradigm. To address this, we combine FSL with hyperparameter optimization to enhance model performance across multiple small-scale datasets. Experimental results demonstrate that our approach improves model performance on the ConstScene dataset over the default training paradigm. The code for this study is available at here.
Maghsood Salimi, Mohammad Loni, Antonio Cicchetti, Marjan Sirjani
IJCNN2
2025 Efficient Torque Prediction for Digital Twins in Quarry Operations: A Data-Driven and Expert-Guided Approach
abstract
Quarry sites present unique operational challenges where the performance of heavy machinery is critical for maintaining efficiency and safety. In such environments, accurate torque prediction is essential for effective engine management and optimal task execution. This work addresses the torque prediction challenge for a wheel loader operating in quarry conditions by proposing a structured three-phase approach to feature selection that reduces model complexity while preserving predictive accuracy. In the first phase, features are selected based on domain expertise to capture the physical and operational realities of quarry machinery. A comprehensive set of features is then employed to establish a robust performance baseline. In the final phase, a data-driven analysis using SHapley Additive Explanations (SHAP) identifies the top five features that most significantly impact torque prediction. Model efficacy was validated via cross-validation, with R-squared and mean-squared error serving as the key performance indicators. Comparative analysis reveals that while SHAP-ranked features yield statistically optimal results, the expert-selected features are more aligned with the practical requirements of quarry operations. These findings support the design of efficient, interpretable digital twins for real-time decisions in challenging environments.
Abdulkarim Habbab, Anas Fattouh, Mohammad Loni, Koteshwar Chirumalla, Bobbie Frank, Markus Bohlin
INDIN3
2025 LS-HAR: Language Supervised Human Action Recognition with Salient Fusion, Construction Sites as a Use-Case
abstract
Detecting human actions is a crucial task for autonomous robots and vehicles, often requiring the integration of various data modalities for improved accuracy. In this study, we introduce a novel approach to Human Action Recognition (HAR) using language supervision named LS-HAR based on skeleton and visual cues. Our method leverages a language model to guide the feature extraction process in the skeleton encoder. Specifically, we employ learnable prompts for the language model conditioned on the skeleton modality to optimize feature representation. Furthermore, we propose a fusion mechanism that combines dual-modality features using a salient fusion module, incorporating attention and transformer mechanisms to address the modalities’ high dimensionality. This fusion process prioritizes informative video frames and body joints, enhancing the recognition accuracy of human actions. Additionally, we introduce a new dataset tailored for real-world robotic applications in construction sites, featuring visual, skeleton, and depth data modalities, named VolvoConstAct. This dataset serves to facilitate the training and evaluation of machine learning models to instruct autonomous construction machines for performing necessary tasks in real-world construction sites. To evaluate our approach, we conduct experiments on our dataset as well as three widely used public datasets: NTU-RGB+D, NTU-RGB+D 120, and NW-UCLA. Results reveal that our proposed method achieves promising performance across all datasets, demonstrating its robustness and potential for various applications. The code, dataset, and demonstration of real-machine experiments are available at: https://mmahdavian.github.io/ls_har/
Mohammad Mahdavian, Mohammad Loni, Ted Samuelsson, Mo Chen 0001
IROS2
2023 DASS: Differentiable Architecture Search for Sparse Neural Networks
abstract
The deployment of Deep Neural Networks (DNNs) on edge devices is hindered by the substantial gap between performance requirements and available computational power. While recent research has made significant strides in developing pruning methods to build a sparse network for reducing the computing overhead of DNNs, there remains considerable accuracy loss, especially at high pruning ratios. We find that the architectures designed for dense networks by differentiable architecture search methods are ineffective when pruning mechanisms are applied to them. The main reason is that the current methods do not support sparse architectures in their search space and use a search objective that is made for dense networks and does not focus on sparsity. This paper proposes a new method to search for sparsity-friendly neural architectures. It is done by adding two new sparse operations to the search space and modifying the search objective. We propose two novel parametric SparseConv and SparseLinear operations in order to expand the search space to include sparse operations. In particular, these operations make a flexible search space due to using sparse parametric versions of linear and convolution operations. The proposed search objective lets us train the architecture based on the sparsity of the search space operations. Quantitative analyses demonstrate that architectures found through DASS outperform those used in the state-of-the-art sparse networks on the CIFAR-10 and ImageNet datasets. In terms of performance and hardware effectiveness, DASS increases the accuracy of the sparse version of MobileNet-v2 from 73.44% to 81.35% (+7.91% improvement) with a 3.87× faster inference time.
Mohammad Loni, Mina Alibeigi, Masoud Daneshtalab
ACM Trans. Embed. Comput. Syst.2
2022 TAS: Ternarized Neural Architecture Search for Resource-Constrained Edge Devices
abstract
Ternary Neural Networks (TNNs) compress network weights and activation functions into 2-bit representation resulting in remarkable network compression and energy efficiency. However, there remains a significant gap in accuracy between TNNs and full-precision counterparts. Recent advances in Neural Architectures Search (NAS) promise opportunities in automated optimization for various deep learning tasks. Unfortunately, this area is unexplored for optimizing TNNs. This paper proposes TAS, a framework that drastically reduces the accuracy gap between TNNs and their full-precision counterparts by integrating quantization into the network design. We experienced that directly applying NAS to the ternary domain provides accuracy degradation as the search settings are customized for full-precision networks. To address this problem, we propose (i) a new cell template for ternary networks with maximum gradient propagation; and (ii) a novel learnable quantizer that adaptively relaxes the ternarization mechanism from the distribution of the weights and activation functions. Experimental results reveal that TAS delivers 2.64% higher accuracy and ≃2.8 ×memory saving over competing methods with the same bit-width resolution on the CIFAR-10 dataset. These results suggest that TAS is an effective method that paves the way for the efficient design of the next generation of quantized neural networks.
Mohammad Loni, Mohammad Riazati, Masoud Daneshtalab, Mikael Sjödin
DATE1
2022 3DLaneNAS: Neural Architecture Search for Accurate and Light-Weight 3D Lane Detection
Ali Zoljodi, Mohammad Loni, Sadegh Abadijou, Mina Alibeigi, Masoud Daneshtalab
ICANN (1)2
2022 FastStereoNet: A Fast Neural Architecture Search for Improving the Inference of Disparity Estimation on Resource-Limited Platforms
abstract
Convolutional neural networks (CNNs) provide the best accuracy for disparity estimation. However, CNNs are computationally expensive, making them unfavorable for resource-limited devices with real-time constraints. Recent advances in neural architectures search (NAS) promise opportunities in automated optimization for disparity estimation. However, the main challenge of the NAS methods is the significant amount of computing time to explore a vast search space [e.g.,$1.6\times 10^{29}$] and costly training candidates. To reduce the NAS computational demand, many proxy-based NAS methods have been proposed. Despite their success, most of them are designed for comparatively small-scale learning tasks. In this article, we propose a fast NAS method, called FastStereoNet, to enable resource-aware NAS within an intractably large search space. FastStereoNet automatically searches for hardware-friendly CNN architectures based on late acceptance hill climbing (LAHC), followed by simulated annealing (SA). FastStereoNet also employs a fine-tuning with a transferred weights mechanism to improve the convergence of the search process. The collection of these ideas provides competitive results in terms of search time and strikes a balance between accuracy and efficiency. Compared to the state of the art, FastStereoNet provides$5.25\times $reduction in search time and$44.4\times $reduction in model size. These benefits are attained while yielding a comparable accuracy that enables seamless deployment of disparity estimation on resource-limited devices. Finally, FastStereoNet significantly improves the perception quality of disparity estimation deployed on field-programmable gate array and Intel Neural Compute Stick 2 accelerator in a significantly less onerous manner.
Mohammad Loni, Ali Zoljodi, Amin Majd, Byung Hoon Ahn, Masoud Daneshtalab, Mikael Sjödin, Hadi Esmaeilzadeh
IEEE Trans. Syst. Man Cybern. Syst.1
2020 DenseDisp: Resource-Aware Disparity Map Estimation by Compressing Siamese Neural Architecture
abstract
Stereo vision cameras are flexible sensors due to providing heterogeneous information such as color, luminance, disparity map (depth), and shape of the objects. Today, Convolutional Neural Networks (CNNs) present the highest accuracy for the disparity map estimation [1]. However, CNNs require considerable computing capacity to process billions of floating-point operations in a real-time fashion. Besides, commercial stereo cameras produce huge size images (e.g., 10 Megapixels [2]), which impose a new computational cost to the system. The problem will be pronounced if we target resource-limited hardware for the implementation. In this paper, we propose DenseDisp, an automatic framework that designs a Siamese neural architecture for disparity map estimation in a reasonable time. DenseDisp leverages a meta-heuristic multi-objective exploration to discover hardware-friendly architectures by considering accuracy and network FLOPS as the optimization objectives. We explore the design space with four different fitness functions to improve the accuracy-FLOPS trade-off and convergency time of the DenseDisp. According to the experimental results, DenseDisp provides up to 39. 1x compression rate while losing around 5% accuracy compared to the state-of-the-art results.
Mohammad Loni, Ali Zoljodi, Daniel Maier 0002, Amin Majd, Masoud Daneshtalab, Mikael Sjödin, Ben H. H. Juurlink, Reza Akbari
CEC1
2019 TOT-Net: An Endeavor Toward Optimizing Ternary Neural Networks
abstract
High computation demands and big memory resources are the major implementation challenges of Convolutional Neural Networks (CNNs) especially for low-power and resource-limited embedded devices. Many binarized neural networks are recently proposed to address these issues. Although they have significantly decreased computation and memory footprint, they have suffered from accuracy loss especially for large datasets. In this paper, we propose TOT-Net, a ternarized neural network with [-1, 0, 1] values for both weights and activation functions that has simultaneously achieved a higher level of accuracy and less computational load. In fact, first, TOT-Net introduces a simple bitwise logic for convolution computations to reduce the cost of multiply operations. To improve the accuracy, selecting proper activation function and learning rate are influential, but also difficult. As the second contribution, we propose a novel piece-wise activation function, and optimized learning rate for different datasets. Our findings first reveal that 0.01 is a preferable learning rate for the studied datasets. Third, by using an evolutionary optimization approach, we found novel piece-wise activation functions customized for TOT-Net. According to the experimental results, TOT-Net achieves 2.15%, 8.77%, and 5.7/5.52% better accuracy compared to XNOR-Net on CIFAR-10, CIFAR-100, and ImageNet top-5/top-1 datasets, respectively.
Najmeh Nazari, Mohammad Loni, Mostafa E. Salehi, Masoud Daneshtalab, Mikael Sjödin
DSD2
2019 NeuroPower: Designing Energy Efficient Convolutional Neural Network Architecture for Embedded Systems
Mohammad Loni, Ali Zoljodi, Sima Sinaei, Masoud Daneshtalab, Mikael Sjödin
ICANN (1)1
2019 SoFA: A Spark-oriented Fog Architecture
abstract
Fog computing offers a wide range of service levels including low bandwidth usage, low response time, support of heterogeneous applications, and high energy efficiency. Therefore, real-time embedded applications could potentially benefit from Fog infrastructure. However, providing high system utilization is an important challenge of Fog computing especially for processing embedded applications. In addition, although Fog computing extends cloud computing by providing more energy efficiency, it still suffers from remarkable energy consumption, which is a limitation for embedded systems. To overcome the above limitations, in this paper, we propose SoFA, a Spark-oriented Fog architecture that leverages Spark functionalities to provide higher system utilization, energy efficiency and scalability. Compared to the common Fog computing platforms where edge devices are only responsible for processing data received from their IoT nodes, SoFA leverages the remaining processing capacity of all other edge devices. To attain this purpose, SoFA provides a distributed processing paradigm by the help of Spark to utilize the whole processing capacity of all the available edge devices leading to increase energy efficiency and system utilization. In other words, SoFA proposes a near-sensor processing solution in which the edge devices act as the Fog nodes. In addition, SoFA provides scalability by taking advantage of Spark functionalities. According to the experimental results, SoFA is a power-efficient and scalable solution desirable for embedded platforms by providing up to 3.1x energy efficiency for the Word-Count benchmark compared to the common Fog processing platform.
Neda Maleki, Mohammad Loni, Masoud Daneshtalab, Mauro Conti, Hossein Fotouhi
IECON2
2018 A Customized Processing-in-Memory Architecture for Biological Sequence Alignment
abstract
Sequence alignment is the most widely used operation in bioinformatics. With the exponential growth of the biological sequence databases, searching a database to find the optimal alignment for a query sequence (that can be at the order of hundreds of millions of characters long) would require excessive processing power and memory bandwidth. Sequence alignment algorithms can potentially benefit from the processing power of massive parallel processors due their simple arithmetic operations, coupled with the inherent fine-grained and coarse-grained parallelism that they exhibit. However, the limited memory bandwidth in conventional computing systems prevents exploiting the maximum achievable speedup. In this paper, we propose a processing-in-memory architecture as a viable solution for the excessive memory bandwidth demand of bioinformatics applications. The design is composed of a set of simple and lightweight processing elements, customized to the sequence alignment algorithm, integrated at the logic layer of an emerging 3D DRAM architecture. Experimental results show that the proposed architecture results in up to 2.4x speedup and 41% reduction in power consumption, compared to a processor-side parallel implementation.
Nasrin Akbari, Mehdi Modarressi, Masoud Daneshtalab, Mohammad Loni
ASAP4
2018 ADONN: Adaptive Design of Optimized Deep Neural Networks for Embedded Systems
abstract
Nowadays, many modern applications, e.g. autonomous system, and cloud data services need to capture and process a big amount of raw data at runtime that ultimately necessitates a high-performance computing model. Deep Neural Network (DNN) has already revealed its learning capabilities in runtime data processing for modern applications. However, DNNs are becoming more deep sophisticated models for gaining higher accuracy which require a remarkable computing capacity. Considering high-performance cloud infrastructure as a supplier of required computational throughput is often not feasible. Instead, we intend to find a near-sensor processing solution which will lower the need for network bandwidth and increase privacy and power efficiency, as well as guaranteeing worst-case response-times. Toward this goal, we introduce ADONN framework, which aims to automatically design a highly robust DNN architecture for embedded devices as the closest processing unit to the sensors. ADONN adroitly searches the design space to find improved neural architectures. Our proposed framework takes advantage of a multi-objective evolutionary approach, which exploits a pruned design space inspired by a dense architecture. Unlike recent works that mainly have tried to generate highly accurate networks, ADONN also considers the network size factor as the second objective to build a highly optimized network fitting with limited computational resource budgets while delivers comparable accuracy level. In comparison with the best result on CIFAR-10 dataset, a generated network by ADONN presents up to 26.4 compression rate while loses only 4% accuracy. In addition, ADONN maps the generated DNN on the commodity programmable devices including ARM Processor, High-Performance CPU, GPU, and FPGA.
Mohammad Loni, Masoud Daneshtalab, Mikael Sjödin
DSD1