Arian Bakhtiarnia

dblp:290/7257 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0001-8624-8661ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Deep learning architectures and training · 50% Probabilistic and Bayesian machine learning · 50%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
online inference
0.712023
Continual Transformers: Redundancy-Free Attention for Online Inference · ICLR 2023
Machine learning › Deep learning architectures and training › attention mechanism
transformer attention
0.712023
Continual Transformers: Redundancy-Free Attention for Online Inference · ICLR 2023

Methods — techniques the papers use, named apart from their topics

transformer · 0.7attention · 0.7
YearPublicationVenuePosition
2026 Continual low-rank scaled dot-product attention
abstract
• A new Continual Inference Transformer that explores low-rank attention is proposed. • We propose two new ways to compute data-driven landmarks for the low-rank attention. • The proposed model leads to up to three orders of magnitude computation reduction. Transformers are widely used for their ability to capture data relations in sequence processing, with great success for a wide range of static tasks. However, the computational and memory footprint of their main component, i.e., the Scaled Dot-product Attention, is commonly overlooked. This makes their adoption infeasible in applications involving stream data processing with constraints in response latency, computational and memory resources. Some works have proposed methods to lower the computational cost of Transformers by using low-rank approximations, sparsity in attention, and efficient formulations for Continual Inference. In this paper, we introduce a new formulation of the Scaled Dot-product Attention based on the Nyström approximation that is suitable for Continual Inference. In experiments on Online Audio Classification and Online Action Detection tasks, the proposed Continual Scaled Dot-product Attention can lower the number of operations by up to three orders of magnitude compared to the original Transformers while retaining the predictive performance of competing models.
Ginés Carreto Picón, Illia Oleksiienko, Lukas Hedegaard, Arian Bakhtiarnia, Alexandros Iosifidis
Neural Networks4
2024 Accurate Gigapixel Crowd Counting by Iterative Zooming and Refinement
abstract
The increasing prevalence of gigapixel resolutions has presented new challenges for crowd counting. Such resolutions are far beyond the memory and computation limits of current GPUs, and available deep neural network architectures and training procedures are not designed for such massive inputs. Although several methods have been proposed to address these challenges, they are either limited to downsampling the input image to a small size, or borrowing from other gigapixel tasks, which are not tailored for crowd counting. In this paper, we propose a novel method called GigaZoom, which iteratively zooms into the densest areas of the image and refines coarser density maps with finer details. We show that GigaZoom obtains the state-of-the-art for gigapixel crowd counting and improves the accuracy of the next best method by 42%.
Arian Bakhtiarnia, Qi Zhang 0013, Alexandros Iosifidis
ICASSP1
2024 MAVAD: Audio-Visual Dataset and Method for Anomaly Detection in Traffic Videos
abstract
This paper introduces the first audio-visual dataset for traffic anomaly detection called MAVAD, taken from real-world scenes, with a diverse range of illumination conditions. In addition, a novel anomaly detection method is proposed which combines visual and audio features extracted from video sequences by means of cross-attention. We demonstrate that the addition of audio improves anomaly detection performance by up to 5.2%. Moreover, the impact of image anonymization is evaluated, showing only a minor decrease in performance averaging at 1.7%.
Blazej Leporowski, Arian Bakhtiarnia, Nicole Bonnici, Adrian Muscat, Luca Zanella, Yiming Wang 0002, Alexandros Iosifidis
ICIP2
2023 Dynamic Split Computing for Efficient Deep EDGE Intelligence
abstract
Deploying deep neural networks (DNNs) on IoT and mobile devices is a challenging task due to their limited computational resources. Thus, demanding tasks are often entirely offloaded to edge servers which can accelerate inference, however, it also causes communication cost and evokes privacy concerns. In addition, this approach leaves the computational capacity of end devices unused. Split computing is a paradigm where a DNN is split into two sections; the first section is executed on the end device, and the output is transmitted to the edge server where the final section is executed. Here, we introduce dynamic split computing, where the optimal split location is dynamically selected based on the state of the communication channel. By using natural bottlenecks that already exist in modern DNN architectures, dynamic split computing avoids retraining and hyperparameter optimization, and does not have any negative impact on the final accuracy of DNNs. Through extensive experiments, we show that dynamic split computing achieves faster inference in edge computing environments where the data rate and server load vary over time.
Arian Bakhtiarnia, Nemanja Milosevic, Qi Zhang 0013, Dragana Bajovic, Alexandros Iosifidis
ICASSP1
2023 Continual Transformers: Redundancy-Free Attention for Online Inference
Lukas Hedegaard, Arian Bakhtiarnia, Alexandros Iosifidis
ICLR2
2023 PromptMix: Text-to-image diffusion models enhance the performance of lightweight networks
abstract
Many deep learning tasks require annotations that are too time consuming for human operators, resulting in small dataset sizes. This is especially true for dense regression problems such as crowd counting which requires the location of every person in the image to be annotated. Techniques such as data augmentation and synthetic data generation based on simulations can help in such cases. In this paper, we introduce PromptMix, a method for artificially boosting the size of existing datasets, that can be used to improve the performance of lightweight networks. First, synthetic images are generated in an end-to-end data-driven manner, where text prompts are extracted from existing datasets via an image captioning deep network, and subsequently introduced to text-to-image diffusion models. The generated images are then annotated using one or more high-performing deep networks, and mixed with the real dataset for training the lightweight network. By extensive experiments on five datasets and two tasks, we show that PromptMix can significantly increase the performance of lightweight networks by up to 26%.
Arian Bakhtiarnia, Qi Zhang 0013, Alexandros Iosifidis
IJCNN1
2022 Single-layer vision transformers for more accurate early exits with less overhead
abstract
Deploying deep learning models in time-critical applications with limited computational resources, for instance in edge computing systems and IoT networks, is a challenging task that often relies on dynamic inference methods such as early exiting. In this paper, we introduce a novel architecture for early exiting based on the vision transformer architecture, as well as a fine-tuning strategy that significantly increase the accuracy of early exit branches compared to conventional approaches while introducing less overhead. Through extensive experiments on image and audio classification as well as audiovisual crowd counting, we show that our method works for both classification and regression problems, and in both single- and multi-modal settings. Additionally, we introduce a novel method for integrating audio and visual modalities within early exits in audiovisual data analysis, that can lead to a more fine-grained dynamic inference.
Arian Bakhtiarnia, Qi Zhang 0013, Alexandros Iosifidis
Neural Networks1
2021 Multi-Exit Vision Transformer for Dynamic Inference
Arian Bakhtiarnia, Qi Zhang 0013, Alexandros Iosifidis
BMVC1
2021 Improving the Accuracy of Early Exits in Multi-Exit Architectures via Curriculum Learning
abstract
Deploying deep learning services for time-sensitive and resource-constrained settings such as IoT using edge computing systems is a challenging task that requires dynamic adjustment of inference time. Multi-exit architectures allow deep neural networks to terminate their execution early in order to adhere to tight deadlines at the cost of accuracy. To mitigate this cost, in this paper we introduce a novel method called Multi-Exit Curriculum Learning that utilizes curriculum learning, a training strategy for neural networks that imitates human learning by sorting the training samples based on their difficulty and gradually introducing them to the network. Experiments on CIFAR-10 and CIFAR-100 datasets and various configurations of multi-exit architectures show that our method consistently improves the accuracy of early exits compared to the standard training approach.
Arian Bakhtiarnia, Qi Zhang 0013, Alexandros Iosifidis
IJCNN1