VLDB 2026 Research / reviewers in the wild / expert
Alireza Tavakkoli
dblp:75/1965
· DBLP profile ↗
24ranked-venue papers
2as first author
14since 2021 · last 2025
0000-0001-9460-1269ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Point-RTD: Replaced Token Denoising for Pretraining Transformer Models on Point CloudsabstractPre-training strategies play a critical role in advancing the performance of transformer-based models for 3D point cloud tasks. In this paper, we introduce Point-RTD (Replaced Token Denoising), a novel pretraining strategy designed to improve token robustness through a corruption-reconstruction framework. Unlike traditional mask-based reconstruction tasks that hide data segments for later prediction, Point-RTD corrupts point cloud tokens and leverages a discriminator-generator architecture for denoising. This shift enables more effective learning of structural priors and significantly enhances model performance and efficiency. On the ShapeNet dataset, Point-RTD reduces reconstruction error by over 93% compared to PointMAE, and achieves more than 14x lower Chamfer Distance on the test set. Our method also converges faster and yields higher classification accuracy on ShapeNet, ModelNet10, and ModelNet40 benchmarks, clearly outperforming the baseline Point-MAE framework in every case. Code is available at https://github.com/GunnerStone/PointRTD. Gunner Stone, Youngsook Choi, Alireza Tavakkoli, Ankita Shukla |
ICMLA | 3 |
| 2025 | Guided and Unguided Conditional Diffusion Mechanisms for Structured and Semantically-Aware 3D Point Cloud GenerationabstractGenerating realistic 3D point clouds is a fundamental problem in computer vision with applications in remote sensing, robotics, and digital object modeling. Existing generative approaches primarily capture geometry, and when semantics are considered, they are typically imposed post hoc through external segmentation or clustering rather than integrated into the generative process itself. We propose a diffusion-based framework that embeds per-point semantic conditioning directly within generation. Each point is associated with a conditional variable corresponding to its semantic label, which guides the diffusion dynamics and enables the joint synthesis of geometry and semantics. This design produces point clouds that are both structurally coherent and segmentation-aware, with object parts explicitly represented during synthesis. Through a comparative analysis of guided and unguided diffusion processes, we demonstrate the significant impact of conditional variables on diffusion dynamics and generation quality. Extensive experiments validate the efficacy of our approach, producing detailed and accurate 3D point clouds tailored to specific parts and features. Gunner Stone, Sushmita Sarker, Alireza Tavakkoli |
ICMLA | 3 |
| 2024 | RaceGAN: A Framework for Preserving Individuality While Converting Racial Information for Image-to-Image TranslationabstractGenerative adversarial networks (GANs) have demonstrated significant progress in unpaired image-to-image translation in recent years for several applications. CycleGAN was the first to lead the way, although it was restricted to a pair of domains. StarGAN overcame this constraint by tackling image-to-image translation across various domains, although it was not able to map in-depth low-level style changes for these domains. Style mapping via reference-guided image synthesis has been made possible by the innovations of StarGANv2 and StyleGAN. However, these models do not maintain individuality and need an extra reference image in addition to the input. Our study aims to translate racial traits by means of multi-domain image-to-image translation. We present RaceGAN, a novel framework capable of mapping style codes over several domains during racial attribute translation while maintaining individuality and high level semantics without relying on a reference image. RaceGAN outperforms other models in translating racial features (i.e., Asian, White, and Black) when tested on Chicago Face Dataset. We also give quantitative findings utilizing InceptionReNetv2-based classification to demonstrate the effectiveness of our racial translation. Moreover, we investigate how well the model partitions the latent space into distinct clusters of faces for each ethnic group. Tasnim Pervin, George Bebis, Alireza Tavakkoli |
ICMLA | 4 |
| 2024 | SpeciServe. a gRPC Infrastructure ConceptabstractSmart city projects require data to be transferred from one destination to the next using a number of different network protocols. The data pipelines involved in these smart city projects often have limited bandwidth or compute resources due to the low power nature of most embedded hardware. The data transferred between devices in these types of embedded systems are often structured in non-standard data schemata. Remote procedure calls (RPC) are implemented to transfer data between devices and switching between RPC implementations can be tricky due to the lack of standardization. There is no guarantee that an existing data schema will work with a different RPC implementation. This makes it difficult for a researcher or system developer to benchmark and compare different RPC im-plementations. In this paper, a conceptual infrastructure named SpeciServe is introduced where gRPC is used as a communication backbone due its support for flatbuffers and multiple server modes. Multiple software services are described to allow for dissimilar RPC implementations to be run in parallel. This system is intended to allow for researchers in machine learning, smart cities, and Internet of Things (loT) to be able test different versions of RPCs and provide support for system developers to define the functions of an edge service. Chase D. Carthen, Araam Zaremehrjardi, Zachary Estreito, Alireza Tavakkoli, Frederick C. Harris Jr., Sergiu M. Dascalu |
SERA | 4 |
| 2024 | A Spatial Data Pipeline for Streaming Smart City DataabstractPoint cloud data in the form of LiDAR is often utilized for its spatial qualities, especially in smart city projects for tasks involving vehicles and pedestrians. However, the process in which LiDAR data is acquired can be cumbersome to setup and automate. In this paper, we introduce a streaming and an on-demand pipeline for capturing LiDAR data from Velodyne Ultra Pucks placed along northern Nevada intersections known as the Living Lab as part of a smart city project for the city of Reno. The data coming from these intersections consist of the following formats: ROS 2 bag file, PCD, LAZ, Google Draco, and PCAP. A streaming point cloud service with PCD, LAZ, and Draco was implemented to stream any of these formats, as well as to allow the user to capture the current monitored point cloud. Additionally, two on-demand web services were implemented for both the PCAP and ROS 2 bag file to enable a user to start and stop the acquisition of LiDAR data in these formats. Through our analysis, it was discovered that Draco provided the best processing time and had a wider range of options that affected the quality of the point cloud. To evaluate this pipeline, the features of existing software were compared and a discussion was provided with an analysis of the point cloud formats. Chase D. Carthen, Araam Zaremehrjardi, Vinh D. Le, Carlos Cardillo, Scotty Strachan, Alireza Tavakkoli, Sergiu M. Dascalu, Frederick C. Harris Jr. |
SERA | 6 |
| 2024 | Exploring the influence of attention for whole-image mammogram classification
Marc Berghouse, George Bebis, Alireza Tavakkoli |
Image Vis. Comput. | 3 |
| 2024 | A comprehensive overview of deep learning techniques for 3D point cloud classification and semantic segmentation
Sushmita Sarker, Prithul Sarker, Gunner Stone, Ryan Gorman, Alireza Tavakkoli, George Bebis, Javad Sattarvand |
Mach. Vis. Appl. | 5 |
| 2023 | Using Action Cameras to Collect Tree MeasurementsabstractThis study investigated the use of a consumer-grade action camera to generate a 3D recreation of a newly established forest plot using structure from motion (SfM) techniques. We tested a novel approach that used 2 videos collected with a GoPro HERO 9 Black during plot setup. After the plot center and witness trees were identified, we recorded the approach and a short walkthrough of the plot that included the witness trees and the plot center stake. Image frames along with position and orientation metadata were extracted from each video. These images were processed in the Pix4DMapper photogrammetry software and exported as 3D point clouds. These point clouds were processed using PDAL to remove noise, identify ground points, and calculate height above ground. Tree position, DBH, and height were extracted from the point cloud and compared to field measurements to assess the utility and practicality of the terrestrial SfM point cloud. Theodore Hartsook, Sarah Bisbing, Alireza Tavakkoli, Jonathan A. Greenberg |
IGARSS | 3 |
| 2023 | Revolutionizing Space Health (Swin-FSR): Advancing Super-Resolution of Fundus Images for SANS Visual Assessment Technology
Khondker Fariha Hossain, Sharif Amit Kamran, Joshua Ong, Andrew G. Lee, Alireza Tavakkoli |
MICCAI (7) | 5 |
| 2023 | Orchestrating Apache NiFi/MiNiFi within a Spatial Data PipelineabstractIn many smart city projects, a common choice to capture spatial information is the inclusion of LiDAR data, but this decision will often invoke severe growing pains within the existing infrastructure. In this paper, we introduce a data pipeline that orchestrates Apache NiFi (NiFi), Apache MiNiFi (MiNiFi), and several other tools as an automated solution in order to relay and archive LiDAR data captured by deployed edge devices. The LiDAR sensors utilized within this workflow are Velodyne Ultra Pucks sensors that capture at a rate of 10 frames per second and produces 6-7 GB packet capture (PCAP) files per hour. By both compressing the file after capturing it and compressing the file in real-time, we discovered that gzip produced a file of 5 GB and saved about 5 minutes in transmission time to NiFi, as well as saving considerable CPU time when compressing the file in real-time. Alternatively, we chose XZ as the compression algorithm for the ingestion of LiDAR data onto an institution compute cluster due to its high compression ratio. In order to evaluate the capabilities of our system design, the features of this data pipeline were compared against existing third-party services, namely Globus and RSync. Chase D. Carthen, Araam Zaremehrjardi, Vinh D. Le, Carlos Cardillo, Scotty Strachan, Alireza Tavakkoli, Frederick C. Harris Jr., Sergiu M. Dascalu |
SERA | 6 |
| 2022 | A comprehensive survey of the approaches for pathway analysis using multi-omics data integrationabstractPathway analysis has been widely used to detect pathways and functions associated with complex disease phenotypes. The proliferation of this approach is due to better interpretability of its results and its higher statistical power compared with the gene-level statistics. A plethora of pathway analysis methods that utilize multi-omics setup, rather than just transcriptomics or proteomics, have recently been developed to discover novel pathways and biomarkers. Since multi-omics gives multiple views into the same problem, different approaches are employed in aggregating these views into a comprehensive biological context. As a result, a variety of novel hypotheses regarding disease ideation and treatment targets can be formulated. In this article, we review 32 such pathway analysis methods developed for multi-omics and multi-cohort data. We discuss their availability and implementation, assumptions, supported omics types and databases, pathway analysis techniques and integration strategies. A comprehensive assessment of each method's practicality, and a thorough discussion of the strengths and drawbacks of each technique will be provided. The main objective of this survey is to provide a thorough examination of existing methods to assist potential users and researchers in selecting suitable tools for their data and analysis purposes, while highlighting outstanding challenges in the field that remain to be addressed for future development. Zeynab Maghsoudi, Alireza Tavakkoli, Tin Chi Nguyen |
Briefings Bioinform. | 3 |
| 2021 | Classification of RIGID and Non-Rigid Transformations with Autoencoder Representations
Alexis R. Tudor, Gunner Stone, Alireza Tavakkoli, Emily Morgan Hand |
ICIP | 3 |
| 2021 | ECG-Adv-GAN: Detecting ECG Adversarial Examples with Conditional Generative Adversarial NetworksabstractElectrocardiogram (ECG) acquisition requires an automated system and analysis pipeline for understanding specific rhythm irregularities. Deep neural networks have become a popular technique for tracing ECG signals, outperforming human experts. Despite this, convolutional neural networks are susceptible to adversarial examples that can misclassify ECG signals and decrease the model’s precision. Moreover, they do not generalize well on the out-of-distribution dataset. The GAN architecture has been employed in recent works to synthesize adversarial ECG signals to increase existing training data. However, they use a disjointed CNN-based classification architecture to detect arrhythmia. Till now, no versatile architecture has been proposed that can detect adversarial examples and classify arrhythmia simultaneously. To alleviate this, we propose a novel Conditional Generative Adversarial Network to simultaneously generate ECG signals for different categories and detect cardiac abnormalities. Moreover, the model is conditioned on class-specific ECG signals to synthesize realistic adversarial examples. Consequently, we compare our architecture and show how it outperforms other classification models in normal/abnormal ECG signal detection by benchmarking real world and adversarial signals. Khondker Fariha Hossain, Sharif Amit Kamran, Alireza Tavakkoli, Lei Pan 0002, Xingjun Ma, Sutharshan Rajasegarar, Chandan Karmaker |
ICMLA | 3 |
| 2021 | RV-GAN: Segmenting Retinal Vascular Structure in Fundus Photographs Using a Novel Multi-scale Generative Adversarial Network
Sharif Amit Kamran, Khondker Fariha Hossain, Alireza Tavakkoli, Stewart Zuckerbrod, Kenton M. Sanders, Salah A. Baker |
MICCAI (8) | 3 |
| 2020 | Improving Robustness Using Joint Attention Network for Detecting Retinal Degeneration From Optical Coherence Tomography ImagesabstractNoisy data and the similarity in the ocular appearances caused by different ophthalmic pathologies pose significant challenges for an automated expert system to accurately detect retinal diseases. In addition, the lack of knowledge transferability and the need for unreasonably large datasets limit clinical application of current machine learning systems. To increase robustness, a better understanding of how the retinal subspace deformations lead to various levels of disease severity needs to be utilized for prioritizing disease-specific model details. In this paper we propose the use of disease-specific feature representation as a novel architecture comprised of two joint networks -- one for supervised encoding of disease model and the other for producing attention maps in an unsupervised manner to retain disease specific spatial information. Our experimental results on publicly available datasets show the proposed joint-network significantly improves the accuracy and robustness of state-of-the-art retinal disease classification networks on unseen datasets. Sharif Amit Kamran, Alireza Tavakkoli, Stewart Zuckerbrod |
ICIP | 2 |
| 2020 | Attention2AngioGAN: Synthesizing Fluorescein Angiography from Retinal Fundus Images using Generative Adversarial NetworksabstractFluorescein Angiography (FA) is a technique that employs the designated camera for Fundus photography incorporating excitation and barrier filters. FA also requires fluorescein dye that is injected intravenously, which might cause adverse effects ranging from nausea, vomiting to even fatal anaphylaxis. Currently, no other fast and non-invasive technique exists that can generate FA without coupling with Fundus photography. To eradicate the need for an invasive FA extraction procedure, we introduce an Attention-based Generative network that can synthesize Fluorescein Angiography from Fundus images. The proposed gan incorporates multiple attention based skip connections in generators and comprises novel residual blocks for both generators and discriminators. It utilizes reconstruction, feature-matching, and perceptual loss along with adversarial training to produces realistic Angiograms that is hard for experts to distinguish from real ones. Our experiments confirm that the proposed architecture surpasses recent state-of-the-art generative networks for fundus-to-angio translation task. Sharif Amit Kamran, Khondker Fariha Hossain, Alireza Tavakkoli, Stewart Zuckerbrod |
ICPR | 3 |
| 2020 | Control Framework for a Hybrid-steel Bridge Inspection RobotabstractAutonomous navigation of steel bridge inspection robots are essential for proper maintenance. Majority of existing robotic solutions for bridge inspection require human intervention to assist in the control and navigation. In this paper, a control system framework has been proposed for a previously designed ARA robot [1], which facilitates autonomous real-time navigation and minimizes human involvement. The mechanical design and control framework of ARA robot enables two different configurations, namely the mobile and inch-worm transformation. In addition, a switching control was developed with 3D point clouds of steel surfaces as the input which allow the robot to switch between mobile and inch-worm transformation. The surface availability algorithm (considers plane, area and height) of the switching control enables the robot to perform inch-worm jumps autonomously. The mobile transformation allows the robot to move on continuous steel surfaces and perform visual inspection of steel bridge structures. Practical experiments on actual steel bridge structures highlight the effective performance of ARA robot with the proposed control framework for autonomous navigation during visual inspection of steel bridges. Hoang-Dung Bui, Umme Hafsa Billah, Chuong Le, Alireza Tavakkoli, Hung Manh La |
IROS | 5 |
| 2019 | Optic-Net: A Novel Convolutional Neural Network for Diagnosis of Retinal Diseases from Optical Tomography ImagesabstractDiagnosing different retinal diseases from Spectral Domain Optical Coherence Tomography (SD-OCT) images is a challenging task. Different automated approaches such as image processing, machine learning and deep learning algorithms have been used for early detection and diagnosis of retinal diseases. Unfortunately, these are prone to error and computational inefficiency, which requires further intervention from human experts. In this paper, we propose a novel convolution neural network architecture to successfully distinguish between different degeneration of retinal layers and their underlying causes. The proposed novel architecture outperforms other classification models while addressing the issue of gradient explosion. Our approach reaches near perfect accuracy of 99.8% and 100% for two separately available Retinal SD-OCT data-set respectively. Additionally, our architecture predicts retinal diseases in real time while outperforming human diagnosticians. Sharif Amit Kamran, Sourajit Saha, Ali Sabbir, Alireza Tavakkoli |
ICMLA | 4 |
| 2016 | Hand motion calibration and retargeting for intuitive object manipulation in immersive virtual environmentsabstractIn this paper a system is proposed to combine small finger movements with the large scale body movements captured from a motion capture system. The strength of the proposed work over previous research is in the real-time and natural interactions that the virtual hands have with their environment. By being able to conform to physics, the virtual hands feel like virtual extensions of one's own hands. This provides a higher degree of immersion and interactivity when compared to more traditional virtual reality systems. Brandon Wilson, Matthew Bounds, Alireza Tavakkoli |
VR | 3 |
| 2010 | Integrating Context into Intent Recognition Systems
Richard Kelley, Christopher King, Amol Ambardekar, Monica N. Nicolescu, Mircea Nicolescu, Alireza Tavakkoli |
ICINCO (2) | 6 |
| 2009 | Non-parametric statistical background modeling for efficient foreground region detection
Alireza Tavakkoli, Mircea Nicolescu, George Bebis, Monica N. Nicolescu |
Mach. Vis. Appl. | 1 |
| 2008 | Understanding human intentions via hidden markov models in autonomous mobile robotsabstractUnderstanding intent is an important aspect of communication among people and is an essential component of the human cognitive system. This capability is particularly relevant for situations that involve collaboration among agents or detection of situations that can pose a threat. In this paper, we propose an approach that allows a robot to detect intentions of others based on experience acquired through its own sensory-motor capabilities, then using this experience while taking the perspective of the agent whose intent should be recognized. Our method uses a novel formulation of Hidden Markov Models designed to model a robot's experience and interaction with the world. The robot's capability to observe and analyze the current scene employs a novel vision-based technique for target detection and tracking, using a non-parametric recursive modeling approach. We validate this architecture with a physically embedded robot, detecting the intent of several people performing various activities. Richard Kelley, Alireza Tavakkoli, Christopher King, Monica N. Nicolescu, Mircea Nicolescu, George Bebis |
HRI | 2 |
| 2008 | Feature Fusion Hierarchies for gender classificationabstractWe present a hierarchical feature fusion model for image classification that is constructed by an evolutionary learning algorithm. The model has the ability to combine local patches whose location, width and height are automatically determined during learning. The representational framework takes the form of a two-level hierarchy which combines feature fusion and decision fusion into a unified model. The structure of the hierarchy itself is constructed automatically during learning to produce optimal local feature combinations. A comparative evaluation of different classifiers is provided on a challenging gender classification image database. It demonstrates the effectiveness of these Feature Fusion Hierarchies (FFH). Fabien Scalzo, George Bebis, Mircea Nicolescu, Leandro A. Loss, Alireza Tavakkoli |
ICPR | 5 |
| 2008 | Efficient background modeling through incremental Support Vector Data DescriptionabstractBackground modeling is an essential and important part of many high-level video processing applications. Recently, the Support Vector Data Description (SVDD) has been introduced for novelty detection when only one class of data is available, i.e. background pixels. This paper proposes a method to efficiently train an SVDD and compares the performance of this training algorithm with the traditional SVDD training techniques. We compare the performance of our method with traditional SVDD and other classification algorithms on various data sets including real video sequences. Alireza Tavakkoli, Mircea Nicolescu, George Bebis, Monica N. Nicolescu |
ICPR | 1 |