VLDB 2026 Research / reviewers in the wild / expert
Senem Velipasalar
dblp:85/6636 · also Senem Velipasalar Gursoy
· DBLP profile ↗
123ranked-venue papers
8as first author
35since 2021 · last 2025
0000-0002-1430-1555ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 7 first-author · 17 since 2021Computer networks · 41 · 6 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 6 since 2021Human-computer interaction and ubiquitous computing · 2Theory of computation · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 3D-PointZshotS: Geometry-Aware 3D Point Cloud Zero-Shot Semantic Segmentation Narrowing the Visual-Semantic GapabstractExisting zero-shot 3D point cloud segmentation methods often struggle with limited transferability from seen classes to unseen classes and from semantic to visual space. To alleviate this, we introduce 3D-PointZshotS, a geometry-aware zero-shot segmentation framework that enhances both feature generation and alignment using latent geometric prototypes (LGPs). Specifically, we integrate LGPs into a generator via a cross-attention mechanism, enriching semantic features with fine-grained geometric details. To further enhance stability and generalization, we introduce a self-consistency loss, which enforces feature robustness against point-wise perturbations. Additionally, we re-represent visual and semantic features in a shared space, bridging the semantic-visual gap and facilitating knowledge transfer to unseen classes. Experiments on three real-world datasets, namely ScanNet, SemanticKITTI, and S3DIS, demonstrate that our method achieves superior performance over four baselines in terms of harmonic mIoU. Code is available at github. Minmin Yang, Huantao Ren, Senem Velipasalar |
AVSS | 3 |
| 2025 | Asynchronous Decentralized Federated Learning Deconstructs Excessively Large Batches
Mustafa Cenk Gursoy, Senem Velipasalar |
GLOBECOM | 4 |
| 2025 | LaTP: LiDAR-aided multimodal token pruning for efficient trajectory prediction of autonomous driving
Yantao Lu, Ning Liu 0007, Yilan Li, Jinchao Chen, Ying Zhang 0060, Yichen Zhu 0001, Senem Velipasalar |
Neural Networks | 9 |
| 2025 | Cross-task and time-aware adversarial attack framework for perception of autonomous driving
Yantao Lu, Ning Liu 0007, Yilan Li, Jinchao Chen, Senem Velipasalar |
Pattern Recognit. | 5 |
| 2024 | VLP: Vision Language Planning for Autonomous DrivingabstractAutonomous driving is a complex and challenging task that aims at safe motion planning through scene understanding and reasoning. While vision-only autonomous driving methods have recently achieved notable performance, through enhanced scene understanding, several key issues, including lack of reasoning, low generalization performance and long-tail scenarios, still need to be addressed. In this paper, we present VLP, a novel Vision-Language-Planningframework that exploits language models to bridge the gap between linguistic understanding and autonomous driving. VLP enhances autonomous driving systems by strengthening both the source memory foundation and the self-driving car's contextual understanding. VLP achieves state-of-the-art end-to-end planning performance on the challenging NuScenes dataset by achieving 35.9% and 60.5% reduction in terms of average L2 error and collision rates, respectively, compared to the previous best method. Moreover, VLP shows improved performance in challenging long-tail scenarios and strong generalization capabilities when faced with new urban environments. Chenbin Pan, Burhaneddin Yaman, Tommaso Nesti, Abhirup Mallik, Alessandro Allievi, Senem Velipasalar, Liu Ren 0001 |
CVPR | 6 |
| 2024 | CLIP-BEVFormer: Enhancing Multi-View Image-Based BEV Detector with Ground Truth FlowabstractAutonomous driving stands as a pivotal domain in computer vision, shaping the future of transportation. Within this paradigm, the backbone of the system plays a crucial role in interpreting the complex environment. However, a notable challenge has been the loss of clear supervision when it comes to Bird's Eye View elements. To address this limitation, we introduce CLIP-BEVFormer, a novel approach that leverages the power of contrastive learning techniques to enhance the multi-view image-derived BEV backbones with ground truth information flow. We conduct extensive experiments on the challenging nuScenes dataset and showcase significant and consistent improvements over the SOTA. Specifically, CLIP-BEVFormer achieves an impressive 8.5% and 9.2% enhancement in terms of NDS and mAP, respectively, over the previous best BEV model on the 3D object detection task. Chenbin Pan, Burhaneddin Yaman, Senem Velipasalar, Liu Ren 0001 |
CVPR | 3 |
| 2024 | Letting 3D Guide the Way: 3D Guided 2D Few-Shot Image ClassificationabstractExisting few-shot image classification networks aim to perform prediction on images belonging to classes that were not seen during training, with only a few labeled images, which are randomly picked from the same image pool as the support set. However, this traditional approach has two main issues: (i) in real-world applications, since support images are randomly picked, the angle they were captured from can be very different from that of the query image, causing the images to look very different and making it hard to match them; (ii) since support and query images, for both training and testing, are sampled from the same image pool, models can overfit the dataset, especially if the image pool contains images with similar color, texture or view angle. Thus, good performance on a dataset does not reflect a model’s real ability. To address these issues, we propose a novel few-shot learning approach referred to as the 3D guided 2D (3DG2D) few-shot image classification. In our proposed approach, the queries are 2D images, and the support set is composed of 3D mesh data, providing different views of an object, in contrast to randomly picked images providing a single view. From each 3D mesh, 14 projection images are generated from different angles. Thus, these projections have significant variance among themselves. To address this challenge, we also propose the Angle Inference Module (AIM), which is used to infer the view angle of a query image so that more attention is given to projection images corresponding to the same view angle as the query image to achieve better prediction performance. We perform experiments on ModelNet40, Toys4K and ShapeNet datasets with 4-fold cross validation, and show that our 3DG2D few-shot classification approach consistently outperforms the state-of-the-art baselines. Jiajing Chen, Minmin Yang, Senem Velipasalar |
WACV | 3 |
| 2024 | Maximum Knowledge Orthogonality Reconstruction with Gradients in Federated LearningabstractFederated learning (FL) aims at keeping client data local to preserve privacy. Instead of gathering the data itself, the server only collects aggregated gradient updates from clients. Following the popularity of FL, there has been considerable amount of work revealing the vulnerability of FL approaches by reconstructing the input data from gradient updates. Yet, most existing works assume an FL setting with unrealistically small batch size, and have poor image quality when the batch size is large. Other works modify the neural network architectures or parameters to the point of being suspicious, and thus, can be detected by clients. Moreover, most of them can only reconstruct one sample input from a large batch. To address these limitations, we propose a novel and analytical approach, referred to as the maximum knowledge orthogonality reconstruction (MKOR), to reconstruct clients' data. Our proposed method reconstructs a mathematically proven high-quality image from large batches. MKOR only requires the server to send secretly modified parameters to clients and can efficiently and inconspicuously reconstruct images from clients' gradient updates. We evaluate MKOR’s performance on MNIST, CIFAR-100, and ImageNet datasets and compare it with the state-of-the-art baselines. The results show that MKOR outperforms the existing approaches, and draw attention to a pressing need for further research on the privacy protection of FL so that comprehensive defense approaches can be developed. The code is available at: https://github.com/wfwf10/MKOR. Senem Velipasalar, Mustafa Cenk Gursoy |
WACV | 2 |
| 2024 | SimpliMix: A Simplified Manifold Mixup for Few-shot Point Cloud ClassificationabstractFew-shot learning often assumes that base classes are abundant and diverse with plentiful well-labeled samples for each class. This ensures that models can generalize effectively from a small amount of data by leveraging prior knowledge learned from base classes. This assumption holds for 2D few-shot learning since the benchmark datasets are large and diverse. However, 3D point cloud few-shot benchmarks are low in magnitude and diversity. We conduct experiments and show that many existing methods overlook this issue and suffer from overfitting on base classes, which hinders generalization ability and test performance. To alleviate the overfitting issue, we propose a simplified manifold mixup, referred to as the SimpliMix, which mixes hidden representations and forces the models to learn more generalized features. We incorporate SimpliMix into existing prototype-based models, perform experiments on ModelNet40-FS, ModelNet40-C-FS and ScanObjectNN-FS datasets, and improve the models by a significant margin. We further conduct cross-domain few-shot classification experiments and show that networks with SimpliMix learn more generalized and transferable features and achieve better performance. The code is available at https://github.com/LexieYang/SimpliMix Minmin Yang, Weiheng Chai, Jiyang Wang, Senem Velipasalar |
WACV | 4 |
| 2024 | Time-aware and task-transferable adversarial attack for perception of autonomous vehicles
Yantao Lu, Haining Ren, Weiheng Chai, Senem Velipasalar, Yilan Li |
Pattern Recognit. Lett. | 4 |
| 2024 | Rethinking the Evaluation of Driver Behavior Analysis ApproachesabstractCrashes caused by distracted driving result in more than 3000 deaths every year in the U.S. Distracted driver behavior detection is instrumental for driver assist systems. Researchers have focused on autonomously detecting distracted driver behavior so that drivers can be alerted in time to reduce the risk of crashes. Despite the large number of approaches presented in the literature, there are still issues related to proper performance evaluation, reproducibility and lack of or very slow adoption of these approaches by the transportation industry. Most existing approaches either do not provide documented and usable codes or use private datasets, or do not present the experiment details, such as data split, sometimes resulting in inflated accuracy numbers. Moreover, these factors also make many results not reproducible. In addition, the performance metrics should be chosen carefully to measure various aspects of different methods, including their generalizability, and action localization ability in time. In this work, we perform a commensurate comparison of different state-of-the-art methods by using different data splits and performance metrics on the StateFarm distracted driving and AI CITY Challenge datasets. With the data split experiments, we highlight the importance of leave-N-driver-out cross validation, since these models should perform well in real-world testing with never-before-seen drivers. The results show the importance of data splitting and the performance metric for the comparison and evaluation of different methods, and their significant effects on the results. Weiheng Chai, Jiyang Wang, Jiajing Chen, Senem Velipasalar, Anuj Sharma 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Vision-Language Models Can Identify Distracted Driver Behavior From Naturalistic VideosabstractRecognizing the activities causing distraction in real-world driving scenarios is critical for ensuring the safety and reliability of both drivers and pedestrians on the roadways. Conventional computer vision techniques are typically data-intensive and require a large volume of annotated training data to detect and classify various distracted driving behaviors, thereby limiting their generalization ability, efficiency and scalability. We aim to develop a generalized framework that showcases robust performance with access to limited or no annotated training data. Recently, vision-language models have offered large-scale visual-textual pretraining that can be adapted to task-specific learning like distracted driving activity recognition. Vision-language pretraining models like CLIP have shown significant promise in learning natural language-guided visual representations. This paper proposes a CLIP-based driver activity recognition approach that identifies driver distraction from naturalistic driving images and videos. CLIP’s vision embedding offers zero-shot transfer and task-based finetuning, which can classify distracted activities from naturalistic driving video. Our results show that this framework offers state-of-the-art performance on zero-shot transfer, finetuning and video-based models for predicting the driver’s state on four public datasets. We propose frame-based and video-based frameworks developed on top of the CLIP’s visual representation for distracted driving detection and classification tasks and report the results. Our code is available at https://github.com/zahid-isu/DriveCLIP Md. Zahid Hasan, Jiajing Chen, Jiyang Wang, Mohammed Shaiqur Rahman, Ameya Joshi, Senem Velipasalar, Chinmay Hegde, Anuj Sharma 0001, Soumik Sarkar |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Improving robustness and efficiency of edge computing models
Yilan Li, Yantao Lu, Helei Cui, Senem Velipasalar |
Wirel. Networks | 4 |
| 2023 | ViewNet: A Novel Projection-Based Backbone with View Pooling for Few-shot Point Cloud ClassificationabstractAlthough different approaches have been proposed for 3D point cloud-related tasks, few-shot learning (FSL) of 3D point clouds still remains under-explored. In FSL, un-like traditional supervised learning, the classes of training and test data do not overlap, and a model needs to rec-ognize unseen classes from only a few samples. Existing FSL methods for 3D point clouds employ point-based models as their backbone. Yet, based on our extensive experiments and analysis, we first show that using a point-based backbone is not the most suitable FSL approach, since (i) a large number of points' features are discarded by the max pooling operation used in 3D point-based backbones, decreasing the ability of representing shape information; (ii) point-based backbones are sensitive to occlusion. To address these issues, we propose employing a projection-and 2D Convolutional Neural Network-based backbone, referred to as the ViewNet, for FSL from 3D point clouds. Our approach first projects a 3D point cloud onto six different views to alleviate the issue of missing points. Also, to generate more descriptive and distinguishing features, we propose View Pooling, which combines different projected plane combinations into five groups and performs max-pooling on each of them. The experiments performed on the ModelNet40, ScanObjectNN and ModelNet40-C datasets, with cross validation, show that our method consistently outperforms the state-of-the-art baselines. Moreover, compared to traditional image classification backbones, such as ResNet, the proposed ViewNet can extract more distinguishing features from multiple views of a point cloud. We also show that ViewNet can be used as a backbone with different FSL heads and provides improved performance compared to traditionally used backbones. Jiajing Chen, Minmin Yang, Senem Velipasalar |
CVPR | 3 |
| 2023 | ToThePoint: Efficient Contrastive Learning of 3D Point Clouds via RecyclingabstractRecent years have witnessed significant developments in point cloud processing, including classification and segmentation. However, supervised learning approaches need a lot of well-labeled data for training, and annotation is labor-and time-intensive. Self-supervised learning, on the other hand, uses unlabeled data, and pretrains a back-bone with a pretext task to extract latent representations to be used with the downstream tasks. Compared to 2D images, self-supervised learning of 3D point clouds is under-explored. Existing models, for self-supervised learning of 3D point clouds, rely on a large number of data samples, and require significant amount of computational re-sources and training time. To address this issue, we propose a novel contrastive learning approach, referred to as To ThePoint. Different from traditional contrastive learning methods, which maximize agreement between features obtained from a pair of point clouds formed only with different types of augmentation, ToThePoint also maximizes the agreement between the permutation invariant features and features discarded after max pooling. We first perform self-supervised learning on the ShapeNet dataset, and then evaluate the performance of the network on different downstream tasks. In the downstream task experiments, performed on the ModelNet40, ModelNet40C, ScanobjectNN and ShapeNet-Part datasets, our proposed ToThe-Point achieves competitive, if not better results compared to the state-of-the-art baselines, and does so with significantly less training time (200 times faster than baselines). Xinglin Li, Jiajing Chen, Jinhui Ouyang, Hanhui Deng, Senem Velipasalar, Di Wu 0002 |
CVPR | 5 |
| 2023 | Sensitivity of Dynamic Network Slicing to Deep Reinforcement Learning Based Jamming AttacksabstractIn this paper, we consider multi-agent deep reinforcement learning (deep RL) based network slicing agents in a dynamic environment with multiple base stations and multiple users. We develop a deep RL based jammer with limited prior information and limited power budget. The goal of the jammer is to minimize the transmission rates achieved with network slicing and thus degrade the network slicing agents’ performance. We design a jammer with both listening and jamming phases and address jamming location optimization as well as jamming channel optimization via deep RL. We evaluate the jammer at the optimized location, generating interference attacks in the optimized set of channels by switching between the jamming phase and listening phase. We show that the proposed jammer can significantly reduce the victims’ performance without direct feedback or prior knowledge on the network slicing policies. Mustafa Cenk Gursoy, Senem Velipasalar |
PIMRC | 3 |
| 2023 | Cross-Modality Feature Fusion Network for Few-Shot 3D Point Cloud ClassificationabstractRecent years have witnessed significant progress in the field of few-shot image classification while few-shot 3D point cloud classification still remains under-explored. Real-world 3D point cloud data often suffers from occlusions, noise and deformation, which make the few-shot 3D point cloud classification even more challenging. In this paper, we propose a cross-modality feature fusion network, for few-shot 3D point cloud classification, which aims to recognize an object given only a few labeled samples, and provides better performance even with point cloud data with missing points. More specifically, we train two models in parallel. One is a projection-based model with ResNet18 as the backbone and the other one is a point-based model with a DGCNN backbone. Moreover, we design a Support-Query Mutual Attention (sqMA) module to fully exploit the correlation between support and query features. Extensive experiments on three datasets, namely ModelNet40, ModelNet40-C and ScanObjectNN, show the effectiveness of our method, and its robustness to missing points. Our proposed method outperforms different state-of-the-art baselines on all datasets. The margin of improvement is even larger on the ScanObjectNN dataset, which is collected from real-world scenes and is more challenging with objects having missing points. Minmin Yang, Jiajing Chen, Senem Velipasalar |
WACV | 3 |
| 2023 | Driver Head Pose Detection From Naturalistic Driving DataabstractDriver behavior analysis plays an important role in driver assistance systems. A driver’s face and head pose hold the key towards understanding whether the driver’s attention and concentration are on the road while driving. Naturalistic driving studies (NDS) allow observing drivers in real-time under naturalistic traffic conditions. Yet, data collected in NDS often comprise low-resolution videos usually with more challenging camera positions compared to controlled studies. For instance, when the camera is not directly facing the driver, classifying head pose becomes more challenging, since the variation between different classes becomes much smaller. In this paper, we propose three different approaches to classify a driver’s head pose from naturalistic videos, which were captured by a camera providing a side view, instead of directly facing the driver. These approaches employ a sequence of five key points on the driver’s face. We compare these three proposed approaches with each other as well as with three different baselines by using leave-one-driver-out cross-validation on nine different drivers. Results show that our proposed method employing a Bidirectional Gated Recurrent Unit (BiGRU) outperforms the best performing baseline by 11% in terms of overall accuracy. Weiheng Chai, Jiajing Chen, Jiyang Wang, Senem Velipasalar, Archana Venkatachalapathy, Yaw Adu-Gyamfi, Jennifer Merickel, Anuj Sharma 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Why Discard if You can Recycle?: A Recycling Max Pooling Module for 3D Point Cloud AnalysisabstractIn recent years, most 3D point cloud analysis models have focused on developing either new network architectures or more efficient modules for aggregating point features from a local neighborhood. Regardless of the network architecture or the methodology used for improved feature learning, these models share one thing, which is the use of max-pooling in the end to obtain permutation invariant features. We first show that this traditional approach causes only a fraction of 3D points contribute to the permutation-invariant features, and discards the rest of the points. In order to address this issue and improve the performance of any baseline 3D point classification or segmentation model, we propose a new module, referred to as the Recycling Max-Pooling (RMP) module, to recycle and utilize the features of some of the discarded points. We incorporate a refinement loss that uses the recycled features to refine the prediction loss obtained from the features kept by traditional max-pooling. To the best of our knowledge, this is the first work that explores recycling of still useful points that are traditionally discarded by max-pooling. We demonstrate the effectiveness of the proposed RMP module by incorporating it into several milestone baselines and state-of-the-art networks for point cloud classification and indoor semantic segmentation tasks. We show that RPM, without any bells and whistles, consistently improves the performance of all the tested networks by using the same base network implementation and hyper-parameters. The code is provided in the supplementary material. Jiajing Chen, Burak Kakillioglu, Huantao Ren, Senem Velipasalar |
CVPR | 4 |
| 2022 | Communication-Efficient and Privacy-Preserving Feature-based Federated Transfer LearningabstractFederated learning has attracted growing interest as it preserves the clients' privacy. As a variant of federated learning, federated transfer learning utilizes the knowledge from similar tasks and thus has also been intensively studied. However, due to the limited radio spectrum, the communication efficiency of federated learning via wireless links is critical since some tasks may require thousands of Terabytes of uplink payload. In order to improve the communication efficiency, we in this paper propose the feature-based federated transfer learning as an innovative approach to reduce the uplink payload by more than five orders of magnitude compared to that of existing approaches. We first introduce the system design in which the extracted features and outputs are uploaded instead of parameter updates, and then determine the required payload with this approach and provide comparisons with the existing approaches. Subsequently, we analyze the random shuffling scheme that preserves the clients' privacy. Finally, we evaluate the performance of the proposed learning scheme via experiments on an image classification task to show its effectiveness. Mustafa Cenk Gursoy, Senem Velipasalar |
GLOBECOM | 3 |
| 2022 | Controlled Sensing and Anomaly Detection Via Soft Actor-Critic Reinforcement LearningabstractTo address the anomaly detection problem in the presence of noisy observations and to tackle the tuning and efficient exploration challenges that arise in deep reinforcement learning algorithms, we in this paper propose a soft actor-critic deep reinforcement learning framework. To evaluate the proposed framework, we measure its performance in terms of detection accuracy, stopping time, and the total number of samples needed for detection. Via simulation results, we demonstrate the performance when soft actor-critic algorithms are employed, and identify the impact of key parameters, such as the sensing cost, on the performance. In all results, we further provide comparisons between the performances of the proposed soft actor-critic and conventional actor-critic algorithms. Chen Zhong 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
ICASSP | 3 |
| 2022 | Multi-Agent Reinforcement Learning with Pointer Networks for Network Slicing in Cellular SystemsabstractIn this paper, we present a multi-agent deep reinforcement learning (deep RL) framework for network slicing in a dynamic environment with multiple base stations. We first introduce the wireless network virtualization (WNV) and the interference channel model. Then, we formulate the network slicing problem in the dynamic environment in which fading varies, users have mobility, and requests are randomly generated over time. Subsequently, we propose a deep RL framework with multiple actors and centralized critic (MACC) to maximize the reward over all base stations instead of pursuing local optimization. The actors are implemented as pointer networks to fit the varying dimension of input. Finally, we evaluate the performance of the proposed deep RL algorithm via simulations to demonstrate its effectiveness. Mustafa Cenk Gursoy, Senem Velipasalar |
ICC | 3 |
| 2022 | Gaitpoint: A Gait Recognition Network Based on Point Cloud AnalysisabstractWe propose a novel gait recognition method that combines convolutional features with features of human pose key points obtained by a point cloud analysis model. Currently, most state-of-the-art works on gait recognition rely on only images and are purely based on convolutional neural networks. Most of these methods are very sensitive to small variations in the appearance of a walking person. For instance, if a person wears a coat or carries a bag, the accuracy of these methods may drop significantly. To address this problem, we propose to treat a sequence of human key points as a point cloud and combine human key point features and convolution feature map for final prediction. The experimental results show the promise of this approach, which outperforms three state-oft-he-art baselines in all walking scenarios, including the ones involving heavy clothing or carried items. Jiajing Chen, Huantao Ren, Frank Sicong Chen, Senem Velipasalar, Vir V. Phoha |
ICIP | 4 |
| 2022 | Robust Deep Reinforcement Learning Based Network Slicing under Adversarial Jamming AttacksabstractIn this paper, we first present a deep reinforcement learning (deep RL) framework for network slicing in a dynamic environment. We propose three different deep RL algorithms, namely actor-critic, deep Q learning (DQN), and soft DQN, to select slices from the best recorded subset which is updated over time to adapt to the dynamic environment. We evaluate the performances of the proposed deep RL agents for network slicing and provide comparisons. Subsequently, we design intelligent jammers also as deep RL agents that significantly degrade the user's sum reward. Finally, we propose effective defensive measures to mitigate jamming attacks by determining the proper time instants to retrain the network slicing policy. Via simulations, we quantify the improvements in the performance with the defensive retraining. Mustafa Cenk Gursoy, Senem Velipasalar, Yalin E. Sagduyu |
PIMRC | 3 |
| 2022 | Learning-Based Robust Anomaly Detection in the Presence of Adversarial AttacksabstractTo address the anomaly detection problem in the presence of noisy sensor observations and probing costs, we in this paper propose a soft actor-critic deep reinforcement learning framework. Moreover, considering adversarial jamming attacks, we design a generative adversarial network (GAN) based framework to identify the jammed sensors. To evaluate the proposed framework, we measure the performance in terms of detection accuracy, stopping time, and the total number of samples needed for detection. Via simulation results, we demonstrate the performances when soft actor-critic algorithms are sensitive to the probing cost and actively adapt to different environment settings. We analyze the impact of jamming attacks and identify the improvements achieved by GAN-based approach. We further provide comparisons between the performances of the proposed soft actor-critic and conventional actor-critic algorithms. Chen Zhong 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
WCNC | 3 |
| 2022 | Capsule network-based semantic segmentation model for thermal anomaly identification on building envelopes
Chenbin Pan, Jiyang Wang, Weiheng Chai, Burak Kakillioglu, Yasser El Masri, Eleanna Panagoulia, Norhan Bayomi, John E. Fernandez, Tarek Rakha, Senem Velipasalar |
Adv. Eng. Informatics | 11 |
| 2022 | Background-Aware 3-D Point Cloud Segmentation With Dynamic Point Feature AggregationabstractWith the proliferation of Lidar sensors and 3D vision cameras, 3D point cloud analysis has attracted significant attention in recent years. After the success of the pioneer work PointNet, deep learning-based methods have been increasingly applied to various tasks, including 3D point cloud segmentation and 3D object classification. In this paper, we propose a novel 3D point cloud learning network, referred to as Dynamic Point Feature Aggregation Network (DPFA-Net), by selectively performing the neighborhood feature aggregation with dynamic pooling and an attention mechanism. DPFA-Net has two variants for semantic segmentation and classification of 3D point clouds. As the core module of the DPFA-Net, we propose a Feature Aggregation layer, in which features of the dynamic neighborhood of each point are aggregated via a self-attention mechanism. In contrast to other segmentation models, which aggregate features from fixed neighborhoods, our approach can aggregate features from different neighbors in different layers providing a more selective and broader view to the query points, and focusing more on the relevant features in a local neighborhood. In addition, to further improve the performance of the proposed semantic segmentation model, we present two novel approaches, namely Two-Stage BF-Net and BF-Regularization to exploit the background-foreground information. Experimental results show that the proposed DPFA-Net achieves the state-of-the-art overall accuracy score for semantic segmentation on the S3DIS dataset, and provides a consistently satisfactory performance across different tasks of semantic segmentation, part segmentation, and 3D object classification. It is also computationally more efficient compared to other methods. Jiajing Chen, Burak Kakillioglu, Senem Velipasalar |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Taking a Deeper Look at the Brain: Predicting Visual Perceptual and Working Memory Load From High-Density fNIRS DataabstractPredicting workload using physiological sensors has taken on a diffuse set of methods in recent years. However, the majority of these methods train models on small datasets, with small numbers of channel locations on the brain, limiting a model's ability to transfer across participants, tasks, or experimental sessions. In this paper, we introduce a new method of modeling a large, cross-participant and cross-session set of high density functional near infrared spectroscopy (fNIRS) data by using an approach grounded in cognitive load theory and employing a Bi-Directional Gated Recurrent Unit (BiGRU) incorporating attention mechanism and self-supervised label augmentation (SLA). We show that our proposed CNN-BiGRU-SLA model can learn and classify different levels of working memory load (WML) and visual processing load (VPL) across participants. Importantly, we leverage a multi-label classification scheme, where our models are trained to predict simultaneously occurring levels of WML and VPL. We evaluate our model using leave-one-participant-out (LOOCV) as well as 10-fold cross validation. Using LOOCV, for binary classification (off/on), we reached an F1-score of 0.9179 for WML and 0.8907 for VPL across 22 participants (each participant did 2 sessions). For multi-level (off, low, high) classification, we reached an F1-score of 0.7972 for WML and 0.7968 for VPL. Using 10-fold cross validation, for multi-level classification, we reached an F1-score of 0.7742 for WML and 0.7741 for VPL. Jiyang Wang, Trevor Grant, Senem Velipasalar, Baocheng Geng, Leanne M. Hirshfield |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | A Survey on Driver Behavior Analysis From In-Vehicle CamerasabstractDistracted or drowsy driving is unsafe driving behavior responsible for thousands of crashes every year. Studying driver behavior has challenges associated with observing drivers in their natural environment. The naturalistic driving study (NDS) has become the most sought-after approach, since it eliminates the bias of a controlled setup, allowing researchers to understand drivers’ behavior in real-world scenarios. Video recordings collected in NDS research are incredibly insightful in identifying driver errors. Computer vision techniques have been used to autonomously analyze video data and classify drivers’ behavior. While computer vision scientists focus on image analytics, NDS researchers are interested in the factors impacting driver behavior. This survey paper makes a concerted effort to serve both communities by comprehensively reviewing studies, describing their data collection, computer vision techniques implemented, and performance in classifying driver behavior. The scope is limited to studies employing at least one camera observing the driver inside a vehicle. Based on their objective, papers have been classified as detecting low-level (e.g. head orientation) or high-level (e.g. distraction detection) driver information. Papers have been further classified based on the datasets they employ. In addition to twelve public datasets, many private datasets have also been identified, and their data collection design is discussed to highlight any impact on model performance. Across each task, algorithms employed and their performance are discussed to establish a baseline. A comparison of different frameworks for NDS video data analytics throws light on the existing gaps in the state-of-the-art that can be addressed by future computer vision research. Jiyang Wang, Weiheng Chai, Archana Venkatachalapathy, Kai Liang Tan, Arya Haghighat, Senem Velipasalar, Yaw Adu-Gyamfi, Anuj Sharma 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Preclinical Stage Alzheimer's Disease Detection Using Magnetic Resonance Image ScansabstractAlzheimer's disease is one of the diseases that mostly affects older people without being a part of aging. The most common symptoms include problems with communicating and abstract thinking, as well as disorientation. It is important to detect Alzheimer's disease in early stages so that cognitive functioning would be improved by medication and training. In this paper, we propose two attention model networks for detecting Alzheimer's disease from MRI images to help early detection efforts at the preclinical stage. We also compare the performance of these two attention network models with a baseline model. Recently available OASIS-3 Longitudinal Neuroimaging, Clinical, and Cognitive Dataset is used to train, evaluate and compare our models. The novelty of this research resides in the fact that we aim to detect Alzheimer's disease when all the parameters, physical assessments, and clinical data state that the patient is healthy and showing no symptoms. Fatih Altay, Guillermo Ramon Sanchez, Yanli Zhang-James, Stephen V. Faraone, Senem Velipasalar, Asif Salekin |
AAAI | 5 |
| 2021 | PT-CapsNet: A Novel Prediction-Tuning Capsule Network Suitable for Deeper ArchitecturesabstractCapsule Networks (CapsNets) create internal representations by parsing inputs into various instances at different resolution levels via a two-phase process – part-whole transformation and hierarchical component routing. Since both of these internal phases are computationally expensive, CapsNet have not found wider use. Existing variations of CapsNets mainly focus on performance comparison with the original CapsNet, and have not outperformed CNN-based models on complex tasks. To address the limitations of the existing CapsNet structures, we propose a novel Prediction-Tuning Capsule Network (PT-CapsNet), and also introduce fully connected PT-Capsules (FC-PT-Caps) and locally connected PT-Capsules (LC-PT-Caps). Different from existing CapsNet structures, our proposed model (i) allows the use of capsules for more difficult vision tasks and provides wider applicability; and (ii) provides better than or comparable performance to CNN-based baselines on these complex tasks. In our experiments, we show robustness to affine transformations, as well as the lightweight and scalability of PT-CapsNet via constructing larger and deeper networks and performing comparisons on classification, semantic segmentation and object detection tasks. The results show consistent performance improvement and significant parameter reduction compared to various baseline models. Code is available at https://github.com/Christinepan881/PT-CapsNet.git. Chenbin Pan, Senem Velipasalar |
ICCV | 2 |
| 2021 | Weighted Average Precision: Adversarial Example Detection for Visual Perception Of Autonomous VehiclesabstractRecent works have shown that neural networks are vulnerable to carefully crafted adversarial examples (AE). By adding small perturbations to original images, AEs are able to deceive victim models, and result in incorrect outputs. Research work in adversarial machine learning started to focus on the detection of AEs in autonomous driving applications. However, existing studies either use simplifying assumptions on the outputs of object detectors or ignore the tracking system in the perception pipeline. In this paper, we first propose a novel similarity distance metric for object detection outputs in autonomous driving applications. Then, we bridge the gap between the current AE detection research and the real-world autonomous systems by providing a temporal AE detection algorithm, which takes the impact of tracking system into consideration. We perform evaluations on Berkeley Deep Drive and CityScapes datasets, by using different white-box and black-box attacks, which show that our approach outperforms the mean-average-precision and mean intersection over-union based AE detection baselines by significantly increasing the detection accuracy. Weiheng Chai, Yantao Lu, Senem Velipasalar |
ICIP | 3 |
| 2021 | Fabricate-Vanish: An Effective And Transferable Black-Box Adversarial Attack Incorporating Feature DistortionabstractAdversarial examples have emerged as increasingly severe threats for deep neural networks. Recent works have revealed that these malicious samples can transfer across different neural networks, and effectively attack other models. The state-of-the-art methodologies leverage Fast Gradient Sign Method to generate obstructing textures, which can cause neural networks to make incorrect inferences. However, the over-reliance on task-specific loss functions makes the adversarial examples less transferable across networks. Moreover, recent de-noising based adaptive defences provide promising performance against aforementioned attacks. Therefore, to achieve better transferability and attack effectiveness, we propose a novel attack, referred to as the Fabricate-Vanish (FV) attack, which is able to erase benign representations and generate obstruction textures simultaneously. The proposed FV attack treats the adversarial example transferability as latent contribution for each layer of deep neural networks, and maximizes the attack performance by balancing transferability and task specific loss function. Our experimental results on ImageNet show that the proposed FV attack achieves the best attack performance and better transferability by degrading the accuracy of classifiers 3.8% more on average compared to the state-of-the-art attacks. Yantao Lu, Xueying Du, Bingkun Sun, Haining Ren, Senem Velipasalar |
ICIP | 5 |
| 2021 | Part-Based Feature Squeezing To Detect Adversarial Examples in Person Re-Identification NetworksabstractAlthough deep neural networks (DNNs) have achieved top performances in different computer vision tasks, such as object detection, image segmentation and person re-identification (ReID), they can easily be deceived by adversarial examples, which are carefully crafted images with perturbations that are imperceptible to human eyes. Such adversarial examples can significantly degrade the performance of existing DNNs. There are also targeted attacks misleading classifiers into making specific decisions based on attackers’ intentions. In this paper, we propose a new method to effectively detect adversarial examples presented to a person ReID network. The proposed method utilizes parts-based feature squeezing to detect the adversarial examples. We apply two types of squeezing to segmented body parts to better detect adversarial examples. We perform extensive experiments over three major datasets with different attacks, and compare the detection performance of the proposed body part-based approach with a ReID method that is not parts-based. Experimental results show that the proposed method can effectively detect the adversarial examples, and has the potential to avoid significant decreases in person ReID performance caused by adversarial examples. Yu Zheng 0016, Senem Velipasalar |
ICIP | 2 |
| 2021 | Adversarial Reinforcement Learning in Dynamic Channel Access and Power ControlabstractDeep reinforcement learning (DRL) has recently been used to perform efficient resource allocation in wireless communications. In this paper, the vulnerabilities of such DRL agents to adversarial attacks is studied. In particular, we consider multiple DRL agents that perform both dynamic channel access and power control in wireless interference channels. For these victim DRL agents, we design a jammer, which is also a DRL agent. We Propose an adversarial jamming attack scheme that utilizes a listening phase and significantly degrades the users' sum rate. Subsequently, we develop an ensemble policy defense strategy against such a jamming atta.ck.er by reloa.di.ng models (saved during retraining) that have minimum transition correlation. Mustafa Cenk Gursoy, Senem Velipasalar |
WCNC | 3 |
| 2020 | Enhancing Cross-Task Black-Box Transferability of Adversarial Examples With Dispersion ReductionabstractNeural networks are known to be vulnerable to carefully crafted adversarial examples, and these malicious samples often transfer, i.e., they remain adversarial even against other models. Although significant effort has been devoted to the transferability across models, surprisingly little attention has been paid to cross-task transferability, which represents the real-world cybercriminal's situation, where an ensemble of different defense/detection mechanisms need to be evaded all at once. We investigate the transferability of adversarial examples across a wide range of real-world computer vision tasks, including image classification, object detection, semantic segmentation, explicit content detection, and text detection. Our proposed attack minimizes the “dispersion” of the internal feature map, overcoming the limitations of existing attacks, that require task-specific loss functions and/or probing a target model. We conduct evaluation on open-source detection and segmentation models, as well as four different computer vision tasks provided by Google Cloud Vision (GCV) APIs. We demonstrate that our approach outperforms existing attacks by degrading performance of multiple CV tasks by a large margin with only modest perturbations. Yantao Lu, Yunhan Jia, Bai Li 0001, Weiheng Chai, Lawrence Carin, Senem Velipasalar |
CVPR | 7 |
| 2020 | Anomaly Detection via Controlled Sensing and Deep Active InferenceabstractIn this paper, we address the anomaly detection problem where the objective is to find the anomalous processes among a given set of processes. To this end, the decision-making agent probes a subset of processes at every time instant and obtains a potentially erroneous estimate of the binary variable which indicates whether or not the corresponding process is anomalous. The agent continues to probe the processes until it obtains a sufficient number of measurements to reliably identify the anomalous processes. In this context, we develop a sequential selection algorithm that decides which processes to be probed at every instant to detect the anomalies with an accuracy exceeding a desired value while minimizing the delay in making the decision and the total number of measurements taken. Our algorithm is based on active inference which is a general framework to make sequential decisions in order to maximize the notion of free energy. We define the free energy using the objectives of the selection policy and implement the active inference framework using a deep neural network approximation. Using numerical experiments, we compare our algorithm with the state-of-the-art method based on deep actor-critic reinforcement learning and demonstrate the superior performance of our algorithm. Geethu Joseph, Chen Zhong 0007, Mustafa Cenk Gursoy, Senem Velipasalar, Pramod K. Varshney |
GLOBECOM | 4 |
| 2020 | Anomaly Detection and Sampling Cost Control via Hierarchical GANsabstractAnomaly detection incurs certain sampling and sensing costs and therefore it is of great importance to strike a balance between the detection accuracy and these costs. In this work, we study anomaly detection by considering the detection of threshold crossings in a stochastic time series without the knowledge of its statistics. To reduce the sampling cost in this detection process, we propose the use of hierarchical generative adversarial networks (GANs) to perform non-uniform sampling. In order to improve the detection accuracy and reduce the delay in detection, we introduce a butter zone in the operation of the proposed GANbased detector. In the experiments, we analyze the performance of the proposed hierarchical GAN detector considering the metrics of detection delay, miss rates, average cost of error, and sampling ratio. We identify the tradeoffs in the performance as the butter zone sizes and the number of GAN levels in the hierarchy vary. We also compare the performance with that of a sampling policy that approximately minimizes the sum of average costs of sampling and error given the parameters of the stochastic process. We demonstrate that the proposed GAN-based detector can have significant performance improvements in terms of detection delay and average cost of error with a larger butter zone but at the cost of increased sampling rates. Chen Zhong 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
GLOBECOM | 3 |
| 2020 | Adversarial Jamming Attacks on Deep Reinforcement Learning Based Dynamic Multichannel AccessabstractAdversarial attack strategies have been widely studied in machine learning applications, and now are increasingly attracting interest in wireless communications as the application of machine learning methods to wireless systems grows along with security concerns. In this paper, we propose two adversarial policies, one based on feed-forward neural networks (FNNs) and the other based on deep reinforcement learning (DRL) policies. Both attack strategies aim at minimizing the accuracy of a DRL-based dynamic channel access agent. We first present the two frameworks and the dynamic attack procedures of the two adversarial policies. Then we demonstrate and compare their performances. Finally, the advantages and disadvantages of the two frameworks are identified. Chen Zhong 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
WCNC | 4 |
| 2020 | Autonomous Selective Parts-Based TrackingabstractObject tracking from videos is still a challenging task due to various changes throughout a video sequence including occlusions, motion blur, scale and other deformation changes. In this paper, we propose a selective parts-based approach, using correlation filters, that makes choices based on a consensus of the parts and global tracking. Moreover, we further enhance our parts-based approach by introducing a segmentation-assisted parts initialization. In addition, we present a genetic algorithmbased method to autonomously select various parameters of the tracking algorithm, as opposed to the common practice of manually tuning those parameters. In contrast to existing partbased methods, the proposed method does not dilute accurate tracking by averaging results over multiple parts at every frame. Instead, we take a selective approach based on the relative weight of the responses across parts. Moreover, we only make location corrections when a part diverges, and rely on these location corrections to maintain an accurate appearance model. In the case of occlusions, which are among the main reasons for using a parts-based approach, our proposed approach consistently achieves the best performance. It is due to the ability to handle occlusion and not dilute decisions with incorrect parts, that our proposed approach enables state-of-the-art performance. The proposed approach was evaluated on videos from three different challenging benchmark datasets. Our approach has resulted in better overall precision and success rates for three different base tracking approaches. Maria Scalzo-Cornacchia, Senem Velipasalar |
IEEE Trans. Image Process. | 2 |
| 2019 | Deep Actor-Critic Reinforcement Learning for Anomaly DetectionabstractAnomaly detection is widely applied in a variety of domains, involving for instance, smart home systems, network traffic monitoring, IoT applications and sensor networks. In this paper, we study deep reinforcement learning based active sequential testing for anomaly detection. We assume that there is an unknown number of abnormal processes at a time and the agent can only check with one sensor in each sampling step. To maximize the confidence level of the decision and minimize the stopping time concurrently, we propose a deep actor-critic reinforcement learning framework that can dynamically select the sensor based on the posterior probabilities. We provide simulation results for both the training phase and testing phase, and compare the proposed framework with the Chernoff test in terms of claim delay and loss. Chen Zhong 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
GLOBECOM | 3 |
| 2019 | Deep Multi-Agent Reinforcement Learning Based Cooperative Edge Caching in Wireless NetworksabstractThe growing demand on high-quality and low-latency multimedia services has led to much interest in edge caching techniques. Motivated by this, we in this paper consider edge caching at the base stations with unknown content popularity distributions. To solve the dynamic control problem of making caching decisions, we propose a deep actor-critic reinforcement learning based multi-agent framework with the aim to minimize the overall average transmission delay. To evaluate the proposed framework, we compare the learning-based performance with three other caching policies, namely least recently used (LRU), least frequently used (LFU), and first-in-first-out (FIFO) policies. Through simulation results, performance improvements of the proposed framework over these three caching algorithms have been identified and its superior ability to adapt to varying environments is demonstrated. Chen Zhong 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
ICC | 3 |
| 2019 | Efficient Human Activity Classification from Egocentric Videos Incorporating Actor-Critic Reinforcement LearningabstractIn this paper, we introduce a novel framework to significantly reduce the computational cost of human temporal activity recognition from egocentric videos while maintaining the accuracy at the same level. We propose to apply the actor-critic model of reinforcement learning to optical flow data to locate a bounding box around region of interest, which is then used for clipping a sub-image from a video frame. We also propose to use one shallow and one deeper 3D convolutional neural network to process the original image and the clipped image region, respectively. We compared our proposed method with another approach using 3D convolutional networks on the recently released Dataset of Multimodal Semantic Egocentric Video. Experimental results show that the proposed method reduces the processing time by 36.4% while providing comparable accuracy at the same time. Yantao Lu, Yilan Li, Senem Velipasalar |
ICIP | 3 |
| 2019 | Autonomous Choice of Deep Neural Network Parameters by a Modified Generative Adversarial NetworkabstractThe choice of parameters, and the design of the network architecture are important factors affecting the performance of deep neural networks. However, this task still heavily depends on trial and error, and empirical results. Considering that there are many design and parameter choices, it is very hard to cover every configuration, and find the optimal structure. In this paper, we propose a novel method that autonomously and simultaneously optimizes multiple parameters of any given deep neural network by using a modified generative adversarial network (GAN). In our approach, two different models compete and improve each other progressively. Without loss of generality, the proposed method has been tested with three different neural network architectures, and three very different datasets and applications. The results show that the presented approach can simultaneously and successfully optimize multiple neural network parameters, and achieve increased accuracy in all three scenarios. Yantao Lu, Senem Velipasalar |
ICIP | 2 |
| 2019 | Throughput-Delay Tradeoffs With Finite Blocklength Coding Over Multiple Coherence BlocksabstractThis paper investigates the performance of wireless systems that employ finite-blocklength channel codes for transmission and operate under queuing constraints in the form of limitations on buffer overflow or delay violation probabilities. A block fading model, in which fading stays constant in each coherence block and changes independently between blocks, is considered. It is assumed that channel coding is performed over multiple coherence blocks. A simple ARQ scheme with error-free feedback without any delay is considered. The channel coding rate with given maximal error probability is considered as the service rate and is incorporated into the effective capacity formulation, which characterizes the maximum constant arrival rate that can be supported under statistical queuing constraints. Performances of variable-rate and fixed-rate transmissions are studied. The optimum error probability for variable-rate and fixed-rate transmissions is shown to be unique. The limiting performance as the number of blocks increases is characterized. The tradeoffs and the interactions between the throughput, the number of coherence blocks over which channel coding is performed, error probabilities, channel coherence duration, and queuing constraints are identified. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
IEEE Trans. Commun. | 3 |
| 2019 | Power Control for Wireless VBR Video Streaming: From Optimization to Reinforcement LearningabstractIn this paper, we investigate the problem of power control for streaming variable bit rate (VBR) videos over wireless links. A system model involving a transmitter (e.g., a base station) that sends VBR video data to a receiver (e.g., a mobile user) equipped with a playout buffer is adopted, as used in dynamic adaptive streaming video applications. In this setting, we analyze power control policies considering the following two objectives: 1) the minimization of the transmit power consumption and 2) the minimization of the transmission completion time of the communication session. In order to play the video without interruptions, the power control policy should also satisfy the requirement in which the VBR video data is delivered to the mobile user without causing playout buffer underflow or overflows. A directional water-filling algorithm, which provides a simple and concise interpretation of the necessary optimality conditions, is identified as the optimal offline policy. Following this, two online policies are proposed for power control based on channel side information (CSI) prediction within a short time window. Dynamic programming is employed to implement the optimal offline and the initial online power control policies that minimize the transmit power consumption in the communication session. Subsequently, reinforcement learning (RL)-based approach is employed for the second online power control policy. Through the simulation results, we show that the optimal offline power control policy that minimizes the overall power consumption leads to substantial energy savings compared with the strategy of minimizing the time duration of video streaming. We also demonstrate that the RL algorithm performs better than the dynamic programming-based online grouped water-filling (GWF) strategy unless the channel is highly correlated. Chuang Ye, Mustafa Cenk Gursoy, Senem Velipasalar |
IEEE Trans. Commun. | 3 |
| 2018 | Power control and mode selection for VBR video streaming in D2D networksabstractIn this paper, we investigate the problem of power control for streaming variable-bit-rate (VBR) videos in a device-to-device (D2D) wireless network. A VBR video traffic model that considers video frame sizes and playout buffers at the mobile users is adopted. A setup with one pair of D2D users (DUs) and one cellular user (CU) is considered and three modes, namely cellular mode, dedicated mode and reuse mode, are employed. Mode selection for the data delivery is determined and the transmit powers of the base station (BS) and device transmitter are optimized with the goal of maximizing the overall transmission rate while VBR video data can be delivered to the CU and DU without causing playout buffer underflows or overflows. A low-complexity algorithm is proposed. Through simulations with VBR video traces over fading channels, we demonstrate that video delivery with mode selection and power control achieves a better performance than just using a single mode throughout the transmission. Chuang Ye, Mustafa Cenk Gursoy, Senem Velipasalar |
WCNC | 3 |
| 2018 | Building predictive models of emotion with functional near-infrared spectroscopy
Danushka Bandara, Senem Velipasalar, Sarah Bratt, Leanne M. Hirshfield |
Int. J. Hum. Comput. Stud. | 2 |
| 2018 | Quality-Driven Resource Allocation for Full-Duplex Delay-Constrained Wireless Video TransmissionsabstractIn this paper, wireless video transmission over full-duplex channels under total bandwidth and minimum required quality constraints is studied. In order to provide the desired performance levels to the end-users in real-time video transmissions, quality of service requirements such as statistical delay constraints are also considered. Effective capacity is used as the throughput metric in the presence of such statistical delay constraints since deterministic delay bounds are difficult to guarantee due to the time-varying nature of wireless fading channels. A communication scenario with multiple pairs of users in which different users have different delay requirements is addressed. Following characterizations from the rate-distortion theory, a logarithmic model of the quality-rate relation is used for predicting the quality of the reconstructed video in terms of the peak signal-to-noise ratio at the receiver side. Since the optimization problem is not concave or convex, the optimal bandwidth and power allocation policies that maximize the weighted sum video quality subject to total bandwidth, maximum transmission power level and minimum required quality constraints are derived by using monotonic optimization theory. Chuang Ye, Mustafa Cenk Gursoy, Senem Velipasalar |
IEEE Trans. Commun. | 3 |
| 2017 | Joint Mode Selection and Resource Allocation for D2D Communications via Vertex ColoringabstractDevice-to-device (D2D) communication underlaid with cellular networks is a new paradigm, proposed to enhance the performance of cellular networks. By allowing a pair of D2D users to communicate directly and share the same spectral resources with the cellular users, D2D communication can achieve higher spectral efficiency, improve the energy efficiency, and lower the traffic delay. In this paper, we propose a novel joint mode selection and channel resource allocation algorithm via the vertex coloring approach. We decompose the problem into three subproblems and design algorithms for each of them. In the first step, we divide the users into groups using a vertex coloring algorithm. In the second step, we solve the power optimization problem using the interior-point method for each group and conduct mode selection between the cellular mode and D2D mode for D2D users, and we assign channel resources to these groups in the final step. Numerical results show that our algorithm achieves higher sum rate and serves more users with relatively small time consumption compared with other algorithms. Also, the influence of system parameters and the tradeoff between sum rate and the number of served users are studied through simulation results. Yi Li 0007, Mustafa Cenk Gursoy, Senem Velipasalar, Jian Tang 0008 |
GLOBECOM | 3 |
| 2017 | Throughput of HARQ-IR with finite blocklength codes and QoS constraintsabstractIn this paper, throughput of hybrid automatic repeat request (HARQ) schemes with finite blocklength codes is studied for both constant-rate and ON-OFF discrete-time Markov arrivals under statistical queuing constraints and deadline limits. After analyzing the decoding error probability and outage probability, the distribution of transmission period is characterized, and the throughput expressions are obtained for both arrival models. Analytical results are verified via Monte Carlo simulations. In the numerical results, the impact of deadline constraints, fixed transmission rate, coding blocklength, and queuing constraints on the throughput is analyzed. Yi Li 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
ISIT | 3 |
| 2017 | Intercell Interference-Aware Scheduling for Delay Sensitive Applications in C-RANabstractCloud radio access network (C-RAN) architecture is a new mobile network architecture that enables cooperative baseband processing and information sharing among multiple cells and achieves high adaptability to nonuniform traffic by centralizing the baseband processing resources in a virtualized baseband unit (BBU) pool. In this work, we formulate the utility of each user using a convex delay cost function, and design a two-step scheduling algorithm with good delay performance for the C-RAN architecture. In the first step, all users in multiple cells are grouped into small user groups, according to their interference levels and estimated utilities. In the second step, channels are matched to the user groups to maximize the system utility. The performance of our algorithm is further studied via simulations, and the advantages of C-RAN architecture is verified. Yi Li 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
VTC Fall | 3 |
| 2017 | Optimal Resource Allocation for Full-Duplex Wireless Video Transmissions under Delay ConstraintsabstractIn this paper, wireless video transmission over full-duplex channels is studied. In order to provide the desired performance levels to the end-users in real-time video transmissions, quality of service (QoS) requirements such as statistical delay constraints are also considered. Effective capacity (EC) is used as the throughput metric in the presence of such statistical delay constraints since deterministic delay bounds are difficult to guarantee due to the time-varying nature of wireless fading channels. A communication scenario with a pair of users and multiple subchannels in which users can have different delay requirements is addressed. Following characterizations from the rate-distortion (R-D) theory, a logarithmic model of the quality-rate relation is used for predicting the quality of the reconstructed video in terms of the peak signal-to-noise ratio (PSNR) at the receiver side. Since the optimization problem is not concave or convex, the optimal power allocation policy that maximizes the weighted sum video quality subject to total transmission power constraint is derived by using monotonic optimization (MO) theory. The optimal scheme is compared with two suboptimal strategies. Chuang Ye, Mustafa Cenk Gursoy, Senem Velipasalar |
WCNC | 3 |
| 2017 | Autonomous Fall Detection With Wearable Cameras by Using Relative Entropy Distance MeasureabstractTimely, precise, and reliable detection of fall events is very important for systems monitoring activities of elderly people, especially the ones living independently. In this paper, we propose an autonomous fall detection system by taking a completely different view compared with existing vision-based activity monitoring systems and applying a reverse approach. In our system, in contrast with static sensors installed at fixed locations, the camera is worn by the subject, and thus, monitoring is not limited only to areas where the sensors are located and extends to wherever the subject may travel. Moreover, the camera provides a richer set of data and helps lower the false positive rates compared with accelerometer-only systems. We employ a modified version of the histograms of oriented gradients (HOG) approach together with the gradient local binary patterns (GLBP). It is shown that, with the same training set, the GLBP feature is more descriptive and discriminative than HOG, histograms of template, and semantic local binary patterns. Moreover, we autonomously compute a threshold, for the detection of fall events, from the training data based on relative entropy, which is a member of Ali-Silvey distance measures. Experiments are performed with ten different people and a total of around 300 associated fall events indoors and outdoors. Experimental results show that, with the autonomously computed threshold, the proposed method provides 93.78% and 89.8% accuracy for detecting falls with indoor and outdoor experiments, respectively. Koray Ozcan, Senem Velipasalar, Pramod K. Varshney |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2016 | Autonomous altitude measurement and landing area detection for indoor UAV applicationsabstractFully autonomous navigation of unmanned vehicles, without relying on pre-installed tags or markers, still remains a challenge especially for GPS-denied areas and complex indoor environments. Robust altitude control and safe landing zone detection are two important tasks for indoor unmanned aerial vehicle (UAV) applications. In this paper, a novel approach is proposed for indoor UAVs to control their altitudes, and autonomously detect safe landing zones without relying on any markers, special setups, or assuming that the environment is known. The proposed method employs both depth data and RGB images to detect and also track the safe landing zones. Burak Kakillioglu, Senem Velipasalar |
AVSS | 2 |
| 2016 | Throughput of Hybrid-ARQ Chase Combining with ON-OFF Markov Arrivals under QoS ConstraintsabstractIn this paper, throughput of hybrid automatic repeat request (HARQ) schemes is studied in the presence of Markovian data arrivals and statistical queuing constraints. In particular, two queuing models are considered. Specifically, when outage occurs, the transmitter keeps the packet, lowers its priority, and attempts to retransmit it later in the first queue model while the packet is discarded and removed from the buffer in the second queue model. The throughput is investigated when outage constraints, statistical queuing constraints and deadline constraints are imposed. The deadline constraint provides a limitation on the number of retransmissions. Under these assumptions, throughput characterizations are obtained for HARQ chase combining (CC) scheme with three types of Markovian sources, namely the ON-OFF discrete-time and fluid Markov sources and Markov modulated Poisson source (MMPS). Our analytical results are verified via Monte Carlo simulations. In the numerical results, the impact of source randomness, deadline constraints, outage probability and queuing constraints on the throughput is analyzed. Yi Li 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
GLOBECOM | 3 |
| 2016 | Device-to-device communication in cellular networks under statistical queueing constraintsabstractDevice-to-device (D2D) communication underlaid with cellular networks is a new paradigm, proposed to enhance the performance of cellular networks. By allowing a pair of D2D users to communicate directly and share the same spectral resources with the cellular users, D2D communication can achieve higher spectral efficiency, improve the energy efficiency, and lower the traffic delay. In this paper, transmission mode selection and resource allocation in a time-division multiplexed (TDM) cellular network with one cellular user, one base station, and a pair of D2D users is investigated under rate and queueing constraints. In particular, four possible modes are considered, namely the cellular mode, dedicated mode, uplink reuse mode, and downlink reuse mode. Using tools from stochastic network calculus, the system throughput under statistical queueing constraints is formulated, efficient resource allocation algorithms for all possible modes are proposed, and the influence of the positions of each node and the queueing constraints is analyzed via numerical results. Scenarios and conditions for different modes to be optimal in the sense of maximizing the sum-throughput are identified. Yi Li 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
ICC | 3 |
| 2016 | Doorway detection for autonomous indoor navigation of unmanned vehiclesabstractFully autonomous navigation of unmanned vehicles, without relying on pre-installed tags or markers, still remains a challenge for GPS-denied areas and complex indoor environments. Doors are important for navigation as the entry/exit points. A novel approach is proposed to autonomously detect™ doorways by using the Project Tango platform. We first detect the candidate door openings from the 3D point cloud, and then use a pre-trained detector on corresponding RGB image regions to verify if these openings are indeed doors. We employ Aggregate Channel Features for detection, which are computationally efficient for real-time applications. Since detection is only performed on candidate regions, the system is more robust against false positives. The approach can be generalized to recognize windows, some architectural structures and obstacles. Experiments show that the proposed method can detect open doors in a robust and efficient manner. Burak Kakillioglu, Koray Ozcan, Senem Velipasalar |
ICIP | 3 |
| 2016 | Robust footstep counting and traveled distance calculation by mobile phones incorporating camera geometryabstractMost available approaches for step counting rely on accelerometer data, and thus are prone to over-counting. In addition, most existing devices calculate the traveled distance based on the counted number of steps and a preset stride length. We present a robust and autonomous method for counting steps and tracking and calculating stride length by using accelerometer, gravity sensor and camera data from smart phones. To provide higher precision, instead of using a preset step and/or stride length, the proposed method calculates the distance traveled with each step by using the camera data. If camera is tilted significantly, the angle data obtained from the gravity sensor is used to account for camera geometry and increase the precision of the calculated step length. Experiments are performed with different subjects and the proposed method is compared with accelerometer-based step counter apps. The results show that incorporating camera geometry increases the accuracy, and the proposed method provides the lowest average error rate in number of steps taken and the calculated traveled distance. Yantao Lu, Senem Velipasalar |
ICIP | 2 |
| 2016 | Multimedia transmission over device-to-device wireless linksabstractThis paper studies the performance of hierarchical modulation-based image transmission in device-to-device (D2D) cellular wireless networks under constraints on both transmit and interference power levels. Hierarchical quadrature amplitude modulation (HQAM) is considered in which high priority (HP) data is protected more than low priority (LP) data. In this setting, closed-form bit error rate (BER) expressions for HP data and LP data are derived over multiple Rayleigh fading subchannels in 3 different transmission modes. The optimal power control that minimizes weighted sum of average BERs of HP bits and LP bits or its upper bound subject to average transmit power and average interference power constraints is derived. Performance comparisons of image transmission in 3 different modes are carried out, and the proposed power control strategies are evaluated in terms of the BERs and received data quality. Chuang Ye, Mustafa Cenk Gursoy, Senem Velipasalar |
ICME | 3 |
| 2016 | Throughput of two-hop wireless channels with queueing constraints and finite blocklength codesabstractIn this paper, throughput of two-hop wireless relay channels is studied in the finite blocklength regime. Half-duplex relay operation, in which the source node initially sends information to the intermediate relay node and the relay node subsequently forwards the messages to the destination, is considered. It is assumed that all messages are stored in buffers before being sent through the channel, and both the source node and the relay operate under statistical queueing constraints. After characterizing the transmission rates in the finite blocklength regime, the system throughput is formulated via queueing analysis. Subsequently, several properties of the throughput function in terms of system parameters are identified, and an efficient algorithm is proposed to maximize the throughput. Interplay between throughput, queueing constraints, relay location, time allocation, and code blocklength is investigated through numerical results. Yi Li 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
ISIT | 3 |
| 2016 | Scheduling in D2D Underlaid Cellular Networks with Deadline ConstraintsabstractIn this paper, we develop a scheduling algorithm for device-to-device (D2D) cellular networks with deadline constraints via the convex delay cost approach. At the beginning of each time slot, the algorithm allocates all available channels to the users, and each user can choose to transmit in different modes. After characterizing the transmission rates and defining the utility for each possible scheduling decision, we propose power optimization algorithms to maximize the utility for each type of decision. Our scheduling algorithm allocates each channel according to the decision that provides the maximum utility value, and it manages mode selection, channel allocation and power optimization. Via simulation results, we discuss the parameter selection for our algorithm and verify the performance improvements by allowing D2D users to share channels with other users. Yi Li 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
VTC Fall | 3 |
| 2016 | Energy Efficiency of Hybrid-ARQ Under Statistical Queuing ConstraintsabstractIn this paper, energy efficiency of hybrid automatic repeat request (HARQ) schemes with statistical queuing constraints is studied for both constant-rate and random Markov arrivals by characterizing the minimum energy per bit and wideband slope. In particular, two queuing models are considered. Specifically, when outage occurs, the transmitter keeps the packet, lowers its priority, and attempts to retransmit it later in the first queue model, while the packet is discarded and removed from the buffer in the second queue model. For both models, energy efficiency is investigated when outage constraints, statistical queuing constraints, and deadline constraints are imposed. The deadline constraint provides a limitation on the number of retransmissions or equivalently the number of HARQ rounds. Under these assumptions, closed-form expressions are obtained for the minimum energy per bit and wideband slope for HARQ with chase combining, and comparisons among different arrival models are made. For instance, it is shown that stricter queuing constraints and more bursty sources degrade the energy efficiency by lowering the wideband slope. In the numerical results, analytical characterizations are verified through simulations. Moreover, the impact of source variations/burstiness, deadline constraints, outage probability, and queuing constraints on the energy efficiency is analyzed. Yi Li 0007, Gozde O. Sahinoglu, Mustafa Cenk Gursoy, Senem Velipasalar |
IEEE Trans. Commun. | 4 |
| 2016 | On the Throughput of Multi-Source Multi-Destination Relay Networks With Queueing ConstraintsabstractIn this paper, the throughput of relay networks with multiple source-destination pairs under queueing constraints has been investigated for both variable-rate and fixed-rate schemes. When channel side information (CSI) is available at the transmitter side, transmitters can adapt their transmission rates according to the channel conditions, and achieve the instantaneous channel capacities. In this case, the departure rates at each node have been characterized for different system parameters, which control the power allocation, time allocation, and decoding order. In the other case of no CSI at the transmitters, a simple automatic repeat request (ARQ) protocol with fixed rate transmission is used to provide reliable communication. Under this ARQ assumption, the instantaneous departure rates at each node can be modeled as an ON-OFF process, and the probabilities of ON and OFF states are identified. With the characterization of the arrival and departure rates at each buffer, stability conditions are identified, and an effective capacity analysis is conducted for both cases to determine the system throughput under statistical queueing constraints. In addition, for the variable-rate scheme, the concavity of the sum rate is shown for certain parameters, helping to improve the efficiency of parameter optimization. Finally, through numerical results, the influence of system parameters and the behavior of the system throughput are identified. Yi Li 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
IEEE Trans. Wirel. Commun. | 3 |
| 2015 | Throughput and mode selection in two-way MIMO systems under queuing constraintsabstractIn this paper, the throughput of and mode selection between half-duplex and full-duplex modes are studied in two-way multiple-input multiple-output (MIMO) systems operating under statistical queuing constraints. In particular, the effective capacity of these systems is determined in order to identify the throughput under constraints on the buffer overflow probability. In the low signal-to-noise ratio (SNR) regime, the optimal input covariance matrices that achieve the minimum energy per bit of the system are investigated. Full-duplex mode is found to have better performance at low SNRs and short distances, while half-duplex mode outperforms full-duplex operation at high SNR levels and long distances. Additionally, in the numerical results, the influence of the self-interference cancelation parameter and QoS exponent on the throughput is analyzed. Yi Li 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
ICC | 3 |
| 2015 | Scale estimation with difference of ordered residualsabstractMultiple model estimation is an important problem in computer vision. Through estimation, one can detect important structural information in an image. A crucial step in multiple model estimation is the ability to dichotomize inliers of a model from outliers. This paper proposes a novel technique for estimating the scale of a model. In contrast to previous adaptive scale estimate works, our method removes the need for user provided input. We achieve accurate scale estimation through consecutive inspection of the ordered residuals. Our results show the ability of the proposed scale estimate metric to maintain accurate scale estimation even with over 90% outliers present in the data. Likewise, we also apply our scale estimator with multiple model estimation problems for detecting planes and two-view motions, demonstrating the ability of our approach to accurately estimate scale in real application oriented scenarios. Maria Scalzo-Cornacchia, Senem Velipasalar |
ICIP | 2 |
| 2015 | On the throughput of ARQ over multiple-access relay fading channels with queueing constraintsabstractIn this paper, the throughput under queuing constraints in multiple-access relay fading channels achieved by fixed-rate transmissions and a simple automatic repeat request (ARQ) protocol is investigated in the absence of channel side information (CSI) at the transmitters. Transmission is considered to be either in the ON or OFF state, depending on the reliability of the reception, and retransmissions are triggered by the ARQ protocol in the OFF state. The probabilities of these transmission states are identified. Through stability and queuing analysis, feasible set of system parameters is determined, and the maximum constant arrival rates at source nodes, which can be supported by the system under queuing constraints, are characterized in terms of the fixed transmission rates, state probabilities, and quality of service parameters. Yi Li 0007, Mustafa Cenk Gursoy, Senem Velipasalar |
WCNC | 3 |
| 2015 | Image and video transmission in cognitive radio systems under sensing uncertaintyabstractThis paper studies the performance of hierarchical-modulation-based image and video transmission in cognitive radio systems with imperfect channel sensing results under constraints on both transmit and interference power. Data intended for transmission is first compressed via source coding techniques and then divided into two priority classes, namely high priority (HP) data and low priority (LP) data, by taking into consideration the unequal importance of bits in the output codestream. After dividing the compressed data into packets of equal size, turbo coding is applied. Finally, the resulting packets are modulated using hierarchical quadrature amplitude modulation (HQAM). In this setting, closed-form bit error probability expressions for HP data and LP data are derived over Nakagami-m fading channels in the presence of sensing errors. Subsequently, the effects of probabilities of detection and false alarm on error rate performance of cognitive transmissions are evaluated. In addition, tradeoffs between the number of retransmissions and peak signal-to-noise ratio (PSNR) quality are analyzed numerically. Moreover, performance comparisons of multimedia transmission with conventional QAM and hierarchical QAM are carried out in terms of the received data quality and number of retransmissions. Chuang Ye, Gozde O. Sahinoglu, Mustafa Cenk Gursoy, Senem Velipasalar |
WCNC | 4 |
| 2015 | Mobile Standards-Based Traffic Light Detection in Assistive Devices for Individuals with Color-Vision DeficiencyabstractConsidering the substantial population affected by some form of color-vision deficiency (CVD), reliable traffic control signal head light detection is an important problem for driver-assistance systems. While a large number of technologies can be used to localize traffic lights, without drastic changes in infrastructure, only visual information can be used in identifying the status of the light. In addition, traffic light detection is not currently integrated into any driver-assistance systems, making driving for individuals with CVD (where permitted) dangerous to other drivers, pedestrians, and themselves. This paper presents a robust, traffic-standards-based, and computationally efficient method for detecting the status of the traffic lights without relying on Global Positioning System, lidar, radar information, or prior (map-based) knowledge. To the extent of our knowledge, this is the first work to use official Institute of Transportation Engineers (U.S.) and British Standards Institute (European Union) standards for defining traffic light colors, as well as integrating a number of fail-safe mechanisms designed to prevent erroneous detection. The algorithm can be easily ported over to an embedded smart camera platform and used as a windshield-mounted driver-assistance device by individuals with CVD. The system can accurately identify the status of the light at 400 ft away from the intersection, reliably detecting solid, faulty, arrow, and high-visibility signal lights. Over 50 h of video (over 2000 intersections) were tested with the system, containing intersections with one to four traffic lights, governing different lanes of traffic, with 97.5% accuracy of solid light detection. Akhan Almagambetov, Senem Velipasalar, Assel Baitassova |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2014 | Camera motion detection for mobile smart cameras using segmented edge-based optical flowabstractDetermining camera motion is a challenging task in applications involving mobile smart cameras. With widespread use of cameras in mobile applications, analyzing motion-based information have become important. Optical flow has been a popular technique in determining camera motion. However, the use of traditional optical flow techniques can be computationally quite expensive and impractical for embedded smart cameras with limited processing power, The aim of this paper is to provide an effective and computationally efficient optical flow technique to determine the camera motion direction. This technique is based on the segmentation of edge features, and has been implemented on an actual embedded platform. We will show that the systematic segmentation of edge features not only reduces computation time drastically, but also provides sufficient details in determining basic camera motion patterns. Anvith Katte Mahabalagiri, Koray Ozcan, Senem Velipasalar |
AVSS | 3 |
| 2014 | Energy efficient image transmission using wireless embedded smart camerasabstractWireless embedded smart cameras provide significant advantages for surveillance applications, due to their mobility and flexibility in deployment. Many surveillance applications rely on data exchange and image transmission through the network. Moreover, with the widespread use of smart phones and social media, image exchange through crowdsourcing can provide invaluable information for surveillance and investigation purposes. However, transmission of images consumes significant amount of energy, and energy is a limited resource for wireless embedded devices. So far, not much attention has been paid to the energy consumption of a static or mobile node during image transmission for different scenarios. In this paper, we present a new approach for image transmission, which employs a relay node and decreases the overall energy consumption of the sender and the relay nodes. We have performed the experiments with actual wireless embedded smart cameras (CITRIC), with TelosB motes, for a static as well as a mobile scenario. We have compared the energy consumption of the proposed method with that of a traditional relay approach, and that of not using a relay node at all. Experimental results show that, with the proposed relay approach, the energy consumption of image transmission can be reduced by up to 16% in a mobile scenario. Yu Zheng 0016, Chuang Ye, Senem Velipasalar, Mustafa Cenk Gursoy |
AVSS | 3 |
| 2014 | Autonomous multi-scale object detection with hough forestsabstractThe objective of this work is to detect a class of objects in images or video using multi-scale voting with random Hough Forests. Hough Forests have several nice properties, including that an implicit shape model is automatically learned from cropped images of a particular class of object and the voting induced by a Hough technique allows the detection method to handle partial occlusions. Typical Hough Forest voting is however scale sensitive when it comes to both training and testing. Currently, searching for multiple scales for an object size is achieved by re-running the detection routine for a given image at numerous manually provided input scales. This work will demonstrate that manually input scale parameters can lower detection rates if all scales in the test set are not accounted for. The novelty of our proposed work is in the creation of an autonomous scale estimation and multi-scalar detection Hough Forest voting technique. The technique proposed to accomplish the automatic scale estimation is to view votes as not votes for discrete locations, but rather as voting rays. The intersection of these rays can then be used to automatically determine the estimated object's center and scale. Maria Scalzo-Cornacchia, Senem Velipasalar |
ICIP | 2 |
| 2014 | Agglomerative clustering for feature point groupingabstractThe objective of this paper is to group feature points on different planes as a means of semantic image segmentation and understanding. The methodology is based on the ability to estimate planar homographies from grouped feature points spanning different unknown number of planes. This paper proposes an alternative to the J-linkage method, which was shown to have benefits in terms of accuracy over other multiple model estimation techniques. J-linkage is an agglomerative clustering technique that uses a set representation of support for a set of possible planar homographies and the Jaccard measure to determine the distance between support sets. The technique proposed in this paper uses a frequency vector to represent the support for a model. This formulation promotes clustering even in the presence of noise and prevents the order in which agglomerative clustering is performed from influencing the results. The feature vector representation requires an alternative distance measure to Jaccard to be exercised, that of cosine similarity. Hence, the method proposed here is called C-linkage. The results show that, compared to the J-linkage method, the proposed technique correctly classifies more points on each plane, and results in less over-segmentation while providing higher Normalized Mutual Information scores for a range of multiple model estimation problems on different datasets. Maria Scalzo-Cornacchia, Senem Velipasalar |
ICIP | 2 |
| 2014 | Energy efficiency in fading relay channels under secrecy and QoS constraintsabstractTransmission over a half-duplex relay channel with secrecy and quality-of-service (QoS) constraints is studied. It is assumed that data is stored in buffers prior to transmission and transmitters operate with a constraint on the buffer overflow probability. Using the effective capacity formulation, secrecy throughput is derived for the half duplex two-hop fading relay system operating in the presence of an eavesdropper. The scenario without security considerations is also addressed. For both scenarios, minimum energy per bit is obtained. The impact of QoS exponents and channel correlation on the throughput is investigated via numerical analysis. Taking into account the path loss effect, the effect of the relay and eavesdropper locations on the throughput and energy efficiency is studied in a simple linear network. Mustafa Ozmen, Chuang Ye, Mustafa Cenk Gursoy, Senem Velipasalar |
PIMRC | 4 |
| 2014 | Distributed wide-area multi-object tracking with non-overlapping camera views
Youlu Wang, Senem Velipasalar, Mustafa Cenk Gursoy |
Multim. Tools Appl. | 2 |
| 2013 | An improved evolutionary algorithm for fundamental matrix estimationabstractThe estimation of the fundamental matrix is an important problem in epipolar geometry. Many estimation methods have been proposed before, including the eight-point algorithm, Simple Evolutionary Agent (SEA) and RANSAC. In this paper, we investigate the evolutionary agent-based algorithm for fundamental matrix estimation, and present a new algorithm that improves the existing evolutionary algorithm both accuracy- and efficiency-wise. The model focuses on selecting a best combination of input points to compute the fundamental matrix via the eight-point algorithm. To improve the existing algorithm, our new model holds competition over all agents for population control and evolutionary experience accumulation. In addition to a larger competition scope, we add the outlier elimination mechanism, which greatly accelerates the algorithm. New parameters are introduced to control the convergence more efficiently. The improved algorithm achieves lower computation load and more accurate results. A general analysis about parameter selection is also provided. Yi Li 0007, Senem Velipasalar, Mustafa Cenk Gursoy |
AVSS | 2 |
| 2013 | Energy-aware and robust task (re)assignment in embedded smart camera networksabstractMulti-camera multi-object tracking problem can be regarded as a multi-player game by adopting a game theoretical approach. In embedded vision sensor networks, energy, processing power and bandwidth are limited, and should be efficiently used. In this paper, in addition to dynamic grouping of the camera nodes, we focus on the (re)assignment of object tracking tasks by simultaneously considering energy levels, processing loads and accuracy/reliability of nodes in utility calculation. Instead of using a predetermined period to perform auctions, nodes trigger the reassignment process in an event-driven manner. Four scenarios are used for triggering reassignment, namely (i) new-object entry, (ii) object lost or exit, (iii) critical energy level or energy decrease rate, and (iv) critical target location and resolution. We also analyze the communication cost in terms of the number of messages sent between the cameras. We have performed experiments with different number of cameras and targets, and varying target trajectories and camera topology. We have computed the lifetime of the network with and without consideration of the energy levels in the task (re)assignment. We have also compared the number of messages sent with periodic reassignment and with the proposed event-driven triggering mechanism. The simulation results show a significant increase in the lifetime of the network as well as a decrease in the number of messages that are sent when the proposed approach is employed. Chuang Ye, Yu Zheng 0016, Senem Velipasalar, Mustafa Cenk Gursoy |
AVSS | 3 |
| 2013 | On the throughput of two-way relay systems under queueing constraintsabstractIn this paper, throughput of two-way relaying under buffer constraints is studied. In the two-way relay system, source nodes initially send their messages to the relay in the multiple-access phase. Relay decodes and stores the messages from different sources in different buffers and subsequently broadcasts a superimposed signal. It is assumed that both source nodes and the relay operate in the presence of statistical queueing constraints. Under these assumptions, arrival rates that can be supported in this system are investigated through the logarithmic moment generating functions of the arrival and service processes. In particular, after identifying the service rates in the multiple-access and broadcast phases and addressing the stability conditions, characterizations of the maximum arrival rates are provided in terms of system resource allocation parameters, signal-to-noise ratios, and quality-of-service exponents. Impact of different parameters on the performance is investigated through numerical results. Yi Li 0007, Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
GLOBECOM | 4 |
| 2013 | Fall detection and activity classification using a wearable smart cameraabstractRobust detection of events and activities, such as falling, sitting and lying down, is a key to a reliable elderly activity monitoring system. While fast and precise detection of falls is critical in providing immediate medical attention, other activities like sitting and lying down can provide valuable information for early diagnosis of potential health problems. In this paper, we present a fall detection and activity classification system using wearable cameras. Since the camera is worn by the subject, monitoring extends to wherever the subject may go. Furthermore, since the captured frames are not of the subject, privacy is preserved. We present an improved fall detection algorithm employing histograms of edge orientations and strengths, and propose an optical flow-based method for activity classification. Trials were performed on five different subjects wearing a camera on their waist, each performing 40 different activities. Experimental results show the success of the proposed method. Koray Ozcan, Anvith Katte Mahabalagiri, Senem Velipasalar |
ICME | 3 |
| 2013 | Achievable Throughput Regions of Fading Broadcast and Interference Channels under QoS ConstraintsabstractTransmission over fading broadcast and interference channels in the presence of quality of service (QoS) constraints is studied. Effective capacity, which provides the maximum constant arrival rate that a given service process can support while satisfying statistical QoS constraints, is employed as the performance metric. In the broadcast scenario, the effective capacity region achieved with superposition coding and successive interference cancellation is identified and is shown to be convex. Subsequently, optimal power control policies that achieve the boundary points of the effective capacity region are investigated, and an algorithm for the numerical computation of the optimal power adaptation schemes for the two-user case is provided. In the interference channel model, achievable throughput regions are determined for three different strategies, namely treating interference as noise, time division with power control and simultaneous decoding. It is demonstrated that as in Gaussian interference channels, simultaneous decoding expectedly performs better (i.e., supports higher arrival rates) when interfering links are strong, and treating interference as noise leads to improved performance when the interfering cross links are weak while time-division strategy should be preferred in between. When the QoS constraints become more stringent, it is observed that the sum-rates achieved by different schemes all diminish and approach each other, and time division with power control interestingly starts outperforming others over a wider range of cross-link strengths. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
IEEE Trans. Commun. | 3 |
| 2013 | Effective Capacity of Two-Hop Wireless Communication SystemsabstractA two-hop wireless communication link in which a source sends data to a destination with the aid of an intermediate relay node is studied. It is assumed that there is no direct link between the source and the destination, and the relay forwards the information to the destination by employing the decode-and-forward scheme. Both the source and intermediate relay nodes are assumed to operate under statistical quality of service (QoS) constraints imposed as limitations on the buffer overflow probabilities. The maximum constant arrival rates that can be supported by this two-hop link in the presence of QoS constraints are characterized by determining the effective capacity of such links as a function of the QoS parameters and signal-to-noise ratios at the source and relay, and the fading distributions of the links. The analysis is performed for both full-duplex and half-duplex relaying. Through this study, the impact upon the throughput of having buffer constraints at the source and intermediate relay nodes is identified. The interactions between the buffer constraints in different nodes and how they affect the performance are studied. The optimal time-sharing parameter in half-duplex relaying is determined, and performance with half-duplex relaying is investigated. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
IEEE Trans. Inf. Theory | 3 |
| 2012 | A Robust Algorithm for the Detection of Vehicle Turn Signals and Brake LightsabstractRobust and lightweight detection of alert signals of front vehicle, such as turn signals and brake lights, is extremely critical, especially in autonomous vehicle applications. Even with cars that are driven by human beings, automatic detection of these signals can aid in the prevention of otherwise deadly accidents. This paper presents a novel, robust and lightweight algorithm for detecting brake lights and turn signals both at night and during the day. The proposed method employs a Kalman filter to reduce the processing load. Much research is focused only on the detection of brake lights at night, but our algorithm is able to detect turn signals as well as brake lights under any lighting conditions with high accuracy rates. Mauricio Casares, Akhan Almagambetov, Senem Velipasalar |
AVSS | 3 |
| 2012 | Lightweight and Robust Shadow Removal for Foreground DetectionabstractBackground subtraction is a commonly used method to detect moving objects from videos captured by static cameras. However, shadows and reflections significantly affect the output of background subtraction algorithms, and distort the shape of the objects obtained as a result. Thus, shadow detection and removal is a crucial post-processing step to perform accurate object tracking required by different applications. We present a lightweight method to detect and remove shadows as well as reflection effects in indoor and outdoor environments by using spatial and spectral features. This method incorporates an adaptive way to set thresholds to avoid preset numbers. We present a comparison of the outputs we obtained with those of several other methods. The experimental results demonstrate the success of the proposed algorithm. Anuja Gawde, Kedar Joshi, Senem Velipasalar |
AVSS | 3 |
| 2012 | Autonomous tracking of vehicle rear lights and detection of brakes and turn signalsabstractAutomatic detection of vehicle alert signals is extremely critical in autonomous vehicle applications and collision avoidance systems, as these detection systems can help in the prevention of deadly and costly accidents. In this paper, we present a novel and lightweight algorithm that uses a Kalman filter and a codebook to achieve a high level of robustness. The algorithm is able to detect braking and turning signals of the vehicle in front both during the daytime and at night (daytime detection being a major advantage over current research), as well as correctly track a vehicle despite changing lanes or encountering periods of no or low-visibility of the vehicle in front. We demonstrate that the proposed algorithm is able to detect the signals accurately and reliably under different lighting conditions. Akhan Almagambetov, Mauricio Casares, Senem Velipasalar |
CISDA | 3 |
| 2012 | Throughput regions for fading interference channels under statistical QoS constraintsabstractCommunication over fading interference channels in the presence of statistical quality of service (QoS) constraints is considered. Effective capacity, which provides the maximum constant arrival rate that a given service process can support while satisfying statistical queueing constraints, is employed as the performance metric. In a two-user and buffer constrained setting, arrival rate regions that can be supported in the fading interference channel are studied. More specifically, for three different strategies, namely treating interference as noise, time division with power control and simultaneous decoding, achievable throughput regions are determined. It is demonstrated that as in Gaussian interference channels, simultaneous decoding expectedly performs better (i.e., supports higher arrival rates) when interfering links are strong, and treating interference as noise leads to improved performance when the interfering cross links are weak while time-division strategy should be preferred in between. When the QoS constraints become more stringent, it is observed that the sum-rates achieved by different schemes all diminish and approach each other, and time division with power control interestingly starts outperforming others over a wider range of cross-link strengths. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
GLOBECOM | 3 |
| 2012 | Energy efficiency in multiaccess fading channels under QoS constraintsabstractIn this paper, transmission over multiaccess fading channels under QoS constraints is studied in the low power regime. QoS constraints are imposed as limitations on the buffer violation probability. The effective capacity, which characterizes the maximum constant arrival rates in the presence of such statistical QoS constraints, is employed as the performance metric. The minimum received bit energy levels and wideband slope regions are characterized for different transmission and reception strategies, namely time-division multiple-access (TDMA), superposition coding with fixed decoding order, and superposition coding with variable decoding order. It is shown that the minimum received bit energies achieved by these different strategies are the same and independent of the QoS constraints. When wideband slope regions are considered, the suboptimality of TDMA with respect to superposition schemes is shown. For the case of superposition coding, it is proven that varying the decoding order at the receiver with the fading realizations does not enlarge the wideband slope region. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
ICC | 3 |
| 2012 | Transmission Strategies in Multiple-Access Fading Channels With Statistical QoS ConstraintsabstractEffective capacity, which provides the maximum constant arrival rate that a given service process can support while satisfying statistical queueing constraints, is analyzed in a multiuser scenario. In particular, the effective capacity region of fading multiple-access channels in the presence of quality of service (QoS) constraints is studied. Perfect channel side information is assumed to be available at both the transmitters and the receiver. It is initially assumed that the transmitters send the information at a fixed power level and, hence, do not employ power control policies. Under this assumption, the performance achieved by superposition coding with successive decoding techniques is investigated. It is shown that varying the decoding order with respect to the channel states can significantly increase the achievable throughput region. In the two-user case, the optimal decoding strategy is determined for the scenario in which the users have the same QoS constraints. The performance of orthogonal transmission strategies is also analyzed. It is shown that for certain QoS constraints, time-division multiple access can achieve better performance than superposition coding if fixed successive decoding order is used at the receiver side. In the subsequent analysis, power control policies are incorporated into the transmission strategies. The optimal power allocation policies for any fixed decoding order over all channel states are identified. For a given variable decoding-order strategy, the conditions that the optimal power control policies must satisfy are determined, and an algorithm that can be used to compute these optimal policies is provided. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
IEEE Trans. Inf. Theory | 3 |
| 2011 | Channel Coding over Multiple Coherence Blocks with Queueing ConstraintsabstractThis paper investigates the performance of wireless systems that employ finite-blocklength channel codes for transmission and operate under queueing constraints in the form of limitations on buffer overflow probabilities. A block fading model, in which fading stays constant in each coherence block and change independently between blocks, is considered. It is assumed that channel coding is performed over multiple coherence blocks. An approximate lower bound on the transmission rate is obtained from Feintein's Lemma. This lower bound is considered as the service rate and is incorporated into the effective capacity formulation, which characterizes the maximum constant arrival rate that can be supported under statistical queuing constraints. Performances of variable-rate and fixed-rate transmissions are studied. The optimum error probability for variable rate transmission and the optimum coding rate for fixed rate transmission are shown to be unique. Moreover, the tradeoff between the throughput and the number of blocks over which channel coding is performed is identified. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
ICC | 3 |
| 2011 | On the Effective Capacity of Two-Hop Communication SystemsabstractIn this paper, two-hop communication between a source and a destination with the aid of an intermediate relay node is considered. Both the source and intermediate relay node are assumed to operate under statistical quality of service (QoS) constraints imposed as limitations on the buffer overflow probabilities. It is further assumed that the nodes send the information at fixed power levels and have perfect channel side information. In this scenario, the maximum constant arrival rates that can be supported by this two-hop link are characterized by finding the effective capacity. Through this analysis, the impact upon the throughput of having buffer constraints at the source and intermediate-hop nodes is identified. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
ICC | 3 |
| 2011 | Wide-area multi-object tracking with non-overlapping camera viewsabstractWe present a system for wide-area multi-object tracking across disjoint camera views. We employ a probabilistic Petri Net-based approach to account for the uncertainties of the vision algorithms (such as unreliable background subtraction, and tracking failure) and to incorporate the available domain knowledge. We combine appearance features of objects as well as the travel-time evidence for target matching and consistent labeling across disjoint camera views. 3D color histogram, Histogram of Oriented Gradients, object size and aspect ratio are used as the appearance features. The distribution of the travel time is modeled by a Gaussian Mixture Model. By incorporating the domain knowledge about the camera configurations and the information about the received packets from other cameras, certain transitions are fired in the probabilistic Petri net. The system is trained to learn different parameters of the matching process. We present wide-area tracking of vehicles as an example where we used three non-overlapping cameras. The first and the third cameras are approximately 150 meters apart from each other with two intersections in the blind region. The results show the success of the proposed method. Youlu Wang, Senem Velipasalar, Mustafa Cenk Gursoy |
ICME | 2 |
| 2011 | Effective capacity region and optimal power control for fading broadcast channelsabstract1Transmission over fading broadcast channels in the presence of quality of service (QoS) constraints is studied. Effective capacity, which provides the maximum constant arrival rate that a given service process can support while satisfying statistical QoS constraints, is employed as the performance metric. The effective capacity region achieved with superposition coding and successive interference cancellation is identified and is shown to be convex. Subsequently, optimal power control policies that achieve the boundary points of the effective capacity region are investigated, and an algorithm for the numerical computation of the optimal power adaptation schemes for the two-user case is provided. Additionally, performance attained with time-division multiplexing (TDM) of messages is studied for comparison with the optimal schemes. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
ISIT | 3 |
| 2011 | Analysis of the accuracy-latency-energy tradeoff for wireless embedded camera networksabstractWireless embedded smart cameras provide flexibility in camera deployment in terms of the locations and number of the cameras. However, these battery-powered embedded vision sensors have very limited energy, memory, and processing power. Energy consumption and latency are two major concerns in wireless embedded camera networks. In multi-camera tracking applications, the amount of data exchanged between cameras has an effect on the tracking accuracy, the energy consumption of the camera nodes and the latency. In this paper, we provide a detailed quantitative analysis of the accuracy-latency-energy tradeoff for overlapping and non-overlapping camera setups when different-sized data packets are transferred in a wireless manner. The experiments have been performed with an actual wireless embedded smart camera network employing CITRIC motes, and performing tracking of objects. Alvaro Pinto, Zhe Zhang 0003, Xin Dong 0008, Senem Velipasalar, Mehmet Can Vuran, Mustafa Cenk Gursoy |
WCNC | 4 |
| 2011 | Energy Efficiency in the Low-SNR Regime under Queueing Constraints and Channel UncertaintyabstractEnergy efficiency of fixed-rate transmissions is studied in the presence of queueing constraints and channel uncertainty. It is assumed that neither the transmitter nor the receiver has channel side information prior to transmission. The channel coefficients are estimated at the receiver via minimum mean-square-error (MMSE) estimation with the aid of training symbols. It is further assumed that the system operates under statistical queueing constraints in the form of limitations on buffer violation probabilities. The optimal fraction of power allocated to training is identified. Spectral efficiency-bit energy tradeoff is analyzed in the low-power and wideband regimes by employing the effective capacity formulation. In particular, it is shown that the bit energy increases without bound in the low-power regime as the average power vanishes. A similar conclusion is reached in the wideband regime if the number of noninteracting subchannels grow without bound with increasing bandwidth. On the other hand, it is proven that if the number of resolvable independent paths and hence the number of noninteracting subchannels remain bounded as the available bandwidth increases, the bit energy diminishes to its minimum value in the wideband regime. For this case, expressions for the minimum bit energy and wideband slope are derived. Overall, energy costs of channel uncertainty and queueing constraints are identified, and the impact of multipath richness and sparsity is determined. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
IEEE Trans. Commun. | 3 |
| 2011 | Adaptive Methodologies for Energy-Efficient Object Detection and Tracking With Battery-Powered Embedded Smart CamerasabstractBattery-powered wireless embedded smart cameras have limited processing power, memory and energy. Since video processing tasks consume considerable amount of energy, it is essential to have lightweight algorithms to increase the energy efficiency of camera nodes. Moreover, just grabbing and buffering a frame require significant amount of energy. Thus, it is not sufficient to only focus on the vision algorithms. Methodologies are needed to determine when and how long a camera can be idle. In this paper, we first present a feedback method for detection and tracking, which provides significant savings in processing time. We take advantage of these savings by sending the microprocessor to idle state at the end of processing a frame. Then, we present an adaptive methodology that can send the camera to idle state not only when the scene is empty but also when there are target objects. Idle state duration is adaptively changed based on the speeds of tracked objects. We then introduce a combined method that employs the feedback method and the adaptive methodology together, and provides further savings in energy consumption. We provide a detailed comparison of these methods, and present experimental results showing the gains in processing time as well as the significant savings in energy consumption and increase in battery life. Mauricio Casares, Senem Velipasalar |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Secure Wireless Communication and Optimal Power Control Under Statistical Queueing ConstraintsabstractIn this paper, secure transmission of information over fading broadcast channels is studied in the presence of statistical queueing constraints. Effective capacity is employed as a performance metric to identify the secure throughput of the system, i.e., effective secure throughput. It is assumed that perfect channel side information (CSI) is available at both the transmitter and the receivers. Initially, the scenario in which the transmitter sends common messages to two receivers and confidential messages to one receiver is considered. For this case, the effective secure throughput region, which is the region of constant arrival rates of common and confidential messages that can be supported by the buffer-constrained transmitter and fading broadcast channel, is defined. It is proven that this effective throughput region is convex implying that time-sharing between any two viable transmission and power control strategies results in effective throughput values inside the region. Then, the optimal power control policies that achieve the boundary points of the effective secure throughput region are investigated and an algorithm for the numerical computation of the optimal power adaptation schemes is provided. Additionally, the throughput region achieved by time-division multiplexing of common and confidential messages is explored. Subsequently, the special case in which the transmitter sends only confidential messages to one receiver is addressed in more detail. For this case, effective secure throughput is formulated and two different power adaptation policies are studied. These power adaptation policies are compared with the opportunistic ones that are optimal in the absence of quality of service (QoS) constraints. It is shown that opportunistic schemes, in which data transmission with high rates and high power occurs only when the main channel is much better than the eavesdropper channel, are no longer optimal under buffer constraints, and the transmitter should send the data at a certain moderate rate and power even when the main channel strength is comparable to that of the eavesdropper channel to avoid buffer overflows. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2010 | Resource-Efficient Salient Foreground Detection for Embedded Smart Cameras br Tracking FeedbackabstractBattery-powered wireless embedded smart cameras have limited processing power, memory and energy. Since video processing tasks consume significant amount of power,the problem of limited resources becomes even more pronounced, and necessitates designing light-weight algorithms suitable for embedded platforms. In this paper, we present a resource-efficient salient foreground detection and tracking algorithm. Contrary to traditional methods that implement foreground object detection and tracking independently and in a sequential manner, the proposed method uses the feedback from the tracking stage in the foreground object detection. We compare the proposed method with a sequential method on the microprocessor of an embedded smart camera, and present the savings in the processing time and energy consumption and the gain in the lifetime of a battery-powered camera for different scenarios. The presented method provides significant savings in terms of the processing time of a frame. We take advantage of these savings by sending the microprocessor to idle state at the end of processing a frame, and when the scene is empty. Mauricio Casares, Senem Velipasalar |
AVSS | 2 |
| 2010 | Energy Consumption and Latency Analysis for Wireless Multimedia Sensor NetworksabstractEnergy and bandwidth are limited resources in wireless sensor networks, and communication consumes significant amount of energy. When wireless vision sensors are used to capture and transfer image and video data, the problems of limited energy and bandwidth become even more pronounced. Thus, message traffic should be decreased to reduce the communication cost. In many applications, the interest is to detect composite and semantically higher-level events based on information from multiple sensors. Rather than sending all the information to the sinks and performing composite event detection at the sinks or control-center, it is much more efficient to push the detection of semantically high-level events within the network, and perform composite event detection in a peer-to-peer and energy-efficient manner across embedded smart cameras. In this paper, three different operation scenarios are analyzed for a wireless vision sensor network. A detailed quantitative comparison of these operation scenarios are presented in terms of energy consumption and latency. This quantitative analysis provides the motivation for, and emphasizes (1) the importance of performing high-level local processing and decision making at the embedded sensor level and (2) need for peer-to-peer communication solutions for wireless multimedia sensor networks. Alvaro Pinto, Zhe Zhang 0003, Xin Dong 0008, Senem Velipasalar, Mehmet Can Vuran, Mustafa Cenk Gursoy |
GLOBECOM | 4 |
| 2010 | Secure Broadcasting over Fading Channels with Statistical QoS ConstraintsabstractIn this paper, the fading broadcast channel with confidential messages is studied in the presence of statistical quality of service (QoS) constraints in the form of limitations on the buffer length. We employ the effective capacity formulation to measure the throughput of the confidential and common messages. We assume that the channel side information (CSI) is available at both the transmitter and the receivers. Considering average power constraints at the transmitter side, we first define the effective secure throughput region, and prove that the throughput region is convex. Then, we obtain the optimal power control policies that achieve the boundary points of the effective secure throughput region. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
GLOBECOM | 3 |
| 2010 | On the Achievable Throughput Region of Multiple-Access Fading Channels with QoS ConstraintsabstractEffective capacity, which provides the maximum constant arrival rate that a given service process can support while satisfying statistical delay constraints, is analyzed in a multiuser scenario. In particular, we study the achievable effective capacity region of the users in multiaccess fading channels (MAC) in the presence of quality of service (QoS) constraints. We assume that channel side information (CSI) is available at both the transmitters and the receiver, and superposition coding technique with successive decoding is used. When the power is fixed at the transmitters, we show that varying the decoding order with respect to the channel state can significantly increase the achievable throughput region. For a two-user case, we obtain the optimal decoding strategy when the users have the same QoS constraints. Meanwhile, it is shown that time-division multiple-access (TDMA) can achieve better performance than superposition coding with fixed successive decoding order at the receiver side for certain QoS constraints. For power and rate adaptation, we determine the optimal power allocation policy with fixed decoding order at the receiver side. Numerical results are provided to demonstrate our results. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
ICC | 3 |
| 2010 | Real-time distributed tracking with non-overlapping camerasabstractWe present a real-time distributed system for tracking with non-overlapping camera views. Each camera performs multi-object tracking, and cameras communicate with each other in a peer-to-peer manner for consistent labeling. To match objects across non-overlapping views, we employ multiple features, namely color histogram, height, travel time and speed. First, camera configuration and reference values of different features are learned in the training phase. Then, we combine multiple evidences by computing an overall similarity score, which is a weighted sum of the similarity scores of different features. Communication and frame processing run in parallel and share memory. Experimental results show the success of the presented system in real-time tracking with non-overlapping cameras and in handling merge cases. Youlu Wang, Li He 0001, Senem Velipasalar |
ICIP | 3 |
| 2010 | Secure communication over fading channels with statistical QoS constraintsabstractIn this paper, secure transmission of information over an ergodic fading channel is studied in the presence of statistical quality of service (QoS) constraints.We employ effective capacity to measure the secure throughput of the system, i.e., effective secure throughput. We assume that the channel side information (CSI) of the main and the eavesdropper channels is available at the transmitter side. Under this assumption, we investigate the optimal power control policies that maximize the effective secure throughput. In particular, it is noted that opportunistic transmission is no longer optimal and the transmitter should not wait to send the data at a high rate until the main channel is much better than the eavesdropper channel. Moreover, it is shown that the benefits of adapting the power with respect to the CSI of both the eavesdropper and main channels rather than the CSI of only the main channel diminish as QoS constraints become more stringent. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
ISIT | 3 |
| 2010 | A Noncooperative Power Control Game in Multiple-Access Fading Channels with QoS ConstraintsabstractIn this paper, a game-theoretic analysis for the resource allocation policies in fading multiple-access channels (MAC) in the presence of quality of service (QoS) constraints is performed. Effective capacity, which provides the maximum constant arrival rate, or throughput, that a given service process can support while satisfying statistical delay constraints, is considered in a multiuser scenario. We assume that the channel side information (CSI) is available at both the receiver and transmitters, and the transmitters are selfish, rational with certain QoS constraints and average power limitations. Without the aid of the receiver, we prove that there is always a unique admissible Nash equilibrium of the noncooperative power control game. The Nash equilibrium of the power control game is proved to be always inside the rate region where successive decoding techniques are used at the receiver. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
WCNC | 3 |
| 2010 | Light-weight salient foreground detection for embedded smart cameras
Mauricio Casares, Senem Velipasalar, Alvaro Pinto |
Comput. Vis. Image Underst. | 2 |
| 2010 | Detection of user-defined, semantically high-level, composite events, and retrieval of event queries
Senem Velipasalar, Lisa M. Brown, Arun Hampapur |
Multim. Tools Appl. | 1 |
| 2010 | Cooperative Object Tracking and Composite Event Detection With Wireless Embedded Smart CamerasabstractEmbedded smart cameras have limited processing power, memory, energy, and bandwidth. Thus, many system- and algorithm-wise challenges remain to be addressed to have operational, battery-powered wireless smart-camera networks. We present a wireless embedded smart-camera system for cooperative object tracking and detection of composite events spanning multiple camera views. Each camera is a CITRIC mote consisting of a camera board and wireless mote. Lightweight and robust foreground detection and tracking algorithms are implemented on the camera boards. Cameras exchange small-sized data wirelessly in a peer-to-peer manner. Instead of transferring or saving every frame or trajectory, events of interest are detected. Simpler events are combined in a time sequence to define semantically higher-level events. Event complexity can be increased by increasing the number of primitives and/or number of camera views they span. Examples of consistently tracking objects across different cameras, updating location of occluded/lost objects from other cameras, and detecting composite events spanning two or three camera views, are presented. All the processing is performed on camera boards. Operating current plots of smart cameras, obtained when performing different tasks, are also presented. Power consumption is analyzed based upon these measurements. Youlu Wang, Senem Velipasalar, Mauricio Casares |
IEEE Trans. Image Process. | 2 |
| 2009 | Cooperative Object Tracking and Event Detection with Wireless Smart CamerasabstractWireless embedded smart cameras not only capture images, but also can perform processing and communication. However,many system- and algorithm-wise challenges remain to be addressed to have operational, battery-powered wireless smart-camera networks, since they have limited processing power, memory, energy and bandwidth. In this paper, we present a wireless, embedded smart camera system for cooperative object tracking and event detection, wherein each camera platform consists of a camera board and a wireless mote. Light-weight background subtraction and tracking algorithms are implemented and run on the camera boards. Cameras communicate in a peer-to-peer manner over wireless links to exchange data, and thus to consistently track objects. In a wireless smart camera system, transferring large amounts of data between cameras should be avoided, since it requires more power, and incurs more communication delay. In the presented system, cameras exchange small-size packets for communication. Also, with wireless smart cameras, it is not viable to transfer all the captured frames to a base station due to limited resources. Instead, we define events of interest beforehand, and embedded smart cameras save only those portions of the live video capture where the defined event scenario occurs. We present results of tracking and detecting objects entering a region of interest, all of which are performed on the microprocessor of camera boards. We also show examples of consistently tracking objects, moving across different camera views, by wireless data exchange. Youlu Wang, Mauricio Casares, Senem Velipasalar |
AVSS | 3 |
| 2009 | Energy Efficiency of Fixed-Rate Wireless Transmissions under Queueing Constraints and Channel UncertaintyabstractEnergy efficiency of fixed-rate transmissions is studied in the presence of queueing constraints and channel uncertainty. It is assumed that neither the transmitter nor the receiver has channel side information prior to transmission. The channel coefficients are estimated at the receiver via minimum mean-square-error (MMSE) estimation with the aid of training symbols. It is further assumed that the system operates under statistical queueing constraints in the form of limitations on buffer violation probabilities. The optimal fraction of power allocated to training is identified. Spectral efficiency-bit energy tradeoff is analyzed in the low-power and wideband regimes by employing the effective capacity formulation. In particular, it is shown that the bit energy increases without bound in the low-power regime as the average power vanishes. On the other hand, it is proven that if sparse multipath fading with bounded number of independent resolvable paths is experienced, the bit energy diminishes to its minimum value in the wideband regime as the available bandwidth increases. For this case, expressions for the minimum bit energy and wideband slope are derived. Overall, energy costs of channel uncertainty and queueing constraints are identified. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
GLOBECOM | 3 |
| 2009 | Light-weight salient foreground detection with adaptive memory requirementabstractDesigning algorithms, which require less memory and consume less power, is very important for the portability to embedded smart cameras, which have limited resources. We present a light-weight and efficient algorithm for salient foreground detection that is highly robust against lighting variations and non-static backgrounds such as scenes with swaying trees. Contrary to traditional methods, memory requirement for the data saved for each pixel is very small in the proposed algorithm. Moreover, the total memory requirement is adaptive, and is decreased even more depending on the amount of activity in the scene. As opposed to existing methods, we treat each pixel differently based on its history. Instead of requiring the same amount of memory for every pixel, we allocate less memory for stable background pixels. The plot of the required memory at each frame also serves as a tool to find the video portions with high activity. Mauricio Casares, Senem Velipasalar |
ICASSP | 2 |
| 2009 | Frame-level temporal calibration of unsynchronized cameras by using Longest Consecutive Common SubsequenceabstractWe present a computationally efficient and robust method for temporally calibrating video sequences from unsynchronized cameras by using object trajectories. Existing methods remain restricted in terms of their assumptions, and/or they are computationally expensive. To match and align the object trajectories, and thus to recover the frame offset between video sequences, we present an algorithm that is based on the Longest Consecutive Common Subsequence. The candidate frame offsets are obtained from each matched trajectory pair, and then a confidence check is performed. The algorithm is robust against possible errors due to background subtraction and location extraction, and can handle large frame offsets. We present experimental results for different frame offset values on different video sequences, which show the robustness of the algorithm in recovering the frame offsets. We also compare the presented algorithm with our previous work to demonstrate the computational efficiency provided. Youlu Wang, Senem Velipasalar |
ICASSP | 2 |
| 2009 | Energy Efficiency of Fixed-Rate Wireless Transmissions under QoS ConstraintsabstractTransmission over wireless fading channels under quality of service (QoS) constraints is studied when only the receiver has perfect channel side information. Being unaware of the channel conditions, transmitter is assumed to send the information at a fixed rate. Under these assumptions, a two-state (ON-OFF) transmission model is adopted, where information is transmitted reliably at a fixed rate in the ON state while no reliable transmission occurs in the OFF state. QoS limitations are imposed as constraints on buffer violation probabilities, and effective capacity formulation is used to identify the maximum arrival rate that a wireless channel can sustain while satisfying statistical QoS constraints. Energy efficiency is investigated by obtaining the minimum bit energy and wideband slope expressions in both low-power and wideband regimes. The increased energy requirements due to the presence of QoS constraints are quantified. Comparisons with variable-rate/fixed-power and variable-rate/variable-power cases are given. Overall, an energy-delay tradeoff for fixed-rate transmission systems is provided. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
ICC | 3 |
| 2009 | Analysis of Energy Efficiency in Fading Channels under QoS ConstraintsabstractEnergy efficiency in fading channels in the presence of Quality of Service (QoS) constraints is studied. Effective capacity, which provides the maximum arrival rate that a wireless channel can sustain while satisfying statistical QoS constraints, is considered. Spectral efficiency-bit energy tradeoff is analyzed in the low-power and wideband regimes by employing the effective capacity formulation, rather than the Shannon capacity. Through this analysis, energy requirements under QoS constraints are identified. The analysis is conducted under two assumptions: perfect channel side information (CSI) available only at the receiver and perfect CSI available at both the receiver and transmitter. In particular, it is shown in the low-power regime that the minimum bit energy required under QoS constraints is the same as that attained when there are no such limitations. However, this performance is achieved as the transmitted power vanishes. Through the wideband slope analysis, the increased energy requirements at low but nonzero power levels in the presence of QoS constraints are determined. A similar analysis is also conducted in the wideband regime. The minimum bit energy and wideband slope expressions are obtained. In this regime, the required bit energy levels are found to be strictly greater than those achieved when Shannon capacity is considered. Overall, a characterization of the energy-bandwidth-delay tradeoff is provided. Mustafa Cenk Gursoy, Deli Qiao, Senem Velipasalar |
IEEE Trans. Wirel. Commun. | 3 |
| 2009 | The impact of QoS constraints on the energy efficiency of fixed-rate wireless transmissionsabstractTransmission over wireless fading channels under quality of service (QoS) constraints is studied when only the receiver has channel side information. Being unaware of the channel conditions, transmitter is assumed to send the information at a fixed rate. Under these assumptions, a two-state (ON-OFF) transmission model is adopted, where information is transmitted reliably at a fixed rate in the ON state while no reliable transmission occurs in the OFF state. QoS limitations are imposed as constraints on buffer violation probabilities, and effective capacity formulation is used to identify the maximum throughput that a wireless channel can sustain while satisfying statistical QoS constraints. Energy efficiency is investigated by obtaining the bit energy required at zero spectral efficiency and the wideband slope in both wideband and low-power regimes assuming that the receiver has perfect channel side information (CSI). Initially, the wideband regime with multipath sparsity is investigated, and the minimum bit energy and wideband slope expressions are found. It is shown that the minimum bit energy requirements increase as the QoS constraints become more stringent. Subsequently, the low-power regime, which is also equivalent to the wideband regime with rich multipath fading, is analyzed. In this case, bit energy requirements are quantified through the expressions of bit energy required at zero spectral efficiency and wideband slope. It is shown for a certain class of fading distributions that the bit energy required at zero spectral efficiency is indeed the minimum bit energy for reliable communications. Moreover, it is proven that this minimum bit energy is attained in all cases regardless of the strictness of the QoS limitations. The impact upon the energy efficiency of multipath sparsity and richness is quantified, and comparisons with variable-rate/fixed-power and variable-rate/variable-power cases are provided. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
IEEE Trans. Wirel. Commun. | 3 |
| 2008 | Continuous Background Update and Object Detection with Non-static CamerasabstractDetecting moving objects is an important part of tracking. Most of the previous work on moving object detection concentrates on fixed cameras. Methods using moving cameras seldom deal with the problem of robustly and continuously updating the background model during all times including the periods when the camera is not static. We propose a method to build and continuously update a background model, and to detect foreground objects not only when the camera is static but also when it is zooming in/out or panning/tilting. For instance, the model built for the zoomed in (out) portion of a video is warped to the reference frame of the model of the zoomed out (in) portion to immediately incorporate changes that occurred in the background, such as objects that are placed or removed. This way, changes are incorporated to the model without requiring a learning period each time camera zooms in/out. This method addresses the problems of detecting moving objects during the zooming in and zooming out periods, detecting objects that are placed in the scene while the camera is non-static and gradually incorporating an overall illumination change to the scene model. We present different experiments covering three different scenarios to demonstrate the success of the proposed method in addressing these issues. Yijia Zhao, Mauricio Casares, Senem Velipasalar |
AVSS | 3 |
| 2008 | Analysis of Energy Efficiency in Fading Channels under QoS ConstraintsabstractEnergy efficiency in fading channels in the presence of QoS constraints is studied. Effective capacity, which provides the maximum constant arrival rate that a given process can support while satisfying statistical delay constraints, is considered. Spectral efficiency-bit energy tradeoff is analyzed in the low-power and wideband regimes by employing the effective capacity formulation, rather than the Shannon capacity, and energy requirements under QoS constraints are identified. The analysis is conducted for the case in which perfect channel side information (CSI) is available at the receiver and also for the case in which perfect CSI is available at both the receiver and transmitter. In particular, it is shown in the low-power regime that the minimum bit energy required in the presence of QoS constraints is the same as that attained when there are no such limitations. However, this performance is achieved as the transmitted power vanishes. Through the wideband slope analysis, the increased energy requirements at low but nonzero power levels are determined. A similar analysis is also conducted in the wideband regime, and minimum bit energy and wideband slope expressions are obtained. In this regime, the required bit energy levels are found to be strictly greater than those achieved when Shannon capacity is considered. Overall, an energy-delay tradeoff is characterized. Deli Qiao, Mustafa Cenk Gursoy, Senem Velipasalar |
GLOBECOM | 3 |
| 2008 | Frame-level temporal calibration of video sequences from unsynchronized cameras
Senem Velipasalar, Marilyn Wolf |
Mach. Vis. Appl. | 1 |
| 2007 | Real-Time Distributed TrackingabstractDistributed smart cameras use distributed computing architectures to analyze imagery from physically distributed cameras. Performing real-time distributed analysis of video introduces substantial new challenges, but also provides substantial benefits over server-based approaches. This work describes some of the algorithms and architectures we have developed for tracking using distributed smart camera systems, including fault-tolerance, synchronization, and multi-band fusion. Marilyn Wolf, Senem Velipasalar, Jason Schlessman, Cheng-Yao Chen, Chang Hong Lin |
ICASSP (4) | 2 |
| 2006 | Design and Verification of Communication Protocols for Peer-to-Peer Multimedia SystemsabstractThis paper addresses issues pertaining to the necessity of utilizing formal verification methods in the design of protocols for peer-to-peer multimedia systems. These systems require sophisticated communication protocols, and these protocols require verification. We discuss two sample protocols designed for two distinct peer-to-peer computer vision applications, namely multi-object multi-camera tracking and distributed gesture recognition. We present simulation and verification results for these protocols, obtained by using the SPIN verification tool, and discuss the importance of verifying the protocols used in peer-to-peer multimedia systems Senem Velipasalar, Chang Hong Lin, Jason Schlessman, Marilyn Wolf |
ICME | 1 |
| 2006 | SCCS: A Scalable Clustered Camera System for Multiple Object Tracking Communicating Via Message Passing InterfaceabstractWe introduce the scalable clustered camera system, a peer-to-peer multi-camera system for multi-object tracking, where different CPUs are used to process inputs from distinct cameras. Instead of transferring control of tracking jobs from one camera to another, each camera in our system performs its own tracking and keeps its own tracks for each target object, thus providing fault tolerance. A fast and robust tracking method is proposed to perform tracking on each camera view, while maintaining consistent labeling. In addition, we introduce a new communication protocol, where the decisions about when and with whom to communicate are made such that frequency and size of transmitted messages are minimized. This protocol incorporates variable synchronization capabilities, so as to allow flexibility with accuracy tradeoffs. We discuss our implementation, consisting of a parallel computing cluster, with communication between the cameras performed by MPI. We present experimental results which demonstrate the success of the proposed peer-to-peer multi-camera tracking system, with accuracy of 95% for a high frequency of synchronization, as well as a worst-case of 15 frames of latency in recovering correct labels at low synchronization frequencies Senem Velipasalar, Jason Schlessman, Cheng-Yao Chen, Marilyn Wolf |
ICME | 1 |
| 2006 | Automatic Counting of Interacting People by using a Single Uncalibrated CameraabstractAutomatic counting of people, entering or exiting a region of interest, is very important for both business and security applications. This paper introduces an automatic and robust people counting system which can count multiple people who interact in the region of interest, by using only one camera. Two-level hierarchical tracking is employed. For cases not involving merges or splits, a fast blob tracking method is used. In order to deal with interactions among people in a more thorough and reliable way, the system uses the mean shift tracking algorithm. Using the first-level blob tracker in general, and employing the mean shift tracking only in the case of merges and splits saves power and makes the system computationally efficient. The system setup parameter can be automatically learned in a new environment from a 3 to 5 minute-video with people going in or out of the target region one at a time. With a 2 GHz Pentium machine, the system runs at about 33 fps on 320times240 images without code optimization. Average accuracy rates of 98.5% and 95% are achieved on videos with normal traffic flow and videos with many cases of merges and splits, respectively Senem Velipasalar, Yingli Tian, Arun Hampapur |
ICME | 1 |
| 2005 | Frame-level temporal calibration of video sequences from unsynchronized cameras by using projective invariantsabstractThis paper describes a new method for temporally calibrating multiple cameras by image processing operations. Existing multi-camera algorithms assume that the input sequences are synchronized either by genlock or by time stamp information and a centralized server. Yet, hardware-based synchronization increases installation cost. Hence, using image information is necessary to align frames from the cameras whose clocks are not synchronized. Our method uses image processing to find the frame offset between sequences so that they can be aligned. We track foreground objects, extract a point of interest for each object as its current location, and find the corresponding location of the object in the other sequence by using projective invariants in P/sup 2/. Our algorithm recovers the frame offset by matching the tracks in different views, and finding the most reliable match out of the possible track pairs. This method does not require information about intrinsic or extrinsic camera parameters, and thanks to information obtained from multiple tracks, is robust to possible errors in background subtraction or location extraction. We present results on different sequences from the PETS2001 database, which show the robustness of the algorithm in recovering the frame offset. Senem Velipasalar, Marilyn Wolf |
AVSS | 1 |
| 2005 | Multiple object tracking and occlusion handling by information exchange between uncalibrated camerasabstractWe introduce a novel and robust method for multi-object tracking from multiple uncalibrated cameras. This method improves consistent labeling by incorporating the field of view lines and location information exchange between cameras by using the projective invariants in P/sup 2/. Each camera keeps its own tracks for each target object. This provides improved tracking as well as distributed processing, in which each camera is operated by a separate CPU that performs its own tracking and labeling. The tracking in each camera view is performed by using a two-level hierarchical structure. The main novelties of the proposed method include: a) the ability to communicate between the cameras at any time to improve and update the tracks of an object instead of tracking in each view independently, and to perform this without camera calibration; b) updating the track of an object without interruption and without any need for an estimation of the moving speed and direction, even if the object is totally invisible. The proposed method recovered 90% of the full occlusion cases. The hierarchical tracking structure makes the algorithm computationally efficient and, after background elimination, the first-level tracking runs at about 62 fps on a 2 GHz Celeron machine without code optimization. We present results obtained from the PETS2001 database, which show the success of the camera communication in partial and complete occlusions. Senem Velipasalar, Marilyn Wolf |
ICIP (2) | 1 |
| 2004 | Recovering field of view lines by using projective invariantsabstractEstablishing correspondences between moving objects is an important problem in multiple camera tracking, and field of view (FOV) lines have been introduced in literature as an efficient tool to resolve the consistent labeling issue. We introduce a new and robust method to find the FOV lines, which uses projective invariants that does not rely on the object movement in the scene and does not require information about the camera parameters. As the labeling scheme suggested before is based on the distance of an object to an FOV line, accurate recovery of the FOV lines provides reliable labeling, and makes less errors in the process. We present results on different sequences, obtained from the PETS200I database, which show the robustness of the algorithm in recovering all visible FOV lines of another camera in the current camera view. Senem Velipasalar, Marilyn Wolf |
ICIP | 1 |
| 2003 | A real-time prototype for small-vocabulary audio-visual ASRabstractWe present a prototype for the automatic recognition of audio-visual speech, developed to augment the IBM ViaVoice/spl trade/ speech recognition system. Frontal face, full frame video is captured through a USB 2.0 interface by means of an inexpensive PC camera, and processed to obtain appearance-based visual features. Subsequently, these are combined with audio features, synchronously extracted from the acoustic signal, using a simple discriminant feature fusion technique. On the average, the required computations utilize approximately 67% of a Pentium/spl trade/ 4, 1.8 GHz processor, leaving the remaining resources available to hidden Markov model based speech recognition. Real-time performance is there- fore achieved for small-vocabulary tasks, such as connected-digit recognition. In the paper, we discuss the prototype architecture based on the ViaVoice engine, the basic algorithms employed, and their necessary modifications to ensure real-time performance and causality of the visual front end processing. We benchmark the resulting system performance on stored videos against prior research experiments, and we report a close match between the two. Jonathan H. Connell, Norman Haas, Etienne Marcheret, Chalapathy Neti, Gerasimos Potamianos, Senem Velipasalar |
ICME | 6 |