VLDB 2026 Research / reviewers in the wild / expert
Okan Köpüklü
dblp:218/6295
· DBLP profile ↗
15ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0001-5281-9462ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 6 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SoundTRC: DNN-based Acoustic Target Region ControlabstractWe propose a deep neural network based automatic acoustic target region control framework, where the goal is to maintain all the relevant speech in the designated target region while muting all speech outside the target region in a multi-speaker conferencing room. We discuss three target regions: angle, angle-distance and distance that reflect common region-based speech control requests, and further propose a unified target region encoding strategy to encode the three different target regions into discriminative and compact target region vector. We propose a unified Cross-Attention Transformer based deep neural network, which takes the mixed speech and corresponding target region description as input and outputs ideal ratio mask that is responsible of masking out speech that is outside of the target region while suppressing the noise simultaneously. We run experiments on both simulated shoe-box like 3D room scenes and photo-realistic and complex 3D room scenes, showing the advantage of our proposed framework. Andrew Markham, Okan Köpüklü |
ICASSP | 3 |
| 2025 | Inference-Adaptive Steering of Neural Networks for Real-Time Area-Based Sound Source SeparationabstractWe propose a novel adaptive steering technique that changes the target area of a spatial-aware multi-microphone sound source separation algorithm during inference without the necessity of retraining the deep neural network (DNN). To achieve this, we first train a DNN aiming to retain speech within a target region, defined by an angular span, while suppressing sound sources stemming from other directions. Afterward, a phase shift is applied to the microphone signals, allowing us to shift the center of the target area during inference at negligible additional cost in computational complexity. Further, we show that the proposed approach performs well in a wide variety of acoustic scenarios, including several speakers inside and outside the target area and additional noise. More precisely, the proposed approach performs on par with DNNs trained explicitly for the steered target area in terms of DNSMOS and SI-SDR. Martin Strauss 0003, Wolfgang Mack, Maria Luis Valero, Okan Köpüklü |
IEEE Signal Process. Lett. | 4 |
| 2023 | From anomaly detection to open set recognition: Bridging the gapabstractThe classifiers that return compact acceptance regions are crucial for the success in anomaly detection and open set recognition settings since we have to determine and reject the anomalies and samples coming from the unknown classes. This paper introduces novel methods that approximate the class acceptance regions with compact hypersphere models for anomaly detection and open set recognition. As opposed to the other deep hypersphere classifiers, we treat the hypersphere centers as learnable parameters and update them based on the changing deep feature representations. In addition, we propose novel loss terms that are more robust to the noisy labels within the outlier exposure and background datasets. The proposed methods bear similarity to the deep distance metric learning classifiers using the triplet loss function with the exception that the anchors are set to the hypersphere centers which are updated dynamically. The experimental results show that the proposed methods achieve the state-of-the-art accuracies on the majority of the tested datasets in the context of anomaly detection and open set recognition. Hakan Çevikalp, Bedirhan Uzun, Yusuf Salk, Hasan Saribas, Okan Köpüklü |
Pattern Recognit. | 5 |
| 2022 | ResectNet: An Efficient Architecture for Voice Activity Detection on Mobile Devices
Okan Köpüklü, Maja Taseska |
INTERSPEECH | 1 |
| 2022 | Dissected 3D CNNs: Temporal skip connections for efficient online video processing
Okan Köpüklü, Stefan Hörmann 0001, Fabian Herzog, Hakan Çevikalp, Gerhard Rigoll |
Comput. Vis. Image Underst. | 1 |
| 2022 | TRAT: Tracking by attention using spatio-temporal features
Hasan Saribas, Hakan Çevikalp, Okan Köpüklü, Bedirhan Uzun |
Neurocomputing | 3 |
| 2021 | How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the WildabstractSuccessful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers within each frame, and (iii) temporal modeling for the reference speaker. Each stage of this pipeline plays an important role for the final performance of the created architecture. Based on a series of controlled experiments, this work presents several practical guidelines for audio-visual active speaker detection. Correspondingly, we present a new architecture called ASDNet, which achieves a new state-of-the-art on the AVA-ActiveSpeaker dataset with a mAP of 93.5% outperforming the second best with a large margin of 4.7%. Our code and pretrained models are publicly available1. Okan Köpüklü, Maja Taseska, Gerhard Rigoll |
ICCV | 1 |
| 2021 | Driver Anomaly Detection: A Dataset and Contrastive Learning ApproachabstractDistracted drivers are more likely to fail to anticipate hazards, which result in car accidents. Therefore, detecting anomalies in drivers' actions (i.e., any action deviating from normal driving) contains the utmost importance to reduce driver-related accidents. However, there are unbounded many anomalous actions that a driver can do while driving, which leads to an `open set recognition' problem. Accordingly, instead of recognizing a set of anomalous actions that are commonly defined by previous dataset providers, in this work, we propose a contrastive learning approach to learn a metric to differentiate normal driving from anomalous driving. For this task, we introduce a new video-based benchmark, the Driver Anomaly Detection (DAD) dataset, which contains normal driving videos together with a set of anomalous actions in its training set. In the test set of the DAD dataset, there are unseen anomalous actions that still need to be winnowed out from normal driving. Our method reaches 0.9673 AUC on the test set, demonstrating the effectiveness of the contrastive learning approach on the anomaly detection task. Our dataset, codes and pre-trained models are publicly available1. Okan Köpüklü, Jiapeng Zheng, Gerhard Rigoll |
WACV | 1 |
| 2021 | Deep compact polyhedral conic classifier for open and closed set recognition
Hakan Çevikalp, Bedirhan Uzun, Okan Köpüklü, Gürkan Öztürk |
Pattern Recognit. | 3 |
| 2020 | DriverMHG: A Multi-Modal Dataset for Dynamic Recognition of Driver Micro Hand Gestures and a Real-Time Recognition FrameworkabstractThe use of hand gestures provides a natural alternative to cumbersome interface devices for Human-Computer Interaction (HCI) systems. However, real-time recognition of dynamic micro hand gestures from video streams is challenging for in-vehicle scenarios since (i) the gestures should be performed naturally without distracting the driver, (ii) micro hand gestures occur within very short time intervals at spatially constrained areas, (iii) the performed gesture should be recognized only once, and (iv) the entire architecture should be designed lightweight as it will be deployed to an embedded system. In this work, we propose an HCI system for dynamic recognition of driver micro hand gestures, which can have a crucial impact in automotive sector especially for safety related issues. For this purpose, we initially collected a dataset named Driver Micro Hand Gestures (DriverMHG), which consists of RGB, depth and infrared modalities. The challenges for dynamic recognition of micro hand gestures have been addressed by proposing a lightweight convolutional neural network (CNN) based architecture which operates online efficiently with a sliding window approach. For the CNN model, several 3-dimensional resource efficient networks are applied and their performances are analyzed. Online recognition of gestures has been performed with 3D-MobileNetV2, which provided the best offline accuracy among the applied networks with similar computational complexities. The final architecture is deployed on a driver simulator operating in real-time. We make DriverMHG dataset and our source code publicly available1. Okan Köpüklü, Thomas Ledwon, Yao Rong 0001, Neslihan Kose, Gerhard Rigoll |
FG | 1 |
| 2019 | Real-time Hand Gesture Detection and Classification Using Convolutional Neural NetworksabstractReal-time recognition of dynamic hand gestures from video streams is a challenging task since (i) there is no indication when a gesture starts and ends in the video, (ii) performed gestures should only be recognized once, and (iii) the entire architecture should be designed considering the memory and power budget. In this work, we address these challenges by proposing a hierarchical structure enabling offline-working convolutional neural network (CNN) architectures to operate online efficiently by using sliding window approach. The proposed architecture consists of two models: (1) A detector which is a lightweight CNN architecture to detect gestures and (2) a classifier which is a deep CNN to classify the detected gestures. In order to evaluate the single-time activations of the detected gestures, we propose to use Levenshtein distance as an evaluation metric since it can measure misclassifications, multiple detections, and missing detections at the same time. We evaluate our architecture on two publicly available datasets - EgoGesture and NVIDIA Dynamic Hand Gesture Datasets - which require temporal detection and classification of the performed hand gestures. ResNeXt-101 model, which is used as a classifier, achieves the state-of-the-art offline classification accuracy of 94.04% and 83.82% for depth modality on EgoGesture and NVIDIA benchmarks, respectively. In real-time detection and classification, we obtain considerable early detections while achieving performances close to offline operation. The codes and pretrained models used in this work are publicly available1. Okan Köpüklü, Ahmet Gunduz, Neslihan Kose, Gerhard Rigoll |
FG | 1 |
| 2019 | Gait Energy Image Restoration Using Generative Adversarial NetworksabstractGait is a biometric property that can be used for human identification in video surveillance. Basically, different gait features require motion of a person walking over one complete gait cycle. For example, in Gait Energy Image (GEI), average of silhouette images over one complete gait cycle is computed. However, in reality, there might be a partial gait cycle data available due to occlusion. In this paper, we propose a Generative Adversarial Network (GAN) in order to address the problem of gait recognition from incomplete gait cycle. Precisely, the network is able to reconstruct complete GEIs from incomplete GEIs. The proposed architecture is composed of (i) a generator which is an auto-encoder network to construct complete GEIs out of incomplete GEIs and (ii) two discriminators, one of which discriminates whether a given image is a full GEI while the other discriminates whether two GEIs belong to the same subject. We evaluate our approach on the OULP large gait dataset confirming that the proposed architecture successfully reconstructs complete GEIs from even extreme incomplete gait cycles. Maryam Babaee, Okan Köpüklü, Stefan Hörmann 0001, Gerhard Rigoll |
ICIP | 3 |
| 2019 | Outlier-Robust Neural Aggregation Network for Video Face IdentificationabstractCurrent approaches for video face recognition rely on image sets containing faces of exclusively one identity. However, as image sets are created by unsupervised methods, it is necessary to consider outlier-afflicted sets for real-life applications. In this paper, we propose an Outlier-Robust Neural Aggregation Network (ORNAN). First, we embed each image into a feature space using a Convolutional Neural Network (CNN). With the help of two cascaded attention blocks, we predict outliers within the image set. By integrating this knowledge into our aggregation network, we adaptively aggregate all feature vectors to form a single feature, mitigating the influence of outliers and noisy features. We show that our network is robust against outliers using outlier-afflicted IJB-B and IJB-C benchmarks while maintaining similar performance without outliers. Stefan Hörmann 0001, Martin Knoche, Maryam Babaee, Okan Köpüklü, Gerhard Rigoll |
ICIP | 4 |
| 2019 | Convolutional Neural Networks with Layer ReuseabstractA convolutional layer in a Convolutional Neural Network (CNN) consists of many filters which apply convolution operation to the input, capture some special patterns and pass the result to the next layer. If the same patterns also occur at the deeper layers of the network, why wouldn't the same convolutional filters be used also in those layers? In this paper, we propose a CNN architecture, Layer Reuse Network (LruNet), where the convolutional layers are used repeatedly without the need of introducing new layers to get a better performance. This approach introduces several advantages: (i) Considerable amount of parameters are saved since we are reusing the layers instead of introducing new layers, (ii) the Memory Access Cost (MAC) can be reduced since reused layer parameters can be fetched only once, (iii) the number of nonlinearities increases with layer reuse, and (iv) reused layers get gradient updates from multiple parts of the network. The proposed approach is evaluated on CIFAR-10, CIFAR-100 and Fashion-MNIST datasets for image classification task, and layer reuse improves the performance by 5.14%, 5.85% and 2.29%, respectively. The source code and pretrained models are publicly available1. Okan Köpüklü, Maryam Babaee, Stefan Hörmann 0001, Gerhard Rigoll |
ICIP | 1 |
| 2018 | Analysis on Temporal Dimension of Inputs for 3D Convolutional Neural Networksabstract3D ConvNets provide a dedicated spatiotemporal representation in order to incorporate motion patterns within video frames. However, compared to 2D convolutions, the 3D convolution kernels increase the number of parameters in the architecture and the floating point operations during inference time, which are of critical importance for real-time applications requiring faster runtime. In this paper, we show a sparse sampling and stacking strategy to span large time intervals for 3D ConvNet architectures that can attain multiple times less inference time by relinquishing little amount of classification accuracy. The proposed approach is validated on action and gesture recognition tasks using two recent video datasets: Jester and Something-Something datasets. Okan Köpüklü, Gerhard Rigoll |
IPAS | 1 |