Nazli Ikizler-Cinbis

dblp:06/8611 · DBLP profile ↗
← Back
33ranked-venue papers
4as first author
10since 2021 · last 2024
0000-0002-8644-2875ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2024 Cross-lingual few-shot sign language recognition
Yunus Can Bilge, Nazli Ikizler-Cinbis, Ramazan Gokberk Cinbis
Pattern Recognit.2
2023 Towards Zero-Shot Sign Language Recognition
abstract
This paper tackles the problem of zero-shot sign language recognition (ZSSLR), where the goal is to leverage models learned over the seen sign classes to recognize the instances of unseen sign classes. In this context, readily available textual sign descriptions and attributes collected from sign language dictionaries are utilized as semantic class representations for knowledge transfer. For this novel problem setup, we introduce three benchmark datasets with their accompanying textual and attribute descriptions to analyze the problem in detail. Our proposed approach builds spatiotemporal models of body and hand regions. By leveraging the descriptive text and attribute embeddings along with these visual representations within a zero-shot learning framework, we show that textual and attribute based class definitions can provide effective knowledge for the recognition of previously unseen sign classes. We additionally introduce techniques to analyze the influence of binary attributes in correct and incorrect zero-shot predictions. We anticipate that the introduced approaches and the accompanying datasets will provide a basis for further exploration of zero-shot learning in sign language recognition.
Yunus Can Bilge, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 MAC: Mask-Augmentation for Motion-Aware Video Representation Learning
Arif Akar, Ufuk Umut Senturk, Nazli Ikizler-Cinbis
BMVC3
2022 TripleDNet: Exploring Depth Estimation with Self-Supervised Representation Learning
Ufuk Umut Senturk, Arif Akar, Nazli Ikizler-Cinbis
BMVC3
2022 MaskSplit: Self-supervised Meta-learning for Few-shot Semantic Segmentation
abstract
Just like other few-shot learning problems, few-shot segmentation aims to minimize the need for manual annotation, which is particularly costly in segmentation tasks. Even though the few-shot setting reduces this cost for novel test classes, there is still a need to annotate the training data. To alleviate this need, we propose a self-supervised training approach for learning few-shot segmentation models. We first use unsupervised saliency estimation to obtain pseudo-masks on images. We then train a simple prototype based model over different splits of pseudo masks and augmentations of images. Our extensive experiments show that the proposed approach achieves promising results, highlighting the potential of self-supervised training. To the best of our knowledge this is the first work that addresses unsupervised few-shot segmentation problem on natural images.
Mustafa Sercan Amac, Ahmet Sencan, Orhun Bugra Baran, Nazli Ikizler-Cinbis, Ramazan Gokberk Cinbis
WACV4
2022 Top-down and bottom-up attentional multiple instance learning for still image action recognition
Cagdas Bas, Nazli Ikizler-Cinbis
Signal Process. Image Commun.2
2021 Red Carpet to Fight Club: Partially-supervised Domain Transfer for Face Recognition in Violent Videos
abstract
In many real-world problems, there is typically a large discrepancy between the characteristics of data used in training versus deployment. A prime example is the analysis of aggression videos: in a criminal incidence, typically suspects need to be identified based on their clean portraitlike photos, instead of their prior video recordings. This results in three major challenges; large domain discrepancy between violence videos and ID-photos, the lack of video examples for most individuals and limited training data availability. To mimic such scenarios, we formulate a realistic domain-transfer problem, where the goal is to transfer the recognition model trained on clean posed images to the target domain of violent videos, where training videos are available only for a subset of subjects. To this end, we introduce the "WildestFaces" dataset, tailored to study cross-domain recognition under a variety of adverse conditions. We divide the task of transferring a recognition model from the domain of clean images to the violent videos into two sub-problems and tackle them using (i) stacked affine-transforms for classifier-transfer, (ii) attention-driven pooling for temporal-adaptation. We additionally formulate a self-attention based model for domain-transfer. We establish a rigorous evaluation protocol for this "clean-to-violent" recognition task, and present a detailed analysis of the proposed dataset and the methods. Our experiments highlight the unique challenges introduced by the WildestFaces dataset and the advantages of the proposed approach.
Yunus Can Bilge, Mehmet Kerim Yucel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis, Pinar Duygulu
WACV4
2021 Using independently recurrent networks for reinforcement learning based unsupervised video summarization
Gokhan Yaliniz, Nazli Ikizler-Cinbis
Multim. Tools Appl.2
2021 Leveraging auxiliary image descriptions for dense video captioning
Emre Boran, Aykut Erdem, Nazli Ikizler-Cinbis, Erkut Erdem, Pranava Swaroop Madhyastha, Lucia Specia
Pattern Recognit. Lett.3
2021 Multi-stream pose convolutional neural networks for human interaction recognition in images
Gokhan Tanisik, Cemil Zalluhoglu, Nazli Ikizler-Cinbis
Signal Process. Image Commun.3
2020 Collective Sports: A multi-task dataset for collective activity recognition
Cemil Zalluhoglu, Nazli Ikizler-Cinbis
Image Vis. Comput.2
2019 Zero-Shot Sign Language Recognition: Can Textual Data Uncover Sign Languages?
Yunus Can Bilge, Nazli Ikizler-Cinbis, Ramazan Gokberk Cinbis
BMVC2
2019 Image Captioning with Unseen Objects
Berkan Demirel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis
BMVC3
2019 Learning Visually Consistent Label Embeddings for Zero-Shot Learning
abstract
In this work, we propose a zero-shot learning method to effectively model knowledge transfer between classes via jointly learning visually consistent word vectors and label embedding model in an end-to-end manner. The main idea is to project the vector space word vectors of attributes and classes into the visual space such that word representations of semantically related classes become more closer, and use the projected vectors in the proposed embedding model to identify unseen classes. We evaluate the proposed approach on two benchmark datasets and the experimental results show that our method yields significant improvements in recognition accuracy.
Berkan Demirel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis
ICIP3
2019 Region based multi-stream convolutional neural networks for collective activity recognition
Cemil Zalluhoglu, Nazli Ikizler-Cinbis
J. Vis. Commun. Image Represent.2
2018 Zero-Shot Object Detection by Hybrid Region Embedding
Berkan Demirel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis
BMVC3
2018 RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes
abstract
Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text.In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes.It comprises of approximately 20K instructional recipes with multiple modalities such as titles, descriptions and aligned set of images.With over 36K automatically generated question-answer pairs, we design a set of comprehension and reasoning tasks that require joint understanding of images and text, capturing the temporal flow of events and making sense of procedural knowledge.Our preliminary results indicate that RecipeQA will serve as a challenging test bed and an ideal benchmark for evaluating machine comprehension systems.The data and leaderboard are available at http://hucvl.github.io/recipeqa.
Semih Yagcioglu, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis
EMNLP4
2018 Histograms of sequences: a novel representation for human interaction recognition
abstract
This study presents a novel representation based on hierarchical histogram of local feature sequences for human interaction recognition. The authors’ method basically combines the power of discriminative sequence mining and histogram representation for the effective recognition of human interactions. Our framework involves extracting visual features from the videos first, and then mining sequences of the visual features that occur consequently in space and time. After the mining step, we represent each video with a histogram pyramid of such sequences. We also propose to use soft clustering in the visual word construction step, such that more information‐rich histograms can be obtained. The authors’ experimental results on challenging human interaction recognition data sets indicate that the proposed algorithm performs on par with the state‐of‐the‐art methods.
Aytac Cavent, Nazli Ikizler-Cinbis
IET Comput. Vis.2
2018 Space-Time Tree Ensemble for Action Recognition and Localization
Shugao Ma, Jianming Zhang 0001, Stan Sclaroff, Nazli Ikizler-Cinbis, Leonid Sigal
Int. J. Comput. Vis.4
2017 Re-evaluating Automatic Metrics for Image Captioning
abstract
Mert Kilickaya, Aykut Erdem, Nazli Ikizler-Cinbis, Erkut Erdem. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Mert Kilickaya, Aykut Erdem, Nazli Ikizler-Cinbis, Erkut Erdem
EACL (1)3
2017 Attributes2Classname: A Discriminative Model for Attribute-Based Unsupervised Zero-Shot Learning
abstract
We propose a novel approach for unsupervised zero-shot learning (ZSL) of classes based on their names. Most existing unsupervised ZSL methods aim to learn a model for directly comparing image features and class names. However, this proves to be a difficult task due to dominance of non-visual semantics in underlying vector-space embeddings of class names. To address this issue, we discriminatively learn a word representation such that the similarities between class and combination of attribute names fall in line with the visual similarity. Contrary to the traditional zero-shot learning approaches that are built upon attribute presence, our approach bypasses the laborious attribute-class relation annotations for unseen classes. In addition, our proposed approach renders text-only training possible, hence, the training can be augmented without the need to collect additional image data. The experimental results show that our method yields state-of-the-art results for unsupervised ZSL in three benchmark datasets.
Berkan Demirel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis
ICCV3
2017 Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures (Extended Abstract)
abstract
Automatic image description generation is a challenging problem that has recently received a large amount of interest from the computer vision and natural language processing communities. In this survey, we classify the known approaches based on how they conceptualise this problem and provide a review of existing models, highlighting their advantages and disadvantages. Moreover, we give an overview of the benchmark image-text datasets and the evaluation measures that have been developed to assess the quality of machine-generated descriptions. Finally we explore future directions in the area of automatic image description.
Raffaella Bernardi, Ruken Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, Barbara Plank
IJCAI6
2017 Data-driven image captioning via salient region discovery
abstract
In the past few years, automatically generating descriptions for images has attracted a lot of attention in computer vision and natural language processing research. Among the existing approaches, data‐driven methods have been proven to be highly effective. These methods compare the given image against a large set of training images to determine a set of relevant images, then generate a description using the associated captions. In this study, the authors propose to integrate an object‐based semantic image representation into a deep features‐based retrieval framework to select the relevant images. Moreover, they present a novel phrase selection paradigm and a sentence generation model which depends on a joint analysis of salient regions in the input and retrieved images within a clustering framework. The authors demonstrate the effectiveness of their proposed approach on Flickr8K and Flickr30K benchmark datasets and show that their model gives highly competitive results compared with the state‐of‐the‐art models.
Mert Kilickaya, Burak Kerim Akkus, Ruken Cakici, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis
IET Comput. Vis.6
2016 Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures
abstract
Automatic description generation from natural images is a challenging problem that has recently received a large amount of interest from the computer vision and natural language processing communities. In this survey, we classify the existing approaches based on how they conceptualize this problem, viz., models that cast description as either generation problem or as a retrieval problem over a visual or multimodal representational space. We provide a detailed review of existing models, highlighting their advantages and disadvantages. Moreover, we give an overview of the benchmark image datasets and the evaluation measures that have been developed to assess the quality of machine-generated image descriptions. Finally we extrapolate future directions in the area of automatic image description generation.
Raffaella Bernardi, Ruken Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, Barbara Plank
J. Artif. Intell. Res.6
2016 Low-level features for visual attribute recognition: An evaluation
Emine Gul Danaci, Nazli Ikizler-Cinbis
Pattern Recognit. Lett.2
2016 Facial descriptors for human interaction recognition in still images
Gokhan Tanisik, Cemil Zalluhoglu, Nazli Ikizler-Cinbis
Pattern Recognit. Lett.3
2015 Two-person interaction recognition via spatial multiple instance embedding
Fadime Sener, Nazli Ikizler-Cinbis
J. Vis. Commun. Image Represent.2
2014 Ensemble of multiple instance classifiers for image re-ranking
Fadime Sener, Nazli Ikizler-Cinbis
Image Vis. Comput.2
2013 Action Recognition and Localization by Hierarchical Space-Time Segments
abstract
We propose Hierarchical Space-Time Segments as a new representation for action recognition and localization. This representation has a two-level hierarchy. The first level comprises the root space-time segments that may contain a human body. The second level comprises multi-grained space-time segments that contain parts of the root. We present an unsupervised method to generate this representation from video, which extracts both static and non-static relevant space-time segments, and also preserves their hierarchical and temporal relationships. Using simple linear SVM on the resultant bag of hierarchical space-time segments representation, we attain better than, or comparable to, state-of-the-art action recognition performance on two challenging benchmark datasets and at the same time produce good action localization results.
Shugao Ma, Jianming Zhang 0001, Nazli Ikizler-Cinbis, Stan Sclaroff
ICCV3
2012 Web-Based Classifiers for Human Action Recognition
abstract
Action recognition in uncontrolled videos is a challenging task, where it is relatively hard to find the large amount of required training videos to model all the variations of the domain. This paper addresses this challenge and proposes a generic method for action recognition. The idea is to use images collected from the Web to learn representations of actions and leverage this knowledge to automatically annotate actions in videos. For this purpose, we first use an incremental image retrieval procedure to collect and clean up the necessary training set for building the human pose classifiers. Our approach is unsupervised in the sense that it requires no human intervention other than the text querying to an internet search engine. Its benefits are two-fold: 1) we can improve retrieval of action images, and 2) we can collect a large generic database of action poses, which can then be used in tagging videos. We present experimental evidence that using action images collected from the Web, annotating actions in the videos is possible. Additionally, we explore how the Web-based pose classifiers can be utilized in conjunction with limited labelled videos. We propose to use “ordered pose pairs” (OPP) for encoding the temporal ordering of poses in our action model, and show that considering the temporal ordering of pose pairs can increase the action recognition accuracy. We also show that by selecting the keyposes with the help of Web-based classifiers, the classification time can be reduced. Our experiments demonstrate that, with or without available video data, the pose models learned from the Web can improve the performance of the action recognition systems.
Nazli Ikizler-Cinbis, Stan Sclaroff
IEEE Trans. Multim.1
2010 Object, Scene and Actions: Combining Multiple Features for Human Action Recognition
Nazli Ikizler-Cinbis, Stan Sclaroff
ECCV (1)1
2010 Object Recognition and Localization Via Spatial Instance Embedding
abstract
We propose an approach for improving object recognition and localization using spatial kernels together with instance embedding. Our approach treats each image as a bag of instances (image features) within a multiple instance learning framework, where the relative locations of the instances are considered as well as the appearance similarity of the localized image features. The introduced spatial kernel augments the recognition power of the instance embedding in an intuitive and effective way, providing increased localization performance. We test our approach over two object datasets and present promising results.
Nazli Ikizler-Cinbis, Stan Sclaroff
ICPR1
2009 Learning actions from the Web
abstract
This paper proposes a generic method for action recognition in uncontrolled videos. The idea is to use images collected from the Web to learn representations of actions and use this knowledge to automatically annotate actions in videos. Our approach is unsupervised in the sense that it requires no human intervention other than the text querying. Its benefits are two-fold: 1) we can improve retrieval of action images, and 2) we can collect a large generic database of action poses, which can then be used in tagging videos. We present experimental evidence that using action images collected from the Web, annotating actions is possible.
Nazli Ikizler-Cinbis, Ramazan Gokberk Cinbis, Stan Sclaroff
ICCV1