Kenan E. Ak

dblp:182/2932 · also Kenan Emir Ak · DBLP profile ↗
← Back
13ranked-venue papers
10as first author
5since 2021 · last 2024
0000-0001-5863-3685ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 first-author
YearPublicationVenuePosition
2024 Controllable Video Generation With Text-Based Instructions
abstract
Most of the existing studies on controllable video generation either transfer disentangled motion to an appearance without detailed control over motion or generate videos of simple actions such as the movement of arbitrary objects conditioned on a control signal from users. In this study, we introduce Controllable Video Generation with text-based Instructions (CVGI) framework that allows text-based control over action performed on a video. CVGI generates videos where hands interact with objects to perform the desired action by generating hand motions with detailed control through text-based instruction from users. By incorporating the motion estimation layer, we divide the task into two sub-tasks: (1) control signal estimation and (2) action generation. In control signal estimation, an encoder models actions as a set of simple motions by estimating low-level control signals for text-based instructions with given initial frames. In action generation, generative adversarial networks (GANs) generate realistic hand-based action videos as a combination of hand motions conditioned on the estimated low control level signal. Evaluations on several datasets (EPIC-Kitchens-55, BAIR robot pushing, and Atari Breakout) show the effectiveness of CVGI in generating realistic videos and in the control over actions.
Ali Koksal, Kenan E. Ak, Ying Sun 0001, Deepu Rajan, Joo-Hwee Lim
IEEE Trans. Multim.2
2023 Leveraging Efficient Training and Feature Fusion in Transformers for Multimodal Classification
abstract
People navigate a world that involves many different modalities and make decision on what they observe. Many of the classification problems that we face in the modern digital world are also multimodal in nature, where textual information on the web rarely occurs alone, and is often accompanied by images, sounds, or videos. The use of transformers in deep learning tasks has proven to be highly effective. However, the relationship between different modalities remains unclear. This paper investigates ways to simultaneously utilize self-attention over both text and vision modalities. We propose a novel architecture that combines the strengths of both modalities. We show that combining a text model with a fixed image model leads to the best classification performance. Additionally, we incorporate a late fusion technique to enhance the architecture’s ability to capture multiple modalities. Our experiments demonstrate that our proposed method outperforms state-of-the-art baselines on Food101, MM-IMDB, and FashionGen datasets.
Kenan E. Ak, Gwang-Gook Lee, Mingwei Shen 0001
ICIP1
2023 Learning by Imagination: A Joint Framework for Text-Based Image Manipulation and Change Captioning
abstract
Image and text are dual modalities of our semantic interpretation. Changing images based on text descriptions allows us to imagine and visualize the world (a.k.a. text-based image manipulation (TIM)). In this paper, we introduce a framework that combines TIM with change captioning (CC) and utilizes the benefits of co-training. CC aims to describe what has changed in a scene and can be regarded as the inverse version of TIM where both tasks rely on generative networks. These generative networks can be regarded as data producers of each other and unlike previous methods, we discover that integrating their learning procedures can benefit both. Since the CC module describes differences between two images as text, the CC module can be used as evaluation criteria and provide feedback. Furthermore, we utilize a shared attention mechanism in TIM and CC modules to localize towards prominent regions as well as enabling a change-aware discriminator. In the opposite direction, the output image synthesized by the TIM module can be assessed with the CC module, by checking whether the ground truth text description can be redescribed. Following this insight, not only do we boost the training of the TIM module, but we also utilize the TIM module as additional supervision for the CC training. Experimental results show that our framework outperforms existing TIM methods on several datasets substantially and we achieve marginal improvements in the CC module. To our best knowledge, this is the first study dedicated to the joint training of TIM and CC tasks.
Kenan E. Ak, Ying Sun 0001, Joo-Hwee Lim
IEEE Trans. Multim.1
2022 $A^3$-FKG: Attentive Attribute-Aware Fashion Knowledge Graph for Outfit Preference Prediction
abstract
With the booming development of the online fashion industry, effective personalized recommender systems have become indispensable for the convenience they brought to the customers and the profits to the e-commercial platforms. Estimating the user’s preference towards the outfit is at the core of a personalized recommendation system. Existing works on fashion recommendation are largely centering on modelling the clothing compatibility without considering the user factor or characterizing the user’s preference over the single item. However, how to effectively model the outfits with either few or even none interactions, is yet under-explored. In this paper, we address the task of personalized outfit preference prediction via a novelAttentiveAttribute-AwareFashionKnowledgeGraph ($A^3$-FKG), which is incorporated to build the association between different outfits with both outfit- and item- level attributes. Additionally, a two-level attention mechanism is developed to capture the user’s preference: 1) User-specific relation-aware attention layer, which captures the user’s fine-grained preferences with different focus on relations for learning outfit representation; 2) Target-aware attention layer, which characterizes the user’s latent diverse interests from his/her behavior sequences for learning user representation. Extensive experiments conducted on a large-scale fashion outfit dataset demonstrate significant improvements over other methods, which verify the excellence of our proposed framework.
Huijing Zhan, Jie Lin 0001, Kenan E. Ak, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Multim.3
2021 Robust Multi-Frame Future Prediction By Leveraging View Synthesis
abstract
In this paper, we focus on the problem of video prediction, i.e., future frame prediction. Most state-of-the-art techniques focus on synthesizing a single future frame at each step. However, this leads to utilizing the model’s own predicted frames when synthesizing multi-step prediction, resulting in gradual performance degradation due to accumulating errors in pixels. To alleviate this issue, we propose a model that can handle multi-step prediction. Additionally, we employ techniques to leverage from view synthesis for future frame prediction, where both problems are treated independently in the literature. Our proposed method employs multiview camera pose prediction and depth-prediction networks to project the last available frame to desired future frames via differentiable point cloud renderer. For the synthesis of moving objects, we utilize an additional refinement stage. In experiments, we show that the proposed framework outperforms state-of-theart methods in both KITTI and Cityscapes datasets.
Kenan E. Ak, Ying Sun 0001, Joo-Hwee Lim
ICIP1
2020 Incorporating Reinforced Adversarial Learning in Autoregressive Image Generation
Kenan E. Ak, Ning Xu 0007, Zhe Lin 0002, Yilin Wang 0002
ECCV (21)1
2020 Learning Cross-Modal Representations for Language-Based Image Manipulation
abstract
In this paper, we propose a generative architecture for manipulating images/scenes with natural language descriptions. This is a challenging task as the generative network is expected to perform the given text instruction without changing the non-affiliating contents of the input image. Two main drawbacks of the existing methods are their limitation of performing changes that would affect only a limited region and the inability of handling complex instructions. The proposed approach, designed to address these limitations initially uses two sets of networks to extract the image and text features respectively. Rather than a simple combination of these two modalities during the image manipulation process, we use an improved technique to compose image and text features. Additionally, the generative network utilizes similarity learning to improve text manipulation which also enforces only the text-relevant changes on the input image. Our experiments on CSS and Fashion Synthesis datasets show that the proposed approach performs remarkably well and outperforms the baseline frameworks in terms of R-precision and FID.
Kenan E. Ak, Ying Sun 0001, Joo-Hwee Lim
ICIP1
2020 Edge-Gan: Edge Conditioned Multi-View Face Image Generation
abstract
Reconstructing photorealistic multi-view images from an image with an arbitrary view has a wide range of applications in the field of face generation. However, most current pixel-based generation models cannot generate sufficiently realistic enough images. To address this problem, we propose an edge-conditioned multi-view image generation model called Edge-GAN. Edge-GAN utilizes edge information to guide the image generation based on the perspective of the target view while the details of the input image are used to influence the target image. Edge-GAN combines the input image with the target pose information to generate a coarse image with an approximate target outline which is then refined to a better quality using adversanal training. Experiments conducted show that our Edge-GAN is able to generate high-quality images of people with convincing details.
Heqing Zou, Kenan E. Ak, Ashraf A. Kassim
ICIP2
2020 Semantically consistent text to fashion image synthesis with an enhanced attentional generative adversarial network
Kenan E. Ak, Joo-Hwee Lim, Jo Yew Tham, Ashraf A. Kassim
Pattern Recognit. Lett.1
2019 Attribute Manipulation Generative Adversarial Networks for Fashion Images
abstract
Recent advances in Generative Adversarial Networks (GANs) have made it possible to conduct multi-domain image-to-image translation using a single generative network. While recent methods such as Ganimation and SaGAN are able to conduct translations on attribute-relevant regions using attention, they do not perform well when the number of attributes increases as the training of attention masks mostly rely on classification losses. To address this and other limitations, we introduce Attribute Manipulation Generative Adversarial Networks (AMGAN) for fashion images. While AMGAN's generator network uses class activation maps (CAMs) to empower its attention mechanism, it also exploits perceptual losses by assigning reference (target) images based on attribute similarities. AMGAN incorporates an additional discriminator network that focuses on attribute-relevant regions to detect unrealistic translations. Additionally, AMGAN can be controlled to perform attribute manipulations on specific regions such as the sleeve or torso regions. Experiments show that AMGAN outperforms state-of-the-art methods using traditional evaluation metrics as well as an alternative one that is based on image retrieval.
Kenan E. Ak, Ashraf A. Kassim, Joo-Hwee Lim, Jo Yew Tham
ICCV1
2018 Learning Attribute Representations With Localization for Flexible Fashion Search
abstract
In this paper, we investigate ways of conducting a detailed fashion search using query images and attributes. A credible fashion search platform should be able to (1) find images that share the same attributes as the query image, (2) allow users to manipulate certain attributes, e.g. replace collar attribute from round to v-neck, and (3) handle region-specific attribute manipulations, e.g. replacing the color attribute of the sleeve region without changing the color attribute of other regions. A key challenge to be addressed is that fashion products have multiple attributes and it is important for each of these attributes to have representative features. To address these challenges, we propose the FashionSearchNet which uses a weakly supervised localization method to extract regions of attributes. By doing so, unrelated features can be ignored thus improving the similarity learning. Also, FashionSearchNet incorporates a new procedure that enables region awareness to be able to handle region-specific requests. FashionSearchNet outperforms the most recent fashion search techniques and is shown to be able to carry out different search scenarios using the dynamic queries.
Kenan E. Ak, Ashraf A. Kassim, Joo-Hwee Lim, Jo Yew Tham
CVPR1
2018 Efficient Multi-attribute Similarity Learning Towards Attribute-Based Fashion Search
abstract
In this paper, we propose an attribute-based query & retrieval system designed for fashion products. Our system addresses the problem of carrying out fashion searches by the query image and attribute manipulation, e.g. replacing long sleeve attribute of a dress to sleeveless. We present the attributes in two groups: (1) general attributes (category, gender etc.) and (2) special attributes (sleeve length, collar etc.). The special attributes are more suitable for the attribute manipulation and thus conducting searches. In order to solve the mentioned fashion search problem, it is crucial for the deep neural networks to understand attribute similarities. To facilitate more specific similarity learning, clothing items are represented by their structural subcomponents or "parts". The parts are estimated using an unsupervised segmentation method and used inside the proposed Convolutional Neural Network (CNN) as an attention mechanism. Meaning, different parts are connected to the special attributes, e.g. sleeve part is connected with sleeve length attribute. With this mechanism, part-based triplet ranking constraint is applied to learn similarity of each special attribute independently from one another in a single network. In the end, the well-defined features are used to conduct the fashion search. Additionally, an adaptive relevance feedback module is used to personalize the fashion search process with the feature descriptions. For our experiments, a new dataset is constructed containing 101,021 images which consist of pure clothing items. Besides achieving decent retrieval results in our dataset, the experiments show that proposed technique outperforms different baselines and is able to adapt towards user's requests.
Kenan E. Ak, Joo-Hwee Lim, Jo Yew Tham, Ashraf A. Kassim
WACV1
2018 Which shirt for my first date? Towards a flexible attribute-based fashion query system
Kenan E. Ak, Joo-Hwee Lim, Jo Yew Tham, Ashraf A. Kassim
Pattern Recognit. Lett.1