Pinar Yanardag Delul

dblp:130/0344 · also Pinar Yanardag · DBLP profile ↗
← Back
29ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0003-0193-7818ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 2 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Plot'n Polish: Zero-Shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
abstract
Text-to-image diffusion models have demonstrated significant capabilities to generate diverse and detailed visuals in various domains, and story visualization is emerging as a particularly promising application. However, as their use in real-world creative domains increases, the need for providing enhanced control, refinement, and the ability to modify images post-generation in a consistent manner becomes an important challenge. Existing methods often lack the flexibility to apply fine or coarse edits while maintaining visual and narrative consistency across multiple frames, preventing creators from seamlessly crafting and refining their visual stories. To address these challenges, we introduce Plot'n Polish, a zero-shot framework that enables consistent story generation and provides fine-grained control over story visualizations at various levels of detail.
Kiymet Akdemir, Jing Shi 0005, Kushal Kafle, Brian L. Price, Pinar Yanardag Delul
AAAI5
2026 MotionFlow: Attention-Driven Motion Transfer in Video Diffusion Models
abstract
Text-to-video models have demonstrated impressive capabilities in producing diverse video content, yet often lack fine-grained control over motion. We address the problem of motion transfer: given a source video and a target text prompt, generate a new video that preserves the source motion while matching the target semantics and allowing large changes in appearance and scene layout. We introduce MotionFlow, a training-free framework that performs test-time latent optimization guided by attention-derived motion cues. MotionFlow first extracts cross-attention maps from a pre-trained video diffusion model and converts them into spatio-temporal motion masks for the source subject. During generation, it optimizes the target latents so that their evolving attention patterns align with these masks, while the target text controls appearance. This avoids direct attention-map replacement and any model-specific fine-tuning, reducing artifacts and improving flexibility. Qualitative and quantitative experiments, including a user study, show that MotionFlow outperforms existing methods in motion fidelity, temporal consistency, and versatility, even under drastic scene changes.
Tuna Han Salih Meral, Hidir Yesiltepe, Connor Dunlop, Pinar Yanardag Delul
AAAI4
2025 FluxSpace: Disentangled Semantic Editing in Rectified Flow Models
abstract
Rectified flow models have emerged as a dominant approach in image generation, showcasing impressive capabilities in high-quality image synthesis. However, despite their effectiveness in visual generation, rectified flow models often struggle with disentangled editing of images. This limitation prevents the ability to perform precise, attribute-specific modifications without affecting unrelated aspects of the image. In this paper, we introduce FluxSpace, a domain-agnostic image editing method leveraging a representation space with the ability to control the semantics of images generated by rectified flow transformers, such as Flux. By leveraging the representations learned by the transformer blocks within the rectified flow models, we propose a set of semantically interpretable representations that enable a wide range of image editing tasks, from fine-grained image editing to artistic creation. This work offers a scalable and effective image editing approach, along with its disentanglement capabilities.
Yusuf Dalva, Kavana Venkatesh, Pinar Yanardag Delul
CVPR3
2025 Explaining in Diffusion: Explaining a Classifier with Diffusion Semantics
abstract
Classifiers are important components in many computer vision tasks, serving as the foundational backbone of a wide variety of models employed across diverse applications. However, understanding the decision-making process of classifiers remains a significant challenge. We propose DiffEx, a novel method that leverages the capabilities of text-to-image diffusion models to explain classifier decisions. Unlike traditional GAN-based explainability models, which are limited to simple, single-concept analyses and typically require training a new model for each classifier, our approach can explain classifiers that focus on single concepts (such as faces or animals) as well as those that handle complex scenes involving multiple concepts. DiffEx employs vision-language models to create a hierarchical list of semantics, allowing users to identify not only the overarching semantic influences on classifiers (e.g., the ‘beard’ semantic in a facial classifier) but also their sub-types, such as ‘goatee’ or ‘Balbo’ beard. Our experiments demonstrate that DiffEx is able to cover a significantly broader spectrum of semantics compared to its GAN counterparts, providing a hierarchical tool that delivers a more detailed and fine-grained understanding of classifier decisions.
Tahira Kazimi, Ritika Allada, Pinar Yanardag Delul
CVPR3
2025 LoRACLR: Contrastive Adaptation for Customization of Diffusion Models
abstract
Recent advances in text-to-image customization have enabled high-fidelity, context-rich generation of personalized images, allowing specific concepts to appear in a variety of scenarios. However, current methods struggle with combining multiple personalized models, often leading to attribute entanglement or requiring separate training to preserve concept distinctiveness. We present LoRACLR, a novel approach for multi-concept image generation that merges multiple LoRA models, each fine-tuned for a distinct concept, into a single, unified model without additional individual fine-tuning. LoRACLR uses a contrastive objective to align and merge the weight spaces of these models, ensuring compatibility while minimizing interference. By enforcing distinct yet cohesive representations for each concept, LoRACLR enables efficient, scalable model composition for high-quality, multi-concept image synthesis. Our results highlight the effectiveness of LoRACLR in accurately merging multiple concepts, advancing the capabilities of personalized image generation.
Enis Simsar, Thomas Hofmann 0001, Federico Tombari, Pinar Yanardag Delul
CVPR4
2025 Contrastive Test-Time Composition of Multiple LoRA Models for Image Generation
Tuna Han Salih Meral, Enis Simsar, Federico Tombari, Pinar Yanardag Delul
ICCV4
2025 LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion Models
abstract
Low-rank Adaptation (LoRA) models have revolutionized the personalization of pre-trained diffusion models by enabling fine-tuning through low-rank, factorized weight matrices specifically optimized for attention layers. These models facilitate the generation of highly customized content across a variety of objects, individuals, and artistic styles without the need for extensive retraining. Despite the availability of over 100K LoRA adapters on platforms like Civit.ai, users often face challenges in navigating, selecting, and effectively utilizing the most suitable adapters due to their sheer volume, diversity, and lack of structured organization. This paper addresses the problem of selecting the most relevant and diverse LoRA models from this vast database by framing the task as a combinatorial optimization problem and proposing a novel submodular framework. Our quantitative and qualitative experiments demonstrate that our method generates diverse outputs across a wide range of domains.
Mert Sonmezer, Matthew Zheng, Pinar Yanardag Delul
ICCV3
2025 ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
abstract
Do the rich representations of multi-modal diffusion transformers (DiTs) exhibit unique properties that enhance their interpretability? We introduce ConceptAttention, a novel method that leverages the expressive power of DiT attention layers to generate high-quality saliency maps that precisely locate textual concepts within images. Without requiring additional training, ConceptAttention repurposes the parameters of DiT attention layers to produce highly contextualized *concept embeddings*, contributing the major discovery that performing linear projections in the output space of DiT attention layers yields significantly sharper saliency maps compared to commonly used cross-attention maps. ConceptAttention even achieves state-of-the-art performance on zero-shot image segmentation benchmarks, outperforming 15 other zero-shot interpretability methods on the ImageNet-Segmentation dataset. ConceptAttention works for popular image models and even seamlessly generalizes to video generation. Our work contributes the first evidence that the representations of multi-modal DiTs are highly transferable to vision tasks like segmentation.
Alec Helbling, Tuna Han Salih Meral, Benjamin Hoover, Pinar Yanardag Delul, Polo Chau
ICML4
2025 LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers
abstract
We introduce LoRAShop, the first framework for multi-concept image generation and editing with LoRA models. LoRAShop builds on a key observation about the feature interaction patterns inside Flux-style diffusion transformers: concept-specific transformer features activate spatially coherent regions early in the denoising process. We harness this observation to derive a disentangled latent mask for each concept in a prior forward pass and blend the corresponding LoRA weights only within regions bounding the concepts to be personalized. The resulting edits seamlessly integrate multiple subjects or styles into the original scene while preserving global context, lighting, and fine details. Our experiments demonstrate that LoRAShop delivers better identity preservation compared to baselines. By eliminating retraining and external constraints, LoRAShop turns personalized diffusion models into a practical `photoshop-with-LoRAs' tool and opens new avenues for compositional visual storytelling and rapid creative iteration.
Yusuf Dalva, Hidir Yesiltepe, Pinar Yanardag Delul
NeurIPS3
2025 Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization
abstract
Text-to-image (T2I) diffusion models have made remarkable strides in generating and editing high-fidelity images from text. Yet, these models remain fundamentally generic, failing to adapt to the nuanced aesthetic preferences of individual users. In this work, we present the first framework for personalized image editing in diffusion models, introducing Collaborative Direct Preference Optimization (C-DPO), a novel method that aligns image edits with user-specific preferences while leveraging collaborative signals from like-minded individuals. Our approach encodes each user as a node in a dynamic preference graph and learns embeddings via a lightweight graph neural network, enabling information sharing across users with overlapping visual tastes. We enhance a diffusion model's editing capabilities by integrating these personalized embeddings into a novel DPO objective, which jointly optimizes for individual alignment and neighborhood coherence. Comprehensive experiments, including user studies and quantitative benchmarks, demonstrate that our method consistently outperforms baselines in generating edits that are aligned with user preferences.
Connor Dunlop, Matthew Zheng, Kavana Venkatesh, Pinar Yanardag Delul
NeurIPS4
2025 CREA: A Collaborative Multi-Agent Framework for Creative Image Editing and Generation
abstract
Creativity in AI imagery remains a fundamental challenge, requiring not only the generation of visually compelling content but also the capacity to add novel, expressive, and artistically rich transformations to images. Unlike conventional editing tasks that rely on direct prompt-based modifications, creative image editing demands an autonomous, iterative approach that balances originality, coherence, and artistic intent. To address this, we introduce CREA, a novel multi-agent collaborative framework that mimics the human creative process. Our framework leverages a team of specialized AI agents who dynamically collaborate to conceptualize, generate, critique, and enhance images. Through extensive qualitative and quantitative evaluations, we demonstrate that CREA significantly outperforms state-of-the-art methods in diversity, semantic alignment, and creative transformation. To the best of our knowledge, this is the first work to introduce the task of creative editing.
Kavana Venkatesh, Connor Dunlop, Pinar Yanardag Delul
NeurIPS3
2025 Dynamic View Synthesis as an Inverse Problem
abstract
In this work, we address dynamic view synthesis from monocular videos as an inverse problem in a training-free setting. By redesigning the noise initialization phase of a pre-trained video diffusion model, we enable high-fidelity dynamic view synthesis without any weight updates or auxiliary modules. We begin by identifying a fundamental obstacle to deterministic inversion arising from zero-terminal signal-to-noise ratio (SNR) schedules and resolve it by introducing a novel noise representation, termed K-order Recursive Noise Representation. We derive a closed form expression for this representation, enabling precise and efficient alignment between the VAE-encoded and the DDIM inverted latents. To synthesize newly visible regions resulting from camera motion, we introduce Stochastic Latent Modulation, which performs visibility aware sampling over the latent space to complete occluded regions. Comprehensive experiments demonstrate that dynamic view synthesis can be effectively performed through structured latent manipulation in the noise initialization phase.
Hidir Yesiltepe, Pinar Yanardag Delul
NeurIPS2
2024 NoiseCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions in Diffusion Models
abstract
Generative models have been very popular in the recent years for their image generation capabilities. GAN-based models are highly regarded for their disentangled latent space, which is a key feature contributing to their success in controlled image editing. On the other hand, diffusion models have emerged as powerful tools for generating high-quality images. However, the latent space of diffusion models is not as thoroughly explored or understood. Existing methods that aim to explore the latent space of diffusion models usually relies on text prompts to pinpoint specific semantics. However, this approach may be restrictive in ar-eas such as art, fashion, or specialized fields like medicine, where suitable text prompts might not be available or easy to conceive thus limiting the scope of existing work. In this paper, we propose an unsupervised method to discover la-tent semantics in text-to-image diffusion models without relying on text prompts. Our method takes a small set of un la-beled images from specific domains, such as faces or cats, and a pre-trained diffusion model, and discovers diverse se-mantics in unsupervised fashion using a contrastive learning objective. Moreover, the learned directions can be ap-plied simultaneously, either within the same domain (such as various types of facial edits) or across different domains (such as applying cat and face edits within the same image) without interfering with each other. Our extensive experi-ments show that our method achieves highly disentangled edits, outperforming existing approaches in both diffusion-based and GAN-based latent space editing methods.
Yusuf Dalva, Pinar Yanardag Delul
CVPR2
2024 RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models
abstract
Recent advancements in diffusion-based models have demonstrated significant success in generating images from text. However, video editing models have not yet reached the same level of visual quality and user control. To address this, we introduce RAVE, a zero-shot video editing method that leverages pre-trained text-to-image diffusion models without additional training. RAVE takes an input video and a text prompt to produce high-quality videos while preserving the original motion and semantic structure. It employs a novel noise shuffling strategy, leveraging spatio-temporal interactions between frames, to produce temporally consis- tent videos faster than existing methods. It is also efficient in terms of memory requirements, allowing it to handle longer videos. RAVE is capable of a wide range of edits, from local attribute modifications to shape transformations. In order to demonstrate the versatility of RAVE, we create a com- prehensive video evaluation dataset ranging from object- focused scenes to complex human activities like dancing and typing, and dynamic scenes featuring swimming fish and boats. Our qualitative and quantitative experiments highlight the effectiveness of RAVE in diverse video editing scenarios compared to existing methods. Our code, dataset and videos can be found in our project webpage.
Özgür Kara, Bariscan Kurtkaya, Hidir Yesiltepe, James M. Rehg, Pinar Yanardag Delul
CVPR5
2024 CONFORM: Contrast is All You Need For High-Fidelity Text-to-Image Diffusion Models
abstract
Images produced by text-to-image diffusion models might not always faithfully represent the semantic intent of the provided text prompt, where the model might overlook or entirely fail to produce certain objects. Existing solutions often require customly tailored functions for each of these problems, leading to sub-optimal results, especially for complex prompts. Our work introduces a novel perspective by tackling this challenge in a contrastive context. Our approach intuitively promotes the segregation of objects in attention maps while also maintaining that pairs of related attributes are kept close to each other. We conduct extensive experiments across a wide variety of scenarios, each involving unique combinations of objects, attributes, and scenes. These experiments effectively showcase the versatil-ity, efficiency, and flexibility of our method in working with both latent and pixel-based diffusion models, including Sta-ble Diffusion and Imagen. Moreover, we publicly share our source code to facilitate further research.
Tuna Han Salih Meral, Enis Simsar, Federico Tombari, Pinar Yanardag Delul
CVPR4
2024 ORACLE: Leveraging Mutual Information for Consistent Character Generation with LoRAs in Diffusion Models
Kiymet Akdemir, Pinar Yanardag Delul
ICCC2
2024 Stylebreeder: Exploring and Democratizing Artistic Styles through Text-to-Image Models
abstract
Text-to-image models are becoming increasingly popular, revolutionizing the landscape of digital art creation by enabling highly detailed and creative visual content generation. These models have been widely employed across various domains, particularly in art generation, where they facilitate a broad spectrum of creative expression and democratize access to artistic creation. In this paper, we introduce STYLEBREEDER, a comprehensive dataset of 6.8M images and 1.8M prompts generated by 95K users on Artbreeder, a platform that has emerged as a significant hub for creative exploration with over 13M users. We introduce a series of tasks with this dataset aimed at identifying diverse artistic styles, generating personalized content, and recommending styles based on user interests. By documenting unique, user-generated styles that transcend conventional categories like 'cyberpunk' or 'Picasso,' we explore the potential for unique, crowd-sourced styles that could provide deep insights into the collective creative psyche of users worldwide. We also evaluate different personalization methods to enhance artistic expression and introduce a style atlas, making these models available in LoRA format for public use. Our research demonstrates the potential of text-to-image diffusion models to uncover and promote unique artistic expressions, further democratizing AI in art and fostering a more diverse and inclusive artistic community. The dataset, code, and models are available at https://stylebreeder.github.io under a Public Domain (CC0) license.
Matthew Zheng, Enis Simsar, Hidir Yesiltepe, Federico Tombari, Joel Simon, Pinar Yanardag Delul
NeurIPS6
2023 Text and Image Guided 3D Avatar Generation and Manipulation
abstract
The manipulation of latent space has recently become an interesting topic in the field of generative models. Recent research shows that latent directions can be used to manipulate images towards certain attributes. However, controlling the generation process of 3D generative models remains a challenge. In this work, we propose a novel 3D manipulation method that can manipulate both the shape and texture of the model using text or image-based prompts such as ’a young face’ or ’a surprised face’. We leverage the power of Contrastive Language-Image Pre-training (CLIP) model and a pre-trained 3D GAN model designed to generate face avatars and create a fully differentiable rendering pipeline to manipulate meshes. More specifically, our method takes an input latent code and modifies it such that the target attribute specified by a text or image prompt is present or enhanced while leaving other attributes largely unaffected. Our method requires only 5 minutes per manipulation, and we demonstrate the effectiveness of our approach with extensive results and comparisons.
Zehranaz Canfes, M. Furkan Atasoy, Alara Dirik, Pinar Yanardag Delul
WACV4
2023 Fantastic Style Channels and Where to Find Them: A Submodular Framework for Discovering Diverse Directions in GANs
abstract
The discovery of interpretable directions in the latent spaces of pre-trained GAN models has recently become a popular topic. In particular, StyleGAN2 has enabled various image generation and manipulation tasks due to its rich and disentangled latent spaces. However, the discovery of such directions is typically made either in a supervised manner, which requires annotated data for each desired manipulation, or in an unsupervised manner, which requires a manual effort to identify the directions. As a result, existing work typically finds only a handful of directions in which controllable edits can be made. In this study, we design a novel submodular framework that finds the most representative and diverse subset of directions in the latent space of StyleGAN2. Our approach takes advantage of the latent space of channel-wise style parameters, so-called stylespace, in which we cluster channels that perform similar manipulations into groups. Our framework promotes diversity by using the notion of clusters and can be efficiently solved with a greedy optimization scheme. We evaluate our framework with qualitative and quantitative experiments and show that our method finds more diverse and disentangled directions.
Enis Simsar, Umut Kocasari, Ezgi Gülperi Er, Pinar Yanardag Delul
WACV4
2022 MIDISpace: Finding Linear Directions in Latent Space for Music Generation
abstract
While recent work has shown that it is possible to find disentangled directions in the latent space of image generative networks, finding directions in the latent space of sequential models for music generation remains a largely unexplored topic. In this work, we propose a method for discovering linear directions in the latent space of a musicgenerating Variational Auto-Encoder (VAE). We use PCA, a statistical method, to transform the input data such that the variation along the new axes is maximized. We apply PCA to the latent space activations of our model and find largely disentangled directions that change the style and characteristics of the input music. Our experiments show that the found directions are often monotonic, global and encode fundamental musical characteristics such as colorfulness, speed, and repetitiveness. Moreover, we propose a set of quantitative metrics to describe different musical styles and characteristics to evaluate our results. We show that the found directions decouple content and can be utilized for style transfer and conditional music generation tasks. Our project page can be found at http://catlab-team.github.io/midispace.
Meliksah Turker, Alara Dirik, Pinar Yanardag Delul
Creativity & Cognition3
2022 FairStyle: Debiasing StyleGAN2 with Style Channel Manipulations
Cemre Karakas, Alara Dirik, Eylul Yalcinkaya, Pinar Yanardag Delul
ECCV (13)4
2022 StyleMC: Multi-Channel Based Fast Text-Guided Image Generation and Manipulation
abstract
Discovering meaningful directions in the latent space of GANs to manipulate semantic attributes typically requires large amounts of labeled data. Recent work aims to overcome this limitation by leveraging the power of Contrastive Language-Image Pre-training (CLIP), a joint text-image model. While promising, these methods require several hours of preprocessing or training to achieve the desired manipulations. In this paper, we present StyleMC, a fast and efficient method for text-driven image generation and manipulation. StyleMC uses a CLIP-based loss and an identity loss to manipulate images via a single text prompt without significantly affecting other attributes. Unlike prior work, StyleMC requires only a few seconds of training per text prompt to find stable global directions, does not require prompt engineering and can be used with any pre-trained StyleGAN2 model. We demonstrate the effectiveness of our method and compare it to state-of-the-art methods. Our code can be found at http://catlab-team.github.io/stylemc.
Umut Kocasari, Alara Dirik, Mert Tiftikci, Pinar Yanardag Delul
WACV4
2021 Shelley: A Crowd-sourced Collaborative Horror Writer
abstract
Fear induction in the form of stories and visual images pervades the history of human culture. Creating a visceral emotion such as fear remains one of the cornerstones of human creativity. As artificial intelligence makes strides in solving challenging analytical problems like chess and Go, an important question still remains: can machines induce extreme human emotions, such as fear? In this work, we propose a deep-learning based collaborative horror writer that collaboratively writes scary stories with people on Twitter. We deploy our system as a bot on Twitter that regularly generates and posts new stories on Twitter, and invites users to participate. Users who interact with the stories produce multiple storylines originating from the same tweet, thereby creating a tree-based story structure. We further perform a validation study on n = 105 subjects to verify whether the generated stories psychologically move people on psychometrically validated measures of effect and anxiety such as I-PANAS-SF [43] and STAI-SF [26]. Our experiments show that 1) stories generated by our bot as well as the stories generated collaboratively between our bot and Twitter users produced statistically significant increases in negative affect and state anxiety compared to the control condition, and 2) collaborated stories are more successful in terms of increasing negative affect and state anxiety than the machine-generated ones. Furthermore, we make three novel datasets used in our framework publicly available at https://github.com/catlab-team/shelley for encouraging further research on this topic.
Pinar Yanardag Delul, Manuel Cebrián, Iyad Rahwan
Creativity & Cognition1
2021 Nightmare Machine: A Large-Scale Study to Induce Fear Using Artificial Intelligence
Pinar Yanardag Delul, Nick Obradovich, Manuel Cebrián, Iyad Rahwan
ICCC1
2021 LatentCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions
abstract
Recent research has shown that it is possible to find interpretable directions in the latent spaces of pre-trained Generative Adversarial Networks (GANs). These directions enable controllable image generation and support a wide range of semantic editing operations, such as zoom or rotation. The discovery of such directions is often done in a supervised or semi-supervised manner and requires manual annotations which limits their use in practice. In comparison, unsupervised discovery allows finding subtle directions that are difficult to detect a priori. In this work, we propose a contrastive learning-based approach to discover semantic directions in the latent space of pre-trained GANs in a self-supervised manner. Our approach finds semantically meaningful dimensions compatible with state-of-the-art methods.
Oguz Kaan Yüksel, Enis Simsar, Ezgi Gülperi Er, Pinar Yanardag Delul
ICCV4
2016 WordRank: Learning Word Embeddings via Robust Ranking
abstract
Embedding words in a vector space has gained a lot of attention in recent years.While stateof-the-art methods provide efficient computation of word similarities via a low-dimensional matrix embedding, their motivation is often left unclear.In this paper, we argue that word embedding can be naturally viewed as a ranking problem due to the ranking nature of the evaluation metrics.Then, based on this insight, we propose a novel framework Wor-dRank that efficiently estimates word representations via robust ranking, in which the attention mechanism and robustness to noise are readily achieved via the DCG-like ranking losses.The performance of WordRank is measured in word similarity and word analogy benchmarks, and the results are compared to the state-of-the-art word embedding techniques.Our algorithm is very competitive to the state-of-the-arts on large corpora, while outperforms them by a significant margin when the training set is limited (i.e., sparse and noisy).With 17 million tokens, WordRank performs almost as well as existing methods using 7.2 billion tokens on a popular word similarity benchmark.Our multi-node distributed implementation of WordRank is publicly available for general usage.
Shihao Ji 0001, Hyokun Yun, Pinar Yanardag Delul, Shin Matsushima, S. V. N. Vishwanathan
EMNLP3
2015 Deep Graph Kernels
abstract
In this paper, we present Deep Graph Kernels, a unified framework to learn latent representations of sub-structures for graphs, inspired by latest advancements in language modeling and deep learning. Our framework leverages the dependency information between sub-structures by learning their latent representations. We demonstrate instances of our framework on three popular graph kernels, namely Graphlet kernels, Weisfeiler-Lehman subtree kernels, and Shortest-Path graph kernels. Our experiments on several benchmark datasets show that Deep Graph Kernels achieve significant improvements in classification accuracy over state-of-the-art graph kernels.
Pinar Yanardag Delul, S. V. N. Vishwanathan
KDD1
2015 A Structural Smoothing Framework For Robust Graph Comparison
abstract
In this paper, we propose a general smoothing framework for graph kernels by taking \textit{structural similarity} into account, and apply it to derive smoothed variants of popular graph kernels. Our framework is inspired by state-of-the-art smoothing techniques used in natural language processing (NLP). However, unlike NLP applications which primarily deal with strings, we show how one can apply smoothing to a richer class of inter-dependent sub-structures that naturally arise in graphs. Moreover, we discuss extensions of the Pitman-Yor process that can be adapted to smooth structured objects thereby leading to novel graph kernels. Our kernels are able to tackle the diagonal dominance problem, while respecting the structural similarity between sub-structures, especially under the presence of edge or label noise. Experimental evaluation shows that not only our kernels outperform the unsmoothed variants, but also achieve statistically significant improvements in classification accuracy over several other graph kernels that have been recently proposed in literature. Our kernels are competitive in terms of runtime, and offer a viable option for practitioners.
Pinar Yanardag Delul, S. V. N. Vishwanathan
NIPS1
2014 Crowdsourced Resource-Sizing of Virtual Appliances
abstract
Using a population of VMware Virtual Center Virtual Ap- pliances (VCVA) and their respective workloads we de- scribe techniques for constructing a model of their resource consumption and performance, speci cally memory require- ments, and average operation-latency by mining logs of ap- plication (VCVA) performance. We use our model to provide sizing recommendations for the virtual appliance and iden- tify features that can be used to provide rough estimates of expected memory consumption. We show results of bet- ter than 70% prediction accuracy (recall) for predicting Physical Memory Usage and better than 80% prediction accuracy (recall) for predicting the average latency of work- load operations. We describe modeling techniques from sta- tistical machine learning that are amenable to representing complex, non-linear systems. Further, via the choice of tech- niques, we present an approach for reasoning about the lim- itations of our model, i.e., identifying when (and why) our model is expected to perform well and poorly.
Pinar Yanardag Delul, Rean Griffith, Anne Holler, K. Shankari, Xiaoyun Zhu, Ravi Soundararajan, Adarsh Jagadeeshwaran, Pradeep Padala
IEEE CLOUD1