Ugur Güdükbay

dblp:26/4312 · DBLP profile ↗
← Back
73ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0003-2462-6959ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 49 · 9 first-author · 12 since 2021Artificial intelligence and machine learning · 13 · 4 since 2021Databases, data management, data science and information retrieval · 5Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Data-driven Inverse Kinematics using Laban Movement Analysis
abstract
Inverse Kinematics (IK) provides control over animation, facilitating the creation of full-body poses by utilizing target end-effector locations. Many approaches address the physical aspects of arranging limb configurations; however, systems that consider the psychological aspects of human motion are lacking. To this end, we introduce a quantitative translation of the qualitative concepts of Laban Movement Analysis (LMA) into computable, continuous style descriptors. Building upon this formulation, we also propose a data-driven Inverse Kinematics (IK) method that directly utilizes these LMA parameters to refine generated animations. Specifically, we refer to LMA Shape Qualities and the attitude towards the Kinesphere to control the orientation of the generated pose along the vertical and horizontal axes. Our Interpolator upsamples sparse end-effector keyframes into dense paths and modulates Time Effort at the trajectory level. Flow Effort is controlled by a pose-similarity objective that deliberately reduces pose similarity to the dataset examples. Through a perception user study, we show that the system can successfully apply LMA-based changes to the motion to express different personality traits. This data-driven system can ease the process of controlling the psychological aspect of generative animation.
Mehmet Akif Sahin, Sinan Sonlu, Ugur Güdükbay
Comput. Graph.3
2026 Direct volume rendering of tree-based tetrahedral adaptive mesh refinement data
Musa Ege Ünalan, Serkan Demirci, Stefan Zellmann, Ugur Güdükbay
Comput. Graph.4
2025 Effects of Embodiment and Personality in LLM-Based Conversational Agents
abstract
This work investigates the effects of personality expression and embodiment in conversational agents. We extend a personality-driven conversational agent framework by integrating LLM-based conversation support to provide information about contemporary scientific topics. We describe a user study built on this system to evaluate two opposing personality styles using three models: a dialogue-only model that conveys personality verbally, an animated human model that expresses personality only through dialogue, and an animated human model expressing personality through dialogue and expressive animations. The users perceive all models positively regarding personality and learning outcomes; however, models with high personality traits are perceived as more engaging than those with low personality traits. We provide an analysis of personality perception, learning, and user experience.
Sinan Sonlu, Bennie Bendiksen, Funda Durupinar, Ugur Güdükbay
VR4
2025 Talk With Socrates: Relation Between Perceived Agent Personality and User Personality in LLM-Based Natural Language Dialogue Using Virtual Reality
abstract
ABSTRACT Large Language Models (LLMs) offer almost immediate human‐like quality responses to user queries. Conversational agent systems support natural language dialogues utilizing LLM backends in combination with Text‐to‐Speech (TTS) and Automatic Speech Recognition (ASR) technologies, enabling life‐like characters in virtual environments. This study investigates the relationship between user personality and perceived agent personality in LLM‐based natural language dialogue. We adopt a Virtual Reality (VR) setting where the user can talk with the agent that assumes the role of Socrates, the famous philosopher. To this end, we utilize a three‐dimensional (3D) avatar model resembling Socrates and use specific LLM prompts to get stylistic answers from OpenAI's Chat Completions Application Programming Interface (API). Our user study measures the agent's personality and the system's ease of use, quality, realism, and immersion concerning the user's self‐reported personality. The results suggest that the user's conscientiousness, extraversion, and emotional stability have a moderate effect on certain personality factors and system qualities. User conscientiousness affects the perceived ease of use, quality, and realism, while user extraversion affects perceived agent conscientiousness, system realism, and immersion. Additionally, the user's emotional stability correlates with perceived extraversion and agreeableness.
Mehmet Efe Sak, Sinan Sonlu, Ugur Güdükbay
Comput. Animat. Virtual Worlds3
2025 Personality Expression Using Co-Speech Gesture
abstract
We express our personality through verbal and nonverbal behavior. While verbal cues are mostly related to the semantics of what we say, nonverbal cues include our posture, gestures, and facial expressions. Appropriate expression of these behavioral elements improves conversational virtual agents’ communication capabilities and realism. Although previous studies focus on co-speech gesture generation, they do not consider the personality aspect of the synthesized animations. We show that automatically generated co-speech gestures naturally express personality traits, and heuristics-based adjustments for such animations can further improve personality expression. To this end, we present a framework for enhancing co-speech gestures with the different personalities of the Five-Factor model. Our experiments suggest that users perceive increased realism and improved personality expression when combining heuristics-based motion adjustments with co-speech gestures.
Sinan Sonlu, Halil Özgür Demir, Ugur Güdükbay
ACM Trans. Appl. Percept.3
2025 Visualization of Large Non-Trivially Partitioned Unstructured Data With Native Distribution on High-Performance Computing Systems
abstract
Interactively visualizing large finite element simulation data on High-Performance Computing (HPC) systems poses several difficulties. Some of these relate to unstructured data, which, even on a single node, is much more expensive to render compared to structured volume data. Worse yet, in the data parallel rendering context, such data with highly non-convex spatial domain boundaries will cause rays along its silhouette to enter and leave a given rank's domains at different distances. This straddling, in turn, poses challenges for both ray marching, which usually assumes successive elements to share a face, and compositing, which usually assumes a single fragment per pixel per rank. We holistically address these issues using a combination of three inter-operating techniques: first, we use a highly optimized GPU ray marching technique that, given an entry point, can march a ray to its exit point with high-performance by exploiting an exclusive-or (XOR) based compaction scheme. Second, we use hardware-accelerated ray tracing to efficiently find the proper entry points for these marching operations. Third, we use a "deep" compositing scheme to properly handle cases where different ranks' ray segments interleave in depth. We use GPU-to-GPU remote direct memory access (RDMA) to achieve interactive frame rates of 10-15 frames per second and higher for our motivating use case, the Fun3D NASA Mars Lander.
Alper Sahistan, Serkan Demirci, Ingo Wald, Stefan Zellmann, João Barbosa, Nathan Morrical, Ugur Güdükbay
IEEE Trans. Vis. Comput. Graph.7
2024 Personality perception in human videos altered by motion transfer networks
Ayda Yurtoglu, Sinan Sonlu, Yalim Dogan, Ugur Güdükbay
Comput. Graph.4
2024 Learning visual similarity for image retrieval with global descriptors and capsule networks
Duygu Durmus, Ugur Güdükbay, Özgür Ulusoy
Multim. Tools Appl.2
2024 Point cloud registration with quantile assignment
abstract
Abstract Point cloud registration is a fundamental problem in computer vision. The problem encompasses critical tasks such as feature estimation, correspondence matching, and transformation estimation. The point cloud registration problem can be cast as a quantile matching problem. We refined the quantile assignment algorithm by integrating prevalent feature descriptors and transformation estimation methods to enhance the correspondence between the source and target point clouds. We evaluated the performances of these descriptors and methods with our approach through controlled experiments on a dataset we constructed using well-known 3D models. This systematic investigation led us to identify the most suitable methods for complementing our approach. Subsequently, we devised a new end-to-end, coarse-to-fine pairwise point cloud registration framework. Finally, we tested our framework on indoor and outdoor benchmark datasets and compared our results with state-of-the-art point cloud registration methods.
Ecenur Oguz, Yalim Dogan, Ugur Güdükbay, Oya Ekin Karasan, Mustafa Ç. Pinar
Mach. Vis. Appl.3
2024 Refining 3D Human Texture Estimation From a Single Image
abstract
Estimating 3D human texture from a single image is essential in graphics and vision. It requires learning a mapping function from input images of humans with diverse poses into the parametric (uv) space and reasonably hallucinating invisible parts. To achieve a high-quality 3D human texture estimation, we propose a framework that adaptively samples the input by a deformable convolution where offsets are learned via a deep neural network. Additionally, we describe a novel cycle consistency loss that improves view generalization. We further propose to train our framework with an uncertainty-based pixel-level image reconstruction loss, which enhances color fidelity. We compare our method against the state-of-the-art approaches and show significant qualitative and quantitative improvements.
Said Fahri Altindis, Adil Meric, Yusuf Dalva, Ugur Güdükbay, Aysegul Dundar
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 State-of-the-art in Large-Scale Volume Visualization Beyond Structured Data
abstract
Abstract Volume data these days is usually massive in terms of its topology, multiple fields, or temporal component. With the gap between compute and memory performance widening, the memory subsystem becomes the primary bottleneck for scientific volume visualization. Simple, structured, regular representations are often infeasible because the buses and interconnects involved need to accommodate the data required for interactive rendering. In this state‐of‐the‐art report, we review works focusing on large‐scale volume rendering beyond those typical structured and regular grid representations. We focus primarily on hierarchical and adaptive mesh refinement representations, unstructured meshes, and compressed representations that gained recent popularity. We review works that approach this kind of data using strategies such as out‐of‐core rendering, massive parallelism, and other strategies to cope with the sheer size of the ever‐increasing volume of data produced by today's supercomputers and acquisition devices. We emphasize the data management side of large‐scale volume rendering systems and also include a review of tools that support the various volume data types discussed.
Jonathan Sarton, Stefan Zellmann, Serkan Demirci, Ugur Güdükbay, Welcome Alexandre-Barff, Laurent Lucas, Jean-Michel Dischler, Stefan Wesner, Ingo Wald
Comput. Graph. Forum4
2023 Personality expression in cartoon animal characters using Sasang typology
abstract
Abstract The movement style is an adequate descriptor of different personalities. While many studies investigate the relationship between apparent personality and high‐level motion qualities in humans, similar research for animal characters still needs to be done. The variety in animals' skeletal configurations and texture complicates their pose estimation process. Our affect analysis framework includes a workflow for pose extraction in animal characters and a parameterization of the high‐level animal motion descriptors inspired by Laban movement analysis. Using a data set of quadruped walk cycles, we prove the display of typologies in cartoon animal characters, reporting the point‐biserial correlation between our motion parameters and the Sasang categories that reflect different personalities.
Hamila Mailee, Sinan Sonlu, Ugur Güdükbay
Comput. Animat. Virtual Worlds3
2023 Quick Clusters: A GPU-Parallel Partitioning for Efficient Path Tracing of Unstructured Volumetric Grids
abstract
We propose a simple yet effective method for clustering finite elements to improve preprocessing times and rendering performance of unstructured volumetric grids without requiring auxiliary connectivity data. Rather than building bounding volume hierarchies (BVHs) over individual elements, we sort elements along with a Hilbert curve and aggregate neighboring elements together, improving BVH memory consumption by over an order of magnitude. Then to further reduce memory consumption, we cluster the mesh on the fly into sub-meshes with smaller indices using a series of efficient parallel mesh re-indexing operations. These clusters are then passed to a highly optimized ray tracing API for point containment queries and ray-cluster intersection testing. Each cluster is assigned a maximum extinction value for adaptive sampling, which we rasterize into non-overlapping view-aligned bins allocated along the ray. These maximum extinction bins are then used to guide the placement of samples along the ray during visualization, reducing the number of samples required by multiple orders of magnitude (depending on the dataset), thereby improving overall visualization interactivity. Using our approach, we improve rendering performance over a competitive baseline on the NASA Mars Lander dataset from 6× (1 frame per second (fps) and 1.0 M rays per second (rps) up to now 6 fps and 12.4 M rps, now including volumetric shadows) while simultaneously reducing memory consumption by 3×(33 GB down to 11 GB) and avoiding any offline preprocessing steps, enabling high-quality interactive visualization on consumer graphics cards. Then by utilizing the full 48 GB of an RTX 8000, we improve the performance of Lander by 17 × (1 fps up to 17 fps, 1.0 M rps up to 35.6 M rps).
Nathan Morrical, Alper Sahistan, Ugur Güdükbay, Ingo Wald, Valerio Pascucci
IEEE Trans. Vis. Comput. Graph.3
2021 Privacy-preserving and robust watermarking on sequential genome data using belief propagation and local differential privacy
abstract
MOTIVATION: Genome data is a subject of study for both biology and computer science since the start of the Human Genome Project in 1990. Since then, genome sequencing for medical and social purposes becomes more and more available and affordable. Genome data can be shared on public websites or with service providers (SPs). However, this sharing compromises the privacy of donors even under partial sharing conditions. We mainly focus on the liability aspect ensued by the unauthorized sharing of these genome data. One of the techniques to address the liability issues in data sharing is the watermarking mechanism. RESULTS: To detect malicious correspondents and SPs-whose aim is to share genome data without individuals' consent and undetected-, we propose a novel watermarking method on sequential genome data using belief propagation algorithm. In our method, we have two criteria to satisfy. (i) Embedding robust watermarks so that the malicious adversaries cannot temper the watermark by modification and are identified with high probability. (ii) Achieving ϵ-local differential privacy in all data sharings with SPs. For the preservation of system robustness against single SP and collusion attacks, we consider publicly available genomic information like Minor Allele Frequency, Linkage Disequilibrium, Phenotype Information and Familial Information. Our proposed scheme achieves 100% detection rate against the single SP attacks with only 3% watermark length. For the worst case scenario of collusion attacks (50% of SPs are malicious), 80% detection is achieved with 5% watermark length and 90% detection is achieved with 10% watermark length. For all cases, the impact of ϵ on precision remained negligible and high privacy is ensured. AVAILABILITY AND IMPLEMENTATION: https://github.com/acoksuz/PPRW\_SGD\_BPLDP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Abdullah Çaglar Öksüz, Erman Ayday, Ugur Güdükbay
Bioinform.3
2021 An augmented crowd simulation system using automatic determination of navigable areas
Yalim Dogan, Sinan Sonlu, Ugur Güdükbay
Comput. Graph.3
2021 Multimodal assessment of apparent personality using feature attention and error consistency constraint
Süleyman Aslan, Ugur Güdükbay, Hamdi Dibeklioglu
Image Vis. Comput.2
2021 Multi-level tetrahedralization-based accelerator for ray-tracing animated scenes
abstract
Abstract We describe a hybrid acceleration structure for ray tracing. The hybrid structure is a Bounding Volume Hierarchy (BVH) where the leaf nodes are tetrahedralized for a decent ray‐surface intersection performance. We use the hybrid acceleration structure (BTH) in a two‐level acceleration structure for rendering animated scenes. There is a BVH at the top level in this two‐level structure and the proposed hybrid structure (BTH) at the bottom level. We test the proposed two‐level structure (BVH‐BTH) for various animated scenes and obtained promising results against other acceleration structures in terms of rendering times. The two‐level BVH‐BTH structure outperforms the two‐level BVH structure for the tested dynamic scenes.
Aytek Aman, Serkan Demirci, Ugur Güdükbay, Ingo Wald
Comput. Animat. Virtual Worlds3
2020 Deep Convolutional Generative Adversarial Networks for Flame Detection in Video
Süleyman Aslan, Ugur Güdükbay, B. Ugur Töreyin, A. Enis Çetin
ICCCI2
2020 Guido: Augmented Reality for Indoor Navigation Using Commodity Hardware
abstract
Indoor positioning is one of the difficult problems in current navigation systems. There is an increasing demand for detecting the locations of objects and humans inside closed environments in various fields including surveillance, robotics, and entertainment. Recent works focus on indoor navigation systems using different technologies including Wireless Local Area Network (WLAN), Radio Frequency Identification (RFID), Inertial Measurement Unit (IMU), and Simultaneous Localization and Mapping (SLAM). Research reveals using these technologies alone is inefficient in terms of accuracy and cost. To address this issue, we propose a marker-based Augmented Reality (AR) indoor navigation system with integrated SLAM and IMU. We use Unity's AR Foundation Framework for highly accurate results with minimum hardware requirements.
Zafer Tan Çankiri, Erdem Ege Marasli, Sait Aktürk, Sinan Sonlu, Ugur Güdükbay
IV5
2020 Recognition of occupational therapy exercises and detection of compensation mistakes for Cerebral Palsy
Mehmet Faruk Ongun, Ugur Güdükbay, Selim Aksoy
J. Vis. Commun. Image Represent.2
2019 Early Wildfire Smoke Detection Based on Motion-based Geometric Image Transformation and Deep Convolutional Generative Adversarial Networks
abstract
Early detection of wildfire smoke in real-time is essentially important in forest surveillance and monitoring systems. We propose a vision-based method to detect smoke using Deep Convolutional Generative Adversarial Neural Networks (DC-GANs). Many existing supervised learning approaches using convolutional neural networks require substantial amount of labeled data. In order to have a robust representation of sequences with and without smoke, we propose a two-stage training of a DCGAN. Our training framework includes, the regular training of a DCGAN with real images and noise vectors, and training the discriminator separately using the smoke images without the generator. Before training the networks, the temporal evolution of smoke is also integrated with a motion-based transformation of images as a pre-processing step. Experimental results show that the proposed method effectively detects the smoke images with negligible false positive rates in real-time.
Süleyman Aslan, Ugur Güdükbay, B. Ugur Töreyin, A. Enis Çetin
ICASSP2
2019 Augmentation of Virtual Agents in Real Crowd Videos
abstract
Augmentation of virtual agents in real crowd videos is an important task for different applications from design simulations of social environments to modeling abnormalities in crowd behavior. We propose a framework for this task, namely for augmenting virtual agents in real crowd videos. Our framework utilizes homography-based video stabilization, Dalal-Triggs detector [1] for pedestrian detection and state-based tracking algorithms to automatically locate the pedestrians in video frames and project them into our 3D simulated environment, where the navigable area of the simulated environment is available as a manually designed and located navigation mesh. We represent the real pedestrians in the video as simple three-dimensional (3D) models in our simulation environment. 3D models representing real, projected agents and the augmented virtual agents are simulated using local path planning coupled with a collision detection and avoidance algorithm, called Reciprocal Velocity Obstacles (RVO) [2]. The virtual agents augmented into the video move plausibly without colliding with static and dynamic obstacles, including other virtual agents and real pedestrians. We provide an extensive graphical user interface for controlling the virtual agents in the scene, including collision avoidance parameters, adjusting the camera in the scene and some standard video player options.
Yalim Dogan, Serkan Demirci, Ugur Güdükbay
VR3
2018 Using real life incidents for creating realistic virtual crowds with data-driven emotion contagion
Ahmet Eren Basak, Ugur Güdükbay, Funda Durupinar
Comput. Graph.2
2018 A group-based approach for gaze behavior of virtual crowds incorporating personalities
abstract
Abstract Predicting interest points of virtual characters and accurately simulating their gaze behavior play a significant role for realistic crowd simulations. We propose a saliency model that enables virtual agents to produce plausible gaze behavior. The model measures the effects of distinct saliency features implemented by examining the state‐of‐the‐art perception studies. When predicting an agent's interest point, we compute the saliency scores by using a weighted sum function for other agents and environment objects in the field of view of the agent for each frame. Then, we determine the most salient entity for each agent in the scene; thus, agents gain a visual understanding of their environment. Besides, our model introduces new aspects to crowd perception, such as perceiving characters as groups of people and applying social norms on crowd gaze behavior, effects of agent personality on gaze, gaze copy phenomena, and effects of agent velocity on attention. For evaluation, we compare the resulting saliency gaze model with real‐world crowd behavior in captured videos. In the experiments, we simulate the gaze behavior in real crowds. The results show that the proposed approach generates plausible gaze behaviors and is easily adaptable to varying scenarios for virtual crowds.
Umut Agil, Ugur Güdükbay
Comput. Animat. Virtual Worlds2
2017 ACMICS: an agent communication model for interacting crowd simulation
Kurtulus Kullu, Ugur Güdükbay, Dinesh Manocha
Auton. Agents Multi Agent Syst.2
2017 PETAL: A fully distributed location service for wireless ad hoc networks
Amir Rahimzadeh Ilkhechi, Ibrahim Korpeoglu, Ugur Güdükbay, Özgür Ulusoy
J. Netw. Comput. Appl.3
2017 Mobile multi-view object image search
Fatih Çalisir, Muhammet Bastan, Özgür Ulusoy, Ugur Güdükbay
Multim. Tools Appl.4
2016 Psychological Parameters for Crowd Simulation: From Audiences to Mobs
abstract
In the social psychology literature, crowds are classified as audiences and mobs. Audiences are passive crowds, whereas mobs are active crowds with emotional, irrational and seemingly homogeneous behavior. In this study, we aim to create a system that enables the specification of different crowd types ranging from audiences to mobs. In order to achieve this goal we parametrize the common properties of mobs to create collective misbehavior. Because mobs are characterized by emotionality, we describe a framework that associates psychological components with individual agents comprising a crowd and yields emergent behaviors in the crowd as a whole. To explore the effectiveness of our framework we demonstrate two scenarios simulating the behavior of distinct mob types.
Funda Durupinar, Ugur Güdükbay, Aytek Aman, Norman I. Badler
IEEE Trans. Vis. Comput. Graph.2
2015 A hand gesture recognition technique for human-computer interaction
Nurettin Çagri Kiliboz, Ugur Güdükbay
J. Vis. Commun. Image Represent.2
2014 Real-time virtual fitting with body measurement and motion smoothing
Umut Gültepe, Ugur Güdükbay
Comput. Graph.2
2014 Application-Specific Heterogeneous Network-on-Chip Design
abstract
As a result of increasing communication demands, application-specific and scalable Network-on-Chips (NoCs) have emerged to connect processing cores and subsystems in Multiprocessor System-on-Chips. A challenge in application-specific NoC design is to find the right balance among different tradeoffs, such as communication latency, power consumption and chip area. We propose a novel approach that generates latency-aware heterogeneous NoC topology. Experimental results show that our approach improves the total communication latency up to 27% with modest power consumption.
Dilek Demirbas, Ismail Akturk, Ozcan Ozturk 0001, Ugur Güdükbay
Comput. J.4
2014 A hybrid representation for modeling, interactive editing, and real-time visualization of terrains with volumetric features
abstract
Terrain rendering is a crucial part of many real-time applications. The easiest way to process and visualize terrain data in real time is to constrain the terrain model in several ways. This decreases the amount of data to be processed and the amount of processing power needed, but at the cost of expressivity and the ability to create complex terrains. The most popular terrain representation is a regular 2D grid, where the vertices are displaced in a third dimension by a displacement map, called a heightmap. This is the simplest way to represent terrain, and although it allows fast processing, it cannot model terrains with volumetric features. Volumetric approaches sample the 3D space by subdividing it into a 3D grid and represent the terrain as occupied voxels. They can represent volumetric features but they require computationally intensive algorithms for rendering, and their memory requirements are high. We propose a novel representation that combines the voxel and heightmap approaches, and is expressive enough to allow creating terrains with caves, overhangs, cliffs, and arches, and efficient enough to allow terrain editing, deformations, and rendering in real time.
Çetin Koca, Ugur Güdükbay
Int. J. Geogr. Inf. Sci.2
2014 Direct volume rendering of unstructured tetrahedral meshes using CUDA and OpenMP
Erhan Okuyan, Ugur Güdükbay
J. Supercomput.2
2013 Dynamic point-region quadtrees for particle simulations
Oguzcan Oguz, Funda Durupinar, Ugur Güdükbay
Inf. Sci.3
2011 Nearest-Neighbor based Metric Functions for indoor scene recognition
Fatih Çakir, Ugur Güdükbay, Özgür Ulusoy
Comput. Vis. Image Underst.2
2010 Emergency crowd simulation for outdoor environments
Oguzcan Oguz, Ates Akaydin, Türker Yilmaz, Ugur Güdükbay
Comput. Graph.4
2010 Fuzzy color histogram-based video segmentation
Onur Küçüktunç, Ugur Güdükbay, Özgür Ulusoy
Comput. Vis. Image Underst.2
2010 Scenario-based query processing for video-surveillance archives
Ediz Saykol, Ugur Güdükbay, Özgür Ulusoy
Eng. Appl. Artif. Intell.2
2010 3D Model compression using Connectivity-Guided Adaptive Wavelet Transform built into 2D SPIHT
Kivanç Köse, A. Enis Çetin, Ugur Güdükbay, Levent Onural
J. Vis. Commun. Image Represent.3
2010 Video copy detection using multiple visual cues and MPEG-7 descriptors
Onur Küçüktunç, Muhammet Bastan, Ugur Güdükbay, Özgür Ulusoy
J. Vis. Commun. Image Represent.3
2009 Special issue on advances in three-dimensional television and video: Guest editorial
Ugur Güdükbay, A. Aydin Alatan
Signal Process. Image Commun.1
2009 Rate-Distortion Efficient Piecewise Planar 3-D Scene Representation From 2-D Images
abstract
In any practical application of the 2-D-to-3-D conversion that involves storage and transmission, representation efficiency has an undisputable importance that is not reflected in the attention the topic received. In order to address this problem, a novel algorithm, which yields efficient 3-D representations in the rate distortion sense, is proposed. The algorithm utilizes two views of a scene to build a mesh-based representation incrementally, via adding new vertices, while minimizing a distortion measure. The experimental results indicate that, in scenes that can be approximated by planes, the proposed algorithm is superior to the dense depth map and, in some practical situations, to the block motion vector-based representations in the rate-distortion sense.
Evren Imre, A. Aydin Alatan, Ugur Güdükbay
IEEE Trans. Image Process.3
2008 Segmentation-based extraction of important objects from video for object-based indexing
abstract
We describe a method to automatically extract important video objects for object-based indexing. Most of the existing salient object detection approaches detect visually conspicuous structures in images, while our method aims to find regions that may be important for indexing in a video database system. Our method works on a shot basis. We first segment each frame to obtain homogeneous regions in terms of color and texture. Then, we extract a set of regional and inter-regional color, shape, texture and motion features for all regions, which are classified as being important or not using SVMs trained on a few hundreds of example regions. Finally, each important region is tracked within each shot for trajectory generation and consistency check. Experimental results from news video sequences show that the proposed approach is effective.
Muhammet Bastan, Ugur Güdükbay, Özgür Ulusoy
ICME2
2008 A video-based text and equation editor for LaTeX
Özcan Öksüz, Ugur Güdükbay, A. Enis Çetin
Eng. Appl. Artif. Intell.2
2008 Automatic detection of salient objects and spatial relations in videos for a video database system
Tarkan Sevilmis, Muhammet Bastan, Ugur Güdükbay, Özgür Ulusoy
Image Vis. Comput.3
2007 Rate-Distortion Based Piecewise Planar 3D Scene Geometry Representation
abstract
This paper proposes a novel 3D piecewise planar reconstruction algorithm, to build a 3D scene representation that minimizes the intensity error between a particular frame and its prediction. 3D scene geometry is exploited to remove the visual redundancy between frame pairs for any predictive coding scheme. This approach associates the rate increase with the quality of representation, and is shown to be rate-distortion efficient by the experiments.
Evren Imre, A. Aydin Alatan, Ugur Güdükbay
ICIP (5)3
2007 A Virtual Garment Design and Simulation System
abstract
In this paper, a 3D graphics environment for virtual garment design and simulation is presented. The proposed system enables the three dimensional construction of a garment from its cloth panels, for which the underlying structure is a mass-spring model. The garment construction process is performed through automatic pattern generation, posterior correction, and seaming. Afterwards, it is possible to do fitting on virtual mannequins as if in a real life tailor's workshop. The system provides the users with the flexibility to design their own garment patterns and make changes on the garment even after the dressing of the model. Furthermore, rendering alternatives for the visualization of knitted and woven fabric are presented.
Funda Durupinar, Ugur Güdükbay
IV2
2007 Procedural visualization of knitwear and woven cloth
Funda Durupinar, Ugur Güdükbay
Comput. Graph.2
2007 Conservative occlusion culling for urban visualization using a slice-wise data structure
Türker Yilmaz, Ugur Güdükbay
Graph. Model.2
2007 Scene Representation Technologies for 3DTV - A Survey
abstract
3-D scene representation is utilized during scene extraction, modeling, transmission and display stages of a 3DTV framework. To this end, different representation technologies are proposed to fulfill the requirements of 3DTV paradigm. Dense point-based methods are appropriate for free-view 3DTV applications, since they can generate novel views easily. As surface representations, polygonal meshes are quite popular due to their generality and current hardware support. Unfortunately, there is no inherent smoothness in their description and the resulting renderings may contain unrealistic artifacts. NURBS surfaces have embedded smoothness and efficient tools for editing and animation, but they are more suitable for synthetic content. Smooth subdivision surfaces, which offer a good compromise between polygonal meshes and NURBS surfaces, require sophisticated geometry modeling tools and are usually difficult to obtain. One recent trend in surface representation is point-based modeling which can meet most of the requirements of 3DTV, however the relevant state-of-the-art is not yet mature enough. On the other hand, volumetric representations encapsulate neighborhood information that is useful for the reconstruction of surfaces with their parallel implementations for multiview stereo algorithms. Apart from the representation of 3-D structure by different primitives, texturing of scenes is also essential for a realistic scene rendering. Image-based rendering techniques directly render novel views of a scene from the acquired images, since they do not require any explicit geometry or texture representation. 3-D human face and body modeling facilitate the realistic animation and rendering of human figures that is quite crucial for 3DTV that might demand real-time animation of human bodies. Physically based modeling and animation techniques produce impressive results, thus have potential for use in a 3DTV framework for modeling and animating dynamic scenes. As a concluding remark, it can be argued that 3-D scene and texture representation techniques are mature enough to serve and fulfill the requirements of 3-D extraction, transmission and display sides in a 3DTV scenario.
A. Aydin Alatan, Yücel Yemez, Ugur Güdükbay, Xenophon Zabulis, Karsten Müller 0001, Çigdem Eroglu Erdem, C. Weigel, Aljoscha Smolic
IEEE Trans. Circuits Syst. Video Technol.3
2006 Realistic Rendering and Animation of a Multi-Layered Human Body Model
abstract
A framework for realistic rendering of a multi-layered human body model is proposed in this paper. The human model is composed of three layers: skeleton, muscle, and skin. The skeleton layer, represented by a set of joints and bones, controls the animation of the human body using inverse kinematics. Muscles are represented with action lines that are defined by a set of control points. An action line applies the force produced by a muscle on the bones and on the skin mesh. The skin layer is modeled as a 3D mesh and deformed during animation by binding the skin layer to both the skeleton and muscle layers. The skin is deformed by a two-step algorithm according to the current state of the skeleton and muscle layers. Performance experiments show that it is possible to obtain real-time frame rates for a moderately complex human model containing approximately 33,000 triangles on the skin layer.
Mehmet Sahin Yesil, Ugur Güdükbay
IV2
2006 Computer vision based method for real-time fire and flame detection
B. Ugur Töreyin, Yigithan Dedeoglu, Ugur Güdükbay, A. Enis Çetin
Pattern Recognit. Lett.3
2005 Real-Time Fire and Flame Detection in Video
abstract
The paper proposes a novel method to detect fire and/or flame by processing the video data generated by an ordinary camera monitoring a scene. In addition to ordinary motion and color clues, flame and fire flicker are detected by analyzing the video in the wavelet domain. Periodic behavior in flame boundaries is detected by performing a temporal wavelet transform. Color variations in fire are detected by computing the spatial wavelet transform of moving fire-colored regions. Other clues used in the fire detection algorithm include irregularity of the boundary of the fire-colored region and the growth of such regions in time. All of the above clues are combined to reach a final decision.
Yigithan Dedeoglu, B. Ugur Töreyin, Ugur Güdükbay, A. Enis Çetin
ICASSP (2)3
2005 A histogram-based approach for object-based query-by-shape-and-color in image and video databases
Ediz Saykol, Ugur Güdükbay, Özgür Ulusoy
Image Vis. Comput.2
2005 BilVideo: Design and Implementation of a Video Database Management System
Mehmet Emin Dönderler, Ediz Saykol, Umut Arslan, Özgür Ulusoy, Ugur Güdükbay
Multim. Tools Appl.5
2005 PHR: A Parallel Hierarchical Radiosity System with Dynamic Load Balancing
Ali Kemal Sinop, Tolga Abaci, Ümit Akkus, Attila Gürsoy, Ugur Güdükbay
J. Supercomput.5
2004 Computer vision based text and equation editor for LATEX
abstract
We present a computer vision based text and equation editor for LATEX. The user writes text and equations on paper and a camera attached to a computer records the actions of the user. In particular, positions of the pen-tip in consecutive image frames are detected. Next, directional and positional information about characters are calculated using these positions. Then, this information is used for on-line character classification. After characters and symbols are found, the corresponding LATEX code is generated.
Özcan Öksüz, Ugur Güdükbay, A. Enis Çetin
ICME2
2004 BilVideo Video Database Management System
Özgür Ulusoy, Ugur Güdükbay, Mehmet Emin Dönderler, Ediz Saykol, Cemil Alper
VLDB2
2004 An efficient query optimization strategy for spatio-temporal queries in video databases
Gulay Ünel, Mehmet Emin Dönderler, Özgür Ulusoy, Ugur Güdükbay
J. Syst. Softw.4
2004 Content-based retrieval of historical Ottoman documents stored as textual images
abstract
There is an accelerating demand to access the visual content of documents stored in historical and cultural archives. Availability of electronic imaging tools and effective image processing techniques makes it feasible to process the multimedia data in large databases. In this paper, a framework for content-based retrieval of historical documents in the Ottoman Empire archives is presented. The documents are stored as textual images, which are compressed by constructing a library of symbols occurring in a document, and the symbols in the original image are then replaced with pointers into the codebook to obtain a compressed representation of the image. The features in wavelet and spatial domain based on angular and distance span of shapes are used to extract the symbols. In order to make content-based retrieval in historical archives, a query is specified as a rectangular region in an input image and the same symbol-extraction process is applied to the query region. The queries are processed on the codebook of documents and the query images are identified in the resulting documents using the pointers in textual images. The querying process does not require decompression of images. The new content-based retrieval framework is also applicable to many other document archives using different scripts.
Ediz Saykol, Ali Kemal Sinop, Ugur Güdükbay, Özgür Ulusoy, A. Enis Çetin
IEEE Trans. Image Process.3
2004 Rule-based spatiotemporal query processing for video databases
Mehmet Emin Dönderler, Özgür Ulusoy, Ugur Güdükbay
VLDB J.3
2003 Direct volume rendering of unstructured grids
Hakan Berk, Cevdet Aykanat, Ugur Güdükbay
Comput. Graph.3
2002 Visualizer: a mesh visualization system using view-dependent refinement
Ugur Güdükbay, Okan Arikan, Bülent Özgüç
Comput. Graph.1
2002 A rule-based video database system architecture
Mehmet Emin Dönderler, Özgür Ulusoy, Ugur Güdükbay
Inf. Sci.3
2002 Stereoscopic View-Dependent Visualization of Terrain Height Fields
abstract
Visualization of large geometric environments has always been an important problem of computer graphics. We present a framework for the stereoscopic view-dependent visualization of large scale terrain models. We use a quadtree based multiresolution representation for the terrain data. This structure is queried to obtain the view-dependent approximations of the terrain model at different levels of detail. In order not to lose depth information, which is crucial for the stereoscopic visualization, we make use of a different simplification criterion, namely, distance-based angular error threshold. We also present an algorithm for the construction of stereo pairs in order to speed up the view-dependent stereoscopic visualization. The approach we use is the simultaneous generation of the triangles for two stereo images using a single draw-list so that the view frustum culling and vertex activation is done only once for each frame. The cracking problem is solved using the dependency information stored for each vertex. We eliminate the popping artifacts that can occur while switching between different resolutions of the data using morphing. We implemented the proposed algorithms on personal computers and graphics workstations. Performance experiments show that the second eye image can be produced approximately 45 percent faster than drawing the two images separately and a smooth stereoscopic visualization can be achieved at interactive frame rates using continuous multiresolution representation of height fields.
Ugur Güdükbay, Türker Yilmaz
IEEE Trans. Vis. Comput. Graph.1
1998 Realistic Speech Animation of Synthetic Faces
abstract
We combined physically based modeling and parameterization to generate realistic speech animation on synthetic faces. We used physically based modeling for muscles. Muscles are modeled as forces deforming the mesh of polygons. A parameterization technique is used for generating mouth shapes for speech animation. Each meaningful part of a text, which is a letter in our case, corresponds to a specific mouth shape and the mouth shape is generated by setting a set of parameters used for representing the muscles and jaw rotation. We also developed a mechanism to generate and synchronize facial expressions while speaking. Some tags specifying the facial expressions are inserted into the input text together with the degree of the expression. In this way, the facial expression with the specified degree is generated and synchronized with speech animation.
B. Uz, Ugur Güdükbay, Bülent Özgüç
CA2
1998 Right-triangular subdivision for texture mapping ray-traced objects
Ugur Akdemir, Bülent Özgüç, Ugur Güdükbay, Alper Selçuk
Vis. Comput.3
1997 A movable jaw model for the human face
Ugur Güdükbay
Comput. Graph.1
1997 A spring force formulation for elastically deformable models
abstract
Continuous deformable models are generally represented using a grid of control points. The elastic properties are then modeled using the interactions between these points. The formulations based on elasticity theory express these interactions using stiffness matrices. These matrices store the elastic properties of the models and they should be evolved in time according to changing elastic properties of the models. However, forming the stiffness matrices at any step of an animation is very difficult and sometimes the differential equations that should be solved to produce animation become ill-conditioned. Instead of modeling the elasticities using stiffness matrices, the interactions between model points could be expressed in terms of external spring forces. In this paper, a spring force formulation for animating elastically deformable models is presented. In this formulation, elastic properties of the materials are represented as external spring forces as opposed to forming complicated stiffness matrices. (C) 1997 Elsevier Science Ltd.
Ugur Güdükbay, Bülent Özgüç, Yilmaz Tokad
Comput. Graph.1
1995 Animating deformable models: different approaches
abstract
Physically-based modeling remedies the problem of producing realistic animation by including forces, masses, strain energies, and other physical quantities. The behavior of physically-based models is governed by the laws of rigid and nonrigid dynamics expressed through a set of equations of motion. This paper discusses various formulations for animating deformable models. The formulations based on elasticity theory express the interactions between discrete deformable model points using the stiffness matrices. These matrices store the elastic properties of the models and they should be evolved in time according to changing elastic properties of the models. An alternative to these formulations seems to be external force formulations of different types. In these types of formulations, elastic properties of the materials are represented as external spring or other tensile forces as opposed to forming complicated stiffness matrices.>
Ugur Güdükbay, Bülent Özgüç
CA1
1994 Animation of deformable models
Ugur Güdükbay, Bülent Özgüç
Comput. Aided Des.1
1993 An animation system for rigid and deformable models
Ugur Güdükbay, Bülent Özgüç, Yilmaz Tokad
Comput. Graph.1
1990 Free-form solid modeling using deformations
Ugur Güdükbay, Bülent Özgüç
Comput. Graph.1