VLDB 2026 Research / reviewers in the wild / expert
Alper Yilmaz 0001
dblp:11/1315-1
· DBLP profile ↗
47ranked-venue papers
11as first author
11since 2021 · last 2026
0000-0003-0755-2628ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 11 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CascadeFormer: A Family of Two-Stage Cascading Transformers for Skeleton-Based Human Action Recognition
Yusen Peng, Alper Yilmaz 0001 |
ICPR (2) | 2 |
| 2025 | HyperGLM: HyperGraph for Video Scene Graph Generation and AnticipationabstractMultimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes. To bridge this gap, Video Scene Graph Generation (VidSGG) has emerged to capture multi-object relationships across video frames. However, prior methods rely on pairwise connections, limiting their ability to handle complex multi-object interactions and reasoning. To this end, we propose the Multimodal Large Language Models (LLMs) on a Scene HyperGraph (HyperGLM), promoting reasoning about multi-way interactions and higher-order relationships. Our approach uniquely integrates entity scene graphs, which capture spatial relationships between objects, with a procedural graph that models their causal transitions, forming a unified HyperGraph. Significantly, HyperGLM enables reasoning by injecting this unified HyperGraph into LLMs. Additionally, we introduce a new Video Scene Graph Reasoning (VSGR) dataset featuring 1.9M frames from third-person, egocentric, and drone views and support five tasks. Empirically, HyperGLM consistently outperforms state-of-the-art methods, effectively modeling and reasoning complex relationships in diverse scenes. Pha Nguyen, Jackson David Cothren, Alper Yilmaz 0001, Khoa Luu |
CVPR | 4 |
| 2025 | Rapid Object Modeling Initialization for Vector Quantized-Variational AutoEncoderabstractThis paper presents a novel approach for codebook initialization in the well-established Vector Quantized-Variational AutoEncoder (VQ-VAE). Rapid Object Modeling (ROM), inspired by few-shot metric learning, leverages the knowledge base of pretrained deep neural networks to rapidly generate object appearance models that are aggregated into reference concept dictionaries. The results of this paper explore the effects of codebook size and embedding vector length on reconstruction error, followed by a qualitative comparison between the reconstructed images. Initializing the VQ-VAE with ROM generated concept dictionaries showed several desirable improvements, including 47% lower reconstruction errors, faster convergence, and more predictable performance. Michael Karnes, Alper Yilmaz 0001 |
ICIP | 2 |
| 2025 | Disrupting explicit encoding paradigms: property-interactive transformers decode T-cell receptor specificity beyond dataset biasesabstractThe human immune response relies on the unique ability of T-cell receptors (TCRs) to specifically bind to peptides, a process essential for immune surveillance and response. Although deep learning methods for prediction of TCR-peptide binding have proliferated, many encoder-based approaches learn dataset biases, greatly overestimating the model results, and ignoring the biochemical mechanisms and spatial properties affecting binding. Through our analysis, we found that interaction pairs generated by cross-mapping the amino acid properties between TCR and peptide implicitly simulate spatial structure, enabling machine learning models to capture information more effectively. Based on this insight, we developed T-cell receptor cross (TCRoss), a transformer-based model for large-scale learning. In addition, we observed that incorporating environmental information into the dataset not only mitigates learning biases but also improves performance. Experiments show that TCRoss consistently outperforms existing models in both observed contexts and de novo peptide scenarios. Wet-lab validation using T-cell activation assays confirmed the model's predictions for nonbinding peptides and provided critical experimental evidence for model assessment. Biophysical validation confirms that high-attention residue pairs correspond to crystallographically observed binding interfaces. Luming Yang, Haoxian Liu, Alec Calanche, Sohret M. Gokcek, Nicholas Sansoterra, Munir Akkaya, Billur Akkaya, Alper Yilmaz 0001 |
Briefings Bioinform. | 9 |
| 2024 | DINTR: Tracking via Diffusion-based InterpolationabstractObject tracking is a fundamental task in computer vision, requiring the localization of objects of interest across video frames. Diffusion models have shown remarkable capabilities in visual generation, making them well-suited for addressing several requirements of the tracking problem. This work proposes a novel diffusion-based methodology to formulate the tracking task. Firstly, their conditional process allows for injecting indications of the target object into the generation process. Secondly, diffusion mechanics can be developed to inherently model temporal correspondences, enabling the reconstruction of actual frames in video. However, existing diffusion models rely on extensive and unnecessary mapping to a Gaussian noise domain, which can be replaced by a more efficient and stable interpolation process. Our proposed interpolation mechanism draws inspiration from classic image-processing techniques, offering a more interpretable, stable, and faster approach tailored specifically for the object tracking task. By leveraging the strengths of diffusion models while circumventing their limitations, our Diffusion-based INterpolation TrackeR (DINTR) presents a promising new paradigm and achieves a superior multiplicity on seven benchmarks across five indicator representations. Pha A. Nguyen, T. Hoang Ngan Le, Jackson David Cothren, Alper Yilmaz 0001, Khoa Luu |
NeurIPS | 4 |
| 2024 | CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosabstractVideo scene graph generation (VidSGG) has emerged as a transformative approach to capturing and interpreting the intricate relationships among objects and their temporal dynamics in video sequences. In this paper, we introduce the new AeroEye dataset that focuses on multi-object relationship modeling in aerial videos. Our AeroEye dataset features various drone scenes and includes a visually comprehensive and precise collection of predicates that capture the intricate relationships and spatial arrangements among objects. To this end, we propose the novel Cyclic Graph Transformer (CYCLO) approach that allows the model to capture both direct and long-range temporal dependencies by continuously updating the history of interactions in a circular manner. The proposed approach also allows one to handle sequences with inherent cyclical patterns and process object relationships in the correct sequential order. Therefore, it can effectively capture periodic and overlapping relationships while minimizing information loss. The extensive experiments on the AeroEye dataset demonstrate the effectiveness of the proposed CYCLO model, demonstrating its potential to perform scene understanding on drone videos. Finally, the CYCLO method consistently achieves State-of-the-Art (SOTA) results on two in-the-wild scene graph generation benchmarks, i.e., PVSG and ASPIRe. Pha A. Nguyen, Xin Li 0005, Jackson David Cothren, Alper Yilmaz 0001, Khoa Luu |
NeurIPS | 5 |
| 2024 | Improving model performance of shortest-path-based centrality measures in network models through scale spaceabstractSummary The quality of the solution in resolving a complex network depends on either the speed or accuracy of the results. While some health studies prioritize high performance, fast algorithms are favored in scenarios requiring rapid decision‐making. A comprehensive understanding of the problem necessitates a detailed analysis of the network and its individual components. Betweenness Centrality (BC) and Closeness Centrality (CC) are commonly employed measures in network studies. This study introduces a new strategy to compute BC and CC that assesses their sensitivity in the scale space while measuring the shortest path. The scale space is generated by incorporating a scale parameter that is shown to achieve up to 60% performance improvements for various datasets. The study provides in‐depth insights into the importance of the scale space analysis. Finally, a flexible measurement tool is provided that is suitable for various types of problems. To demonstrate the flexibility and applicability, we experimented with two methods for 10 different graphs using the proposed approach. Kenan Mengüç, Alper Yilmaz 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Learning to Drive Using Sparse Imitation Reinforcement LearningabstractIn this paper, we propose Sparse Imitation Reinforcement Learning (SIRL), a hybrid end-to-end control policy that combines the sparse expert driving knowledge with reinforcement learning (RL) policy for autonomous driving (AD) task in CARLA simulation environment. The sparse expert is designed based on hand-crafted rules which is suboptimal but provides a risk-averse strategy by enforcing experience for critical scenarios such as pedestrian and vehicle avoidance, and traffic light detection. As it has been demonstrated, training a RL agent from scratch is data-inefficient and time consuming particularly for the urban driving task, due to the complexity of situations stemming from the vast size of state space. Our SIRL strategy provides a solution to solve these problems by fusing the output distribution of the sparse expert policy and the RL policy to generate a composite driving policy. With the guidance of the sparse expert during the early training stage, SIRL strategy accelerates the training process and keeps the RL exploration from causing a catastrophe outcome, and ensures safe exploration. To some extent, the SIRL agent is imitating the driving expert’s behavior. At the same time, it continuously gains knowledge during training therefore it keeps making improvement beyond the sparse expert, and can surpass both the sparse expert and a traditional RL agent. We experimentally validate the efficacy of proposed SIRL approach in a complex urban scenario within the CARLA simulator. Besides, we compare the SIRL agent’s performance for risk-averse exploration and high learning efficiency with the traditional RL approach. We additionally demonstrate the SIRL agent’s generalization ability to transfer the driving skill to unseen environment. The supplementary material is available at https://superhan2611.github.io/. Yuci Han, Alper Yilmaz 0001 |
ICPR | 2 |
| 2022 | HVIOnet: A deep learning based hybrid visual-inertial odometry approach for unmanned aerial system position estimation
Muhammet Fatih Aslan, Akif Durdu, Abdullah Yusefi, Alper Yilmaz 0001 |
Neural Networks | 4 |
| 2021 | Pocformer: A Lightweight Transformer Architecture For Detection Of Covid-19 Using Point Of Care UltrasoundabstractThe rapid and seemingly endless expansion of COVID-19 can be traced back to the inefficiency and shortage of testing kits that offer accurate results in a timely manner. An emerging popular technique, which adopts improvements made in mobile ultrasound technology, allows for healthcare professionals to conduct rapid screenings on a large scale. We present an image-based solution that aims at automating the testing process which allows for rapid mass testing to be conducted with or without a trained medical professional that can be applied to rural environment and third world countries. Our contributions towards rapid large-scale testing includes a novel deep learning architecture capable of analyzing ultrasound data that can run in real time and significantly improve the current state-of-the-art detection accuracies using image based COVID-19 detection. Shehan Perera, Srikar Adhikari, Alper Yilmaz 0001 |
ICIP | 3 |
| 2021 | A spacetime model for one-shot active contour extraction scheme for human detection in image sequences
Nima A. Gard, Colin Bunker, Alper Yilmaz 0001 |
Comput. Vis. Image Underst. | 3 |
| 2020 | End-to-end Deep Learning Methods for Automated Damage Detection in Extreme Events at Various ScalesabstractRobust Mask R-CNN (Mask Regional Convolutional Neural Network) methods are proposed and tested for automatic detection of cracks on structures or their components that may be damaged during extreme events, such as earthquakes. We curated a new dataset with 2,021 labeled images for training and validation and aimed to find end-to-end deep neural networks for crack detection in the field. With data augmentation and parameters fine-tuning, Path Aggregation Network (PANet) with spatial attention mechanisms and High- resolution Network (HRNet) are introduced into Mask R-CNNs. The tests on three public datasets with low- or high-resolution images demonstrate that the proposed methods can achieve a big improvement over alternative networks, so the proposed method may be sufficient for crack detection for a variety of scales in real applications. Yongsheng Bai, Halil Sezen, Alper Yilmaz 0001 |
ICPR | 3 |
| 2020 | Map-Based Temporally Consistent Geolocalization through Learning Motion TrajectoriesabstractIn this paper, we propose a novel trajectory learning method that exploits motion trajectories on topological map using recurrent neural network for temporally consistent ge-olocalization of object. Inspired by human's ability to both be aware of distance and direction of self-motion in navigation, our trajectory learning method learns a pattern representation of trajectories encoded as a sequence of distances and turning angles to assist self-localization. We pose the learning process as a conditional sequence prediction problem in which each output locates the object on a traversable edge in a map. Considering the prediction sequence ought to be topologically connected in the graph-structured map, we adopt two different hypotheses generation and elimination strategies to eliminate disconnected sequence prediction. We demonstrate our approach on the KITTI stereo visual odometry dataset which is a city-scale environment. The key benefits of our approach to geolocalization are that 1) we take advantage of powerful sequence modeling ability of recurrent neural network and its robustness to noisy input, 2) only require a map in the form of a graph and 3) simply use an affordable sensor that generates motion trajectory. The experiments show that the motion trajectories can be learned by training an recurrent neural network, and temporally consistent geolocation can be predicted with both of the proposed strategies. Bing Zha, Alper Yilmaz 0001 |
ICPR | 2 |
| 2019 | Multiple Hypothesis Testing Approach to Pedestrian INS with Map-MatchingabstractInertial sensors became wearable/portable with the advances in sensing and computing technologies in the last two decades. Captured motion data can be used to build a pedestrian inertial navigation system (INS); however, time-variant bias and noise characteristics of low-cost sensors cause severe errors in positioning. To overcome the growing errors of so-called dead-reckoning (DR) solution, this research adopts a robust pedestrian INS based on Kalman Filter (KF) with zero-velocity update (ZUPT) aid. Despite accurate traveled distance estimates, obtained trajectories diverge from actual paths because of the heading estimation errors. In the absence of external corrections (e.g., GPS, UWB), map information is commonly employed to eliminate position drift; therefore, INS solution is fed into a higher level map-matching filter for further corrections. Unlike common Particle Filter (PF) map-matching, we implicitly model map by rasterizing the floor plans that reduces the computational burden of PF. Second major usage of the introduced approach, which makes the Bayesian estimation cycle non-recursive by serving as a constant spatial prior in the designed filter, is to generate probabilities for a self-initialization method referred to as the Multiple Hypothesis Testing (MHT). Extracted scores update hypothesis probabilities in a cumulative manner and hypothesis with maximum probability provides the correct initial location and heading. Qualitative and quantitative experimental results show feasible trajectories, and negligible return-to-start and stride errors that validates the success of the proposed approach. Muhammed Taha Köroglu, Mehmet Korkmaz, Alper Yilmaz 0001, Akif Durdu |
IPIN | 3 |
| 2018 | DCF-BoW: Build Match Graph Using Bag of Deep Convolutional Features for Structure From MotionabstractMatching a large number of images is quite time-consuming for structure from motion (SfM) due to the image matching by comparing features between all image pairs. In this letter, a bag of deep convolutional features (DCF-BoW) model is proposed to create match graph to reduce the number of matches. First, the convolutional feature map of an image is extracted using the VGG-16 convolutional neural network trained on ImageNet. Then, each local region in the original image can be represented by a feature vector in the feature map. The feature vectors are normalized and used to construct a bag of words model, which could convert each image into a DCF-BoW representation. Finally, the match graph is constructed by selecting the top 10 images with the highest similarities, which are calculated by computing the distances between those DCF-BoW representations. The experiment results show that the proposed DCF-BoW can create the match graph effectively in short time and find the potential overlapping image pairs. The match graph created by the proposed DCF-BoW is better than those built by DBoW3 and VocabTree, which is clearly showed in precision-recall curve on the Urban data set. The results of the SfM reconstruction based on the match graph created by the proposed DCF-BoW are slightly worse than those of the exhaustive matching, while the number of matches is reduced by 97.4% and 92.1%, respectively, on the Urban and South Building data sets. Jie Wan 0002, Alper Yilmaz 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | 4D ISIP: 4D Implicit Surface Interest Point Detection
Shirui Li, Alper Yilmaz 0001, Changlin Xiao, Hua Li 0009 |
ICIG (1) | 2 |
| 2017 | A unique target representation and voting mechanism for visual trackingabstractOne of the main problems in visual tracking is how to robustly represent the target. In this paper, we propose a simple yet efficient method that provides a unique “target representation” by generating discriminative non-uniform subspaces from the feature space which we refer to as cells. Each cell is attributed with a measure that highlights how likely it describes the target or the background. In addition, we keep a codebook of spatial locations of the features which are mapped to the cell similar to that of the R-Table in generalized Hough transform. Using the uniqueness measure as weight, the target center is estimated by using a modified Hough voting scheme to address non-rigid deformations. In the experiment, we use color as the pixel's descriptor and demonstrate comparable performance on the Online Tracking Benchmark (OTB) dataset respect to the other state-of-the-art. Changlin Xiao, Alper Yilmaz 0001 |
ICIP | 2 |
| 2016 | Efficient tracking with distinctive target colors and silhouetteabstractTarget tracking using color based appearance models is very popular in visual tracking. However, trackers based only on color are fragile and often drift to the background when it has similar appearances. In this paper, we propose an efficient way to use distinctive target colors to track the target and eliminate the drift problem. Colors are sampled from the target and its immediate surrounding region. And color samples coming from target result in more distinctive target color. In our approach, we use a short and a long time color histogram to represent the target color. The short time color histogram is used to calculate the distinctiveness of colors while the long time color histogram is used to keep the target color that is consistent over time. In our approach, the target is not marked as a rectangle or other geometric primitives, instead, we track it with its own silhouette. Using silhouette to mark target significantly reduces the false positive information during online learning. Also, the color models are updated with a dynamic learning factor which is based on the tracking result. After testing with many tracking sequences and comparison with other state-of-art trackers, the proposed tracking algorithm shows comparably better performance with very high tracking rate. Changlin Xiao, Alper Yilmaz 0001 |
ICPR | 2 |
| 2016 | An Integrated Active Contour Approach to Shoreline Mapping Using HSI and DEMabstractThis paper introduces a new approach to shoreline mapping, which provides critical input to nautical charting, coastal zone management, and legal boundary determination. We extract shorelines by fusing AVIRIS hyperspectral imagery (HSI) with a LiDAR-generated DEM using a multiphase active contour segmentation technique. Our approach employs a study of object spectra and a knowledge-based segmentation scheme for generating an initial solution followed by a contour evolution technique to achieve a subpixel level of accuracy, while maintaining low computational complexity. Introducing the DEM into shoreline delineation from HSI proves to be a useful tool in eliminating misclassifications and in increasing localization accuracy. Experimental results show that our integrated approach to mapping shorelines has promising outcomes that provide a way to exploit the rich information found in HSI. Anuchit Sukcharoenpong, Alper Yilmaz 0001, Ron Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | A Multi-transformational Model for Background Subtraction with Moving Cameras
Daniya Zamalieva, Alper Yilmaz 0001, James W. Davis |
ECCV (1) | 2 |
| 2014 | Exploiting Temporal Geometry for Moving Camera Background SubtractionabstractIn this paper, we introduce a new method for background subtraction for freely moving cameras and arbitrary scene geometry. Instead of relying on frame-to-frame estimation, we simultaneously estimate the epipolar geometries induced by a moving camera in a temporally consistent manner for a number of frames using the temporal fundamental matrix (TFM). The TFM is robustly estimated from the track lets generated by dense point tracking and used to compute the probability of each track let belonging to the background. In order to ensure the color, spatial, and temporal consistencies of track let labeling, we minimize spatiotemporal labeling cost in locality of track lets. Extensive experiments with challenging videos show that the proposed method is comparable and, in most cases, outperforms the state-of-the-art. Daniya Zamalieva, Alper Yilmaz 0001, James W. Davis |
ICPR | 2 |
| 2014 | Persistent tracking of static scene features using geometry
Jinwei Jiang, Alper Yilmaz 0001 |
Comput. Vis. Image Underst. | 2 |
| 2014 | Background subtraction for the moving camera: A geometric approach
Daniya Zamalieva, Alper Yilmaz 0001 |
Comput. Vis. Image Underst. | 2 |
| 2014 | Line Matching in Wide-Baseline Stereo: A Top-Down ApproachabstractThis paper introduces a new algorithm for matching lines across images that exploit the epipolar geometry and the coplanarity constraints between pairs of lines. In contrast to common treatment in matching of interest points, we use the epipolar geometry to constrain coplanarity conditions between line-pairs. This treatment eliminates the potential matching problems due to the incomplete line observations with nonmatching endpoints. This observation is used to detect a set of candidate line-pair correspondences. These matching pairs are then verified via local homography transforms derived from the neighboring interest point correspondences. This step results in a line affinity matrix which is processed to obtain matching lines. During this process, we do not use appearance models and show that the proposed treatment performs better than the state-ofthe- art appearance and geometry based methods especially for images with wide-baseline. Mohammed Al-Shahri, Alper Yilmaz 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Estimating driving behavior by a smartphoneabstractIn this paper, we propose an approach to understand the driver behavior using smartphone sensors. The aim for analyzing the sensory data acquired using a smartphone is to design a car-independent system which does not need vehicle mounted sensors measuring turn rates, gas consumption or tire pressure. The sensory data utilized in this paper includes the accelerometer, gyroscope and the magnetometer. Using these sensors we obtain position, speed, acceleration, deceleration and deflection angle sensory information and estimate commuting safety by statistically analyzing driver behavior. In contrast to state of the art, this work uses no external sensors, resulting in a cost efficient, simplistic and user-friendly system. Haluk Eren, Semiha Makinist, Erhan Akin, Alper Yilmaz 0001 |
Intelligent Vehicles Symposium | 4 |
| 2012 | Interactive Image Segmentation Using Dirichlet Process Multiple-View LearningabstractSegmenting semantically meaningful whole objects from images is a challenging problem, and it becomes especially so without higher level common sense reasoning. In this paper, we present an interactive segmentation framework that integrates image appearance and boundary constraints in a principled way to address this problem. In particular, we assume that small sets of pixels, which are referred to as seed pixels, are labeled as the object and background. The seed pixels are used to estimate the labels of the unlabeled pixels using Dirichlet process multiple-view learning, which leverages 1) multiple-view learning that integrates appearance and boundary constraints and 2) Dirichlet process mixture-based nonlinear classification that simultaneously models image features and discriminates between the object and background classes. With the proposed learning and inference algorithms, our segmentation framework is experimentally shown to produce both quantitatively and qualitatively promising results on a standard dataset of images. In particular, our proposed framework is able to segment whole objects from images given insufficient seeds. Lei Ding 0002, Alper Yilmaz 0001 |
IEEE Trans. Image Process. | 2 |
| 2011 | Inferring social relations from visual conceptsabstractIn this paper, we study the problem of social relational inference using visual concepts which serve as indicators of actors' social interactions. While social network analysis from videos has started to gain attention in the recent years, the existing work either uses proximity or co-occurrence statistics, or exploit a holistic model of the scene content where the relations are assumed to stay constant throughout the video. This work permits changing relations and argues that there exists a relationship between the visual concepts and the social relations among actors, which is a fundamentally new concept in computer vision. Specifically, we leverage the existing large-scale concept detectors to generate concept score vectors to represent the video content, and we further map them to grouping cues that are used to detect the social structure. In our framework, a probabilistic graphical model with temporal smoothing provides a means to analyze social relations among actors and detect communities. Experiments on Youtube videos and theatrical movies validate the proposed framework. Lei Ding 0002, Alper Yilmaz 0001 |
ICCV | 2 |
| 2011 | Planar shape representation and matching under projective transformation
Panu Srestasathiern, Alper Yilmaz 0001 |
Comput. Vis. Image Underst. | 2 |
| 2011 | Kernel-based object tracking using asymmetric kernels with adaptive scale and orientation selection
Alper Yilmaz 0001 |
Mach. Vis. Appl. | 1 |
| 2010 | Learning Relations among Movie Characters: A Social Network Perspective
Lei Ding 0002, Alper Yilmaz 0001 |
ECCV (4) | 2 |
| 2010 | Enhancing Interactive Image Segmentation with Automatic Label Set Augmentation
Lei Ding 0002, Alper Yilmaz 0001 |
ECCV (6) | 2 |
| 2010 | Social Network Approach to Analysis of Soccer GameabstractVideo understanding has been an active area of research, where many articles have been published on how to detect and track objects in videos, and how to analyze their trajectories. These methods, however, only provided heuristic low level information without providing a higher level understanding of global relations within the whole context. This paper presents a new way to provide such understanding using social network approach in soccer videos. Our approach considers representing interactions between the objects in the video as a social network. This network is then analyzed by detecting small communities using modularity, which relates social interaction. Additionally, we analyze the centrality of nodes which provides importance of individuals composing the network. In particular, we introduce five centralities exploiting directed and weighted social network. The partitions of the resulting social network are shown to relate to clusters of soccer players with respect to their role in the game. Kyoung-Jin Park, Alper Yilmaz 0001 |
ICPR | 2 |
| 2010 | Interactive image segmentation using probabilistic hypergraphs
Lei Ding 0002, Alper Yilmaz 0001 |
Pattern Recognit. | 2 |
| 2008 | Efficient object shape recovery via slicing planesabstractRecovering the three-dimensional (3D) object shape remains an unresolved area of research on the cross-section of computer vision, photogrammetry and bioinformatics. Although various techniques have been developed, the computational complexity and the constraints introduced to overcome the problems have limited their applicability in the real world scenarios. In this paper, we propose a method that is based on the projective geometry between the object space and the silhouette-images taken from multiple view-points. The approach eliminates the problems related to dense feature point matching and camera calibration that are generally adopted by many state of the art shape reconstruction methods. The object shape is reconstructed by establishing a set of hypothetical planes slicing the object volume and estimating the projective geometric relations between the images of these planes. The experimental results show that the 3D object shape can be recovered by applying minimal constraints. Po-Lun Lai, Alper Yilmaz 0001 |
CVPR | 2 |
| 2008 | Image Segmentation as Learning on HypergraphsabstractIn this paper, we propose to use hypergraphs as the model for images and pose image segmentation as a machine learning problem in which some pixels (called seeds) are labeled as the objects and background. Using the seed pixels, our method predicts the labels for all unlabeled pixels. We present the relations of the proposed method to other hypergraph based learning techniques. We give an adaptive procedure for constructing image hypergraphs and achieve promising results on a real image dataset. Lei Ding 0002, Alper Yilmaz 0001 |
ICMLA | 2 |
| 2008 | View invariant object recognitionabstractThis paper introduces a method for the recognition planar objects under projective geometry. Our method is based on a similarity measure invariant to projective transform. The proposed similarity measure utilizes the distribution of the projective relations between the conic section pairs, which are estimated from the object¿s shape. We conjecture that given two objects of the same type, which are viewed from different viewpoints generate similar histograms, such that their difference is smaller than the histograms generated from other object types. The proposed measure has shown promising performance on the Brown shape database. Panu Srestasathiern, Alper Yilmaz 0001 |
ICPR | 2 |
| 2008 | A differential geometric approach to representing the human actions
Alper Yilmaz 0001, Mubarak Shah |
Comput. Vis. Image Underst. | 1 |
| 2007 | Object Tracking by Asymmetric Kernel Mean Shift with Automatic Scale and Orientation SelectionabstractTracking objects using the mean shift method is performed by iteratively translating a kernel in the image space such that the past and current object observations are similar. Traditional mean shift method requires a symmetric kernel, such as a circle or an ellipse, and assumes constancy of the object scale and orientation during the course of tracking. In a tracking scenario, it is not uncommon to observe objects with complex shapes whose scale and orientation constantly change due to the camera and object motions. In this paper, we present an object tracking method based on the asymmetric kernel mean shift, in which the scale and orientation of the kernel adaptively change depending on the observations at each iteration. Proposed method extends the traditional mean shift tracking, which is performed in the image coordinates, by including the scale and orientation as additional dimensions and simultaneously estimates all the unknowns in a few number of mean shift iterations. The experimental results show that the proposed method is superior to the traditional mean shift tracking in the following aspects: 1) it provides consistent object tracking throughout the video; 2) it is not effected by the scale and orientation changes of the tracked objects; 3) it is less prone to the background clutter. Alper Yilmaz 0001 |
CVPR | 1 |
| 2006 | Matching actions in presence of camera motion
Alper Yilmaz 0001, Mubarak Shah |
Comput. Vis. Image Underst. | 1 |
| 2005 | Actions Sketch: A Novel Action RepresentationabstractIn this paper, we propose to model an action based on both the shape and the motion of the performing object. When the object performs an action in 3D, the points on the outer boundary of the object are projected as 2D (x, y) contour in the image plane. A sequence of such 2D contours with respect to time generates a spatiotemporal volume (STV) in (x, y, t), which can be treated as 3D object in the (x, y, t) space. We analyze STV by using the differential geometric surface properties to identify action descriptors capturing both spatial and temporal properties. A set of action descriptors is called an action sketch. The first step in our approach is to generate STV by solving the point correspondence problem between consecutive frames. The correspondences are determined using a two-step graph theoretical approach. After the STV is generated, actions descriptors are computed by analyzing the differential geometric properties of STV. Finally, using these descriptors, we perform action recognition, which is also formulated as graph theoretical problem. Several experimental results are presented to demonstrate our approach. Alper Yilmaz 0001, Mubarak Shah |
CVPR (1) | 1 |
| 2005 | Recognizing Human Actions in Videos Acquired by Uncalibrated Moving CamerasabstractMost work in action recognition deals with sequences acquired by stationary cameras with fixed viewpoints. Due to the camera motion, the trajectories of the body parts contain not only the motion of the performing actor but also the motion of the camera. In addition to the camera motion, different viewpoints of the same action in different environments result in different trajectories, which can not be matched using standard approaches. In order to handle these problems, we propose to use the multi-view geometry between two actions. However, well known epipolar geometry of the static scenes where the cameras are stationary is not suitable for our task. Thus, we propose to extend the standard epipolar geometry to the geometry of dynamic scenes where the cameras are moving. We demonstrate the versatility of the proposed geometric approach for recognition of actions in a number of challenging sequences. Alper Yilmaz 0001, Mubarak Shah |
ICCV | 1 |
| 2004 | Contour-Based Object Tracking with Occlusion Handling in Video Acquired Using Mobile CamerasabstractWe propose a tracking method which tracks the complete object regions, adapts to changing visual features, and handles occlusions. Tracking is achieved by evolving the contour from frame to frame by minimizing some energy functional evaluated in the contour vicinity defined by a band. Our approach has two major components related to the visual features and the object shape. Visual features (color, texture) are modeled by semiparametric models and are fused using independent opinion polling. Shape priors consist of shape level sets and are used to recover the missing object regions during occlusion. We demonstrate the performance of our method on real sequences with and without object occlusions. Alper Yilmaz 0001, Xin Li 0022, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Target tracking in airborne forward looking infrared imagery
Alper Yilmaz 0001, Khurram Shafique, Mubarak Shah |
Image Vis. Comput. | 1 |
| 2002 | Estimation of Arbitrary Albedo and Shape from Shading for Symmetric ObjectsabstractIn this paper, we propose a shape from shading (SFS) approach to recover both the shape and the reflectance properties of symmetric objects using a single image. The common constraint of constant or piece-wise constant albedo for lambertJan surfaces is relaxed to arbitrary albedo. The proposed method can be categorized as a linear shape from shading method, which linearizes the reflectance function for symmetric objects using the symmetry cues of the shape and the albedo, and iteratively computes the depth values. Estimated depth values are then used to recover pixel-wise surface albedo. To show the usefulness of the proposed method, we present experimental results for both synthetic and real images. Alper Yilmaz 0001, Mubarak Shah |
BMVC | 1 |
| 2002 | View-Invariant Representation and Recognition of Actions
Cen Rao, Alper Yilmaz 0001, Mubarak Shah |
Int. J. Comput. Vis. | 2 |
| 2001 | Eigenhill vs. eigenface and eigenedge
Alper Yilmaz 0001, Muhittin Gökmen |
Pattern Recognit. | 1 |
| 2000 | Eigenhill vs. Eigenface and EigenedgeabstractIn this study, we present a new approach to overcome the problems in face recognition associated with illumination changes by utilizing the edge images rather than intensity values. However, using edges directly has its problems. To combine the advantages of algorithms based on shading and edges while overcoming their drawbacks, we introduced "hills" which are obtained by covering edges with a membrane. Each hill image is then described as a combination of most descriptive eigenvectors, called "eigenhills", spanning hills space. We compare the recognition performances of eigenface, eigenedge and eigenhills methods by considering illumination and orientation changes on Purdue A&R face database and showed experimentally that our approach has the best recognition performance. Alper Yilmaz 0001, Muhittin Gökmen |
ICPR | 1 |