VLDB 2026 Research / reviewers in the wild / expert
Birkan Tunç
dblp:45/8105
· DBLP profile ↗
17ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-2294-4024ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
3D vision · 56% Face, body and person analysis · 44% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Theoretical computer science
2 papers |
Graph algorithms and graph theory · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d face reconstruction |
0.8 | 1 | 2024 | Inequality-Constrained 3D Morphable Face Model Fitting · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › 3D vision › 3d face reconstruction
3d face model fitting |
0.4 | 1 | 2020 | Inequality-Constrained and Robust 3D Face Model Fitting · ECCV (9) 2020 |
Computer vision › 3D vision › camera calibration
camera model |
0.4 | 1 | 2020 | Can Facial Pose and Expression Be Separated With Weak Perspective Camera? · CVPR 2020 |
Computer vision › Face, body and person analysis
facial action unit recognition |
0.4 | 1 | 2020 | Can Facial Pose and Expression Be Separated With Weak Perspective Camera? · CVPR 2020 |
Computer vision › Face, body and person analysis
facial expression analysis |
0.4 | 1 | 2020 | Can Facial Pose and Expression Be Separated With Weak Perspective Camera? · CVPR 2020 |
Computer vision › 3D vision › camera calibration › camera model
weak perspective camera |
0.4 | 1 | 2020 | Can Facial Pose and Expression Be Separated With Weak Perspective Camera? · CVPR 2020 |
Data mining › clustering
approximate clustering |
0.4 | 1 | 2020 | Discovering Synchronized Subsets of Sequences: A Large Scale Solution · CVPR 2020 |
Data mining
clustering |
0.4 | 1 | 2020 | Discovering Synchronized Subsets of Sequences: A Large Scale Solution · CVPR 2020 |
Data mining
pattern mining |
0.4 | 1 | 2020 | Discovering Synchronized Subsets of Sequences: A Large Scale Solution · CVPR 2020 |
Graph algorithms and graph theory › graph theory
clique |
0.4 | 1 | 2020 | Discovering Synchronized Subsets of Sequences: A Large Scale Solution · CVPR 2020 |
Graph algorithms and graph theory › graph theory › clique
maximal clique |
0.4 | 1 | 2020 | Discovering Synchronized Subsets of Sequences: A Large Scale Solution · CVPR 2020 |
Methods — techniques the papers use, named apart from their topics
log-barrier method · 1.5gradient correlation · 1.5jung's theorem · 0.9euclidean ball fitting · 0.9epsilon-expanded clusters · 0.9point distribution model · 0.43d-to-2d mapping · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Micro-DualNet: Dual-Path Spatio-Temporal Network for Micro-Action Recognition
Naga VS Raviteja Chappa, Evangelos Sariyanidi, Lisa Yankowitz, Gokul Nair, Casey Zampella, Robert T. Schultz, Birkan Tunç |
FG | 7 |
| 2025 | 3D Face Reconstruction Error Decomposed: A Modular Benchmark for Fair and Fast Method EvaluationabstractComputing the standard benchmark metric for 3D face reconstruction, namely geometric error, requires a number of steps, such as mesh cropping, rigid alignment, or point correspondence. Current benchmark tools are monolithic (they implement a specific combination of these steps), even though there is no consensus on the best way to measure error. We present a toolkit for a Modularized 3D Face reconstruction Benchmark (M3DFB), where the fundamental components of error computation are segregated and interchangeable, allowing one to quantify the effect of each. Furthermore, we propose a new component, namely correction, and present a computationally efficient approach that penalizes for mesh topology inconsistency. Using this toolkit, we test 16 error estimators with 10 reconstruction methods on two real and two synthetic datasets. Critically, the widely used ICP-based estimator provides the worst benchmarking performance, as it significantly alters the true ranking of the top- 5 reconstruction methods. Notably, the correlation of ICP with the true error can be as low as 0.41. Moreover, non-rigid alignment leads to significant improvement (correlation larger than 0.90), highlighting the importance of annotating 3D landmarks on datasets. Finally, the proposed correction scheme, together with non-rigid warping, leads to an accuracy on a par with the best non-rigid ICP-based estimators, but runs an order of magnitude faster. Our open-source codebase is designed for researchers to easily compare alternatives for each component, thus helping accelerating progress in benchmarking for 3D face reconstruction and, furthermore, supporting the improvement of learned reconstruction methods, which depend on accurate error estimation for effective training. Evangelos Sariyanidi, Claudio Ferrari, Federico Nocentini, Stefano Berretti, Andrea Cavallaro, Birkan Tunç |
FG | 6 |
| 2025 | Beyond FACS: Data-driven Facial Expression Dictionaries, with Application to Predicting AutismabstractThe Facial Action Coding System (FACS) has been used by numerous studies to investigate the links between facial behavior and mental health. The laborious and costly process of FACS coding has motivated the development of machine learning frameworks for Action Unit (AU) detection. Despite intense efforts spanning three decades, the detection accuracy for many AUs is considered to be below the threshold needed for behavioral research. Also, many AUs are excluded altogether, making it impossible to fulfill the ultimate goal of FACSthe representation of any facial expression in its entirety. This paper considers an alternative approach. Instead of creating automated tools that mimic FACS experts, we propose to use a new coding system that mimics the key properties of FACS. Specifically, we construct a data-driven coding system called the Facial Basis, which contains units that correspond to localized and interpretable 3D facial movements, and overcomes three structural limitations of automated FACS coding. First, the proposed method is completely unsupervised, bypassing costly, laborious and variable manual annotation. Second, Facial Basis reconstructs all observable movement, rather than relying on a limited repertoire of recognizable movements (as in automated FACS). Finally, the Facial Basis units are additive, whereas AUs may fail detection when they appear in a non-additive combination. The proposed method outperforms the most frequently used AU detector in predicting autism diagnosis from in-person and remote conversations, highlighting the importance of encoding facial behavior comprehensively. To our knowledge, Facial Basis is the first alternative to FACS for deconstructing facial expressions in videos into localized movements. We provide an open source implementation of the method at github.com/sariyanidi/FacialBasis. Evangelos Sariyanidi, Lisa Yankowitz, Robert T. Schultz, John D. Herrington, Birkan Tunç, Jeffrey Cohn |
FG | 5 |
| 2024 | Detecting Autism from Head Movements using KinesicsabstractHead movements play a crucial role in social interactions. The quantification of communicative movements such as nodding, shaking, orienting, and backchanneling is significant in behavioral and mental health research. However, automated localization of such head movements within videos remains challenging in computer vision due to their arbitrary start and end times, durations, and frequencies. In this work, we introduce a novel and efficient coding system for head movements, grounded in Birdwhistell’s kinesics theory, to automatically identify basic head motion units such as nodding and shaking. Our approach first defines the smallest unit of head movement, termed kine, based on the anatomical constraints of the neck and head. We then quantify the location, magnitude, and duration of kines within each angular component of head movement. Through defining possible combinations of identified kines, we define a higher-level construct, kineme, which corresponds to basic head motion units such as nodding and shaking. We validate the proposed framework by predicting autism spectrum disorder (ASD) diagnosis from video recordings of interacting partners. We show that the multi-scale property of the proposed framework provides a significant advantage, as collapsing behavior across temporal scales reduces performance consistently. Finally, we incorporate another fundamental behavioral modality, namely speech, and show that distinguishing between speaking- and listening-time head movements significantly improves ASD classification performance. Muhittin Gökmen, Evangelos Sariyanidi, Lisa Yankowitz, Casey Zampella, Robert T. Schultz, Birkan Tunç |
ICMI | 6 |
| 2024 | Inequality-Constrained 3D Morphable Face Model Fittingabstract3D morphable model (3DMM) fitting on 2D data is traditionally done via unconstrained optimization with regularization terms to ensure that the result is a plausible face shape and is consistent with a set of 2D landmarks. This paper presents inequality-constrained 3DMM fitting as the first alternative to regularization in optimization-based 3DMM fitting. Inequality constraints on the 3DMM's shape coefficients ensure face-like shapes without modifying the objective function for smoothness, thus allowing for more flexibility to capture person-specific shape details. Moreover, inequality constraints on landmarks increase robustness in a way that does not require per-image tuning. We show that the proposed method stands out with its ability to estimate person-specific face shapes by jointly fitting a 3DMM to multiple frames of a person. Further, when used with a robust objective function, namely gradient correlation, the method can work "in-the-wild" even with a 3DMM constructed from controlled data. Lastly, we show how to use the log-barrier method to efficiently implement the method. To our knowledge, we present the first 3DMM fitting framework that requires no learning yet is accurate, robust, and efficient. The absence of learning enables a generic solution that allows flexibility in the input image size, interchangeable morphable models, and incorporation of camera matrix. Evangelos Sariyanidi, Casey Zampella, Robert T. Schultz, Birkan Tunç |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Meta-evaluation for 3D Face Reconstruction Via Synthetic DataabstractThe standard benchmark metric for 3D face reconstruction is the geometric error between reconstructed meshes and the ground truth. Nearly all recent reconstruction methods are validated on real ground truth scans, in which case one needs to establish point correspondence prior to error computation, which is typically done with the Chamfer (i.e., nearest neighbor) criterion. However, a simple yet fundamental question have not been asked: Is the Chamfer error an appropriate and fair benchmark metric for 3D face reconstruction? More generally, how can we determine which error estimator is a better benchmark metric? We present a meta-evaluation framework that uses synthetic data to evaluate the quality of a geometric error estimator as a benchmark metric for face reconstruction. Further, we use this framework to experimentally compare four geometric error estimators. Results show that the standard approach not only severely underestimates the error, but also does so inconsistently across reconstruction methods, to the point of even altering the ranking of the compared methods. Moreover, although non-rigid ICP leads to a metric with smaller estimation bias, it could still not correctly rank all compared reconstruction methods, and is significantly more time consuming than Chamfer. In sum, we show several issues present in the current benchmarking and propose a procedure using synthetic data to address these issues. Evangelos Sariyanidi, Claudio Ferrari, Stefano Berretti, Robert T. Schultz, Birkan Tunç |
IJCB | 5 |
| 2020 | Discovering Synchronized Subsets of Sequences: A Large Scale SolutionabstractFinding the largest subset of sequences (i.e., time series) that are correlated above a certain threshold, within large datasets, is of significant interest for computer vision and pattern recognition problems across domains, including behavior analysis, computational biology, neuroscience, and finance. Maximal clique algorithms can be used to solve this problem, but they are not scalable. We present an approximate, but highly efficient and scalable, method that represents the search space as a union of sets called ϵ-expanded clusters, one of which is theoretically guaranteed to contain the largest subset of synchronized sequences. The method finds synchronized sets by fitting a Euclidean ball on ϵ-expanded clusters, using Jung's theorem. We validate the method on data from the three distinct domains of facial behavior analysis, finance, and neuroscience, where we respectively discover the synchrony among pixels of face videos, stock market item prices, and dynamic brain connectivity data. Experiments show that our method produces results comparable to, but up to 300 times faster than, maximal clique algorithms, with speed gains increasing exponentially with the number of input sequences. Evangelos Sariyanidi, Casey Zampella, G. Keith Bartley, John D. Herrington, Theodore D. Satterthwaite, Robert T. Schultz, Birkan Tunç |
CVPR | 7 |
| 2020 | Can Facial Pose and Expression Be Separated With Weak Perspective Camera?abstractSeparating facial pose and expression within images requires a camera model for 3D-to-2D mapping. The weak perspective (WP) camera has been the most popular choice; it is the default, if not the only option, in state-of-the-art facial analysis methods and software. WP camera is justified by the supposition that its errors are negligible when the subjects are relatively far from the camera, yet this claim has never been tested despite nearly 20 years of research. This paper critically examines the suitability of WP camera for separating facial pose and expression. First, we theoretically show that WP causes pose-expression ambiguity, as it leads to estimation of spurious expressions. Next, we experimentally quantify the magnitude of spurious expressions. Finally, we test whether spurious expressions have detrimental effects on a common facial analysis application, namely Action Unit (AU) detection. Contrary to conventional wisdom, we find that severe pose-expression ambiguity exists even when subjects are not close to the camera, leading to large false positive rates in AU detection. We also demonstrate that the magnitude and characteristics of spurious expressions depend on the point distribution model used to model the expressions. Our results suggest that common assumptions about WP need to be revisited in facial expression modeling, and that facial analysis software should encourage and facilitate the use of the true camera model whenever possible. Evangelos Sariyanidi, Casey Zampella, Robert T. Schultz, Birkan Tunç |
CVPR | 4 |
| 2020 | Inequality-Constrained and Robust 3D Face Model Fitting
Evangelos Sariyanidi, Casey Zampella, Robert T. Schultz, Birkan Tunç |
ECCV (9) | 4 |
| 2017 | Subject-Specific Structural Parcellations Based on Randomized AB-divergences
Nicolas Honnorat, Drew Parker, Birkan Tunç, Christos Davatzikos, Ragini Verma |
MICCAI (1) | 3 |
| 2017 | On characterizing population commonalities and subject variations in brain networks
Yasser Ghanbari, Luke Bloy, Birkan Tunç, Varsha Shankar, Timothy P. L. Roberts, J. Christopher Edgar, Robert T. Schultz, Ragini Verma |
Medical Image Anal. | 3 |
| 2016 | Label-Informed Non-negative Matrix Factorization with Manifold Regularization for Discriminative Subnetwork Detection
Takanori Watanabe, Birkan Tunç, Drew Parker, Junghoon Kim 0006, Ragini Verma |
MICCAI (1) | 2 |
| 2015 | Semantics of object representation in machine learning
Birkan Tunç |
Pattern Recognit. Lett. | 1 |
| 2012 | Local Zernike Moments: A new representation for face recognitionabstractIn this paper, we propose a new image representation called Local Zernike Moments (LZM) for face recognition. In recent years, local image representations such as Gabor and Local Binary Patterns (LBP) have attracted great interest due to their success in handling difficulties of face recognition. In this study, we aim to develop an alternative representation to further improve the face recognition performance. We achieve this by utilizing Zernike Moments which have been successfully used as shape descriptors for character recognition. We modify global Zernike moments to obtain a local representation by computing the moments at every pixel of a face image by considering its local neighborhood, thus decomposing the image into a set of images, moment components, to capture the micro structure around each pixel. Our experiments on FERET face database reveal the superior performance of LZM over Gabor and LBP representations. Evangelos Sariyanidi, Volkan Dagli, Salih Cihan Tek, Birkan Tunç, Muhittin Gökmen |
ICIP | 4 |
| 2012 | Class dependent factor analysis and its application to face recognition
Birkan Tunç, Volkan Dagli, Muhittin Gökmen |
Pattern Recognit. | 1 |
| 2011 | Robust face recognition with class dependent factor analysisabstractA general framework for face recognition under different variations such as illumination and facial expressions is proposed. The model utilizes the class information in a supervised manner to define separate manifolds for each class. Manifold embeddings are achieved by a nonlinear manifold learning technique. Inside each manifold, a mixture of Gaussians is designated to introduce a generative model. By this way, a novel connection between the manifold learning and probabilistic generative models is achieved. The proposed model learns system parameters in a probabilistic framework, allowing a Bayesian decision model. Experimental evaluations with face recognition under illumination changes and facial expressions were performed to realize the ability of the proposed model to handle different types of variations. Our recognition performances were comparable to state-of art results. Birkan Tunç, Volkan Dagli, Muhittin Gökmen |
IJCB | 1 |
| 2009 | Context ranking machine and its application to rigid localization of deformable objectsabstractIn this paper, we exploit the context information embodied in an image to develop a machine learning method called context ranking machine (CRM). Specifically, we leverage two kinds of context information: identity context and metric context. The identity context of an image patch refers to its origin (e.g., from which image it is cropped), and the metric context refers to its distance to the exact surrounding box of the target object inside the image. We use these context information in two ways. First, for object localization, instead of learning classifiers to separate the whole negative pool from all positives, we separate each positive from its own negatives sharing the same identity context. Second, we rank image patches according to their resemblance to the ground truth by establishing a connection between appearance based features and metric properties of the image. The CRM learns an image-based ranking algorithm via boosting and achieves an improved localization accuracy. We performed tests on echocardiogram images to localize heart chambers, and face images for eye band localization. Birkan Tunç, Shaohua Kevin Zhou, Jin Hyeong Park, Muhittin Gökmen |
ICIP | 1 |