VLDB 2026 Research / reviewers in the wild / expert
Itir Önal
dblp:135/8878 · also Itir Onal, Itir Önal Ertugrul
· DBLP profile ↗
23ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0002-0999-8626ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated detection of positive and non-positive shyness in infants from videos
Michail Panagiotis Bofos, Eliala Alice Salvadori, Cristina Colonnesi, Itir Önal |
FG | 4 |
| 2026 | Beyond the Surface: Incorporating 3D Facial Information into Action Unit Detection
Mehmet Kaan Özkan, Itir Önal |
FG | 2 |
| 2026 | RGB-TViT: Multimodal RGB-Thermal Transformers for Early Prediction of Vasovagal Reactions in Blood Donation
Malou Van Der Velde, Itir Önal, Elisabeth Huis in 't Veld, Pieter Spronck, Lee-Ling S. Ong, Judita Rudokaite |
FG | 2 |
| 2025 | Revisiting Representation Learning and Identity Adversarial Training for Facial Behavior UnderstandingabstractFacial Action Unit (AU) detection has gained significant attention as it enables the breakdown of complex facial expressions into individual muscle movements. In this paper, we revisit two fundamental factors in AU detection: diverse and large-scale data and subject identity regularization. Motivated by recent advances in foundation models, we highlight the importance of data and introduce Face 9 M, a diverse dataset comprising 9 million facial images from multiple public sources. Pretraining a masked autoencoder on Face9M yields strong performance in AU detection and facial expression tasks. More importantly, we emphasize that the Identity Adversarial Training (IAT) has not been well explored in AU tasks. To fill this gap, we first show that subject identity in AU datasets creates shortcut learning for the model and leads to suboptimal solutions to AU predictions. Secondly, we demonstrate that strong IAT regularization is necessary to learn identityinvariant features. Finally, we elucidate the design space of IAT and empirically show that IAT circumvents the identity-based shortcut learning and results in a better solution. Our proposed methods, Facial Masked Autoencoder (FMAE) and IAT, are simple, generic and effective. Remarkably, the proposed FMAEIAT approach achieves new state-of-the-art F1 scores on BP4D ($67.1 \%$), BP4D+ ($66.8 \%$), and DISFA ($70.1 \%$) databases, significantly outperforming previous work. We release the code and model at https://github.com/forever208/FMAE-IAT. Mang Ning, Albert Ali Salah, Itir Önal |
FG | 3 |
| 2025 | DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT SpaceabstractThis paper explores image modeling from the frequency space and introduces DCTdiff, an end-to-end diffusion generative paradigm that efficiently models images in the discrete cosine transform (DCT) space. We investigate the design space of DCTdiff and reveal the key design factors. Experiments on different frameworks (UViT, DiT), generation tasks, and various diffusion samplers demonstrate that DCTdiff outperforms pixel-based diffusion models regarding generative quality and training efficiency. Remarkably, DCTdiff can seamlessly scale up to 512$\times$512 resolution without using the latent diffusion paradigm and beats latent diffusion (using SD-VAE) with only 1/4 training cost. Finally, we illustrate several intriguing properties of DCT image modeling. For example, we provide a theoretical proof of why `image diffusion can be seen as spectral autoregression', bridging the gap between diffusion and autoregressive models. The effectiveness of DCTdiff and the introduced properties suggest a promising direction for image modeling in the frequency space. The code is at https://github.com/forever208/DCTdiff. Mang Ning, Mingxiao Li 0002, Jianlin Su, Haozhe Jia, Lanmiao Liu, Martin Benes 0001, Wenshuo Chen, Albert Ali Salah, Itir Önal |
ICML | 9 |
| 2024 | Expanding PyAFAR: A Novel Privacy-Preserving Infant AU DetectorabstractWe enhance PyAFAR11Code will be available on: https:\\affectanalysisgroup.github.io/PyAFARI, an open source, Python-based library for facial action unit detection by introducing a privacy-protected infant AU detector. To prevent reconstruction of the training images, we train the infant AU detector by extracting histogram of gradients (HoG) features and using an efficient Light Gradient Boosting Machine (LightGBM) classifier. Models are trained with two large, well-annotated databases. The performance of our approach is comparable to previously developed deep models that have not been released due to privacy concerns. Our models are available for use and further fine-tuning, contributing to the advancement of facial action unit detection. Itir Önal, Saurabh Hinduja, Maneesh Bilalpur, Daniel S. Messinger, Jeffrey F. Cohn |
FG | 1 |
| 2024 | Face the Needle: Predicting Risk of Fear and Fainting During Blood Donation Through Video AnalysisabstractThere are physiological, hormonal and psychological markers that occur early in a procedure involving needles. These so-called vasovagal reactions range from feeling nauseous, dizzy, to completely passing out. In an early stage, they are difficult to measure and self-report before it is too late to prevent them. This study aims to explore different features from regular video and thermal facial video recordings of blood donors in the waiting room, prior to a blood donation procedure, in order to assess to what extent it is possible to predict whether a donor will experience a low or high level of vasovagal reaction later on during the blood donation. The results showed that the best performance was achieved using pre-trained ResNet152 models with GRU on a continuous video stream, achieving an F1 of 0.69, a PR-AUC score of 0.81, and an MCC score of 0.56. This model also achieved a precision of 0.52, recall of 0.94, F1 score of 0.67, and MCC score of 0.42 on new, previously unseen mobile video data. Although the model requires further improvement, it outperforms self-reported vasovagal reaction scores and shows the potential to predict who is at risk of experiencing vasovagal reactions using facial video data. Judita Rudokaite, Itir Önal, Lee-Ling S. Ong, Mart P. Janssen, Elisabeth Huis in 't Veld |
FG | 2 |
| 2024 | Elucidating the Exposure Bias in Diffusion ModelsabstractDiffusion models have demonstrated impressive generative capabilities, but their exposure bias problem, described as the input mismatch between training and sampling, lacks in-depth exploration. In this paper, we investigate the exposure bias problem in diffusion models by first analytically modelling the sampling distribution, based on which we then attribute the prediction error at each sampling step as the root cause of the exposure bias issue. Furthermore, we discuss potential solutions to this issue and propose an intuitive metric for it. Along with the elucidation of exposure bias, we propose a simple, yet effective, training-free method called Epsilon Scaling to alleviate the exposure bias. We show that Epsilon Scaling explicitly moves the sampling trajectory closer to the vector field learned in the training phase by scaling down the network output, mitigating the input mismatch between training and sampling. Experiments on various diffusion frameworks (ADM, DDIM, EDM, LDM, DiT, PFGM++) verify the effectiveness of our method. Remarkably, our ADM-ES, as a state-of-the-art stochastic sampler, obtains 2.17 FID on CIFAR-10 under 100-step unconditional generation. The code is at https://github.com/forever208/ADM-ES Mang Ning, Mingxiao Li 0002, Jianlin Su, Albert Ali Salah, Itir Önal |
ICLR | 5 |
| 2024 | Multimodal Prediction of Obsessive-Compulsive Disorder and Comorbid Depression Severity and Energy Delivered by Deep Brain ElectrodesabstractTo develop reliable, valid, and efficient measures of obsessive-compulsive disorder (OCD) severity, comorbid depression severity, and total electrical energy delivered (TEED) by deep brain stimulation (DBS), we trained and compared random forests regression models in a clinical trial of participants receiving DBS for refractory OCD. Six participants were recorded during open-ended interviews at pre- and post-surgery baselines and then at 3-month intervals following DBS activation. Ground-truth severity was assessed by clinical interview and self-report. Visual and auditory modalities included facial action units, head and facial landmarks, speech behavior and content, and voice acoustics. Mixed-effects random forest regression with Shapley feature reduction strongly predicted severity of OCD, comorbid depression, and total electrical energy delivered by the DBS electrodes (intraclass correlation, ICC, = 0.83, 0.87, and 0.81, respectively. When random effects were omitted from the regression, predictive power decreased to moderate for severity of OCD and comorbid depression and remained comparable for total electrical energy delivered (ICC = 0.60, 0.68, and 0.83, respectively). Multimodal measures of behavior outperformed ones from single modalities. Feature selection achieved large decreases in features and corresponding increases in prediction. The approach could contribute to closed-loop DBS that would automatically titrate DBS based on affect measures. Saurabh Hinduja, Ali Darzi, Itir Önal, Nicole R. Provenza, Ron Gadot, Eric A. Storch, Sameer A. Sheth, Wayne K. Goodman, Jeffrey F. Cohn |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Automated Emotional Valence Estimation in Infants with Stochastic and Strided Temporal SamplingabstractWe propose the first automated approach to estimate the emotional valence of infants from their facial behavior. We use the state-of-the-art transformer-based video masked autoencoder (VideoMAE) that is pre-trained on a large video dataset as a backbone, and finetune it on two large, well-annotated infant video datasets (SIBSMILE and MODELING). To augment the limited data, we propose a novel video temporal augmentation method called Stochastic and Strided Temporal Sampling (SSTS). We demonstrate the effectiveness of our approach for infant valence estimation by achieving 0.671 Concordance Correlation Coefficient (CCC) on SIBSMILE and MODELING. The experiments show that SSTS remarkably accelerates the training speed by 8 times while gaining the best valence estimation performance. Lastly, we suggest that face detection and cropping (coarse registration) is a promising alternative to landmark-based registration (i.e. fine registration) in data pre-processing when accurate infant facial landmark detectors are inaccessible. Mang Ning, Itir Önal, Daniel S. Messinger, Jeffrey F. Cohn, Albert Ali Salah |
ACII | 2 |
| 2021 | Synthetic Expressions are Better Than Real for Learning to Detect Facial ActionsabstractCritical obstacles in training classifiers to detect facial actions are the limited sizes of annotated video databases and the relatively low frequencies of occurrence of many actions. To address these problems, we propose an approach that makes use of facial expression generation. Our approach reconstructs the 3D shape of the face from each video frame, aligns the 3D mesh to a canonical view, and then trains a GAN-based network to synthesize novel images with facial action units of interest. To evaluate this approach, a deep neural network was trained on two separate datasets: One network was trained on video of synthesized facial expressions generated from FERA17; the other network was trained on unaltered video from the same database. Both networks used the same train and validation partitions and were tested on the test partition of actual video from FERA17. The network trained on synthesized facial expressions outperformed the one trained on actual facial expressions and surpassed current state-of-the-art approaches. Koichiro Niinuma, Itir Önal, Jeffrey F. Cohn, László A. Jeni |
WACV | 2 |
| 2020 | Multimodal Interaction in PsychopathologyabstractThis paper presents an introduction to the Multimodal Interaction in Psychopathology workshop, which is held virtually in conjunction with the 22nd ACM International Conference on Multimodal Interaction on October 25th, 2020. This workshop has attracted submissions in the context of investigating multimodal interaction to reveal mechanisms and assess, monitor, and treat psychopathology. Keynote speakers from diverse disciplines present an overview of the field from different vantages and comment on future directions. Here we summarize the goals and the content of the workshop. Itir Önal, Jeffrey F. Cohn, Hamdi Dibeklioglu |
ICMI | 1 |
| 2019 | FACS3D-Net: 3D Convolution based Spatiotemporal Representation for Action Unit DetectionabstractMost approaches to automatic facial action unit (AU) detection consider only spatial information and ignore AU dynamics. For humans, dynamics improves AU perception. Is same true for algorithms? To make use of AU dynamics, recent work in automated AU detection has proposed a sequential spatiotemporal approach: Model spatial information using a 2D CNN and then model temporal information using LSTM (Long-Short-Term Memory). Inspired by the experience of human FACS coders, we hypothesized that combining spatial and temporal information simultaneously would yield more powerful AU detection. To achieve this, we propose FACS3D-Net that simultaneously integrates 3D and 2D CNN. Evaluation was on the Expanded BP4D+ database of 200 participants. FACS3D-Net outperformed both 2D CNN and 2D CNN-LSTM approaches. Visualizations of learnt representations suggest that FACS3D-Net is consistent with the spatiotemporal dynamics attended to by human FACS coders. To the best of our knowledge, this is the first work to apply 3D CNN to the problem of AU detection. Le Yang 0009, Itir Önal, Jeffrey F. Cohn, Zakia Hammal, Dongmei Jiang, Hichem Sahli |
ACII | 2 |
| 2019 | PAttNet: Patch-attentive deep network for action unit detection
Itir Önal, László A. Jeni, Jeffrey F. Cohn |
BMVC | 1 |
| 2019 | Unmasking the Devil in the Details: What Works for Deep Facial Action Coding?
Koichiro Niinuma, László A. Jeni, Itir Önal, Jeffrey F. Cohn |
BMVC | 3 |
| 2019 | Cross-domain AU Detection: Domains, Learning Approaches, and MeasuresabstractFacial action unit (AU) detectors have performed well when trained and tested within the same domain. Do AU detectors transfer to new domains in which they have not been trained? To answer this question, we review literature on cross-domain transfer and conduct experiments to address limitations of prior research. We evaluate both deep and shallow approaches to AU detection (CNN and SVM, respectively) in two large, well-annotated, publicly available databases, Expanded BP4D+ and GFT. The databases differ in observational scenarios, participant characteristics, range of head pose, video resolution, and AU base rates. For both approaches and databases, performance decreased with change in domain, often to below the threshold needed for behavioral research. Decreases were not uniform, however. They were more pronounced for GFT than for Expanded BP4D+ and for shallow relative to deep learning. These findings suggest that more varied domains and deep learning approaches may be better suited for promoting generalizability. Until further improvement is realized, caution is warranted when applying AU classifiers from one domain to another. Itir Önal, Jeffrey F. Cohn, László A. Jeni, Zheng Zhang 0023, Lijun Yin 0001 |
FG | 1 |
| 2019 | AFAR: A Deep Learning Based Tool for Automated Facial Affect RecognitionabstractAutomated facial affect recognition is crucial to multiple domains (e.g., health, education, entertainment). Commercial tools are available but costly and of unknown validity. Open-source ones [1] lack user-friendly GUI for use by non-programmers. For both types, evidence of domain transfer and options for retraining for use in new domains typically are lacking. Itir Önal, László A. Jeni, Wanqiao Ding, Jeffrey F. Cohn |
FG | 1 |
| 2018 | Automated Affect Detection in Deep Brain Stimulation for Obsessive-Compulsive Disorder: A Pilot StudyabstractAutomated measurement of affective behavior in psychopathology has been limited primarily to screening and diagnosis. While useful, clinicians more often are concerned with whether patients are improving in response to treatment. Are symptoms abating, is affect becoming more positive, are unanticipated side effects emerging? When treatment includes neural implants, need for objective, repeatable biometrics tied to neurophysiology becomes especially pressing. We used automated face analysis to assess treatment response to deep brain stimulation (DBS) in two patients with intractable obsessive-compulsive disorder (OCD). One was assessed intraoperatively following implantation and activation of the DBS device. The other was assessed three months post-implantation. Both were assessed during DBS on and o conditions. Positive and negative valence were quantified using a CNN trained on normative data of 160 non-OCD participants. Thus, a secondary goal was domain transfer of the classifiers. In both contexts, DBS-on resulted in marked positive affect. In response to DBS-off, affect flattened in both contexts and alternated with increased negative affect in the outpatient setting. Mean AUC for domain transfer was 0.87. These findings suggest that parametric variation of DBS is strongly related to affective behavior and may introduce vulnerability for negative affect in the event that DBS is discontinued. Jeffrey F. Cohn, László A. Jeni, Itir Önal, Donald Malone, Michael S. Okun, David A. Borton, Wayne K. Goodman |
ICMI | 3 |
| 2018 | Modeling and synthesis of kinship patterns of facial expressions
Itir Önal, László A. Jeni, Hamdi Dibeklioglu |
Image Vis. Comput. | 1 |
| 2017 | What Will Your Future Child Look Like? Modeling and Synthesis of Hereditary Patterns of Facial DynamicsabstractAnalysis of kinship from facial images or videos is an important problem. Prior machine learning and computer vision studies approach kinship analysis as a verification or recognition task. In this paper, first time in the literature, we propose a kinship synthesis framework, which generates smile videos of (probable) children from the smile videos of parents. While the appearance of a child's smile is learned using a convolutional encoder-decoder network, another neural network models the dynamics of the corresponding smile. The smile video of the estimated child is synthesized by the combined use of appearance and dynamics models. In order to validate our results, we perform kinship verification experiments using videos of real parents and estimated children generated by our framework. The results show that generated videos of children achieve higher correct verification rates than those of real children. Our results also indicate that the use of generated videos together with the real ones in the training of kinship verification models, increases the accuracy, suggesting that such videos can be used as a synthetic dataset. Itir Önal, Hamdi Dibeklioglu |
FG | 1 |
| 2017 | Does the Strength of Sentiment Matter? A Regression Based Approach on Turkish Social Media
Ali Mert Ertugrul, Itir Önal, Cengiz Acartürk |
NLDB | 2 |
| 2014 | Modeling the Brain Connectivity for Pattern AnalysisabstractAn information theoretic approach is proposed to estimate the degree of connectivity for each voxel with its neighboring voxels. The neighborhood system is defined by spatial and functional connectivity metrics. Then, a local mesh of variable size is formed around each voxel using spatial or functional neighborhood. The mesh arc weights, called Mesh Arc Descriptors (MAD), are estimated by a linear regression model fitted to the voxel intensity values of the functional Magnetic Resonance Images (fMRI). Finally, the error term of the linear regression equation is used to estimate the mesh size for a voxel by optimizing Akaike's information Criterion, Bayesian Information Criterion and Rissanen's Minimum Description Length. fMRI measurements are obtained during a memory encoding and retrieval experiment performed on a subject who is exposed to the stimuli from 10 semantic categories. For each sample, a k-NN classifier is trained using the Mesh Arc Descriptors (MAD) having the variable mesh sizes. The classification performances reflect that the suggested variable-size Mesh Arc Descriptors represents the mental states better than the classical multi-voxel pattern representation. Moreover, we observe that the degree of connectivities in the brain greatly varies for each voxel. Itir Önal, Emre Aksan, Burak Velioglu, Orhan Firat, Mete Ozay, Ilke Öztekin, Fatos T. Yarman-Vural |
ICPR | 1 |
| 2013 | An information theoretic approach to classify cognitive states using fMRIabstractIn this study, an information theoretic approach is proposed to model brain connectivity during a cognitive processing task, measured by functional Magnetic Resonance Imaging (fMRI). For this purpose, a local mesh of varying size is formed around each voxel. The arc weights of each mesh are estimated using a linear regression model by minimizing the squared error. Then, the optimal mesh size for each sample, that represents the information distribution in the brain, is estimated by minimizing various information criteria which employ the mean square error of linear regression model. The estimated mesh size shows the degree of locality or degree of connectivity of the voxels for the underlying cognitive process. The samples are generated during an fMRI experiment employing item recognition (IR) and judgment of recency (JOR) tasks. For each sample, estimated arc weights of the local mesh with optimal size are used to classify whether it belongs to IR or JOR tasks. Results indicate that the suggested connectivity model with optimal mesh size for each sample represent the information distribution in the brain better than the state-of-the art methods. Itir Önal, Mete Ozay, Orhan Firat, Ilke Öztekin, Fatos T. Yarman-Vural |
BIBE | 1 |