EDBT 2026 Demo / reviewers in the wild / expert
Fadime Sener
dblp:119/1497
· DBLP profile ↗
22ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0001-5004-6005ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PALM: A Dataset and Baseline for Learning Multi-Subject Hand PriorabstractThe ability to grasp objects, signal with gestures, and share emotion through touch all stem from the unique capabilities of human hands. Yet creating high-quality personalized hand avatars from images remains challenging due to complex geometry, appearance, and articulation, particularly under unconstrained lighting and limited views. Progress has also been limited by the lack of datasets that jointly provide accurate 3D geometry, high-resolution multiview imagery, and a diverse population of subjects. To address this, we present PALM, a large-scale dataset comprising 13k high-quality hand scans from 263 subjects and$90 k$multi-view images, capturing rich variation in skin tone, age, and geometry. To show its utility, we present a baseline PALM-Net, a multi-subject prior over hand geometry and material properties learned via physically based inverse rendering, enabling realistic, relightable single-image hand avatar personalization. PALM's scale and diversity make it a valuable real-world resource for hand modeling and related research. Zicong Fan, Edoardo Remelli, David Dimond, Fadime Sener, Liuhao Ge, Bugra Tekin, Cem Keskin, Shreyas Hampali |
3DV | 4 |
| 2025 | Context-Enhanced Memory-Refined Transformer for Online Action DetectionabstractOnline Action Detection (OAD) detects actions in streaming videos using past observations. State-of-the-art OAD approaches model past observations and their interactions with an anticipated future. The past is encoded using short-and long-term memories to capture immediate and long-range dependencies, while anticipation compensates for missing future context. We identify a training-inference discrepancy in existing OAD methods that hinders learning effectiveness. The training uses varying lengths of short-term memory, while inference relies on a full-length short-term memory. As a remedy, we propose a Context-enhanced Memory-Refined Transformer (CMeRT). CMeRT introduces a context-enhanced encoder to improve frame representations using additional near-past context. It also features a memory-refined decoder to leverage near-future generation to enhance performance. CMeRT1achieves state-of-the-art in online detection and anticipation on THUMOS’14, CrossTask, and EPIC-Kitchens-100. Zhanzhong Pang, Fadime Sener, Angela Yao |
CVPR | 2 |
| 2025 | Streaming Videollms for Real-Time Procedural Video Understanding
Dibyadip Chatterjee, Edoardo Remelli, Yale Song, Bugra Tekin, Abhay Mittal, Bharat Bhatnagar, Necati Cihan Camgöz, Shreyas Hampali, Eric Sauser, Shugao Ma, Angela Yao, Fadime Sener |
ICCV | 12 |
| 2025 | Spatial and temporal beliefs for mistake detection in assembly tasksabstractAssembly tasks, as an integral part of daily routines and activities, involve a series of sequential steps that are prone to error. This paper proposes a novel method for identifying ordering mistakes in assembly tasks based on knowledge-grounded beliefs. The beliefs comprise spatial and temporal aspects, each serving a unique role. Spatial beliefs capture the structural relationships among assembly components and indicate their topological feasibility. Temporal beliefs model the action preconditions and enforce sequencing constraints. Furthermore, we introduce a learning algorithm that dynamically updates and augments the belief sets online. To evaluate, we first test our approach in deducing predefined rules on synthetic data based on industry assembly. We also verify our approach on the real-world Assembly101 dataset, enhanced with annotations of component information. Our framework achieves superior performance in detecting ordering mistakes under both synthetic and real-world settings, highlighting the effectiveness of our approach. • We present two belief sets for the assembly tasks. • We propose a novel mistake detection framework for assembly tasks. • The framework can better detect ordering mistakes and can be integrated with perception modules. Guodong Ding, Fadime Sener, Shugao Ma, Angela Yao |
Comput. Vis. Image Underst. | 2 |
| 2024 | Cost-Sensitive Learning for Long-Tailed Temporal Action Segmentation
Zhanzhong Pang, Fadime Sener, Shrinivas Ramasubramanian, Angela Yao |
BMVC | 2 |
| 2024 | X-MIC: Cross-Modal Instance Conditioning for Egocentric Action GeneralizationabstractLately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recog-nition. However, the adaptation of these models to egocentric videos has been largely unexplored. To address this gap, we propose a simple yet effective cross-modal adaptation framework, which we call X-MIC. Using a video adapter, our pipeline learns to align frozen text embeddings to each egocentric video directly in the shared embedding space. Our novel adapter architecture retains and improves generalization of the pre-trained VLMs by disentangling learnable temporal modeling and frozen visual en-coder. This results in an enhanced alignment of text embeddings to each egocentric video, leading to a significant improvement in cross-dataset generalization. We evaluate our approach on the Epic-Kitchens, Ego4D, and EGTEA datasets for fine-grained cross-dataset action generalization, demonstrating the effectiveness of our method.11https://github.com/annusha/xmic Anna Kukleva, Fadime Sener, Edoardo Remelli, Bugra Tekin, Eric Sauser, Bernt Schiele, Shugao Ma |
CVPR | 2 |
| 2024 | Long-Tail Temporal Action Segmentation with Group-Wise Temporal Logit Adjustment
Zhanzhong Pang, Fadime Sener, Shrinivas Ramasubramanian, Angela Yao |
ECCV (30) | 2 |
| 2024 | On the Utility of 3D Hand Poses for Action Recognition
Md. Salman Shamil, Dibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela Yao |
ECCV (6) | 3 |
| 2024 | DiffH2O: Diffusion-Based Synthesis of Hand-Object Interactions from Textual Descriptions
Sammy Joe Christen, Shreyas Hampali, Fadime Sener, Edoardo Remelli, Tomas Hodan, Eric Sauser, Shugao Ma, Bugra Tekin |
SIGGRAPH Asia | 3 |
| 2024 | Temporal Action Segmentation: An Analysis of Modern TechniquesabstractTemporal action segmentation (TAS) in videos aims at densely identifying video frames in minutes-long videos with multiple action classes. As a long-range video understanding task, researchers have developed an extended collection of methods and examined their performance using various benchmarks. Despite the rapid growth of TAS techniques in recent years, no systematic survey has been conducted in these sectors. This survey analyzes and summarizes the most significant contributions and trends. In particular, we first examine the task definition, common benchmarks, types of supervision, and prevalent evaluation measures. In addition, we systematically investigate two essential techniques of this topic, i.e., frame representation and temporal modeling, which have been studied extensively in the literature. We then conduct a thorough review of existing TAS works categorized by their levels of supervision and conclude our survey by identifying and emphasizing several research gaps. Guodong Ding, Fadime Sener, Angela Yao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | AssemblyHands: Towards Egocentric Activity Understanding via 3D Hand Pose EstimationabstractWe present AssemblyHands, a large-scale benchmark dataset with accurate 3D hand pose annotations, to facilitate the study of egocentric activities with challenging hand-object interactions. The dataset includes synchronized egocentric and exocentric images sampled from the recent Assembly101 dataset, in which participants assemble and disassemble take-apart toys. To obtain high-quality 3D hand pose annotations for the egocentric images, we develop an efficient pipeline, where we use an initial set of manual annotations to train a model to automatically annotate a much larger dataset. Our annotation model uses multi-view feature fusion and an iterative refinement scheme, and achieves an average keypoint error of 4.20 mm, which is 85% lower than the error of the original annotations in Assembly101. AssemblyHands provides 3.0M annotated images, including 490K egocentric images, making it the largest existing benchmark dataset for egocentric 3D hand pose estimation. Using this data, we develop a strong single-view baseline of 3D hand pose estimation from egocentric images. Furthermore, we design a novel action classification task to evaluate predicted 3D hand poses. Our study shows that having higher-quality hand poses directly improves the ability to recognize actions. Takehiko Ohkawa, Fadime Sener, Tomas Hodan, Luan Tran, Cem Keskin |
CVPR | 3 |
| 2023 | Opening the Vocabulary of Egocentric ActionsabstractHuman actions in egocentric videos often feature hand-object interactions composed of a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations — sparsity of action compositions and a closed set of interacting objects. This paper proposes a novel open vocabulary action recognition task. Given a set of verbs and objects observed during training, the goal is to generalize the verbs to an open vocabulary of actions with seen and novel objects. To this end, we decouple the verb and object predictions via an object-agnostic _verb encoder_ and a prompt-based _object encoder_. The prompting leverages CLIP representations to predict an open vocabulary of interacting objects. We create open vocabulary benchmarks on the EPIC-KITCHENS-100 and Assembly101 datasets; whereas closed-action methods fail to generalize, our proposed method is effective. In addition, our object encoder significantly outperforms existing open-vocabulary visual recognition methods in recognizing novel interacting objects. Dibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela Yao |
NeurIPS | 2 |
| 2023 | Transferring Knowledge From Text to Video: Zero-Shot Anticipation for Procedural ActionsabstractCan we teach a robot to recognize and make predictions for activities that it has never seen before? We tackle this problem by learning models for video from text. This paper presents a hierarchical model that generalizes instructional knowledge from large-scale text corpora and transfers the knowledge to video. Given a portion of an instructional video, our model recognizes and predicts coherent and plausible actions multiple steps into the future, all in rich natural language. To demonstrate the capabilities of our model, we introduce the Tasty Videos Dataset V2, a collection of 4022 recipes for zero-shot learning, recognition and anticipation. Extensive experiments with various evaluation metrics demonstrate the potential of our method for generalization, given limited video data for training models. Fadime Sener, Rishabh Saraf, Angela Yao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesabstractAssembly101 is a new procedural activity dataset fea-turing 4321 videos of people assembling and disassembling 101 “take-apart” toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natu-ral variations in action ordering, mistakes, and corrections. Assembly101 is the first multi-view action dataset, with si-multaneous static (8) and egocentric (4) recordings. Se-quences are annotated with more than 100K coarse and 1M fine-grained action segments, and I8M 3D hand poses. We benchmark on three action understanding tasks: recognition, anticipation and temporal segmentation. Ad-ditionally, we propose a novel task of detecting mistakes. The unique recording format and rich set of annotations al-low us to investigate generalization to new toys, cross-view transfer, long-tailed distributions, and pose vs. appearance. We envision that Assemblyl0l will serve as a new challenge to investigate various activity understanding problems. Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Dipika Singhania, Robert Wang 0002, Angela Yao |
CVPR | 1 |
| 2022 | Transformed ROIs for capturing visual transformations in videos
Fadime Sener, Angela Yao |
Comput. Vis. Image Underst. | 2 |
| 2020 | Temporal Aggregate Representations for Long-Range Video Understanding
Fadime Sener, Dipika Singhania, Angela Yao |
ECCV (16) | 1 |
| 2019 | Unsupervised Learning of Action Classes With Continuous Temporal EmbeddingabstractThe task of temporally detecting and segmenting actions in untrimmed videos has seen an increased attention recently. One problem in this context arises from the need to define and label action boundaries to create annotations for training which is very time and cost intensive. To address this issue, we propose an unsupervised approach for learning action classes from untrimmed video sequences. To this end, we use a continuous temporal embedding of framewise features to benefit from the sequential nature of activities. Based on the latent space created by the embedding, we identify clusters of temporal segments across all videos that correspond to semantic meaningful action classes. The approach is evaluated on three challenging datasets, namely the Breakfast dataset, YouTube Instructions, and the 50Salads dataset. While previous works assumed that the videos contain the same high level activity, we furthermore show that the proposed approach can also be applied to a more general setting where the content of the videos is unknown. Anna Kukleva, Hilde Kuehne, Fadime Sener, Juergen Gall |
CVPR | 3 |
| 2019 | Zero-Shot Anticipation for Instructional ActivitiesabstractHow can we teach a robot to predict what will happen next for an activity it has never seen before? We address the problem of zero-shot anticipation by presenting a hierarchical model that generalizes instructional knowledge from large-scale text-corpora and transfers the knowledge to the visual domain. Given a portion of an instructional video, our model predicts coherent and plausible actions multiple steps into the future, all in rich natural language. To demonstrate the anticipation capabilities of our model, we introduce the Tasty Videos dataset, a collection of 2511 recipes for zero-shot learning, recognition and anticipation. Fadime Sener, Angela Yao |
ICCV | 1 |
| 2018 | Unsupervised Learning and Segmentation of Complex Activities From VideoabstractThis paper presents a new method for unsupervised segmentation of complex activities from video into multiple steps, or sub-activities, without any textual input. We propose an iterative discriminative-generative approach which alternates between discriminatively learning the appearance of sub-activities from the videos' visual features to sub-activity labels and generatively modelling the temporal structure of sub-activities using a Generalized Mallows Model. In addition, we introduce a model for background to account for frames unrelated to the actual activities. Our approach is validated on the challenging Breakfast Actions and Inria Instructional Videos datasets and outperforms both unsupervised and weakly-supervised state of the art. Fadime Sener, Angela Yao |
CVPR | 1 |
| 2017 | DRAW: Deep Networks for Recognizing Styles of Artists Who Illustrate Children's BooksabstractThis paper is motivated from a young boy's capability to recognize an illustrator's style in a totally different context. In the book "We are All Born Free" [1], composed of selected rights from the Universal Declaration of Human Rights interpreted by different illustrators, the boy was surprised to see a picture similar to the ones in the "Winnie the Witch" series drawn by Korky Paul (Figure [1]). The style was noticeable in other characters of the same illustrator in different books as well. The capability of a child to easily spot the style was shown to be valid for other illustrators such as Axel Scheffler and Debi Gliori. The boy's enthusiasm let us to start the journey to explore the capabilities of machines to recognize the style of illustrators. Samet Hicsonmez, Nermin Samet, Fadime Sener, Pinar Duygulu |
ICMR | 3 |
| 2015 | Two-person interaction recognition via spatial multiple instance embedding
Fadime Sener, Nazli Ikizler-Cinbis |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | Ensemble of multiple instance classifiers for image re-ranking
Fadime Sener, Nazli Ikizler-Cinbis |
Image Vis. Comput. | 1 |