EDBT 2026 Demo / reviewers in the wild / expert
László A. Jeni
dblp:35/7547 · also László Attila Jeni
· DBLP profile ↗
46ranked-venue papers
4as first author
29since 2021 · last 2026
0000-0002-2830-700XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 3 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 1 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeXabstractWe introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial progress, the field lacks a standardized evaluation framework to thoroughly assess their cross-modality comprehension performance. MAVERIX curates 2,556 questions from 700 videos, in the form of both multiple-choice and open-ended formats, explicitly designed to evaluate multimodal models through questions that necessitate tight integration of video and audio information, spanning a broad spectrum of agentic scenarios. MAVERIX uniquely provides models with questions that closely mimic the multimodal understanding experiences available to humans during decision-making processes. To our knowledge, MAVERIX is the first benchmark aimed explicitly at assessing comprehensive audiovisual integration in such granularity. Experiments with state-of-the-art models, including Qwen 2.5 Omni and Gemini 2.5 Flash-Lite, show performance around 64% accuracy, while human experts reach near-ceiling performance of 92.8%, exposing a substantial gap to human-level comprehension. With standardized evaluation protocols, a rigorously annotated pipeline, and a public toolkit, MAVERIX establishes a challenging testbed for advancing audiovisual multimodal intelligence, with the website publicly available below. Liuyue Xie, Avik Kuthiala, George Z. Wei, Ananya Bal, Mosam Dabhi, Liting Wen, Taru Rustagi, Ethan Lai, Sushil Khyalia, Rohan Choudhury, Morteza Ziyadi, László A. Jeni |
AAAI | 15 |
| 2026 | RAT4D: Rig and Animate Objects without Surface Templates in 4DabstractWe present a surface-template-free method for reconstructing dense, rigged (re-animatable) 3D models from monocular videos. By combining robust pose optimization with differentiable Gaussian splatting, this work bridges the gap between flexible template-free approaches and the visual quality of template-based methods. Starting from noisy 2D keypoints, we refine 3D poses through kinematic and temporal constraints, then attach Gaussian primitives to the optimized skeleton for differentiable supervision and rendering. We demonstrate dense, rigged (re-animatable) 3D models without surface templates across humans, animals, insects, and everyday articulated objects, and we empirically show that closing the rendering–pose loop improves 3D lifting from noisy landmarks. A key enabler of these template-free reconstructions is our kinematic optimization, which reduces 3D pose error by 20–25% relative to template-free baselines; at the same time, our results approach template-based visual metrics (PSNR/SSIM within 5%). We also demonstrate the method’s practical utility by detecting and correcting geometric inconsistencies in AI-generated videos. While limited to articulated subjects with detectable keypoints, the approach provides a practical pipeline that serves as a drop-in refinement to improve 3D lifting in existing pipelines and enables the creation of rigged 3D assets from casual captures when expensive surface templates (e.g., MoCap-derived) are unavailable. Mosam Dabhi, Simon Lucey, László A. Jeni |
WACV | 3 |
| 2026 | Training-free Multi-view 4D Human Motion Reconstruction Virtual Reality System
Yijie He, Joel Julin, Ryosuke Ichikari, Satoki Ogiso, Satoshi Nakae, Akihiro Sato, Takeshi Kurata, László A. Jeni |
WACV | 10 |
| 2025 | DiSRT-In-Bed: Diffusion-Based Sim-to-Real Transfer Framework for In-Bed Human Mesh RecoveryabstractIn-Bed human mesh recovery can be crucial and enabling for several healthcare applications, including sleep pattern monitoring, rehabilitation support, and pressure ulcer prevention. However, it is difficult to collect large real-world visual datasets in this domain, in part due to privacy and expense constraints, which in turn presents significant challenges for training and deploying deep learning models. Existing in-bed human mesh estimation methods often rely heavily on real-world data, limiting their ability to generalize across different in-bed scenarios, such as varying coverings and environmental settings. To address this, we propose a Sim-To-Real Transfer Framework for in-bed human mesh recovery from overhead depth images, which leverages large-scale synthetic data alongside limited or no real-world samples. We introduce a diffusion model that bridges the gap between synthetic data and real data to support generalization in real-world in-bed pose and body inference scenarios. Extensive experiments and ablation studies validate the effectiveness of our framework, demonstrating significant improvements in robustness and adaptability across diverse healthcare scenarios. Project page can be found at https://jing-g2.github.io/DiSRT-In-Bed/. László A. Jeni, Zackory Erickson |
CVPR | 3 |
| 2025 | Custom Condition Generation for Zero-Shot Human-Scene Interactions SynthesisabstractExisting methods for creating human interactions within scenes show promise for common interactions, but often fail with less frequent ones. To overcome this, we introduce a new approach that creates tailored conditions for generating these interactions without previously seen examples. This method leverages the strengths of both large language models (LLMs) and vision-language models (VLMs). Unlike the GenZI, the current state-of-the-art approach, which struggles with rare interactions due to its reliance on VLM inpainting, our method follows a three-step process: first, we generate a preliminary human posture using VLMs and then estimate this posture in three dimensions. Next, we refine the conditions to fit the specific scene and interaction by analyzing the inputs with both LLMs and VLMs. Finally, we fine-tune the placement, orientation, and posture of the human figure using specific optimization techniques. Our experimental results show that this method performs well across a wide range of interactions, including those that are less common. Ryosuke Kawamura, Zoltán Ádám Milacski, Fernando De la Torre, László A. Jeni, Koichiro Niinuma |
FG | 4 |
| 2025 | AlignDiff: Learning Physically-Grounded Camera Alignment via DiffusionabstractAccurate camera calibration is a fundamental task for 3D perception, especially when dealing with real-world, in-the-wild environments where complex optical distortions are common. Existing methods often rely on pre-rectified images or calibration patterns, which limits their applicability and flexibility. In this work, we introduce a novel framework that addresses these challenges by jointly modeling camera intrinsic and extrinsic parameters using a generic ray camera model. Unlike previous approaches, AlignDiff shifts focus from semantic to geometric features, enabling more accurate modeling of local distortions. We propose AlignDiff, a diffusion model conditioned on geometric priors, enabling the simultaneous estimation of camera distortions and scene geometry. To enhance distortion prediction, we incorporate edge-aware attention, focusing the model on geometric features around image edges, rather than semantic content. Furthermore, to enhance generalizability to real-world captures, we incorporate a large database of ray-traced lenses containing over three thousand samples. This database characterizes the distortion inherent in a diverse variety of lens forms. Our experiments demonstrate that the proposed method significantly reduces the angular error of estimated ray bundles by ~8.2 degrees and overall calibration accuracy, outperforming existing approaches on challenging, real-world datasets. Liuyue Xie, Jiancong Guo, Ozan Cakmakci, Andre Araujo, László A. Jeni, Zhiheng Jia |
ICCV | 5 |
| 2025 | Gaussian Splatting Lucas-KanadeabstractGaussian Splatting and its dynamic extensions are effective for reconstructing 3D scenes from 2D images when there is significant camera movement to facilitate motion parallax and when scene objects remain relatively static. However, in many real-world scenarios, these conditions are not met. As a consequence, data-driven semantic and geometric priors have been favored as regularizers, despite their bias toward training data and their neglect of broader movement dynamics.
Departing from this practice, we propose a novel analytical approach that adapts the classical Lucas-Kanade method to dynamic Gaussian splatting. By leveraging the intrinsic properties of the forward warp field network, we derive an analytical velocity field that, through time integration, facilitates accurate scene flow computation. This enables the precise enforcement of motion constraints on warp fields, thus constraining both 2D motion and 3D positions of the Gaussians. Our method excels in reconstructing highly dynamic scenes with minimal camera movement, as demonstrated through experiments on both synthetic and real-world scenes. Liuyue Xie, Joel Julin, Koichiro Niinuma, László A. Jeni |
ICLR | 4 |
| 2025 | GHOST: Grounded Human Motion Generation with Open Vocabulary Scene-and-Text ContextsabstractThe connection between our 3D surroundings and the descriptive language that characterizes them would be well-suited for localizing and generating human motion in context but for one problem. The complexity introduced by multiple modalities makes capturing this connection challenging with a fixed set of descriptors. Specifically, closed vocabulary scene encoders, which require learning text-scene associations from scratch, have been favored in the literature, often resulting in inaccurate motion grounding. In this paper, we propose a method that integrates an open vocabulary scene encoder into the architecture, establishing a robust connection between text and scene. Our two-step approach starts with pretraining the scene encoder through knowledge distillation from an existing open vocabulary semantic image segmentation model, ensuring a shared text-scene feature space. Subsequently, the scene encoder is fine-tuned for conditional motion generation, incorporating two novel regularization losses that regress the category and size of the goal object. Our methodology achieves up to a 30% reduction in the goal object distance metric compared to the prior state-of-the-art baseline model on the HUMANISE dataset. This improvement is demonstrated through evaluations conducted using three implementations of our framework, a perceptual study, and an open vocabulary experiment. Additionally, our method is designed to accommodate future 2D open vocabulary segmentation methods for distillation in a plug-and-play manner. Zoltán Ádám Milacski, Koichiro Niinuma, Ryosuke Kawamura, Fernando De la Torre, László A. Jeni |
WACV | 5 |
| 2025 | Through the Curved Cover: Synthesizing Cover Aberrated Scenes with Refractive FieldabstractRecent extended reality headsets and field robots have adopted covers to protect the front-facing cameras from environmental hazards and falls. The surface irregularities on the cover can lead to optical aberrations like blurring and non-parametric distortions. Novel view synthesis methods like NeRF and 3D Gaussian Splatting are ill-equipped to synthesize from sequences with optical aberrations. To address this challenge, we introduce SynthCover to enable novel view synthesis through protective covers for downstream extended reality applications. SynthCover employs a Refractive Field that estimates the cover's geometry, enabling precise analytical calculation of refracted rays. Experiments on synthetic and real-world scenes demonstrate our method's ability to accurately model scenes viewed through protective covers, achieving a significant improvement in rendering quality compared to prior methods. We also show that the model can adjust well to various cover geometries with synthetic sequences captured with covers of different surface curvatures. To motivate further studies on this problem, we provide the benchmarked dataset containing real and synthetic walkable scenes captured with protective cover optical aberrations. Liuyue Xie, Jiancong Guo, László A. Jeni, Zhiheng Jia, Mingyang Li 0001, Yunwen Zhou |
WACV | 3 |
| 2025 | Exploring image and skeleton-based action recognition approaches for clinical in-bed classification of simulated epileptic seizure movementsabstractEpileptic seizure classification based on seizure semiology requires automated, quantitative approaches to support the diagnosis of epilepsy, which affects 1% of the world’s population. Current approaches address the problem on a seizure level, neglecting the detailed evaluation of the classification of the underlying action features, also known as Movements of Interest (MOIs), which are critical for epileptologists in determining their classifications. Moreover, it hinders objective comparison of these approaches and attribution of performance differences due to datasets, intra-dataset MOI distribution, or architecture variations. Objective evaluation of action recognition techniques is crucial, with MOIs serving as foundational elements of semiology for clinical in-bed applications to facilitate epileptic seizure classification. However, until now, there were no MOI datasets available nor benchmarks comparing different action recognition approaches for this clinical problem. Therefore, as a pilot, we introduced a novel, simulated seizure semiology dataset carried out by 8 experienced epileptologists in an EMU bed, consisting of 7 MOI classes. We compare several computer vision methods for MOI classification, two image-based (I3D and Uniformerv2), and two skeleton-based (ST-GCN++ and PoseC3D) action recognition approaches. This study emphasizes the advantages of a 2-stage skeleton-based action recognition approach in a transfer learning setting (4 classes) and the multi-scale challenge of MOI classification (7 classes), advocating for the integration of skeleton-based methods with hand gesture recognition technologies in the future. The study’s controlled MOI simulation dataset provides us with the opportunity to advance the development of automated epileptic seizure classification systems, paving the way for enhancing their performance and having the potential to contribute to improved patient care. Tamás Karácsony, Nicholas Fearns, Denise Birk, Selina Denise Trapp, Katharina Ernst, Christian Vollmar, Jan Rémi, László A. Jeni, Fernando De la Torre, João Paulo da Silva Cunha |
Expert Syst. Appl. | 8 |
| 2024 | 3D-LFM: Lifting Foundation ModelabstractThe lifting of a 3D structure and camera from 2D land-marks is at the cornerstone of the discipline of computer vision. Traditional methods have been confined to specific rigid objects, such as those in Perspective-n-Point (PnP) problems, but deep learning has expanded our capability to reconstruct a wide range of object classes (e.g. C3DPO [18] and PAUL [24]) with resilience to noise, occlusions, and perspective distortions. However, all these techniques have been limited by the fundamental need to establish correspondences across the 3D training data, significantly limiting their utility to applications where one has an abundance of “in-correspondence” 3D data. Our approach harnesses the inherent permutation equivariance of transformers to manage varying numbers of points per 3D data instance, withstands occlusions, and generalizes to unseen categories. We demonstrate state-of-the-art performance across 2D-3D lifting task benchmarks. Since our approach can be trained across such a broad class of structures, we refer to it simply as a 3D Lifting Foundation Model (3D-LFM) - the first of its kind. Mosam Dabhi, László A. Jeni, Simon Lucey |
CVPR | 2 |
| 2024 | CoGS: Controllable Gaussian SplattingabstractCapturing and re-animating the 3D structure of artic-ulated objects present significant barriers. On one hand, methods requiring extensively calibrated multi-view setups are prohibitively complex and resource-intensive, limiting their practical applicability. On the other hand, while single-camera Neural Radiance Fields (NeRFs) offer a more streamlined approach, they have excessive training and rendering costs. 3D Gaussian Splatting would be a suitable alternative but for two reasons. Firstly, existing methods for 3D dynamic Gaussians require synchronized multi- view cameras, and secondly, the lack of controllability in dynamic scenarios. We present CoGS, a methodfor Controllable Gaussian Splatting, that enables the direct ma-nipulation of scene elements, offering real-time control of dynamic scenes without the prerequisite of pre-computing control signals. We evaluated CoGS using both synthetic and real-world datasets that include dynamic objects that differ in degree of difficulty. In our evaluations, CoGS con-sistently outperformed existing dynamic and controllable neural representations in terms of visual fidelity. Joel Julin, Zoltán Ádám Milacski, Koichiro Niinuma, László A. Jeni |
CVPR | 5 |
| 2024 | Video Question Answering with Procedural Programs
Rohan Choudhury, Koichiro Niinuma, Kris Makoto Kitani, László A. Jeni |
ECCV (38) | 4 |
| 2024 | Enhanced Product Classification Using Learned Prompt Ensembling and Dual Interpolation with CLIP-Based ModelabstractRetail product classification is a crucial technology due to its market size and potential. Several algorithms have been proposed; however, these approaches are not suitable due to their inability to accommodate new data without retraining or their insufficient performance caused by class names reflecting product-specific names. In this study, we adopt a CLIP-based model for retail product classification to overcome the challenges associated with model retraining for new products and the issue of unique class names. Our approach, learned prompt ensembling and dual interpolation (LPEDI), combines prompt learning and its ensembling with encoder fine-tuning, and employs dual interpolation for coefficient adjustment. The method outperforms existing solutions on two retail product datasets, achieving a 5.9% improvement for in-distribution data and a 4.5% gain for out-of-distribution data. These results establish LPEDI as a practical and effective solution for retail product classification. Takahisa Yamamoto, Koichiro Niinuma, László A. Jeni |
MMSP | 3 |
| 2024 | Don't Look Twice: Faster Video Transformers with Run-Length TokenizationabstractVideo transformers are slow to train due to extremely large numbers of input tokens, even though many video tokens are repeated over time. Existing methods to remove uninformative tokens either have significant overhead, negating any speedup, or require tuning for different datasets and examples. We present Run-Length Tokenization (RLT), a simple approach to speed up video transformers inspired by run-length encoding for data compression. RLT efficiently finds and removes `runs' of patches that are repeated over time before model inference, then replaces them with a single patch and a positional encoding to represent the resulting token's new length.
Our method is content-aware, requiring no tuning for different datasets, and fast, incurring negligible overhead.
RLT yields a large speedup in training, reducing the wall-clock time to fine-tune a video transformer by 30% while matching baseline model performance. RLT also works without training, increasing model throughput by 35% with only 0.1% drop in accuracy.
RLT speeds up training at 30 FPS by more than 100%, and on longer video datasets, can reduce the token count by up to 80\%. Our project page is at rccchoudhury.github.io/projects/rlt. Rohan Choudhury, Guanglei Zhu, Koichiro Niinuma, Kris Makoto Kitani, László A. Jeni |
NeurIPS | 6 |
| 2024 | 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion ModelsabstractExisting dynamic scene generation methods mostly rely on distilling knowledge from pre-trained 3D generative models, which are typically fine-tuned on synthetic object datasets.
As a result, the generated scenes are often object-centric and lack photorealism.
To address these limitations, we introduce a novel pipeline designed for photorealistic text-to-4D scene generation, discarding the dependency on multi-view generative models and instead fully utilizing video generative models trained on diverse real-world datasets.
Our method begins by generating a reference video using the video generation model.
We then learn the canonical 3D representation of the video using a freeze-time video, delicately generated from the reference video.
To handle inconsistencies in the freeze-time video, we jointly learn a per-frame deformation to model these imperfections.
We then learn the temporal deformation based on the canonical representation to capture dynamic interactions in the reference video.
The pipeline facilitates the generation of dynamic scenes with enhanced photorealism and structural integrity, viewable from multiple perspectives, thereby setting a new standard in 4D scene generation. Chaoyang Wang 0001, Peiye Zhuang, Willi Menapace, Aliaksandr Siarohin, Junli Cao, László A. Jeni, Sergey Tulyakov, Hsin-Ying Lee 0001 |
NeurIPS | 7 |
| 2024 | Deep learning methods for single camera based clinical in-bed movement action recognition
Tamás Karácsony, László A. Jeni, Fernando De la Torre, João Paulo da Silva Cunha |
Image Vis. Comput. | 2 |
| 2023 | Flow Supervision for Deformable NeRFabstractIn this paper we present a new method for deformable NeRF that can directly use optical flow as supervision. We overcome the major challenge with respect to the computationally inefficiency of enforcing the flow constraints to the backward deformation field, used by deformable NeRFs. Specifically, we show that inverting the backward deformation function is actually not needed for computing scene flows between frames. This insight dramatically simplifies the problem, as one is no longer constrained to deformation functions that can be analytically inverted. Instead, thanks to the weak assumptions required by our derivation based on the inverse function theorem, our approach can be extended to a broad class of commonly used backward deformation field. We present results on monocular novel view synthesis with rapid object motion, and demonstrate significant improvements over baselines without flow supervision. Chaoyang Wang 0001, Lachlan E. MacDonald, László A. Jeni, Simon Lucey |
CVPR | 3 |
| 2023 | DyLiN: Making Light Field Networks DynamicabstractLight Field Networks, the re-formulations of radiance fields to oriented rays, are magnitudes faster than their coordinate network counterparts, and provide higher fidelity with respect to representing 3D structures from 2D observations. They would be well suited for generic scene representation and manipulation, but suffer from one problem: they are limited to holistic and static scenes. In this paper, we propose the Dynamic Light Field Network (DyLiN) method that can handle non-rigid deformations, including topological changes. We learn a deformation field from input rays to canonical rays, and lift them into a higher dimensional space to handle discontinuities. We further introduce CoDyLiN, which augments DyLiN with controllable attribute inputs. We train both models via knowledge distillation from pretrained dynamic radiance fields. We evaluated DyLiN using both synthetic and real world datasets that include various non-rigid deformations. DyLiN qualitatively outperformed and quantitatively matched state-of-the-art methods in terms of visual fidelity, while being 25 – 71× computationally faster. We also tested CoDyLiN on attribute annotated data and it surpassed its teacher model. Project page: https://dylin2023.github.io. Joel Julin, Zoltán Ádám Milacski, Koichiro Niinuma, László A. Jeni |
CVPR | 5 |
| 2023 | Multimodal Feature Selection for Detecting Mothers' Depression in Dyadic Interactions with their Adolescent OffspringabstractDepression is the most common psychological disorder, a leading cause of disability world-wide, and a major contributor to inter-generational transmission of psychopathology within families. To contribute to our understanding of depression within families and to inform modality selection and feature reduction, it is critical to identify interpretable features in developmentally appropriate contexts. Mothers with and without depression were studied. Depression was defined as history of treatment for depression and elevations in current or recent symptoms. We explored two multimodal feature selection strategies in dyadic interaction tasks of mothers with their adolescent children for depression detection. Modalities included face and head dynamics, facial action units, speech-related behavior, and verbal features. The initial feature space was vast and inter-correlated (collinear). To reduce dimensionality and gain insight into the relative contribution of each modality and feature, we explored feature selection strategies using Variance Inflation Factor (VIF) and Shapley values. On an average collinearity correction through VIF resulted in about 4 times feature reduction across unimodal and multimodal features. Collinearity correction was also found to be an optimal intermediate step prior to Shapley analysis. Shapley feature selection following VIF yielded best performance. The top 15 features obtained through Shapley achieved 78% accuracy. The most informative features came from all four modalities sampled, which supports the importance of multimodal feature selection. Maneesh Bilalpur, Saurabh Hinduja, Laura A. Cariola, Lisa Sheeber, Nick Alien, László A. Jeni, Louis-Philippe Morency, Jeffrey F. Cohn |
FG | 6 |
| 2023 | CoNFies: Controllable Neural Face AvatarsabstractNeural Radiance Fields (NeRF) are compelling techniques for modeling dynamic 3D scenes from 2D image collections. These volumetric representations would be well suited for synthesizing novel facial expressions but for two problems. First, deformable NeRFs are object agnostic and model holistic movement of the scene: they can replay how the motion changes over time, but they cannot alter it in an interpretable way. Second, controllable volumetric representations typically require either time-consuming manual annotations or 3D supervision to provide semantic meaning to the scene. We propose a controllable neural representation for face self-portraits (CoNFies), that solves both of these problems within a common framework, and it can rely on automated processing. We use automated facial action recognition (AFAR) to characterize facial expressions as a combination of action units (AU) and their intensities. AUs provide both the semantic locations and control labels for the system. CoNFies outperformed competing methods for novel view and expression synthesis in terms of visual and anatomic fidelity of expressions. Koichiro Niinuma, László A. Jeni |
FG | 3 |
| 2023 | TEMPO: Efficient Multi-View Pose Estimation, Tracking, and ForecastingabstractExisting volumetric methods for predicting 3D human pose estimation are accurate, but computationally expensive and optimized for single time-step prediction. We present TEMPO, an efficient multi-view pose estimation model that learns a robust spatiotemporal representation, improving pose accuracy while also tracking and forecasting human pose. We significantly reduce computation compared to the state-of-the-art by recurrently computing per-person 2D pose features, fusing both spatial and temporal information into a single representation. In doing so, our model is able to use spatiotemporal context to predict more accurate human poses without sacrificing efficiency. We further use this representation to track human poses over time as well as predict future poses. Finally, we demonstrate that our model is able to generalize across datasets without scene-specific fine-tuning. TEMPO achieves 10% better MPJPE with a 33× improvement in FPS compared to TesseTrack on the challenging CMU Panoptic Studio dataset. Our code and demos are available at https://rccchoudhury.github.io/tempo2023/. Rohan Choudhury, Kris Makoto Kitani, László A. Jeni |
ICCV | 3 |
| 2023 | LightSpeed: Light and Fast Neural Light Fields on Mobile DevicesabstractReal-time novel-view image synthesis on mobile devices is prohibitive due to the limited computational power and storage. Using volumetric rendering methods, such as NeRF and its derivatives, on mobile devices is not suitable due to the high computational cost of volumetric rendering. On the other hand, recent advances in neural light field representations have shown promising real-time view synthesis results on mobile devices. Neural light field methods learn a direct mapping from a ray representation to the pixel color. The current choice of ray representation is either stratified ray sampling or Plücker coordinates, overlooking the classic light slab (two-plane) representation, the preferred representation to interpolate between light field views. In this work, we find that using the light slab representation is an efficient representation for learning a neural light field. More importantly, it is a lower-dimensional ray representation enabling us to learn the 4D ray space using feature grids which are significantly faster to train and render. Although mostly designed for frontal views, we show that the light-slab representation can be further extended to non-frontal scenes using a divide-and-conquer strategy. Our method provides better rendering quality than prior light field methods and a significantly better trade-off between rendering quality and speed than prior light field methods. Aarush Gupta, Junli Cao, Chaoyang Wang 0001, Ju Hu, Sergey Tulyakov, Jian Ren 0005, László A. Jeni |
NeurIPS | 7 |
| 2022 | MBW: Multi-view Bootstrapping in the WildabstractLabeling articulated objects in unconstrained settings has a wide variety of applications including entertainment, neuroscience, psychology, ethology, and many fields of medicine. Large offline labeled datasets do not exist for all but the most common articulated object categories (e.g., humans). Hand labeling these landmarks within a video sequence is a laborious task. Learned landmark detectors can help, but can be error-prone when trained from only a few examples. Multi-camera systems that train fine-grained detectors have shown significant promise in detecting such errors, allowing for self-supervised solutions that only need a small percentage of the video sequence to be hand-labeled. The approach, however, is based on calibrated cameras and rigid geometry, making it expensive, difficult to manage, and impractical in real-world scenarios. In this paper, we address these bottlenecks by combining a non-rigid 3D neural prior with deep flow to obtain high-fidelity landmark estimates from videos with only two or three uncalibrated, handheld cameras. With just a few annotations (representing $1-2\%$ of the frames), we are able to produce 2D results comparable to state-of-the-art fully supervised methods, along with 3D reconstructions that are impossible with other existing approaches. Our Multi-view Bootstrapping in the Wild (MBW) approach demonstrates impressive results on standard human datasets, as well as tigers, cheetahs, fish, colobus monkeys, chimpanzees, and flamingos from videos captured casually in a zoo. We release the codebase for MBW as well as this challenging zoo dataset consisting of image frames of tail-end distribution categories with their corresponding 2D and 3D labels generated from minimal human intervention. Mosam Dabhi, Chaoyang Wang 0001, Tim Clifford, László A. Jeni, Ian R. Fasel, Simon Lucey |
NeurIPS | 4 |
| 2022 | 3D Human Pose, Shape and Texture From Low-Resolution Images and Videosabstract3D human pose and shape estimation from monocular images has been an active research area in computer vision. Existing deep learning methods for this task rely on high-resolution input, which however, is not always available in many scenarios such as video surveillance and sports broadcasting. Two common approaches to deal with low-resolution images are applying super-resolution techniques to the input, which may result in unpleasant artifacts, or simply training one model for each resolution, which is impractical in many realistic applications. To address the above issues, this paper proposes a novel algorithm called RSC-Net, which consists of a Resolution-aware network, a Self-supervision loss, and a Contrastive learning scheme. The proposed method is able to learn 3D body pose and shape across different resolutions with one single model. The self-supervision loss enforces scale-consistency of the output, and the contrastive learning scheme enforces scale-consistency of the deep features. We show that both these new losses provide robustness when learning in a weakly-supervised manner. Moreover, we extend the RSC-Net to handle low-resolution videos and apply it to reconstruct textured 3D pedestrians from low-resolution input. Extensive experiments demonstrate that the RSC-Net can achieve consistently better results than the state-of-the-art methods for challenging low-resolution images. Xiangyu Xu 0002, Hao Chen 0102, Francesc Moreno-Noguer, László A. Jeni, Fernando De la Torre |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | High Fidelity 3D Reconstructions with Limited Physical ViewsabstractMulti-view triangulation is the gold standard for 3D reconstruction from 2D correspondences given known calibration and sufficient views. However in practice, expensive multi-view setups – involving tens sometimes hundreds of cameras – are required in order to obtain the high fidelity 3D reconstructions necessary for many modern applications. In this paper we present a novel approach that leverages recent advances in 2D-3D lifting using neural shape priors while also enforcing multi-view equivariance. We show how our method can achieve comparable fidelity to expensive calibrated multi-view rigs using a limited (2-3) number of uncalibrated camera views. Mosam Dabhi, Chaoyang Wang 0001, Kunal Saluja, László A. Jeni, Ian R. Fasel, Simon Lucey |
3DV | 4 |
| 2021 | Facial Action Units and Head Dynamics in Longitudinal Interviews Reveal OCD and Depression severity and DBS EnergyabstractN euromodulation therapy, specifically Deep Brain Stimulation (DBS) of the ventral capsule/ventral striatum (VC/VS), is promising treatment for severe and intractable obsessive-compulsive disorder (OCD). To assess treatment response to DBS, reliable biomarkers are needed. We explored the hypothesis that facial action units and head dynamics in an interview context reveal severity of OCD, related depression, and DBS energy in participants undergoing DBS treatment. Participants were 5 patients (3 females, 2 males) with implanted DBS to VC/VS. They were recorded during brief open-ended interviews by a clinician at pre- and post-surgery baselines and then at 3-month intervals following activation of the DBS electrodes. Facial action units and head dynamics were assessed using AFAR (Automatic Facial Affect Recognition). OCD severity was assessed using clinical interview (YBOCS-II) and depression symptoms were assessed using participant self-report (BDI). After testing for multicollinearity and dropping highly-correlated features, a linear mixed-effects model using chi-square feature selection predicted 61% of the variation in YBOCS-II; 59% of the variation in BDI; and 37% of the variation in delivered energy by DBS to VC/VS. These findings suggest that automatically detected facial action units and head dynamics are potential biomarkers of OCD, depression se verity, and DBS energy. Ali Darzi, Nicole R. Provenza, László A. Jeni, David A. Borton, Sameer A. Sheth, Wayne K. Goodman, Jeffrey F. Cohn |
FG | 3 |
| 2021 | Deep Implicit Surface Point Prediction NetworksabstractDeep neural representations of 3D shapes as implicit functions have been shown to produce high fidelity models surpassing the resolution-memory trade-off faced by the explicit representations using meshes and point clouds. However, most such approaches focus on representing closed shapes. Unsigned distance function (UDF) based approaches have been proposed recently as a promising alternative to represent both open and closed shapes. However, since the gradients of UDFs vanish on the surface, it is challenging to estimate local (differential) geometric properties like the normals and tangent planes which are needed for many downstream applications in vision and graphics. There are additional challenges in computing these properties efficiently with a low-memory footprint. This paper presents a novel approach that models such surfaces using a new class of implicit representations called the closest surface-point (CSP) representation. We show that CSP allows us to represent complex surfaces of any topology (open or closed) with high fidelity. It also allows for accurate and efficient computation of local geometric properties. We further demonstrate that it leads to efficient implementation of downstream algorithms like sphere-tracing for rendering the 3D surface as well as to create explicit mesh-based representations. Extensive experimental evaluation on the ShapeNet dataset validate the above contributions with results surpassing the state-of-the-art. Code and data are available at https://sites.google.com/view/cspnet. Rahul Venkatesh, Tejan Karmali, Sarthak Sharma, Aurobrata Ghosh, Venkatesh Babu Radhakrishnan, László A. Jeni, Maneesh Kumar Singh 0001 |
ICCV | 6 |
| 2021 | Synthetic Expressions are Better Than Real for Learning to Detect Facial ActionsabstractCritical obstacles in training classifiers to detect facial actions are the limited sizes of annotated video databases and the relatively low frequencies of occurrence of many actions. To address these problems, we propose an approach that makes use of facial expression generation. Our approach reconstructs the 3D shape of the face from each video frame, aligns the 3D mesh to a canonical view, and then trains a GAN-based network to synthesize novel images with facial action units of interest. To evaluate this approach, a deep neural network was trained on two separate datasets: One network was trained on video of synthesized facial expressions generated from FERA17; the other network was trained on unaltered video from the same database. Both networks used the same train and validation partitions and were tested on the test partition of actual video from FERA17. The network trained on synthesized facial expressions outperformed the one trained on actual facial expressions and surpassed current state-of-the-art approaches. Koichiro Niinuma, Itir Önal, Jeffrey F. Cohn, László A. Jeni |
WACV | 4 |
| 2020 | 3D Human Shape and Pose from a Single Low-Resolution Image with Self-Supervised Learning
Xiangyu Xu 0002, Hao Chen 0102, Francesc Moreno-Noguer, László A. Jeni, Fernando De la Torre |
ECCV (9) | 4 |
| 2019 | PAttNet: Patch-attentive deep network for action unit detection
Itir Önal, László A. Jeni, Jeffrey F. Cohn |
BMVC | 2 |
| 2019 | Unmasking the Devil in the Details: What Works for Deep Facial Action Coding?
Koichiro Niinuma, László A. Jeni, Itir Önal, Jeffrey F. Cohn |
BMVC | 2 |
| 2019 | Cross-domain AU Detection: Domains, Learning Approaches, and MeasuresabstractFacial action unit (AU) detectors have performed well when trained and tested within the same domain. Do AU detectors transfer to new domains in which they have not been trained? To answer this question, we review literature on cross-domain transfer and conduct experiments to address limitations of prior research. We evaluate both deep and shallow approaches to AU detection (CNN and SVM, respectively) in two large, well-annotated, publicly available databases, Expanded BP4D+ and GFT. The databases differ in observational scenarios, participant characteristics, range of head pose, video resolution, and AU base rates. For both approaches and databases, performance decreased with change in domain, often to below the threshold needed for behavioral research. Decreases were not uniform, however. They were more pronounced for GFT than for Expanded BP4D+ and for shallow relative to deep learning. These findings suggest that more varied domains and deep learning approaches may be better suited for promoting generalizability. Until further improvement is realized, caution is warranted when applying AU classifiers from one domain to another. Itir Önal, Jeffrey F. Cohn, László A. Jeni, Zheng Zhang 0023, Lijun Yin 0001 |
FG | 3 |
| 2019 | AFAR: A Deep Learning Based Tool for Automated Facial Affect RecognitionabstractAutomated facial affect recognition is crucial to multiple domains (e.g., health, education, entertainment). Commercial tools are available but costly and of unknown validity. Open-source ones [1] lack user-friendly GUI for use by non-programmers. For both types, evidence of domain transfer and options for retraining for use in new domains typically are lacking. Itir Önal, László A. Jeni, Wanqiao Ding, Jeffrey F. Cohn |
FG | 2 |
| 2018 | Brute-Force Facial Landmark Analysis With a 140, 000-Way ClassifierabstractWe propose a simple approach to visual alignment, focusing on the illustrative task of facial landmark estimation. While most prior work treats this as a regression problem, we instead formulate it as a discrete K-way classification task, where a classifier is trained to return one of K discrete alignments. One crucial benefit of a classifier is the ability to report back a (softmax) distribution over putative alignments. We demonstrate that this distribution is a rich representation that can be marginalized (to generate uncertainty estimates over groups of landmarks) and conditioned on (to incorporate top-down context, provided by temporal constraints in a video stream or an interactive human user). Such capabilities are difficult to integrate into classic regression-based approaches. We study performance as a function of the number of classes K, including the extreme "exemplar class" setting where K is equal to the number of training examples (140K in our setting). Perhaps surprisingly, we show that classifiers can still be learned in this setting. When compared to prior work in classification, our K is unprecedentedly large, including many "fine-grained" classes that are very similar. We address these issues by using a multi-label loss function that allows for training examples to be non-uniformly shared across discrete classes. We perform a comprehensive experimental analysis of our method on standard benchmarks, demonstrating state-of-the-art results for facial alignment in videos. László A. Jeni, Deva Ramanan |
AAAI | 2 |
| 2018 | Automated Affect Detection in Deep Brain Stimulation for Obsessive-Compulsive Disorder: A Pilot StudyabstractAutomated measurement of affective behavior in psychopathology has been limited primarily to screening and diagnosis. While useful, clinicians more often are concerned with whether patients are improving in response to treatment. Are symptoms abating, is affect becoming more positive, are unanticipated side effects emerging? When treatment includes neural implants, need for objective, repeatable biometrics tied to neurophysiology becomes especially pressing. We used automated face analysis to assess treatment response to deep brain stimulation (DBS) in two patients with intractable obsessive-compulsive disorder (OCD). One was assessed intraoperatively following implantation and activation of the DBS device. The other was assessed three months post-implantation. Both were assessed during DBS on and o conditions. Positive and negative valence were quantified using a CNN trained on normative data of 160 non-OCD participants. Thus, a secondary goal was domain transfer of the classifiers. In both contexts, DBS-on resulted in marked positive affect. In response to DBS-off, affect flattened in both contexts and alternated with increased negative affect in the outpatient setting. Mean AUC for domain transfer was 0.87. These findings suggest that parametric variation of DBS is strongly related to affective behavior and may introduce vulnerability for negative affect in the event that DBS is discontinued. Jeffrey F. Cohn, László A. Jeni, Itir Önal, Donald Malone, Michael S. Okun, David A. Borton, Wayne K. Goodman |
ICMI | 2 |
| 2018 | Recognizing Visual Signatures of Spontaneous Head GesturesabstractHead movements are an integral part of human nonverbal communication. As such, the ability to detect various types of head gestures from video is important for robotic systems that need to interact with people or for assistive technologies that may need to detect conversational gestures to aid communication. To this end, we propose a novel Multi-Scale Deep Convolution-LSTM architecture, capable of recognizing short and long term motion patterns found in head gestures, from video data of natural and unconstrained conversations. In particular, our models use Convolutional Neural Networks (CNNs) to learn meaningful representations from short time windows over head motion data. To capture longer term dependencies, we use Recurrent Neural Networks (RNNs) that extract temporal patterns across the output of the CNNs. We compare against classical approaches using discriminative and generative graphical models and show that our model is able to significantly outperform baseline models. Mohit Sharma 0001, Dragan Ahmetovic, László A. Jeni, Kris Makoto Kitani |
WACV | 3 |
| 2018 | Modeling and synthesis of kinship patterns of facial expressions
Itir Önal, László A. Jeni, Hamdi Dibeklioglu |
Image Vis. Comput. | 2 |
| 2018 | Viewpoint-Consistent 3D Face AlignmentabstractMost approaches to face alignment treat the face as a 2D object, which fails to represent depth variation and is vulnerable to loss of shape consistency when the face rotates along a 3D axis. Because faces commonly rotate three dimensionally, 2D approaches are vulnerable to significant error. 3D morphable models, employed as a second step in 2D+3D approaches are robust to face rotation but are computationally too expensive for many applications, yet their ability to maintain viewpoint consistency is unknown. We present an alternative approach that estimates 3D face landmarks in a single face image. The method uses a regression forest-based algorithm that adds a third dimension to the common cascade pipeline. 3D face landmarks are estimated directly, which avoids fitting a 3D morphable model. The proposed method achieves viewpoint consistency in a computationally efficient manner that is robust to 3D face rotation. To train and test our approach, we introduce the Multi-PIE Viewpoint Consistent database. In empirical tests, the proposed method achieved simple yet effective head pose estimation and viewpoint consistency on multiple measures relative to alternative approaches. Sergey Tulyakov, László A. Jeni, Jeffrey F. Cohn, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Sayette Group Formation Task (GFT) Spontaneous Facial Expression DatabaseabstractDespite the important role that facial expressions play in interpersonal communication and our knowledge that interpersonal behavior is influenced by social context, no currently available facial expression database includes multiple interacting participants. The Sayette Group Formation Task (GFT) database addresses the need for well-annotated video of multiple participants during unscripted interactions. The database includes 172,800 video frames from 96 participants in 32 three-person groups. To aid in the development of automated facial expression analysis systems, GFT includes expert annotations of FACS occurrence and intensity, facial landmark tracking, and baseline results for linear SVM, deep learning, active patch learning, and personalized classification. Baseline performance is quantified and compared using identical partitioning and a variety of metrics (including means and confidence intervals). The highest performance scores were found for the deep learning and active patch learning methods. Learn more at http://osf.io/7wcyz. Jeffrey M. Girard, Wen-Sheng Chu, László A. Jeni, Jeffrey F. Cohn |
FG | 3 |
| 2017 | FERA 2017 - Addressing Head Pose in the Third Facial Expression Recognition and Analysis ChallengeabstractThe field of Automatic Facial Expression Analysis has grown rapidly in recent years. However, despite progress in new approaches as well as benchmarking efforts, most evaluations still focus on either posed expressions, near-frontal recordings, or both. This makes it hard to tell how existing expression recognition approaches perform under conditions where faces appear in a wide range of poses (or camera views), displaying ecologically valid expressions. The main obstacle for assessing this is the availability of suitable data, and the challenge proposed here addresses this limitation. The FG 2017 Facial Expression Recognition and Analysis challenge (FERA 2017) extends FERA 2015 to the estimation of Action Units occurrence and intensity under different camera views. In this paper we present the third challenge in automatic recognition of facial expressions, to be held in conjunction with the 12th IEEE conference on Face and Gesture Recognition, May 2017, in Washington, United States. Two sub-challenges are defined: the detection of AU occurrence, and the estimation of AU intensity. In this work we outline the evaluation protocol, the data used, and the results of a baseline method for both sub-challenges. Michel F. Valstar, Enrique Sánchez-Lozano, Jeffrey F. Cohn, László A. Jeni, Jeffrey M. Girard, Zheng Zhang 0023, Lijun Yin 0001, Maja Pantic |
FG | 4 |
| 2017 | Dense 3D face alignment from 2D video for real-time use
László A. Jeni, Jeffrey F. Cohn, Takeo Kanade |
Image Vis. Comput. | 1 |
| 2016 | Continuous Supervised Descent Method for Facial Landmark Localisation
Marc Oliu, Ciprian A. Corneanu, László A. Jeni, Jeffrey F. Cohn, Takeo Kanade, Sergio Escalera |
ACCV (2) | 3 |
| 2014 | Spatio-temporal Event Classification Using Time-Series Kernel Based Structured Sparsity
László A. Jeni, András Lörincz, Zoltán Szabó 0001, Jeffrey F. Cohn, Takeo Kanade |
ECCV (4) | 1 |
| 2013 | Facing Imbalanced Data-Recommendations for the Use of Performance MetricsabstractRecognizing facial action units (AUs) is important for situation analysis and automated video annotation. Previous work has emphasized face tracking and registration and the choice of features classifiers. Relatively neglected is the effect of imbalanced data for action unit detection. While the machine learning community has become aware of the problem of skewed data for training classifiers, little attention has been paid to how skew may bias performance metrics. To address this question, we conducted experiments using both simulated classifiers and three major databases that differ in size, type of FACS coding, and degree of skew. We evaluated influence of skew on both threshold metrics (Accuracy, F-score, Cohen's kappa, and Krippendorf's alpha) and rank metrics (area under the receiver operating characteristic (ROC) curve and precision-recall curve). With exception of area under the ROC curve, all were attenuated by skewed distributions, in many cases, dramatically so. While ROC was unaffected by skew, precision-recall curves suggest that ROC may mask poor performance. Our findings suggest that skew is a critical factor in evaluating performance metrics. To avoid or minimize skew-biased estimates of performance, we recommend reporting skew-normalized scores along with the obtained ones. László A. Jeni, Jeffrey F. Cohn, Fernando De la Torre |
ACII | 1 |
| 2012 | 3D shape estimation in video sequences provides high precision evaluation of facial expressions
László A. Jeni, András Lörincz, Tamás Nagy, Zsolt Palotai, Judit Sebok, Zoltán Szabó 0001, Dániel Takács |
Image Vis. Comput. | 1 |