EDBT 2026 Demo / reviewers in the wild / expert
Omid Taheri
dblp:34/4091
· DBLP profile ↗
18ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Predicting 4D Hand Trajectory from Monocular VideosabstractWe present HaPTIC, an approach that infers coherent 4D hand trajectories from monocular videos. Current videobased hand pose reconstruction methods primarily focus on improving frame-wise 3D pose using adjacent frames rather than studying consistent 4D hand trajectories in space. Despite the additional temporal cues, they generally underperform compared to image-based methods due to the scarcity of annotated video data. To address these issues, we repurpose a state-of-the-art image-based transformer to take in multiple frames and directly predict a coherent trajectory. We introduce two types of lightweight attention layers: cross-view self-attention to fuse temporal information, and global cross-attention to bring in larger spatial context. Our method infers 4D hand trajectories similar to the ground truth while maintaining strong 2 D reprojection alignment. We apply the method to both egocentric and allocentric videos. It significantly outperforms existing methods in global trajectory accuracy while being comparable to the state-of-the-art in single-image pose estimation. Yufei Ye 0001, Yao Feng 0001, Omid Taheri, Haiwen Feng, Shubham Tulsiani, Michael J. Black |
3DV | 3 |
| 2025 | 3D Whole-Body Grasp Synthesis with Directional ControllabilityabstractSynthesizing 3D whole bodies that realistically grasp objects is useful for animation, mixed reality, and robotics. This is challenging, because the hands and body need to look natural w.r.t. each other, the grasped object, as well as the local scene (i.e., a receptacle supporting the object). Moreover, training data for this task is really scarce, while capturing new data is expensive. Recent work goes beyond finite datasets via a divide-and-conquer approach; it first generates a “guiding” right-hand grasp, and then searches for bodies that match this. However, the guiding-hand synthesis lacks controllability and receptacle awareness, so it likely has an implausible direction (i.e., a body can't match this without penetrating the receptacle) and needs corrections through major post-processing. Moreover, the body search needs exhaustive sampling and is expensive. These are strong limitations. We tackle these with a novel method called CWGrasp. Our key idea is that performing geometry-based reasoning “early on,” instead of “too late,” provides rich “control” signals for inference. To this end, CWGrasp first samples a plausible reaching-direction vector (used later for both the arm and hand) from a probabilistic model built via ray-casting from the object and collision checking. Then, it generates a reaching body with a desired arm direction, as well as a “guiding” grasping hand with a desired palm direction that complies with the arm's one. Eventually, CWGrasp refines the body to match the “guiding” hand, while plausibly contacting the scene. Notably, generating already-compatible “parts” greatly simplifies the “whole”. Moreover, CWGrasp uniquely tackles both right- and left-hand grasps. We evaluate on the GRAB and ReplicaGrasp datasets. CWGrasp outperforms baselines, at lower runtime and budget, while all components help performance. Code and models are available at https://gpaschalidis.github.io/cwgrasp. Georgios Paschalidis, Romana Wilschut, Dimitrije Antic, Omid Taheri, Dimitrios Tzionas |
3DV | 4 |
| 2025 | InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsabstractWe introduce InteractVLM, a novel method to estimate 3D contact points on human bodies and objects from single in- the-wild images, enabling accurate human-object joint reconstruction in 3D. This is challenging due to occlusions, depth ambiguities, and widely varying object shapes. Existing methods rely on 3D contact annotations collected via expensive motion-capture systems or tedious manual labeling, limiting scalability and generalization. To overcome this, InteractVLM harnesses the broad visual knowledge of large Vision-Language Models (VLMs), fine-tuned with limited 3D contact data. However, directly applying these models is non-trivial, as they reason "only" in 2D, while human-object contact is inherently 3D. Thus we introduce a novel "Render-Localize-Lift" module that: (1) embeds 3D body and object surfaces in 2D space via multiview rendering, (2) trains a novel multi-view localization model (MV-Loc) to infer contacts in 2D, and (3) lifts these to 3D. Additionally, we propose a new task called Semantic Human Contact estimation, where human contact predictions are conditioned explicitly on object semantics, enabling richer interaction modeling. InteractVLM outperforms existing work on contact estimation and also facilitates 3D reconstruction from an in-the-wild image. To estimate 3D human and object pose, we infer initial body and object meshes, then infer contacts on both of these via InteractVLM, and last exploit these for fitting the meshes to image evidence. Results show that our approach performs promisingly in the wild. Code and models are available at https://interactvlm.is.tue.mpg.de. Saikumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri, Cordelia Schmid, Michael J. Black, Dimitrios Tzionas |
CVPR | 4 |
| 2025 | A Versatile and Differentiable Hand-Object Interaction RepresentationabstractSynthesizing accurate hands-object interactions (HOI) is critical for applications in Computer Vision, Augmented Reality (AR), and Mixed Reality (MR). Despite recent ad-vances, the accuracy of reconstructed or generated HOI leaves room for refinement. Some techniques have improved the accuracy of dense correspondences by shifting focus from generating explicit contacts to using rich HOI fields. Still, they lack full differentiability or continuity and are tai-lored to specific tasks. In contrast, we present a Coarse Hand-Object Interaction Representation (CHOIR), a novel, versatile and fully differentiable field for HOI modelling. CHOIR leverages discrete unsigned distances for continu-ous shape and pose encoding, alongside multivariate Gaus-sian distributions to represent dense contact maps with few parameters. To demonstrate the versatility of CHOIR we design JointDiffus ion, a diffusion model to learn a grasp distribution conditioned on noisy hand-object interactions or only object geometries, for both refinement and synthe-sis applications. We demonstrate JointDiffus ion, s improve-ments over the SOTA in both applications: it increases the contact F1 score by 5% for refinement and decreases the sim. displacement by 46% for synthesis. Our exper-iments show that JointDiffusion with CHOIR yield supe-rior contact accuracy and physical realism compared to SOTA methods designed for specific tasks. Project page: https://theomorales.com/CHOIR Théo Morales, Omid Taheri, Gerard Lacey |
WACV | 2 |
| 2024 | GRIP: Generating Interaction Poses Using Spatial Cues and Latent ConsistencyabstractHands are dexterous and highly versatile manipulators that are central to how humans interact with objects and their environment. Consequently, modeling realistic hand-object interactions, including the subtle motion of individual fingers, is critical for applications in computer graphics, computer vision, and mixed reality. Prior work on capturing and modeling humans interacting with objects in 3D focuses on the body and object motion, often ignoring hand pose. In contrast, we introduce GRIP, a learning-based method that takes, as input, the 3D motion of the body and the object, and synthesizes realistic motion for both hands before, during, and after object interaction. As a preliminary step before synthesizing the hand motion, we first use a network, ANet, to denoise the arm motion. Then, we leverage the spatio-temporal relationship between the body and the object to extract novel temporal interaction cues, and use them in a two-stage inference pipeline to generate the hand motion. In the first stage, we introduce a new approach to encourage motion temporal consistency in the latent space (LTC) and generate consistent interaction motions. In the second stage, GRIP generates refined hand poses to avoid hand-object penetrations. Given sequences of noisy body and object motion, GRIP “upgrades” them to include hand-object interaction. Quantitative experiments and perceptual studies demonstrate that GRIP outperforms baseline methods and generalizes to unseen objects and motions from different motion-capture datasets. Our models and code are available for research purposes at https://grip.is.tue.mpg.de. Omid Taheri, Yi Zhou 0023, Dimitrios Tzionas, Yang Zhou 0009, Duygu Ceylan, Sören Pirk, Michael J. Black |
3DV | 1 |
| 2024 | WANDR: Intention-guided Human Motion GenerationabstractSynthesizing natural human motions that enable a 3D human avatar to walk and reach for arbitrary goals in 3D space remains an unsolved problem with many applications. Existing methods (data-driven or using reinforcement learning) are limited in terms of generalization and motion naturalness. A primary obstacle is the scarcity of training data that combines locomotion with goal reaching. To ad-dress this, we introduce WANDR, a data-driven model that takes an avatar's initial pose and a goal's 3D position and generates natural human motions that place the end effec-tor (wrist) on the goal location. To solve this, we intro-duce novel intention features that drive rich goal-oriented movement. Intention guides the agent to the goal, and in-teractively adapts the generation to novel situations without needing to define sub-goals or the entire motion path. Cru-cially, intention allows training on datasets that have goal-oriented motions as well as those that do not. WANDR is a conditional Variational Auto-Encoder (c- VAE), which we train using the AMASS and CIRCLE datasets. We evaluate our method extensively and demonstrate its ability to gener-ate natural and long-term motions that reach 3D goals and generalize to unseen goal locations. Our models and code are available for research purposes at wandr.is.tue.mpg.de. Markos Diomataris, Nikos Athanasiou, Omid Taheri, Xi Wang 0021, Otmar Hilliges, Michael J. Black |
CVPR | 3 |
| 2024 | HUMOS: Human Motion Model Conditioned on Body Shape
Shashank Tripathi, Omid Taheri, Christoph Lassner, Michael J. Black, Daniel Holden, Carsten Stoll |
ECCV (16) | 2 |
| 2024 | InterCap: Joint Markerless 3D Tracking of Humans and Objects in Interaction from Multi-view RGB-D ImagesabstractAbstract Humans constantly interact with objects to accomplish tasks. To understand such interactions, computers need to reconstruct these in 3D from images of whole bodies manipulating objects, e.g., for grasping, moving and using the latter. This involves key challenges, such as occlusion between the body and objects, motion blur, depth ambiguities, and the low image resolution of hands and graspable object parts. To make the problem tractable, the community has followed a divide-and-conquer approach, focusing either only on interacting hands, ignoring the body, or on interacting bodies, ignoring the hands. However, these are only parts of the problem. On the contrary, recent work focuses on the whole problem. The GRAB dataset addresses whole-body interaction with dexterous hands but captures motion via markers and lacks video, while the BEHAVE dataset captures video of body-object interaction but lacks hand detail. We address the limitations of prior work with InterCap, a novel method that reconstructs interacting whole-bodies and objects from multi-view RGB-D data, using the parametric whole-body SMPL-X model and known object meshes. To tackle the above challenges, InterCap uses two key observations: (i) Contact between the body and object can be used to improve the pose estimation of both. (ii) Consumer-level Azure Kinect cameras let us set up a simple and flexible multi-view RGB-D system for reducing occlusions, with spatially calibrated and temporally synchronized cameras. With our InterCap method we capture the InterCap dataset, which contains 10 subjects (5 males and 5 females) interacting with 10 daily objects of various sizes and affordances, including contact with the hands or feet. To this end, we introduce a new data-driven hand motion prior, as well as explore simple ways for automatic contact detection based on 2D and 3D cues. In total, InterCap has 223 RGB-D videos, resulting in 67,357 multi-view frames, each containing 6 RGB-D images, paired with pseudo ground-truth 3D body and object meshes. Our InterCap method and dataset fill an important gap in the literature and support many research directions. Data and code are available at https://intercap.is.tue.mpg.de . Yinghao Huang, Omid Taheri, Michael J. Black, Dimitrios Tzionas |
Int. J. Comput. Vis. | 2 |
| 2023 | ARCTIC: A Dataset for Dexterous Bimanual Hand-Object ManipulationabstractHumans intuitively understand that inanimate objects do not move by themselves, but that state changes are typically caused by human manipulation (e.g., the opening of a book). This is not yet the case for machines. In part this is because there exist no datasets with ground-truth 3D annotations for the study of physically consistent and synchronised motion of hands and articulated objects. To this end, we introduce ARCTIC - a dataset of two hands that dexterously manipulate objects, containing 2 .1M video frames paired with accurate 3D hand and object meshes and detailed, dynamic contact information. It contains bi-manual articulation of objects such as scissors or laptops, where hand poses and object states evolve jointly in time. We propose two novel articulated hand-object interaction tasks: (1) Consistent motion reconstruction: Given a monocular video, the goal is to reconstruct two hands and articulated objects in 3D, so that their motions are spatio-temporally consistent. (2) Interaction field estimation: Dense relative hand-object distances must be estimated from images. We introduce two baselines ArcticNet and InterField, respectively and evaluate them qualitatively and quantitatively on ARCTIC. Our code and data are available at https://arctic.is.tue.mpg.de. Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, Otmar Hilliges |
CVPR | 2 |
| 2023 | 3D Human Pose Estimation via Intuitive PhysicsabstractEstimating 3D humans from images often produces implausible bodies that lean, float, or penetrate the floor. Such methods ignore the fact that bodies are typically supported by the scene. A physics engine can be used to enforce physical plausibility, but these are not differentiable, rely on unrealistic proxy bodies, and are difficult to integrate into existing optimization and learning frameworks. In contrast, we exploit novel intuitive-physics (IP) terms that can be inferred from a 3D SMPL body interacting with the scene. Inspired by biomechanics, we infer the pressure heatmap on the body, the Center of Pressure (CoP) from the heatmap, and the SMPL body's Center of Mass (CoM). With these, we develop IPMAN, to estimate a 3D body from a color image in a “stable” configuration by encouraging plausible floor contact and overlapping CoP and CoM. Our IP terms are intuitive, easy to implement, fast to compute, differentiable, and can be integrated into existing optimization and regression methods. We evaluate IPMAN on standard datasets and MoYo, a new dataset with synchronized multi-view images, ground-truth 3D bodies with complex poses, body-floor contact, CoM and pressure. IPMAN produces more plausible results than the state of the art, improving accuracy for static poses, while not hurting dynamic ones. Code and data are available for research at https://ipman.is.tue.mpg.de. Shashank Tripathi, Lea Müller, Chun-Hao P. Huang, Omid Taheri, Michael J. Black, Dimitrios Tzionas |
CVPR | 4 |
| 2022 | GOAL: Generating 4D Whole-Body Motion for Hand-Object GraspingabstractGenerating digital humans that move realistically has many applications and is widely studied, but existing meth-odsfocus on the major limbs of the body, ignoring the hands and head. Hands have been separately studied, but the fo-cus has been on generating realistic static grasps of objects. To synthesize virtual characters that interact with the world, we need to generate full-body motions and realistic hand grasps simultaneously. Both sub-problems are challenging on their own and, together, the state space of poses is sig-nificantly larger, the scales of hand and body motions dif-fer, and the whole-body posture and the hand grasp must agree, satisfy physical constraints, and be plausible. Additionally, the head is involved because the avatar must look at the object to interact with it. For the first time, we ad-dress the problem of generating full-body, hand and head motions of an avatar grasping an unknown object. As in-put, our method, called GOAL, takes a 3D object, its pose, and a starting 3D body pose and shape. GOAL outputs a sequence of whole-body poses using two novel networks. First, GNet generates a goal whole-body grasp with a re-alistic body, head, arm, and hand pose, as well as hand-object contact. Second, MNet generates the motion be-tween the starting and goal pose. This is challenging, as it requires the avatar to walk towards the object with foot-ground contact, orient the head towards it, reach out, and grasp it with a realistic hand pose and hand-object con-tact. To achieve this the networks exploit a representation that combines SMPL-X body parameters and 3D vertex off-sets. We train and evaluate GOAL, both qualitatively and quantitatively, on the GRAB dataset. Results show that GOAL generalizes well to unseen objects, outperforming baselines. A perceptual study shows that GOAL's gener-ated motions approach the realism of GRAB's ground truth. GOAL takes a step towards generating realistic full-body object grasping motion. Our models and code are available at https://goal.is.tue.mpg.de. Omid Taheri, Vasileios Choutas, Michael J. Black, Dimitrios Tzionas |
CVPR | 1 |
| 2020 | GRAB: A Dataset of Whole-Body Human Grasping of Objects
Omid Taheri, Nima Ghorbani, Michael J. Black, Dimitrios Tzionas |
ECCV (4) | 1 |
| 2014 | Reweighted l1-norm penalized LMS for sparse channel estimation and its analysis
Omid Taheri, Sergiy A. Vorobyov |
Signal Process. | 1 |
| 2012 | Decimated least mean squares for frequency sparse channel estimationabstractThe standard least mean squares (LMS) parameter estimation method does not assume any special structure for the parameters being estimated. However, when additional knowledge about the system is available, the performance of LMS can be improved by appropriate modification of the algorithm. We develop such modifications for the case of estimating frequency sparse channels. Such modifications provide either better performance or less complexity when compared to the standard LMS algorithm. Decimated LMS and zero attracting decimated LMS are the two methods proposed in this paper. Simulation results are also provided to compare the performance of the proposed algorithms to the standard LMS and other sparsity aware modifications of LMS. Omid Taheri, Sergiy A. Vorobyov |
ICASSP | 1 |
| 2011 | Sparse channel estimation with lp-norm and reweighted l1-norm penalized least mean squaresabstractThe least mean squares (LMS) algorithm is one of the most popular recursive parameter estimation methods. In its standard form it does not take into account any special characteristics that the parameterized model may have. Assuming that such model is sparse in some domain (for example, it has sparse impulse or frequency response), we aim at developing such LMS algorithms that can adapt to the underlying sparsity and achieve better parameter estimates. Particularly, the example of channel estimation with sparse channel impulse response is considered. The proposed modifications of LMS are the lp-norm and reweighted l1-norm penalized LMS algorithms. Our simulation results confirm the superiority of the proposed algorithms over the standard LMS as well as other sparsity-aware modifications of LMS available in the literature. Omid Taheri, Sergiy A. Vorobyov |
ICASSP | 1 |
| 2008 | Noniterative Joint Channel Equalization and Decoding Based on State Extended Viterbi AlgorithmabstractA new receiver for convolutional coded information, which performs both the equalization and decoding in a joint scheme, based on the Viterbi algorithm is proposed. This method is based on creating a trellis for the joint system assuming some virtual auxiliary registers in the structure of the coder. In this way, the state space of the convolutional coder extends to support different states of the channel, enabling the receiver to include different states of the channel directly during the process of recovering the source information. Simulations indicate that the capability of taking the channel states into account during the decoding increases the performance of the overall receiver in the case of channels with memory. Also, the proposed method has lower complexity than a separated equalizer and hard decoder structure. A similar receiver for a memoryless channel can outperform conventional receivers. Simulation results for channels with and without memory are discussed. Iraj Hosseini, Kaveh Mahdaviani, Omid Taheri, Norman C. Beaulieu |
ICC | 3 |
| 2007 | Hiding Data in Color Halftone Images using Dot Diffusion with Nonlinear ThresholdingabstractImage data hiding is the hiding of invisible patterns in an image without degrading its visual quality. This hidden pattern can be visualized by overlaying the original image with watermarked ones. Also it can be visualized using a simple XNOR operation. In this paper we propose a method called dot diffusion with nonlinear thresholding (DDNT) to embed hidden data in halftone images and its modification to color halftone images which will be called CDDNT. The main advantage of this method is that the hidden pattern's intensity can be adjusted by three parameters. Also it utilizes the inherent parallelism in dot diffusion halftoning method, although it suffers from its lower visual quality than some other halftoning methods such as error diffusion. Omid Taheri, Ahmad Movahedian Attar, Mohammad Mehdi DaneshPanah |
ICASSP (2) | 1 |
| 2007 | A Preprocessing Method for PAPR Reduction in OFDM Systems by Modifying FFT and IFFT MatricesabstractIn this paper we propose an algorithm for PAPR reduction in OFDM systems, based on a modification in the IFFT matrix at the transmitter and using the inverse matrix at the receiver. In this method two columns of the IFFT matrix are replaced with a linear combination of them in order to reduce the PAPR. These columns correspond to the maximum and minimum in the sequence of absolute values for each OFDM symbol. It will be shown that we may choose more than one pair in order to reduce the PAPR more effectively. The proposed method entails less complexity at the transmitter in comparison with other PAPR reduction algorithms. It also requires less increase in SNR for the same BER compared to other methods. A trade-off between complexity and system performance can set the number of pairs of maximum and minimum processed. The only drawback of this method is that the modified IFFT matrix is not the same for all OFDM symbols; therefore some extra information about the modified matrix must be sent. Keyvan Kasiri, Iraj Hosseini, Omid Taheri, M. J. Omidi, P. Glenn Gulak |
PIMRC | 3 |