VLDB 2026 Research / reviewers in the wild / expert
Emre Aksan
dblp:152/9962
· DBLP profile ↗
16ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0002-9836-9011ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Physically Plausible Full-Body Hand-Object Interaction SynthesisabstractWe propose a physics-based method for synthesizing dexterous hand-object interactions in a full-body setting. While recent advancements have addressed specific facets of human-object interactions, a comprehensive physics-based approach remains a challenge. Existing methods often focus on isolated segments of the interaction process and rely on data-driven techniques that may result in artifacts. In contrast, our proposed method embraces reinforcement learning $(R L)$ and physics simulation to mitigate the limitations of data-driven approaches. Through a hierarchical framework, we first learn skill priors for both body and hand movements in a decoupled setting. The generic skill priors learn to decode a latent skill embedding into the motion of the underlying part. A high-level policy then controls hand-object interactions in these pretrained latent spaces, guided by task objectives of grasping and 3D target trajectory following. It is trained using a novel reward function that combines an adversarial style term with a task reward, encouraging natural motions while fulfilling the task incentives. Our method successfully accomplishes the complete interaction task, from approaching an object to grasping and subsequent manipulation. We compare our approach against kinematics-based baselines and show that it leads to more physically plausible motions. Video and code are available at https://eth-ait.github.io/phys-fullbody-grasp/. Jona Braun, Sammy Joe Christen, Muhammed Kocabas, Emre Aksan, Otmar Hilliges |
3DV | 4 |
| 2024 | Optimizing Diffusion Noise Can Serve As Universal Motion PriorsabstractWe propose Diffusion Noise Optimization (DNO), a new method that effectively leverages existing motion diffusion models as motion priors for a wide range of motion-related tasks. Instead of training a task-specific diffusion model for each new task, DNO operates by optimizing the diffusion latent noise of an existing pre-trained text-to-motion model. Given the corresponding latent noise of a human motion, it propagates the gradient from the target criteria defined on the motion space through the whole denoising process to update the diffusion latent noise. As a result, DNO supports any use cases where criteria can be defined as a function of motion. In particular, we show that, for motion editing and control, DNO outperforms existing meth-ods in both achieving the objective and preserving the motion content. DNO accommodates a diverse range of editing modes, including changing trajectory, pose, joint lo-cations, or avoiding newly added obstacles. In addition, DNO is effective in motion denoising and completion, pro-ducing smooth and realistic motion from noisy and partial inputs. DNO achieves these results at inference time with-out the need for model retraining, offering great versatility for any defined reward or loss function on the motion rep-resentation. Korrawe Karunratanakul, Konpat Preechakul, Emre Aksan, Thabo Beeler, Supasorn Suwajanakorn, Siyu Tang 0001 |
CVPR | 3 |
| 2022 | Reconstructing Action-Conditioned Human-Object Interactions Using Commonsense Knowledge PriorsabstractWe present a method for inferring diverse 3D models of human-object interactions from images. Reasoning about how humans interact with objects in complex scenes from a single 2D image is a challenging task given ambiguities arising from the loss of information through projection. In addition, modeling 3D interactions requires the generalization ability towards diverse object categories and interaction types. We propose an action-conditioned modeling of interactions that allows us to infer diverse 3D arrangements of humans and objects without supervision on contact regions or 3D scene geometry. Our method extracts high-level commonsense knowledge from large language models (such as GPT-3), and applies them to perform 3D reasoning of human-object interactions. Our key insight is priors extracted from large language models can help in reasoning about human-object contacts from textural prompts only. We quantitatively evaluate the inferred 3D models on a large human-object interaction dataset and show how our method leads to better 3D reconstructions. We further qualitatively evaluate the effectiveness of our method on real images and demonstrate its generalizability towards interaction types and object categories. Xi Wang 0021, Gen Li 0010, Yen-Ling Kuo, Muhammed Kocabas, Emre Aksan, Otmar Hilliges |
3DV | 5 |
| 2022 | D-Grasp: Physically Plausible Dynamic Grasp Synthesis for Hand-Object InteractionsabstractWe introduce the dynamic grasp synthesis task: given an object with a known 6D pose and a grasp reference, our goal is to generate motions that move the object to a target 6D pose. This is challenging, because it requires reasoning about the complex articulation of the human hand and the intricate physical interaction with the object. We propose a novel method that frames this problem in the reinforcement learning framework and leverages a physics simulation, both to learn and to evaluate such dynamic interactions. A hierarchical approach decomposes the task into low-level grasping and high-level motion synthesis. It can be used to generate novel hand sequences that approach, grasp, and move an object to a desired location, while retaining human-likeness. We show that our approach leads to stable grasps and generates a wide range of motions. Furthermore, even imperfect labels can be corrected by our method to generate dynamic interaction sequences. Video and code are available at: https://eth-ait.github.io/d-grasp/. Sammy Joe Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song 0006, Otmar Hilliges |
CVPR | 3 |
| 2022 | LiP-Flow: Learning Inference-Time Priors for Codec Avatars via Normalizing Flows in Latent Space
Emre Aksan, Shugao Ma, Akin Caliskan, Stanislav Pidhorskyi, Alexander Richard, Shih-En Wei, Jason M. Saragih, Otmar Hilliges |
ECCV (26) | 1 |
| 2021 | A Spatio-temporal Transformer for 3D Human Motion PredictionabstractWe propose a novel Transformer-based architecture for the task of generative modelling of 3D human motion. Previous work commonly relies on RNN-based models considering shorter forecast horizons reaching a stationary and often implausible state quickly. Recent studies show that implicit temporal representations in the frequency domain are also effective in making predictions for a predetermined horizon. Our focus lies on learning spatio-temporal representations autoregressively and hence generation of plausible future developments over both short and long term. The proposed model learns high dimensional embeddings for skeletal joints and how to compose a temporally coherent pose via a decoupled temporal and spatial self-attention mechanism. Our dual attention concept allows the model to access current and past information directly and to capture both the structural and the temporal dependencies explicitly. We show empirically that this effectively learns the underlying motion dynamics and reduces error accumulation over time observed in auto-regressive models. Our model is able to make accurate short-term predictions and generate plausible motion sequences over long horizons. We make our code publicly available at https://github.com/eth-ait/motion-transformer. Emre Aksan, Manuel Kaufmann, Otmar Hilliges |
3DV | 1 |
| 2020 | Convolutional Autoencoders for Human Motion InfillingabstractIn this paper we propose a convolutional autoencoder to address the problem of motion infilling for 3D human motion data. Given a start and end sequence, motion infilling aims to complete the missing gap in between, such that the filled in poses plausibly forecast the start sequence and naturally transition into the end sequence. To this end, we propose a single, end-to-end trainable convolutional autoencoder. We show that a single model can be used to create natural transitions between different types of activities. Furthermore, our method is not only able to fill in entire missing frames, but it can also be used to complete gaps where partial poses are available (e.g. from end effectors), or to clean up other forms of noise (e.g. Gaussian). Also, the model can fill in an arbitrary number of gaps that potentially vary in length. In addition, no further post-processing on the model's outputs is necessary such as smoothing or closing discontinuities at the end of the gap. At the heart of our approach lies the idea to cast motion infilling as an inpainting problem and to train a convolutional de-noising autoencoder on image-like representations of motion sequences. At training time, blocks of columns are removed from such images and we ask the model to fill in the gaps. We demonstrate the versatility of the approach via a number of complex motion sequences and report on thorough evaluations performed to better understand the capabilities and limitations of the proposed approach. Manuel Kaufmann, Emre Aksan, Jie Song 0006, Fabrizio Pece, Remo Ziegler, Otmar Hilliges |
3DV | 2 |
| 2020 | Towards End-to-End Video-Based Eye-Tracking
Seonwook Park, Emre Aksan, Xucong Zhang, Otmar Hilliges |
ECCV (12) | 2 |
| 2020 | CoSE: Compositional Stroke EmbeddingsabstractWe present a generative model for stroke-based drawing tasks which is able to model complex free-form structures. While previous approaches rely on sequence-based models for drawings of basic objects or handwritten text, we propose a model that treats drawings as a collection of strokes that can be composed into complex structures such as diagrams (e.g., flow-charts). At the core of the approach lies a novel auto-encoder that projects variable-length strokes into a latent space of fixed dimension. This representation space allows a relational model, operating in latent space, to better capture the relationship between strokes and to predict subsequent strokes. We demonstrate qualitatively and quantitatively that our proposed approach is able to model the appearance of individual strokes, as well as the compositional structure of larger diagram drawings. Our approach is suitable for interactive use cases such as auto-completing diagrams. We make code and models publicly available at https://eth-ait.github.io/cose. Emre Aksan, Thomas Deselaers, Andrea Tagliasacchi, Otmar Hilliges |
NeurIPS | 1 |
| 2019 | Structured Prediction Helps 3D Human Motion ModellingabstractHuman motion prediction is a challenging and important task in many computer vision application domains. Existing work only implicitly models the spatial structure of the human skeleton. In this paper, we propose a novel approach that decomposes the prediction into individual joints by means of a structured prediction layer that explicitly models the joint dependencies. This is implemented via a hierarchy of small-sized neural networks connected analogously to the kinematic chains in the human body as well as a joint-wise decomposition in the loss function. The proposed layer is agnostic to the underlying network and can be used with existing architectures for motion modelling. Prior work typically leverages the H3.6M dataset. We show that some state-of-the-art techniques do not perform well when trained and tested on AMASS, a recently released dataset 14 times the size of H3.6M. Our experiments indicate that the proposed layer increases the performance of motion forecasting irrespective of the base network, joint-angle representation, and prediction horizon. We furthermore show that the layer also improves motion predictions qualitatively. We make code and models publicly available at https://ait.ethz.ch/projects/2019/spl. Emre Aksan, Manuel Kaufmann, Otmar Hilliges |
ICCV | 1 |
| 2019 | STCN: Stochastic Temporal Convolutional Networks
Emre Aksan, Otmar Hilliges |
ICLR (Poster) | 1 |
| 2018 | DeepWriting: Making Digital Ink Editable via Deep Generative ModelingabstractDigital ink promises to combine the flexibility and aesthetics of handwriting and the ability to process, search and edit digital text. Character recognition converts handwritten text into a digital representation, albeit at the cost of losing personalized appearance due to the technical difficulties of separating the interwoven components of content and style. In this paper, we propose a novel generative neural network architecture that is capable of disentangling style from content and thus making digital ink editable. Our model can synthesize arbitrary text, while giving users control over the visual appearance (style). For example, allowing for style transfer without changing the content, editing of digital ink at the word level and other application scenarios such as spell-checking and correction of handwritten text. We furthermore contribute a new dataset of handwritten text with fine-grained annotations at the character level and report results from an initial user evaluation. Emre Aksan, Fabrizio Pece, Otmar Hilliges |
CHI | 1 |
| 2018 | Deep inertial poser: learning to reconstruct human pose from sparse inertial measurements in real timeabstractWe demonstrate a novel deep neural network capable of reconstructing human full body pose in real-time from 6 Inertial Measurement Units (IMUs) worn on the user's body. In doing so, we address several difficult challenges. First, the problem is severely under-constrained as multiple pose parameters produce the same IMU orientations. Second, capturing IMU data in conjunction with ground-truth poses is expensive and difficult to do in many target application scenarios (e.g., outdoors). Third, modeling temporal dependencies through non-linear optimization has proven effective in prior work but makes real-time prediction infeasible. To address this important limitation, we learn the temporal pose priors using deep learning. To learn from sufficient data, we synthesize IMU data from motion capture datasets. A bi-directional RNN architecture leverages past and future information that is available at training time. At test time, we deploy the network in a sliding window fashion, retaining real time capabilities. To evaluate our method, we recorded DIP-IMU, a dataset consisting of 10 subjects wearing 17 IMUs for validation in 64 sequences with 330 000 time instants; this constitutes the largest IMU dataset publicly available. We quantitatively evaluate our approach on multiple datasets and show results from a real-time implementation. DIP-IMU and the code are available for research purposes. 1 Yinghao Huang, Manuel Kaufmann, Emre Aksan, Michael J. Black, Otmar Hilliges, Gerard Pons-Moll |
ACM Trans. Graph. | 3 |
| 2017 | Learning Human Motion Models for Long-Term PredictionsabstractWe propose a new architecture for the learning of predictive spatio-temporal motion models from data alone. Our approach, dubbed the Dropout Autoencoder LSTM (DAELSTM), is capable of synthesizing natural looking motion sequences over long-time horizons1 without catastrophic drift or motion degradation. The model consists of two components, a 3-layer recurrent neural network to model temporal aspects and a novel autoencoder that is trained to implicitly recover the spatial structure of the human skeleton via randomly removing information about joints during training. This Dropout Autoencoder (DAE) is then used to filter each predicted pose by a 3-layer LSTM network, reducing accumulation of correlated error and hence drift over time. Furthermore to alleviate insufficiency of commonly used quality metric, we propose a new evaluation protocol using action classifiers to assess the quality of synthetic motion sequences. The proposed protocol can be used to assess quality of generated sequences of arbitrary length. Finally, we evaluate our proposed method on two of the largest motion-capture datasets available and show that our model outperforms the state-of-the-art techniques on a variety of actions, including cyclic and acyclic motion, and that it can produce natural looking sequences over longer time horizons than previous methods. Partha Ghosh, Jie Song 0006, Emre Aksan, Otmar Hilliges |
3DV | 3 |
| 2017 | Guiding InfoGAN with Semi-supervision
Adrian Spurr, Emre Aksan, Otmar Hilliges |
ECML/PKDD (1) | 2 |
| 2014 | Modeling the Brain Connectivity for Pattern AnalysisabstractAn information theoretic approach is proposed to estimate the degree of connectivity for each voxel with its neighboring voxels. The neighborhood system is defined by spatial and functional connectivity metrics. Then, a local mesh of variable size is formed around each voxel using spatial or functional neighborhood. The mesh arc weights, called Mesh Arc Descriptors (MAD), are estimated by a linear regression model fitted to the voxel intensity values of the functional Magnetic Resonance Images (fMRI). Finally, the error term of the linear regression equation is used to estimate the mesh size for a voxel by optimizing Akaike's information Criterion, Bayesian Information Criterion and Rissanen's Minimum Description Length. fMRI measurements are obtained during a memory encoding and retrieval experiment performed on a subject who is exposed to the stimuli from 10 semantic categories. For each sample, a k-NN classifier is trained using the Mesh Arc Descriptors (MAD) having the variable mesh sizes. The classification performances reflect that the suggested variable-size Mesh Arc Descriptors represents the mental states better than the classical multi-voxel pattern representation. Moreover, we observe that the degree of connectivities in the brain greatly varies for each voxel. Itir Önal, Emre Aksan, Burak Velioglu, Orhan Firat, Mete Ozay, Ilke Öztekin, Fatos T. Yarman-Vural |
ICPR | 2 |