EDBT 2026 Demo / reviewers in the wild / expert
Maria Parelli
dblp:282/2105
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
3D vision · 59% Face, body and person analysis · 41% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation |
1.6 | 2 | 2025 | MagicHOI: Leveraging 3D Priors for Accurate Hand-Object Reconstruction from Short Monocular Video Clips · ICCV 2025 HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from Video · CVPR 2024 |
Computer vision › 3D vision › 3d reconstruction › object reconstruction
hand-object reconstruction |
1.6 | 2 | 2025 | MagicHOI: Leveraging 3D Priors for Accurate Hand-Object Reconstruction from Short Monocular Video Clips · ICCV 2025 HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from Video · CVPR 2024 |
Computer vision › 3D vision › object pose estimation
hand-object pose estimation |
0.9 | 1 | 2025 | MagicHOI: Leveraging 3D Priors for Accurate Hand-Object Reconstruction from Short Monocular Video Clips · ICCV 2025 |
Computer vision › 3D vision
3d reconstruction |
0.8 | 1 | 2024 | HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from Video · CVPR 2024 |
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation |
0.8 | 1 | 2024 | HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from Video · CVPR 2024 |
Computer vision › 3D vision
hand-object interaction |
0.2 | 1 | 2024 | HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from Video · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
3d priors · 0.9hand-object constraints · 0.8compositional articulated implicit model · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MagicHOI: Leveraging 3D Priors for Accurate Hand-Object Reconstruction from Short Monocular Video Clips
Maria Parelli, Christoph Gebhardt, Zicong Fan, Jie Song 0006 |
ICCV | 3 |
| 2024 | HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from VideoabstractSince humans interact with diverse objects every day, the holistic 3D capture of these interactions is important to understand and model human behaviour. However, most ex-isting methods for hand-object reconstruction from RGB ei-ther assume pre-scanned object templates or heavily rely on limited 3D hand-object data, restricting their ability to scale and generalize to more unconstrained interaction settings. To address this, we introduce HOLD the first category-agnostic method that reconstructs an articulated hand and an object jointly from a monocular interaction video. We develop a compositional articulated implicit model that can reconstruct disentangled 3D hands and ob-jects from 2D images. We also further incorporate hand-object constraints to improve hand-object poses and con-sequently the reconstruction quality. Our method does not rely on any 3D hand-object annotations while significantly outperforming fully-supervised baselines in both in-the-lab and challenging in-the-wild settings. Moreover, we qualita-tively show its robustness in reconstructing from in-the-wild videos. See here for code, data, models, and updates. Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen 0025, Muhammed Kocabas, Michael J. Black, Otmar Hilliges |
CVPR | 2 |
| 2023 | Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes
Alexandros Delitzas, Maria Parelli, Nikolas Hars, Georgios Vlassis, Sotiris Anagnostidis, Gregor Bachmann, Thomas Hofmann 0001 |
BMVC | 2 |
| 2023 | Interpretable Visual Question Answering Via Reasoning SupervisionabstractTransformer-based architectures have recently demonstrated remarkable performance in the Visual Question Answering (VQA) task. However, such models are likely to disregard crucial visual cues and often rely on multimodal shortcuts and inherent biases of the language modality to predict the correct answer, a phenomenon commonly referred to as lack of visual grounding. In this work, we alleviate this shortcoming through a novel architecture for visual question answering that leverages common sense reasoning as a supervisory signal. Reasoning supervision takes the form of a textual justification of the correct answer, with such annotations being already available on large-scale Visual Common Sense Reasoning (VCR) datasets. The model’s visual attention is guided toward important elements of the scene through a similarity loss that aligns the learned attention distributions guided by the question and the correct reasoning. We demonstrate both quantitatively and qualitatively that the proposed approach can boost the model’s visual perception capability and lead to performance increase, without requiring training on explicit grounding annotations. Maria Parelli, Dimitrios Mallis, Markos Diomataris, Vassilis Pitsikalis |
ICIP | 1 |
| 2022 | Spatio-Temporal Graph Convolutional Networks for Continuous Sign Language RecognitionabstractWe address the challenging problem of continuous sign language recognition (CSLR) from RGB videos, proposing a novel deep-learning framework that employs spatio-temporal graph convolutional networks (ST-GCNs), which operate on multiple, appropriately fused feature streams, capturing the signer’s pose, shape, appearance, and motion information. In addition to introducing such networks to the continuous recognition problem, our model’s novelty lies on: (i) the feature streams considered and their blending into three ST-GCN modules; (ii) the combination of such modules with bi-directional long short-term memory networks, thus capturing both short-term embedded signing dynamics and long-range feature dependencies; and (iii) the fusion scheme, where the resulting modules operate in parallel, their posteriors aligned via a guiding connectionist temporal classification method, and fused for sign gloss prediction. Notably, concerning (i), in addition to traditional CSLR features, we investigate the utility of 3D human pose and shape parameterization via the "ExPose" approach, as well as 3D skeletal joint information that is regressed from detected 2D joints. We evaluate the proposed system on two well-known CSLR benchmarks, conducting extensive ablations on its modules. We achieve the new state-of-the-art on one of the two datasets, while reaching very competitive performance on the other. Maria Parelli, Katerina Papadimitriou, Gerasimos Potamianos, Georgios Pavlakos, Petros Maragos |
ICASSP | 1 |