EDBT 2026 Demo / reviewers in the wild / expert
Lilian Calvet
dblp:126/0796
· DBLP profile ↗
18ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-2565-2297ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-view Surface Reconstruction Using Normal and Reflectance Cues
Robin Bruneau, Baptiste Brument, Yvain Quéau, Jean Mélou, François Lauze, Jean-Denis Durou, Lilian Calvet |
Int. J. Comput. Vis. | 7 |
| 2026 | NeuralBoneReg: An instance-specific label-free point cloud-based method for multi-modal bone surface registrationabstractBACKGROUND: In computer- and robot-assisted orthopedic surgery (CAOS), patient-specific surgical plans are generated from preoperative medical imaging data to define target locations and implant trajectories. During surgery, these plans must be precisely transferred to the intraoperative setting to guide accurate execution. The accuracy and success of this transfer rely on cross-registration between preoperative and intraoperative data. However, the substantial heterogeneity across imaging modalities and devices renders this registration process challenging and error-prone, leading to inaccuracies. Consequently, more robust and accurate methods for automatic, modality-agnostic multimodal registration of bone surfaces would have a substantial clinical impact. METHODS: We propose NeuralBoneReg, an instance-specific self-supervised, surface-based framework for bone surface registration using 3D point clouds as an intermediate representation. NeuralBoneReg comprises two key components: an implicit neural unsigned distance field (UDF) module and a multilayer perceptron (MLP)-based registration module. The UDF module learns a neural representation of the preoperative bone model. The registration module solves both global initialization and local refinement by generating a set of transformation hypotheses to register the intraoperative point cloud with the preoperative neural UDF based on a coarse-to-fine strategy. Compared to state-of-the-art (SOTA) supervised registration, NeuralBoneReg operates in an instance-specific self-supervised manner, without requiring inter-subject training data with ground truth transformations. We evaluated NeuralBoneReg against baseline methods on two publicly available multi-modal datasets: a CT-ultrasound dataset of the fibula and tibia (UltraBones100k) and a CT-RGB-D dataset of spinal vertebrae (SpineDepth). The evaluation also includes a newly introduced CT-ultrasound dataset of cadaveric subjects containing femur and pelvis (UltraBonesHip), which will be made publicly available. RESULTS: Quantitative and qualitative results show that NeuralBoneReg achieves competitive performance across anatomies and modalities. On UltraBones100k, it obtains an RRE of 1.83±1.30°, an RTE of 2.02±1.30mm, an RR of 0.89, a CD of 0.82±0.12mm, and an HD95 of 2.06±0.36mm. On UltraBonesHip, it maintains stable performance with an RRE of 1.90±1.56°, an RTE of 2.21±0.86mm, an RR of 0.88, a CD of 2.50±1.08mm, and an HD95 of 8.97±4.08mm, while other methods degrade significantly. On SpineDepth, it achieves an RRE of 3.78±19.34°, an RTE of 2.80±3.75mm, an RR of 0.84, a CD of 1.78±1.61mm, and an HD95 of 4.26±4.58mm. Overall, the method consistently achieves accuracy close to pseudo ground truth across datasets. CONCLUSION: NeuralBoneReg achieves robust, accurate, and modality-agnostic registration of bone surfaces, offering a promising solution for reliable cross-modal alignment in computer- and robot-assisted orthopedic surgery. Luohong Wu, Matthias Seibold, Nicola Cavalcanti, Yunke Ao, Roman Flepp, Aidana Massalimova, Lilian Calvet, Philipp Fürnstahl |
Medical Image Anal. | 7 |
| 2025 | Next-generation surgical navigation: Marker-less multi-view 6DoF pose estimation of surgical instrumentsabstractState-of-the-art research of traditional computer vision is increasingly leveraged in the surgical domain. A particular focus in computer-assisted surgery is to replace marker-based tracking systems for instrument localization with pure image-based 6DoF pose estimation using deep-learning methods. However, state-of-the-art single-view pose estimation methods do not yet meet the accuracy required for surgical navigation. In this context, we investigate the benefits of multi-view setups for highly accurate and occlusion-robust 6DoF pose estimation of surgical instruments and derive recommendations for an ideal camera system that addresses the challenges in the operating room. Our contributions are threefold. First, we present a multi-view RGB-D video dataset of ex-vivo spine surgeries, captured with static and head-mounted cameras and including rich annotations for surgeon, instruments, and patient anatomy. Second, we perform an extensive evaluation of three state-of-the-art single-view and multi-view pose estimation methods, analyzing the impact of camera quantities and positioning, limited real-world data, and static, hybrid, or fully mobile camera setups on the pose accuracy, occlusion robustness, and generalizability. Third, we design a multi-camera system for marker-less surgical instrument tracking, achieving an average position error of 1.01mm and orientation error of 0.89° for a surgical drill, and 2.79mm and 3.33° for a screwdriver under optimal conditions. Our results demonstrate that marker-less tracking of surgical instruments is becoming a feasible alternative to existing marker-based systems. Jonas Hein, Nicola Cavalcanti, Daniel Suter, Lukas Zingg, Fabio Carrillo, Lilian Calvet, Mazda Farshad, Nassir Navab, Marc Pollefeys, Philipp Fürnstahl |
Medical Image Anal. | 6 |
| 2024 | RNb-NeuS: Reflectance and Normal-Based Multi-View 3D ReconstructionabstractThis paper introduces a versatile paradigm for integrating multi-view reflectance (optional) and normal maps acquired through photometric stereo. Our approach employs a pixel-wise joint reparameterization of reflectance and normal, considering them as a vector of radiances rendered under simulated, varying illumination. This re-parameterization enables the seamless integration of reflectance and normal maps as input data in neural volume rendering-based 3D reconstruction while preserving a single optimization objective. In contrast, recent multi-view photometric stereo (MVPS) methods depend on multiple, potentially conflicting objectives. Despite its apparent simplicity, our proposed approach outperforms state-of-the-art approaches in MVPS benchmarks across F-score, Chamfer distance, and mean angular error metrics. Notably, it significantly improves the detailed 3D reconstruction of areas with high curvature or low visibility. Baptiste Brument, Robin Bruneau, Yvain Quéau, Jean Mélou, François Lauze, Jean-Denis Durou, Lilian Calvet |
CVPR | 7 |
| 2024 | Domain adaptation strategies for 3D reconstruction of the lumbar spine using real fluoroscopy data
Sascha Jecklin, Youyang Shen, Amandine Gout, Daniel Suter, Lilian Calvet, Lukas Zingg, Jennifer Straub, Nicola Cavalcanti, Mazda Farshad, Philipp Fürnstahl, Hooman Esfandiari |
Medical Image Anal. | 5 |
| 2021 | Using Multiple Images and Contours for Deformable 3D-2D Registration of a Preoperative CT in Laparoscopic Liver Surgery
Yamid Espinel, Lilian Calvet, Karim Botros, Emmanuel Buc, Christophe Tilmant, Adrien Bartoli |
MICCAI (4) | 2 |
| 2021 | Image-Based Incision Detection for Topological Intraoperative 3D Model Update in Augmented Reality Assisted Laparoscopic Surgery
Tom François, Lilian Calvet, Callyane Sève-d'Erceville, Nicolas Bourdel, Adrien Bartoli |
MICCAI (4) | 2 |
| 2021 | AliceVision Meshroom: An open-source 3D reconstruction pipelineabstractThis paper introduces the Meshroom software and its underlying 3D computer vision framework AliceVision. This solution provides a photogrammetry pipeline to reconstruct 3D scenes from a set of unordered images. It also features other pipelines for fusing multi-bracketing low dynamic range images into high dynamic range, stitching multiple images into a panorama and estimating the motion of a moving camera. Meshroom's node-graph architecture allows the user to customize the different pipelines to adjust them to their domain specific needs. The user can interactively add other processing nodes to modify a pipeline, export intermediate data to analyze the result of the algorithms and easily compare the outputs given by different sets of parameters. The software package is released in open source and relies on open file formats. These features enable researchers to conveniently run the pipelines, access and visualize the data at each step, thus promoting the sharing and the reproducibility of the results. Carsten Griwodz, Simone Gasparini, Lilian Calvet, Pierre Gurdjos, Fabien Castan, Benoit Maujean, Gregoire De Lillo, Yann Lanthony |
MMSys | 3 |
| 2021 | Detection, segmentation, and 3D pose estimation of surgical tools using convolutional neural networks and algebraic geometryabstractBackground and objective: Surgical tool detection, segmentation, and 3D pose estimation are crucial components in Computer-Assisted Laparoscopy (CAL). The existing frameworks have two main limitations. First, they do not integrate all three components. Integration is critical; for instance, one should not attempt computing pose if detection is negative. Second, they have highly specific requirements, such as the availability of a CAD model. We propose an integrated and generic framework whose sole requirement for the 3D pose is that the tool shaft is cylindrical. Our framework makes the most of deep learning and geometric 3D vision by combining a proposed Convolutional Neural Network (CNN) with algebraic geometry. We show two applications of our framework in CAL: tool-aware rendering in Augmented Reality (AR) and tool-based 3D measurement. Methods: We name our CNN as ART-Net (Augmented Reality Tool Network). It has a Single Input Multiple Output (SIMO) architecture with one encoder and multiple decoders to achieve detection, segmentation, and geometric primitive extraction. These primitives are the tool edge-lines, mid-line, and tip. They allow the tool’s 3D pose to be estimated by a fast algebraic procedure. The framework only proceeds if a tool is detected. The accuracy of segmentation and geometric primitive extraction is boosted by a new Full resolution feature map Generator (FrG). We extensively evaluate the proposed framework with the EndoVis and new proposed datasets. We compare the segmentation results against several variants of the Fully Convolutional Network (FCN) and U-Net. Several ablation studies are provided for detection, segmentation, and geometric primitive extraction. The proposed datasets are surgery videos of different patients. Results: In detection, ART-Net achieves 100.0 % in both average precision and accuracy. In segmentation, it achieves 81.0 % in mean Intersection over Union (mIoU) on the robotic EndoVis dataset (articulated tool), where it outperforms both FCN and U-Net, by 4.5 p p and 2.9 p p , respectively. It achieves 88.2 % in mIoU on the remaining datasets (non-articulated tool). In geometric primitive extraction, ART-Net achieves 2.45 ∘ and 2.23 ∘ in mean Arc Length (mAL) error for the edge-lines and mid-line, respectively, and 9.3 pixels in mean Euclidean distance error for the tool-tip. Finally, in terms of 3D pose evaluated on animal data, our framework achieves 1.87 mm, 0.70 mm, and 4.80 mm mean absolute errors on the X , Y , and Z coordinates, respectively, and 5 . 94 ∘ angular error on the shaft orientation. It achieves 2.59 mm and 1.99 mm in mean and median location error of the tool head evaluated on patient data. Conclusions: The proposed framework outperforms existing ones in detection and segmentation. Compared to separate networks, integrating the tasks in a single network preserves accuracy in detection and segmentation but substantially improves accuracy in geometric primitive extraction. Overall, our framework has similar or better accuracy in 3D pose estimation while largely improving robustness against the very challenging imaging conditions of laparoscopy. The source code of our framework and our annotated dataset will be made publicly available at https://github.com/kamruleee51/ART-Net . Md. Kamrul Hasan 0002, Lilian Calvet, Navid Rabbani, Adrien Bartoli |
Medical Image Anal. | 2 |
| 2021 | Augmented Reality Guided Laparoscopic Surgery of the UterusabstractA major research area in Computer Assisted Intervention (CAI) is to aid laparoscopic surgery teams with Augmented Reality (AR) guidance. This involves registering data from other modalities such as MR and fusing it with the laparoscopic video in real-time, to reveal the location of hidden critical structures. We present the first system for AR guided laparoscopic surgery of the uterus. This works with pre-operative MR or CT data and monocular laparoscopes, without requiring any additional interventional hardware such as optical trackers. We present novel and robust solutions to two main sub-problems: the initial registration, which is solved using a short exploratory video, and update registration, which is solved with real-time tracking-by-detection. These problems are challenging for the uterus because it is a weakly-textured, highly mobile organ that moves independently of surrounding structures. In the broader context, our system is the first that has successfully performed markerless real-time registration and AR of a mobile human organ with monocular laparoscopes in the OR. Toby Collins, Daniel Pizarro-Perez, Simone Gasparini, Nicolas Bourdel, Pauline Chauvet, Michel Canis, Lilian Calvet, Adrien Bartoli |
IEEE Trans. Medical Imaging | 7 |
| 2018 | Popsift: a faithful SIFT implementation for real-time applicationsabstractThe keypoint detector and descriptor Scalable Invariant Feature Transform (SIFT) [8] is famous for its ability to extract and describe keypoints in 2D images of natural scenes. It is used in ranging from object recognition to 3D reconstruction. However, SIFT is considered compute-heavy. This has led to the development of many keypoint extraction and description methods that sacrifice the wide applicability of SIFT for higher speed. We present our CUDA implementation named PopSift that does not sacrifice any detail of the SIFT algorithm, achieves a keypoint extraction and description performance that is as accurate as the best existing implementations, and runs at least 100x faster on a high-end consumer GPU than existing CPU implementations on a desktop CPU. Without any algorithmic trade-offs and short-cuts that sacrifice quality for speed, we extract at >25 fps from 1080p images with upscaling to 3840x2160 pixels on a high-end consumer GPU. Carsten Griwodz, Lilian Calvet, Pål Halvorsen |
MMSys | 2 |
| 2016 | Detection and Accurate Localization of Circular Fiducials under Highly Challenging ConditionsabstractUsing fiducial markers ensures reliable detection and identification of planar features in images. Fiducials are used in a wide range of applications, especially when a reliable visual reference is needed, e.g., to track the camera in cluttered or textureless environments. A marker designed for such applications must be robust to partial occlusions, varying distances and angles of view, and fast camera motions. In this paper, we present a robust, highly accurate fiducial system, whose markers consist of concentric rings, along with its theoretical foundations. Relying on projective properties, it allows to robustly localize the imaged marker and to accurately detect the position of the image of the (common) circle center. We demonstrate that our system can detect and accurately localize these circular fiducials under very challenging conditions and the experimental results reveal that it outperforms other recent fiducial systems. Lilian Calvet, Pierre Gurdjos, Carsten Griwodz, Simone Gasparini |
CVPR | 1 |
| 2016 | Immersed gaming in MinecraftabstractThis demonstration will showcase mixed reality technologies that we developed for a series of public art performances in Vienna in October 2015 in a collaboration of performance artists and researchers. The focus of the demonstration is on natural interaction techniques that can be used intuitively to control an avatar in a virtual 3D world. We combine virtual reality devices with optical location tracking, hand gesture recognition and smart devices. Conference attendees will be able to walk around in a Minecraft world by physically moving in the real world and to perform actions on virtual world items using hand gestures. They can also test our initial system for shared avatar control, in which a user in the real world cooperates with a user in the virtual world. Finally, attendees will have the opportunity to give us feedback about their experience with our system. Milan Loviska, Otto Krause, Herman Arnold Engelbrecht, Jason B. Nel, Gregor Schiele, Alwyn Burger, Stephan Schmeißer, Christopher Cichiwskyj, Lilian Calvet, Carsten Griwodz, Pål Halvorsen |
MMSys | 9 |
| 2015 | Exploitation of producer intent in relation to bandwidth and QoE for online video streaming servicesabstractThis paper is the product of recent advances in research on users' intent during multimedia content retrieval. Our goal is to save bandwidth while streaming video clips from a browsable on-demand service, while maintaining or even improving the users' quality of experience (QoE). Understanding user intent allows us to predict whether streaming a particular video in a low quality constitutes a reduced QoE for a user. However, many VoD streaming services today are used by users for a wide variety of reasons, meaning that user intent cannot be inferred from their use of the service alone. However, our investigation demonstrates that user intent does in most cases coincide with producer intent. We can also demonstrate that the latter can be inferred from the content itself as well as associated metadata. By transitivity, we can choose a default video quality that satisfies the users QoE in the majority of cases. Michael Riegler 0001, Lilian Calvet, Amandine Calvet, Pål Halvorsen, Carsten Griwodz |
NOSSDAV | 2 |
| 2014 | 3D Interest Maps From Simultaneous Video RecordingsabstractWe consider an emerging situation where multiple cameras are filming the same event simultaneously from a diverse set of angles. The captured videos provide us with the multiple view geometry and an understanding of the 3D structure of the scene. We further extend this understanding by introducing the concept of 3D interest map in this paper. As most users naturally film what they find interesting from their respective viewpoints, the 3D structure can be annotated with the level of interest, naturally crowdsourced from the users. A 3D interest map can be understood as an extension of saliency maps in the 3D space that captures the semantics of the scene. We evaluate the idea of 3D interest maps on two real datasets, taken from the environment or the cameras that are equipped enough to have an estimation of the poses of cameras and a reasonable synchronization between them. We study two aspects of the 3D interest maps in our evaluation. First, by projecting them into 2D, we compare them to state-of-the-art saliency maps. Second, to demonstrate the usefulness of the 3D interest maps, we apply them to a video mashup system that automatically produces an edited video from one of the datasets. Axel Carlier, Lilian Calvet, Duong-Trung-Dung Nguyen, Wei Tsang Ooi, Pierre Gurdjos, Vincent Charvillat |
ACM Multimedia | 2 |
| 2013 | An Enhanced Structure-from-Motion Paradigm Based on the Absolute Dual Quadric and Images of Circular PointsabstractThis work aims at introducing a new unified Structure from Motion (SfM) paradigm in which images of circular point-pairs can be combined with images of natural points. An imaged circular point-pair encodes the 2D Euclidean structure of a world plane and can easily be derived from the image of a planar shape, especially those including circles. A classical SfM method generally runs two steps: first a projective factorization of all matched image points (into projective cameras and points) and second a camera self calibration that updates the obtained world from projective to Euclidean. This work shows how to introduce images of circular points in these two SfM steps while its key contribution is to provide the theoretical foundations for combining "classical" linear self-calibration constraints with additional ones derived from such images. We show that the two proposed SfM steps clearly contribute to better results than the classical approach. We validate our contributions on synthetic and real images. Lilian Calvet, Pierre Gurdjos |
ICCV | 1 |
| 2012 | Camera tracking using concentric circle markers: Paradigms and algorithmsabstractA C2Tag refers to a set of concentric circles of different radii. C2Tags have been recently introduced in computer vision, in particular for camera calibration, as they offer highly interesting photometric and geometric properties, compared to the classical “checkerboard” tags. In this work, we propose the general paradigm of camera tracking based on a planar marker consisting of at least two C2Tags. All the involved steps are described: detection, identification, 2D reconstruction, calibration and 3D reconstruction. Our contribution is to introduce the following missing steps required to deal with long video sequences under time constraints: key-frame selection, bundle adjustment and intermediate pose refinement. Lilian Calvet, Pierre Gurdjos, Vincent Charvillat |
ICIP | 1 |
| 2012 | Camera tracking based on circular point factorization
Lilian Calvet, Pierre Gurdjos, Vincent Charvillat |
ICPR | 1 |