Jan-Jakob Sonke

dblp:20/4093 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
17since 2021 · last 2025
0000-0001-5155-5274ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021
YearPublicationVenuePosition
2025 Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-Shot Semantic Segmentation
abstract
Generalized Few-Shot Semantic Segmentation (GFSS) aims to extend a segmentation model to novel classes with only a few annotated examples while maintaining performance on base classes. Recently, pretrained vision-language models (VLMs) such as CLIP have been leveraged in GFSS to improve generalization on novel classes through multi-modal prototypes learning. However, existing prototype-based methods are inherently deterministic, limiting the adaptability of learned prototypes to diverse samples, particularly for novel classes with scarce annotations. To address this, we propose FewCLIP, a probabilistic prototype calibration framework over multi-modal prototypes from the pretrained CLIP, thus providing more adaptive prototype learning for GFSS. Specifically, FewCLIP first introduces a prototype calibration mechanism, which refines frozen textual prototypes with learnable visual calibration prototypes, leading to a more discriminative and adaptive representation. Furthermore, unlike deterministic prototype learning techniques, FewCLIP introduces distribution regularization over these calibration prototypes. This probabilistic formulation ensures structured and uncertainty-aware prototype learning, effectively mitigating overfitting to limited novel class data while enhancing generalization. Extensive experimental results on PASCAL-5$^i$ and COCO-20$^i$ datasets demonstrate that our proposed FewCLIP significantly outperforms state-of-the-art approaches across both GFSS and class-incremental setting. The code is available at https://github.com/jliu4ai/FewCLIP.
Jie Liu 0043, Pan Zhou 0002, Jan-Jakob Sonke, Efstratios Gavves
ICCV4
2025 CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent Cooperation
abstract
In this work, we address the cooperation problem among large language model (LLM) based embodied agents, where agents must cooperate to achieve a common goal. Previous methods often execute actions extemporaneously and incoherently, without long-term strategic and cooperative planning, leading to redundant steps, failures, and even serious repercussions in complex tasks like search-and-rescue missions where discussion and cooperative plan are crucial. To solve this issue, we propose Cooperative Plan Optimization (CaPo) to enhance the cooperation efficiency of LLM-based embodied agents. Inspired by human cooperation schemes, CaPo improves cooperation efficiency with two phases: 1) meta plan generation, and 2) progress-adaptive meta plan and execution. In the first phase, all agents analyze the task, discuss, and cooperatively create a meta-plan that decomposes the task into subtasks with detailed steps, ensuring a long-term strategic and coherent plan for efficient coordination. In the second phase, agents execute tasks according to the meta-plan and dynamically adjust it based on their latest progress (e.g., discovering a target object) through multi-turn discussions. This progress-based adaptation eliminates redundant actions, improving the overall cooperation efficiency of agents. Experimental results on the ThreeDworld Multi-Agent Transport and Communicative Watch-And-Help tasks demonstrate CaPo's much higher task completion rate and efficiency compared with state-of-the-arts. The code is released at https://github.com/jliu4ai/CaPo.
Jie Liu 0043, Pan Zhou 0002, Yingjun Du, Ah-Hwee Tan, Cees Snoek, Jan-Jakob Sonke, Efstratios Gavves
ICLR6
2025 Probabilistic Interactive 3D Segmentation with Hierarchical Neural Processes
abstract
Interactive 3D segmentation has emerged as a promising solution for generating accurate object masks in complex 3D scenes by incorporating user-provided clicks. However, two critical challenges remain underexplored: (1) effectively generalizing from sparse user clicks to produce accurate segmentations and (2) quantifying predictive uncertainty to help users identify unreliable regions. In this work, we propose \emph{NPISeg3D}, a novel probabilistic framework that builds upon Neural Processes (NPs) to address these challenges. Specifically, NPISeg3D introduces a hierarchical latent variable structure with scene-specific and object-specific latent variables to enhance few-shot generalization by capturing both global context and object-specific characteristics. Additionally, we design a probabilistic prototype modulator that adaptively modulates click prototypes with object-specific latent variables, improving the model’s ability to capture object-aware context and quantify predictive uncertainty. Experiments on four 3D point cloud datasets demonstrate that NPISeg3D achieves superior segmentation performance with fewer clicks while providing reliable uncertainty estimations.
Jie Liu 0043, Pan Zhou 0002, Zehao Xiao, Wenzhe Yin, Jan-Jakob Sonke, Efstratios Gavves
ICML6
2024 Dynamic Prototype Adaptation with Distillation for Few-shot Point Cloud Segmentation
abstract
Few-shot point cloud segmentation seeks to generate per-point masks for previously unseen categories, using only a minimal set of annotated point clouds as reference. Existing prototype-based methods rely on support prototypes to guide the segmentation of query point clouds, but they encounter challenges when significant object variations exist between the support prototypes and query features. In this work, we present dynamic prototype adaptation (DPA), which explicitly learns task-specific prototypes for each query point cloud to tackle the object variation problem. DPA achieves the adaptation through prototype rectification, aligning vanilla prototypes from support with the query feature distribution, and prototype-to-query attention, extracting task-specific context from query point clouds. Furthermore, we introduce a prototype distillation regularization term, enabling knowledge transfer between early-stage prototypes and their deeper counterparts during adaption. By iteratively applying these adaptations, we generate task-specific prototypes for accurate mask predictions on query point clouds. Extensive experiments on two popular benchmarks show that DPA surpasses state-of-the-art methods by a significant margin, e.g., 7.43% and 6.39% under the 2 -way 1 -shot setting on S3DIS and ScanNet, respectively. Code is available at https://github.com/jliu4ai/DPA.
Jie Liu 0043, Wenzhe Yin, Yunlu Chen, Jan-Jakob Sonke, Efstratios Gavves
3DV5
2024 Kandinsky Conformal Prediction: Efficient Calibration of Image Segmentation Algorithms
abstract
Image segmentation algorithms can be understood as a collection of pixel classifiers, for which the outcomes of nearby pixels are correlated. Classifier models can be cal-ibrated using Inductive Conformal Prediction, but this re-quires holding back a sufficiently large calibration dataset for computing the distribution of non-conformity scores of the model's predictions. If one only requires only marginal calibration on the image level, this calibration set consists of all individual pixels in the images available for calibration. However, if the goal is to attain proper calibration for each individual pixel classifier, the calibration set consists of in-dividual images. In a scenario where data are scarce (such as the medical domain), it may not always be possible to set aside sufficiently many images for this pixel-level calibration. The method we propose, dubbed “Kandinsky calibration ”, makes use of the spatial structure present in the distribution of natural images to simultaneously calibrate the classifiers of “similar” pixels. This can be seen as an intermediate approach between marginal (imagewise) and conditional (pixelwise) calibration, where non-conformity scores are ag-gregated over similar image regions, thereby making more efficient use of the images available for calibration.$We$run experiments on segmentation algorithms trained and calibrated on subsets of the public MS-COCO and Medical Decathlon datasets, demonstrating that Kandinsky calibration method can significantly improve the coverage. When compared to both pixelwise and imagewise calibration on little data, the Kandinsky method achieves much lower coverage errors, indicating the data efficiency of the Kandinsky calibration.
Joren Brunekreef, Eric Marcus, Ray Sheombarsing, Jan-Jakob Sonke, Jonas Teuwen
CVPR4
2024 Task-Driven Wavelets Using Constrained Empirical Risk Minimization
abstract
Deep Neural Networks (DNNs) are widely used for their ability to effectively approximate large classes of functions. This flexibility, however, makes the strict enforcement of con-straints on DNNs a difficult problem. In contexts where it is critical to limit the function space to which certain network components belong, such as wavelets employed in Multi-Resolution Analysis (MRA), naive constraints via additional terms in the loss function are inadequate. To address this, we introduce a Convolutional Neural Network (CNN) wherein the convolutional filters are strictly constrained to be wavelets. This allows the filters to update to task-optimized wavelets during the training procedure. Our primary contri-bution lies in the rigorous formulation of these filters via a constrained empirical risk minimization framework, thereby providing an exact mechanism to enforce these structural constraints. While our work is grounded in theory, we investigate our approach empirically through applications in medical imaging, particularly in the task of contour prediction around various organs, achieving superior performance compared to baseline methods.
Eric Marcus, Ray Sheombarsing, Jan-Jakob Sonke, Jonas Teuwen
CVPR3
2024 How to Train Neural Field Representations: A Comprehensive Study and Benchmark
abstract
Neural fields (NeFs) have recently emerged as a versatile method for modeling signals of various modalities, including images, shapes, and scenes. Subsequently, a number of works have explored the use of NeFs as representations for downstream tasks, e.g. classifying an image based on the parameters of a NeF that has been fit to it. However, the impact of the NeF hyperparameters on their quality as downstream representation is scarcely understood and re-mains largely unexplored. This is in part caused by the large amount of time required to. fit datasets of neuralfields. In this work, we propose a JAX-based library1 that lever-ages parallelization to enable fast optimization of large-scale NeF datasets, resulting in a significant speed-up. With this library, we perform a comprehensive study that inves-tigates the effects of different hyperparameters on fitting NeFs for downstream tasks. In particular, we explore the use of a shared initialization, the effects of overtraining, and the expressiveness of the network architectures used. Our study provides valuable insights on how to train NeFs and offers guidance for optimizing their effectiveness in down-stream applications. Finally, based on the proposed library and our analysis, we propose Neural Field Arena, a bench-mark consisting of neural field variants of popular vision datasets, including MNIST, CIFAR, variants of ImageNet, and ShapeNetv2. Our library and the Neural Field Arena will be open-sourced to introduce standardized benchmarking and promote further research on neural fields.
Samuele Papa, Riccardo Valperga, David M. Knigge, Miltiadis Kofinas, Phillip Lippe, Jan-Jakob Sonke, Efstratios Gavves
CVPR6
2024 Click Prompt Learning with Optimal Transport for Interactive Segmentation
Jie Liu 0043, Wenzhe Yin, Jan-Jakob Sonke, Efstratios Gavves
ECCV (33)4
2024 Space-Time Continuous PDE Forecasting using Equivariant Neural Fields
abstract
Recently, Conditional Neural Fields (NeFs) have emerged as a powerful modelling paradigm for PDEs, by learning solutions as flows in the latent space of the Conditional NeF. Although benefiting from favourable properties of NeFs such as grid-agnosticity and space-time-continuous dynamics modelling, this approach limits the ability to impose known constraints of the PDE on the solutions -- such as symmetries or boundary conditions -- in favour of modelling flexibility. Instead, we propose a space-time continuous NeF-based solving framework that - by preserving geometric information in the latent space of the Conditional NeF - preserves known symmetries of the PDE. We show that modelling solutions as flows of pointclouds over the group of interest $G$ improves generalization and data-efficiency. Furthermore, we validate that our framework readily generalizes to unseen spatial and temporal locations, as well as geometric transformations of the initial conditions - where other NeF-based PDE forecasting methods fail -, and improve over baselines in a number of challenging geometries.
David M. Knigge, David R. Wessels, Riccardo Valperga, Samuele Papa, Jan-Jakob Sonke, Erik J. Bekkers, Efstratios Gavves
NeurIPS5
2024 Domain Adaptation with Cauchy-Schwarz Divergence
abstract
Domain adaptation aims to use training data from one or multiple source domains to learn a hypothesis that can be generalized to a different, but related, target domain. As such, having a reliable measure for evaluating the discrepancy of both marginal and conditional distributions is crucial. We introduce Cauchy-Schwarz (CS) divergence to the problem of unsupervised domain adaptation (UDA). The CS divergence offers a theoretically tighter generalization error bound than the popular Kullback-Leibler divergence. This holds for the general case of supervised learning, including multi-class classification and regression. Furthermore, we illustrate that the CS divergence enables a simple estimator on the discrepancy of both marginal and conditional distributions between source and target domains in the representation space, without requiring any distributional assumptions. We provide multiple examples to illustrate how the CS divergence can be conveniently used in both distance metric- or adversarial training-based UDA frameworks, resulting in compelling performance. The code of our paper is available at \url{https://github.com/ywzcode/CS-adv}.
Wenzhe Yin, Shujian Yu, Yicong Lin, Jie Liu 0043, Jan-Jakob Sonke, Efstratios Gavves
UAI5
2023 Modelling Long Range Dependencies in $N$D: From Task-Specific to a General Purpose CNN
David M. Knigge, David W. Romero, Albert Gu, Efstratios Gavves, Erik J. Bekkers, Jakub M. Tomczak, Mark Hoogendoorn, Jan-Jakob Sonke
ICLR8
2023 Noise2Aliasing: Unsupervised Deep Learning for View Aliasing and Noise Reduction in 4DCBCT
Samuele Papa, Efstratios Gavves, Jan-Jakob Sonke
MICCAI (10)3
2023 PC-Reg: A pyramidal prediction-correction approach for large deformation image registration
abstract
Deformable image registration plays an important role in medical image analysis. Deep neural networks such as VoxelMorph and TransMorph are fast, but limited to small deformations and face challenges in the presence of large deformations. To tackle large deformations in medical image registration, we propose PC-Reg, a pyramidal Prediction and Correction method for deformable registration, which treats multi-scale registration akin to solving an ordinary differential equation (ODE) across scales. Starting with a zero-initialized deformation at the coarse level, PC-Reg follows the predictor-corrector regime and progressively predicts a residual flow and a correction flow to update the deformation vector field through different scales. The prediction in each scale can be regarded as a single step of ODE integration. PC-Reg can be easily extended to diffeomorphic registration and is able to alleviate the multiscale accumulated upsampling and diffeomorphic integration error. Further, to transfer details from full resolution to low scale, we introduce a distillation loss, where the output is used as the target label for intermediate outputs. Experiments on inter-patient deformable registration show that the proposed method significantly improves registration not only for large but also for small deformations.
Wenzhe Yin, Jan-Jakob Sonke, Efstratios Gavves
Medical Image Anal.2
2022 Few-shot Semantic Segmentation with Support-induced Graph Convolutional Network
Jie Liu 0043, Yanqi Bao, Wenzhe Yin, Yang Gao 0001, Jan-Jakob Sonke, Efstratios Gavves
BMVC6
2022 Dynamic Prototype Convolution Network for Few-Shot Semantic Segmentation
abstract
The key challenge for few-shot semantic segmentation (FSS) is how to tailor a desirable interaction among sup-port and query features and/or their prototypes, under the episodic training scenario. Most existing FSS methods im-plement such support/query interactions by solely leveraging plain operations - e.g., cosine similarity and feature concatenation - for segmenting the query objects. How-ever, these interaction approaches usually cannot well capture the intrinsic object details in the query images that are widely encountered in FSS, e.g., if the query object to be segmented has holes and slots, inaccurate segmentation al-most always happens. To this end, we propose a dynamic prototype convolution network (DPCN) to fully capture the aforementioned intrinsic details for accurate FSS. Specifi-cally, in DPCN, a dynamic convolution module (DCM) is firstly proposed to generate dynamic kernels from support foreground, then information interaction is achieved by con-volution operations over query features using these kernels. Moreover, we equip DPCN with a support activation mod-ule (SAM) and a feature filtering module (FFM) to generate pseudo mask and filter out background information for the query images, respectively. SAM and FFM together can mine enriched context information from the query features. Our DPCN is also flexible and efficient under the k-shot FSS setting. Extensive experiments on PASCAL-5iand COCO 20ishow that DPCN yields superior performances under both 1-shot and 5-shot settings.
Jie Liu 0043, Yanqi Bao, Guosen Xie, Huan Xiong, Jan-Jakob Sonke, Efstratios Gavves
CVPR5
2022 Recurrent Variational Network: A Deep Learning Inverse Problem Solver applied to the task of Accelerated MRI Reconstruction
abstract
Magnetic Resonance Imaging can produce detailed images of the anatomy and physiology of the human body that can assist doctors in diagnosing and treating pathologies such as tumours. However, MRI suffers from very long acquisition times that make it susceptible to patient motion artifacts and limit its potential to deliver dynamic treatments. Conventional approaches such as Parallel Imaging and Compressed Sensing allow for an increase in MRI acquisition speed by reconstructing MR images from sub-sampled MRI data acquired using multiple receiver coils. Recent advancements in Deep Learning combined with Parallel Imaging and Compressed Sensing techniques have the potential to produce high-fidelity reconstructions from highly accelerated MRI data. In this work we present a novel Deep Learning-based Inverse Problem solver applied to the task of Accelerated MRI Reconstruction, called the Recurrent Variational Network (RecurrentVarNet), by exploiting the properties of Convolutional Recurrent Neural Networks and unrolled algorithms for solving Inverse Problems. The RecurrentVarNet consists of multiple recurrent blocks, each responsible for one iteration of the unrolled variational optimization scheme for solving the inverse problem of multi-coil Accelerated MRI Reconstruction. Contrary to traditional approaches, the optimization steps are performed in the observation domain (k-space) instead of the image domain. Each block of the RecurrentVarNet refines the observed k-space and comprises a data consistency term and a recurrent unit which takes as input a learned hidden state and the prediction of the previous block. Our proposed method achieves new state of the art qualitative and quantitative reconstruction results on 5-fold and 10-fold accelerated data from a public multi-coil brain dataset, outperforming previous conventional and deep learning-based approaches. Our code is publicly available at https://github.com/NKI-AI/direct.
George Yiasemis, Jan-Jakob Sonke, Clarisa I. Sánchez, Jonas Teuwen
CVPR2
2022 Dynamic Transformer for Few-shot Instance Segmentation
abstract
Few-shot instance segmentation aims to train an instance segmentation model that can fast adapt to novel classes with only a few reference images. Existing methods are usually derived from standard detection models and tackle few-shot instance segmentation indirectly by conducting classification, box regression, and mask prediction on a large set of redundant proposals followed by indispensable post-processing, e.g., Non-Maximum Suppression. Such complicated hand-crafted procedures and hyperparameters lead to degraded optimization and insufficient generalization ability. In this work, we propose an end-to-end Dynamic Transformer Network, DTN for short, to directly segment all target object instances from arbitrary categories given by reference images, relieving the requirements of dense proposal generation and post-processing. Specifically, a small set of Dynamic Queries, conditioned on reference images, are exclusively assigned to target object instances and generate all the instance segmentation masks of reference categories simultaneously. Moreover, a Semantic-induced Transformer Decoder is introduced to constrain the cross-attention between dynamic queries and target images within the pixels of the reference category, which suppresses the noisy interaction with the background and irrelevant categories. Extensive experiments are conducted on the COCO-20 dataset. The experiment results demonstrate that our proposed Dynamic Transformer Network significantly outperforms the state-of-the-arts.
Jie Liu 0043, Yongtuo Liu, Subhransu Maji, Jan-Jakob Sonke, Efstratios Gavves
ACM Multimedia5
2019 Recurrent inference machines for reconstructing heterogeneous MRI data
Kai Lønning, Patrick Putzky, Jan-Jakob Sonke, Liesbeth Reneman, Matthan W. A. Caan, Max Welling
Medical Image Anal.3
2015 Diversifying Multi-Objective Gradient Techniques and their Role in Hybrid Multi-Objective Evolutionary Algorithms for Deformable Medical Image Registration
abstract
Gradient methods and their value in single-objective, real-valued optimization are well-established. As such, they play a key role in tackling real-world, hard optimization problems such as deformable image registration (DIR). A key question is to which extent gradient techniques can also play a role in a multi-objective approach to DIR. We therefore aim to exploit gradient information within an evolutionary-algorithm-based multi-objective optimization framework for DIR. Although an analytical description of the multi-objective gradient (the set of all Pareto-optimal improving directions) is available, it is nontrivial how to best choose the most appropriate direction per solution because these directions are not necessarily uniformly distributed in objective space. To address this, we employ a Monte-Carlo method to obtain a discrete, spatially-uniformly distributed approximation of the set of Pareto-optimal improving directions. We then apply a diversification technique in which each solution is associated with a unique direction from this set based on its multi- as well as single-objective rank. To assess its utility, we compare a state-of-the-art multi-objective evolutionary algorithm with three different hybrid versions thereof on several benchmark problems and two medical DIR problems. Results show that the diversification strategy successfully leads to unbiased improvement, helping an adaptive hybrid scheme solve all problems, but the evolutionary algorithm remains the most powerful optimization method, providing the best balance between proximity and diversity.
Kleopatra Pirpinia, Tanja Alderliesten, Jan-Jakob Sonke, Marcel van Herk, Peter A. N. Bosman
GECCO3
2012 Directional Interpolation for Motion Weighted 4D Cone-Beam CT Reconstruction
Jan-Jakob Sonke
MICCAI (1)2
2008 On-the-Fly Motion-Compensated Cone-Beam CT Using an a Priori Motion Model
Simon Rit, Jochem Wolthaus, Marcel van Herk, Jan-Jakob Sonke
MICCAI (1)4