EDBT 2026 Demo / reviewers in the wild / expert
Jan C. van Gemert
dblp:25/3153 · also Jan van Gemert
· DBLP profile ↗
86ranked-venue papers
7as first author
33since 2021 · last 2025
0000-0002-3913-2786ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 70 · 5 first-author · 25 since 2021Artificial intelligence and machine learning · 51 · 4 first-author · 22 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Physics From Video: Unsupervised Physical Parameter Estimation for Continuous Dynamical SystemsabstractExtracting physical dynamical system parameters from recorded observations is key in natural science. Current methods for automatic parameter estimation from video train supervised deep networks on large datasets. Such datasets require labels, which are difficult to acquire. While some unsupervised techniques–which depend on frame prediction–exist, they suffer from long training times, initialization in-stabilities, only consider motion-based dynamical systems, and are evaluated mainly on synthetic data. In this work, we propose an unsupervised method to estimate the physical parameters of known, continuous governing equations from single videos suitable for different dynamical systems beyond motion and robust to initialization. Moreover, we remove the need for frame prediction by implementing a KL-divergence-based loss function in the latent space, which avoids convergence to trivial solutions and reduces model size and compute. We first evaluate our model on synthetic data, as commonly done. After which, we take the field closer to reality by recording Delfys75: our own real-world dataset of 75 videos for five different types of dynamical systems to evaluate our method and others. Our method compares favorably to others. Code and data are available online: https://github.com/Alejandro-neuro/Learning_physics_from_video. Alejandro Castañeda Garcia, Jan Warchocki, Jan C. van Gemert, Daan Brinks, Nergis Tömen |
CVPR | 3 |
| 2025 | End-to-End Implicit Neural Representations for ClassificationabstractImplicit neural representations (INRs) such as NeRF and SIREN encode a signal in neural network parameters and show excellent results for signal reconstruction. Using INRs for downstream tasks, such as classification, is however not straightforward. Inherent symmetries in the parameters pose challenges and current works primarily focus on designing architectures that are equivariant to these symmetries. However, INR-based classification still significantly under-performs compared to pixel-based methods like CNNs. This work presents an end-to-end strategy for initializing SIRENs together with a learned learning-rate scheme, to yield representations that improve classification accuracy. We show that a simple, straightforward, Transformer model applied to a meta-learned SIREN, without incorporating explicit symmetry equivariances, outperforms the current state-of-the-art. On the CIFAR-10 SIREN classification task, we improve the state-of-the-art without augmentations from 38.8% to 59.6%, and from 63.4% to 64.7% with augmentations. We demonstrate scalability on the high-resolution Imagenette dataset achieving reasonable reconstruction quality with a classification accuracy of 60.8% and are the first to do INR classification on the full ImageNet-1K dataset where we achieve a SIREN classification performance of 23.6%. To the best of our knowledge, no other SIREN classification approach has managed to set a classification baseline for any high-resolution image dataset. Our code is available at https://github.com/SanderGielisse/MWT. Alexander Gielisse, Jan C. van Gemert |
CVPR | 2 |
| 2025 | GazeHTA: End-to-End Gaze Target Detection with Head-Target Association
Zhiyi Lin 0003, Jouh Yeong Chew, Jan C. van Gemert, Xucong Zhang |
ICRA | 3 |
| 2025 | CleanUMamba: A Compact Mamba Network for Speech Denoising using Channel PruningabstractThis paper presents CleanUMamba, a time-domain neural network architecture designed for real-time causal audio denoising directly applied to raw waveforms. CleanUMamba leverages a U-Net encoder-decoder structure, incorporating the Mamba state-space model in the bottleneck layer. By replacing conventional self-attention and LSTM mechanisms with Mamba, our architecture offers superior denoising performance while maintaining a constant memory footprint, enabling streaming operation. To enhance efficiency, we applied structured channel pruning, achieving an 8X reduction in model size without compromising audio quality. Our model demonstrates strong results in the Interspeech 2020 Deep Noise Suppression challenge. Specifically, CleanUMamba achieves a PESQ score of 2.42 and STOI of 95.1% with only 442K parameters and 468M MACs, matching or outperforming larger models in real-time performance. Code will be available at: https://github.com/lab-emi/CleanUMamba Sjoerd Groot, Qinyu Chen, Jan C. van Gemert, Chang Gao 0002 |
ISCAS | 3 |
| 2024 | MSD: A Benchmark Dataset for Floor Plan Generation of Building Complexes
Casper van Engelenburg, Fatemeh Mostafavi, Emanuel Kuhn, Yuntae Jeon, Michael Franzen, Matthias Standfest, Jan C. van Gemert, Seyran Khademi |
ECCV (58) | 7 |
| 2024 | Learn & drop: fast learning of cnns based on layer droppingabstractAbstract This paper proposes a new method to improve the training efficiency of deep convolutional neural networks. During training, the method evaluates scores to measure how much each layer’s parameters change and whether the layer will continue learning or not. Based on these scores, the network is scaled down such that the number of parameters to be learned is reduced, yielding a speed-up in training. Unlike state-of-the-art methods that try to compress the network to be used in the inference phase or to limit the number of operations performed in the back-propagation phase, the proposed method is novel in that it focuses on reducing the number of operations performed by the network in the forward propagation during training. The proposed training strategy has been validated on two widely used architecture families: VGG and ResNet. Experiments on MNIST, CIFAR-10 and Imagenette show that, with the proposed method, the training time of the models is more than halved without significantly impacting accuracy. The FLOPs reduction in the forward propagation during training ranges from 17.83% for VGG-11 to 83.74% for ResNet-152. As for the accuracy, the impact depends on the depth of the model and the decrease is between 0.26% and 2.38% for VGGs and between 0.4 and 3.2% for ResNets. These results demonstrate the effectiveness of the proposed technique in speeding up learning of CNNs. The technique will be especially useful in applications where fine-tuning or online training of convolutional models is required, for instance because data arrive sequentially. Giorgio Cruciata, Luca Cruciata, Liliana Lo Presti, Jan C. van Gemert, Marco La Cascia |
Neural Comput. Appl. | 4 |
| 2023 | Differentiable Transportation PruningabstractDeep learning algorithms are increasingly employed at the edge. However, edge devices are resource constrained and thus require efficient deployment of deep neural networks. Pruning methods are a key tool for edge deployment as they can improve storage, compute, memory bandwidth, and energy usage. In this paper we propose a novel accurate pruning technique that allows precise control over the output network size. Our method uses an efficient optimal transportation scheme which we make end-to-end differentiable and which automatically tunes the exploration-exploitation behavior of the algorithm to find accurate sparse sub-networks. We show that our method achieves state-of-the-art performance compared to previous pruning methods on 3 different datasets, using 5 different models, across a wide range of pruning ratios, and with two types of sparsity budgets and pruning granularities. Yunqiang Li, Jan C. van Gemert, Torsten Hoefler, Bert Moons, Evangelos Eleftheriou, Bram-Ernst Verhoef |
ICCV | 2 |
| 2023 | Objects do not disappear: Video object detection by single-frame object location anticipationabstractObjects in videos are typically characterized by continuous smooth motion. We exploit continuous smooth motion in three ways. 1) Improved accuracy by using object motion as an additional source of supervision, which we obtain by anticipating object locations from a static keyframe. 2) Improved efficiency by only doing the expensive feature computations on a small subset of all frames. Because neighboring video frames are often redundant, we only compute features for a single static keyframe and predict object locations in subsequent frames. 3) Reduced annotation cost, where we only annotate the keyframe and use smooth pseudo-motion between keyframes. We demonstrate computational efficiency, annotation efficiency, and improved mean average precision compared to the state-of-the-art on four datasets: ImageNet VID, EPIC KITCHENS-55, YouTube-BoundingBoxes and Waymo Open dataset. Our source code is available at https://github.com/L-KID/Video-object-detection-by-location-anticipation. Fatemeh Karimi Nejadasl, Jan C. van Gemert, Olaf Booij, Silvia L. Pintea |
ICCV | 3 |
| 2023 | A step towards understanding why classification helps regressionabstractA number of computer vision deep regression approaches report improved results when adding a classification loss to the regression loss. Here, we explore why this is useful in practice and when it is beneficial. To do so, we start from precisely controlled dataset variations and data samplings and find that the effect of adding a classification loss is the most pronounced for regression with imbalanced data. We explain these empirical findings by formalizing the relation between the balanced and imbalanced regression losses. Finally, we show that our findings hold on two real imbalanced image datasets for depth estimation (NYUD2-DIR), and age estimation (IMDB-WIKI-DIR), and on the problem of imbalanced video progress prediction (Breakfast). Our main takeaway is: for a regression task, if the data sampling is imbalanced, then add a classification loss. Silvia L. Pintea, Yancong Lin, Jouke Dijkstra, Jan C. van Gemert |
ICCV | 4 |
| 2023 | Understanding weight-magnitude hyperparameters in training binary networks
Joris Quist, Yunqiang Li, Jan C. van Gemert |
ICLR | 3 |
| 2023 | Color Equivariant Convolutional NetworksabstractColor is a crucial visual cue readily exploited by Convolutional Neural Networks (CNNs) for object recognition. However, CNNs struggle if there is data imbalance between color variations introduced by accidental recording conditions. Color invariance addresses this issue but does so at the cost of removing all color information, which sacrifices discriminative power. In this paper, we propose Color Equivariant Convolutions (CEConvs), a novel deep learning building block that enables shape feature sharing across the color spectrum while retaining important color information. We extend the notion of equivariance from geometric to photometric transformations by incorporating parameter sharing over hue-shifts in a neural network. We demonstrate the benefits of CEConvs in terms of downstream performance to various tasks and improved robustness to color changes, including train-test distribution shifts. Our approach can be seamlessly integrated into existing architectures, such as ResNets, and offers a promising solution for addressing color-based domain shifts in CNNs. Attila Lengyel 0001, Ombretta Strafforello, Robert-Jan Bruintjes, Alexander Gielisse, Jan C. van Gemert |
NeurIPS | 5 |
| 2023 | LAB: Learnable Activation Binarizer for Binary Neural NetworksabstractBinary Neural Networks (BNNs) are receiving an up-surge of attention for bringing power-hungry deep learning towards edge devices. The traditional wisdom in this space is to employ sign(.) for binarizing feature maps. We argue and illustrate that sign(.) is a uniqueness bottleneck, limiting information propagation throughout the network. To alleviate this, we propose to dispense sign(.), replacing it with a learnable activation binarizer (LAB), allowing the network to learn a fine-grained binarization kernel per layer - as opposed to global thresholding. LAB is a novel universal module that can seamlessly be integrated into existing architectures. To confirm this, we plug it into four seminal BNNs and show a considerable accuracy boost at the cost of tolerable increase in delay and complexity. Finally, we build an end-to-end BNN (coined as LAB-BNN) around LAB, and demonstrate that it achieves competitive performance on par with the state-of-the-art on ImageNet. Our code can be found in our repository: https://github.com/sfalkena/LAB. Sieger Falkena, Hadi Jamali Rad, Jan C. van Gemert |
WACV | 3 |
| 2022 | Equal Bits: Enforcing Equally Distributed Binary Network WeightsabstractBinary networks are extremely efficient as they use only two symbols to define the network: {+1, −1}. One can make the prior distribution of these symbols a design choice. The recent IR-Net of Qin et al. argues that imposing a Bernoulli distribution with equal priors (equal bit ratios) over the binary weights leads to maximum entropy and thus minimizes information loss. However, prior work cannot precisely control the binary weight distribution during training, and therefore cannot guarantee maximum entropy. Here, we show that quantizing using optimal transport can guarantee any bit ratio, including equal ratios. We investigate experimentally that equal bit ratios are indeed preferable and show that our method leads to optimization benefits. We show that our quantization method is effective when compared to state-of-the-art binarization methods, even when using binary weight pruning. Our code is available at https://github.com/liyunqianggyn/Equal-Bits-BNN. Yunqiang Li, Silvia L. Pintea, Jan C. van Gemert |
AAAI | 3 |
| 2022 | NeRD++: Improved 3D-mirror symmetry learning from a single image
Yancong Lin, Silvia L. Pintea, Jan C. van Gemert |
BMVC | 3 |
| 2022 | Copy-Pasting Coherent Depth Regions Improves Contrastive Learning for Urban-Scene Segmentation
Liang Zeng 0005, Attila Lengyel 0001, Nergis Tömen, Jan C. van Gemert |
BMVC | 4 |
| 2022 | Deep vanishing point detection: Geometric priors make dataset variations vanishabstractDeep learning has improved vanishing point detection in images. Yet, deep networks require expensive annotated datasets trained on costly hardware and do not generalize to even slightly different domains, and minor problem variants. Here, we address these issues by injecting deep vanishing point detection networks with prior knowledge. This prior knowledge no longer needs to be learned from data, saving valuable annotation efforts and compute, unlocking realistic few-sample scenarios, and reducing the impact of domain changes. Moreover, the interpretability of the priors allows to adapt deep networks to minor problem variations such as switching between Manhattan and non-Manhattan worlds. We seamlessly incorporate two geometric priors: (i) Hough Transform – mapping image pixels to straight lines, and (ii) Gaussian sphere – mapping lines to great circles whose intersections denote vanishing points. Experimentally, we ablate our choices and show comparable accuracy to existing models in the large-data setting. We validate our model's improved data efficiency, robustness to domain changes, adaptability to non-Manhattan settings. Yancong Lin, Ruben Wiersma, Silvia L. Pintea, Klaus Hildebrandt, Elmar Eisemann, Jan C. van Gemert |
CVPR | 6 |
| 2022 | Heart rate estimation in intense exercise videosabstractEstimating heart rate from video allows non-contact health monitoring with applications in patient care, human interaction, and sports. Existing work can robustly measure heart rate under some degree of motion by face tracking. However, this is not always possible in unconstrained settings, as the face might be occluded or even outside the camera. Here, we present IntensePhysio: a challenging video heart rate estimation dataset with realistic face occlusions, severe subject motion, and ample heart rate variation. To ensure heart rate variation in a realistic setting we record each subject for around 1-2 hours. The subject is exercising (at a moderate to high intensity) on a cycling ergometer with an attached video camera and is given no instructions regarding positioning or movement. We have 11 subjects, and approximately 20 total hours of video. We show that the existing remote photo-plethysmography methods have difficulty in estimating heart rate in this setting. In addition, we present IBIS-CNN, a new baseline using spatio-temporal superpixels, which improves on existing models by eliminating the need for a visible face/face tracking. We will make the code and data publically available soon.1 Yeshwanth Napolean, Anwesh Marwade, Nergis Tömen, Puck Alkemade, Thijs Eijsvogels, Jan C. van Gemert |
ICIP | 6 |
| 2022 | Humans Disagree With the IoU for Measuring Object Detector Localization ErrorabstractThe localization quality of automatic object detectors is typically evaluated by the Intersection over Union (IoU) score. In this work, we show that humans have a different view on localization quality. To evaluate this, we conduct a survey with more than 70 participants. Results show that for localization errors with the exact same IoU score, humans might not consider that these errors are equal, and express a preference. Our work is the first to evaluate IoU with humans and makes it clear that relying on IoU scores alone to evaluate localization errors might not be sufficient. Ombretta Strafforello, Vanathi Rajasekart, Osman Semih Kayhan, Oana Inel, Jan C. van Gemert |
ICIP | 5 |
| 2022 | FlexConv: Continuous Kernel Convolutions With Differentiable Kernel Sizes
David W. Romero, Robert-Jan Bruintjes, Jakub M. Tomczak, Erik J. Bekkers, Mark Hoogendoorn, Jan C. van Gemert |
ICLR | 6 |
| 2022 | AmsterTime: A Visual Place Recognition Benchmark Dataset for Severe Domain ShiftabstractWe introduce AmsterTime: a challenging dataset to benchmark visual place recognition (VPR) in presence of a severe domain shift. AmsterTime offers a collection of 2,500 well-curated images matching the same scene from a street view matched to historical archival image data from Amsterdam city. The image pairs capture the same place with different cameras, viewpoints, and appearances. Unlike existing benchmark datasets, AmsterTime is directly crowdsourced in a GIS navigation platform (Mapillary). We evaluate various baselines, including non-learning, supervised and self-supervised methods, pre-trained on different relevant datasets, for both verification and retrieval tasks. Our result credits the best accuracy to the ResNet-101 model pre-trained on the Landmarks dataset for both verification and retrieval tasks by 84% and 24%, respectively. Additionally, a subset of Amsterdam landmarks is collected for feature evaluation in a classification task. Classification labels are further used to extract the visual explanations using Grad-CAM for inspection of the learned similar visuals in a deep metric learning models. Burak Yildiz, Seyran Khademi, Ronny Siebes, Jan C. van Gemert |
ICPR | 4 |
| 2022 | Assessment of Parkinson's Disease Severity From Videos Using Deep ArchitecturesabstractParkinson's disease (PD) diagnosis is based on clinical criteria, i.e., bradykinesia, rest tremor, rigidity, etc. Assessment of the severity of PD symptoms with clinical rating scales, however, is subject to inter-rater variability. In this paper, we propose a deep learning based automatic PD diagnosis method using videos to assist the diagnosis in clinical practices. We deploy a 3D Convolutional Neural Network (CNN) as the baseline approach for the PD severity classification and show the effectiveness. Due to the lack of data in clinical field, we explore the possibility of transfer learning from non-medical dataset and show that PD severity classification can benefit from it. To bridge the domain discrepancy between medical and non-medical datasets, we let the network focus more on the subtle temporal visual cues, i.e., the frequency of tremors, by designing a Temporal Self-Attention (TSA) mechanism. Seven tasks from the Movement Disorders Society - Unified PD rating scale (MDS-UPDRS) part III are investigated, which reveal the symptoms of bradykinesia and postural tremors. Furthermore, we propose a multi-domain learning method to predict the patient-level PD severity through task-assembling. We show the effectiveness of TSA and task-assembling method on our PD video dataset empirically. We achieve the best MCC of 0.55 on binary task-level and 0.39 on three-class patient-level classification. Zhao Yin, Victor Geraedts, Ziqi Wang 0004, Maria Contarino, Hamdi Dibeklioglu, Jan C. van Gemert |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Deep Unsupervised Image Hashing by Maximizing Bit EntropyabstractUnsupervised hashing is important for indexing huge image or video collections without having expensive annotations available. Hashing aims to learn short binary codes for compact storage and efficient semantic retrieval. We propose an unsupervised deep hashing layer called Bi-Half Net that maximizes entropy of the binary codes. Entropy is maximal when both possible values of the bit are uniformly (half-half) distributed. To maximize bit entropy, we do not add a term to the loss function as this is difficult to optimize and tune. Instead, we design a new parameter-free network layer to explicitly force continuous image features to approximate the optimal half-half bit distribution. This layer is shown to minimize a penalized term of the Wasserstein distance between the learned continuous image features and the optimal half-half bit distribution. Experimental results on the image datasets FLICKR25K, NUS-WIDE, CIFAR-10, MS COCO, MNIST and the video datasets UCF-101 and HMDB-51 show that our approach leads to compact codes and compares favorably to the current state-of-the-art. Yunqiang Li, Jan C. van Gemert |
AAAI | 2 |
| 2021 | Frequency learning for structured CNN filters with Gaussian fractional derivatives
Nikhil Saldanha, Silvia L. Pintea, Jan C. van Gemert, Nergis Tömen |
BMVC | 3 |
| 2021 | No Frame Left Behind: Full Video Action RecognitionabstractNot all video frames are equally informative for recognizing an action. It is computationally infeasible to train deep networks on all video frames when actions develop over hundreds of frames. A common heuristic is uniformly sampling a small number of video frames and using these to recognize the action. Instead, here we propose full video action recognition and consider all video frames. To make this computational tractable, we first cluster all frame activations along the temporal dimension based on their similarity with respect to the classification task, and then temporally aggregate the frames in the clusters into a smaller number of representations. Our method is end-to-end trainable and computationally efficient as it relies on temporally localized clustering in combination with fast Hamming distances in feature space. We evaluate on UCF101, HMDB51, Breakfast, and Something-Something V1 and V2, where we compare favorably to existing heuristic frame sampling methods. Silvia L. Pintea, Fatemeh Karimi Nejadasl, Olaf Booij, Jan C. van Gemert |
CVPR | 5 |
| 2021 | Zero-Shot Day-Night Domain Adaptation with a Physics PriorabstractWe explore the zero-shot setting for day-night domain adaptation. The traditional domain adaptation setting is to train on one domain and adapt to the target domain by exploiting unlabeled data samples from the test set. As gathering relevant test data is expensive and sometimes even impossible, we remove any reliance on test data imagery and instead exploit a visual inductive prior derived from physics-based reflection models for domain adaptation. We cast a number of color invariant edge detectors as trainable layers in a convolutional neural network and evaluate their robustness to illumination changes. We show that the color invariant layer reduces the day-night distribution shift in feature map activations throughout the network. We demonstrate improved performance for zero-shot day to night domain adaptation on both synthetic as well as natural datasets in various tasks, including classification, segmentation and place recognition. Attila Lengyel 0001, Sourav Garg, Michael Milford, Jan C. van Gemert |
ICCV | 4 |
| 2021 | Spectral Leakage and Rethinking the Kernel Size in CNNsabstractConvolutional layers in CNNs implement linear filters which decompose the input into different frequency bands. However, most modern architectures neglect standard principles of filter design when optimizing their model choices regarding the size and shape of the convolutional kernel. In this work, we consider the well-known problem of spectral leakage caused by windowing artifacts in filtering operations in the context of CNNs. We show that the small size of CNN kernels make them susceptible to spectral leakage, which may induce performance-degrading artifacts. To address this issue, we propose the use of larger kernel sizes along with the Hamming window function to alleviate leakage in CNN architectures. We demonstrate improved classification accuracy on multiple benchmark datasets including Fashion-MNIST, CIFAR-10, CIFAR-100 and ImageNet with the simple use of a standard window function in convolutional layers. Finally, we show that CNNs employing the Hamming window display increased robustness against various adversarial attacks. Our code is available online1. Nergis Tömen, Jan C. van Gemert |
ICCV | 2 |
| 2021 | The Arm-Swing is Discriminative in Video Gait Recognition for Athlete Re-IdentificationabstractIn this paper we evaluate running gait as an attribute for video person re-identification in a long-distance running event. We show that running gait recognition achieves competitive performance compared to appearance-based approaches in the cross-camera retrieval task and that gait and appearance features are complementary to each other. For gait, the arm swing during running is less distinguishable when using binary gait silhouettes, due to ambiguity in the torso region. We propose to use human semantic parsing to create partial gait silhouettes where the torso is left out. Leaving out the torso improves recognition results by allowing the arm swing to be more visible in the frontal and oblique viewing angles, which offers hints that arm swings are somewhat personal. Experiments show an increase of 3.2% mAP on the CampusRun and increased accuracy with 4.8% in the frontal and rear view on CASIA-B, compared to using the full body silhouettes. Yapkan Choi, Yeshwanth Napolean, Jan C. van Gemert |
ICIP | 3 |
| 2021 | Hallucination In Object Detection - A Study In Visual Part VERIFICATIONabstractWe show that object detectors can hallucinate and detect missing objects; potentially even accurately localized at their expected, but non-existing, position. This is particularly problematic for applications that rely on visual part verification: detecting if an object part is present or absent. We show how popular object detectors hallucinate objects in a visual part verification task and introduce the first visual part verification dataset: DelftBikes1, which has 10,000 bike photographs, with 22 densely annotated parts per image, where some parts may be missing. We explicitly annotated an extra object state label for each part to reflect if a part is missing or intact. We propose to evaluate visual part verification by relying on recall and compare popular object detectors on DelftBikes.1https://github.com/oskyhn/DelftBikes Osman Semih Kayhan, Bart Vredebregt, Jan C. van Gemert |
ICIP | 3 |
| 2021 | Exploiting Learned Symmetries in Group Equivariant ConvolutionsabstractGroup Equivariant Convolutions (GConvs) enable convolutional neural networks to be equivariant to various transformation groups, but at an additional parameter and compute cost. We investigate the filter parameters learned by GConvs and find certain conditions under which they become highly redundant. We show that GConvs can be efficiently decomposed into depthwise separable convolutions while preserving equivariance properties and demonstrate improved performance and data efficiency on two datasets. All code is publicly available at github.com/Attila94/SepGrouPy. Attila Lengyel 0001, Jan C. van Gemert |
ICIP | 2 |
| 2021 | Semi-Supervised Lane Detection With Deep Hough TransformabstractCurrent work on lane detection relies on large manually annotated datasets. We reduce the dependency on annotations by leveraging massive cheaply available unlabelled data. We propose a novel loss function exploiting geometric knowledge of lanes in Hough space, where a lane can be identified as a local maximum. By splitting lanes into separate channels, we can localize each lane via simple global max-pooling. The location of the maximum encodes the layout of a lane, while the intensity indicates the the probability of a lane being present. Maximizing the log-probability of the maximal bins helps neural networks find lanes without labels. On the CULane and TuSimple datasets, we show that the proposed Hough Transform loss improves performance significantly by learning from large amounts of unlabelled images. Yancong Lin, Silvia L. Pintea, Jan C. van Gemert |
ICIP | 3 |
| 2021 | PUNet: Temporal Action Proposal Generation With Positive Unlabeled Learning Using Key Frame AnnotationsabstractPopular approaches to classifying action segments in long, realistic, untrimmed videos start with high quality action proposals. Current action proposal methods based on deep learning are trained on labeled video segments. Obtaining annotated segments for untrimmed videos is time consuming, expensive and error-prone as annotated temporal action boundaries are imprecise, subjective and inconsistent. By embracing this uncertainty we explore to significantly speed up temporal annotations by using just a single key frame label for each action instance instead of the inherently imprecise start and end frames. To tackle the class imbalance by using only a single frame, we evaluate an extremely simple Positive-Unlabeled algorithm (PU-learning). We demonstrate on THUMOS’14 and ActivityNet that using a single key frame label give good results while being significantly faster to annotate. In addition, we show that our simple method, PUNet1, is data-efficient which further reduces the need for expensive annotations.1https://github.com/NoorZia/punet Noor Ul Sehr Zia, Osman Semih Kayhan, Jan C. van Gemert |
ICIP | 3 |
| 2021 | Deep Continuous NetworksabstractCNNs and computational models of biological vision share some fundamental principles, which opened new avenues of research. However, fruitful cross-field research is hampered by conventional CNN architectures being based on spatially and depthwise discrete representations, which cannot accommodate certain aspects of biological complexity such as continuously varying receptive field sizes and dynamics of neuronal responses. Here we propose deep continuous networks (DCNs), which combine spatially continuous filters, with the continuous depth framework of neural ODEs. This allows us to learn the spatial support of the filters during training, as well as model the continuous evolution of feature maps, linking DCNs closely to biological models. We show that DCNs are versatile and highly applicable to standard image classification and reconstruction problems, where they improve parameter and data efficiency, and allow for meta-parametrization. We illustrate the biological plausibility of the scale distributions learned by DCNs and explore their performance in a neuroscientifically inspired pattern completion task. Finally, we investigate an efficient implementation of DCNs by changing input contrast. Nergis Tömen, Silvia L. Pintea, Jan C. van Gemert |
ICML | 3 |
| 2021 | Resolution Learning in Deep Convolutional Networks Using Scale-Space TheoryabstractResolution in deep convolutional neural networks (CNNs) is typically bounded by the receptive field size through filter sizes, and subsampling layers or strided convolutions on feature maps. The optimal resolution may vary significantly depending on the dataset. Modern CNNs hard-code their resolution hyper-parameters in the network architecture which makes tuning such hyper-parameters cumbersome. We propose to do away with hard-coded resolution hyper-parameters and aim to learn the appropriate resolution from data. We use scale-space theory to obtain a self-similar parametrization of filters and make use of the N-Jet: a truncated Taylor series to approximate a filter by a learned combination of Gaussian derivative filters. The parameter σ of the Gaussian basis controls both the amount of detail the filter encodes and the spatial extent of the filter. Since σ is a continuous parameter, we can optimize it with respect to the loss. The proposed N-Jet layer achieves comparable performance when used in state-of-the art architectures, while learning the correct resolution in each layer automatically. We evaluate our N-Jet layer on both classification and segmentation, and we show that learning σ is especially beneficial when dealing with inputs at multiple sizes. Silvia L. Pintea, Nergis Tömen, Stanley F. Goes, Marco Loog, Jan C. van Gemert |
IEEE Trans. Image Process. | 5 |
| 2020 | Black Magic in Deep Learning: How Human Skill Impacts Network Training
Kanav Anand, Ziqi Wang 0004, Marco Loog, Jan C. van Gemert |
BMVC | 4 |
| 2020 | On Translation Invariance in CNNs: Convolutional Layers Can Exploit Absolute Spatial LocationabstractIn this paper we challenge the common assumption that convolutional layers in modern CNNs are translation invariant. We show that CNNs can and will exploit the absolute spatial location by learning filters that respond exclusively to particular absolute locations by exploiting image boundary effects. Because modern CNNs filters have a huge receptive field, these boundary effects operate even far from the image boundary, allowing the network to exploit absolute spatial location all over the image. We give a simple solution to remove spatial location encoding which improves translation invariance and thus gives a stronger visual inductive bias which particularly benefits small data sets. We broadly demonstrate these benefits on several architectures and various applications such as image classification, patch matching, and two video classification datasets. Osman Semih Kayhan, Jan C. van Gemert |
CVPR | 2 |
| 2020 | Deep Hough-Transform Line Priors
Yancong Lin, Silvia L. Pintea, Jan C. van Gemert |
ECCV (22) | 3 |
| 2020 | Tilting at windmills: Data augmentation for deep pose estimation does not help with occlusionsabstractOcclusion degrades the performance of human pose estimation. In this paper, we introduce targeted keypoint and body part occlusion attacks. The effects of the attacks are systematically analyzed on the best performing methods. In addition, we propose occlusion specific data augmentation techniques against keypoint and part attacks. Our extensive experiments show that human pose estimation methods are not robust to occlusion and data augmentation does not solve the occlusion problems. Rafal Pytel, Osman Semih Kayhan, Jan C. van Gemert |
ICPR | 3 |
| 2020 | Zoom-CAM: Generating Fine-grained Pixel Annotations from Image LabelsabstractCurrent weakly supervised object localization and segmentation rely on class-discriminative visualization techniques to generate pseudo-labels for pixel-level training. Such visualization methods, including class activation mapping (CAM) and Grad-CAM, use only the deepest, lowest resolution convolutional layer, missing all information in intermediate layers. We propose Zoom-CAM: going beyond the last lowest resolution layer by integrating the importance maps over all activations in intermediate layers. Zoom-CAM captures fine-grained small-scale objects for various discriminative class instances, which are commonly missed by the baseline visualization methods. We focus on generating pixel-level pseudo-labels from class labels. The quality of our pseudo-labels evaluated on the ImageNet localization task exhibits more than 2.8% improvement on top-1 error. For weakly supervised semantic segmentation our generated pseudo-labels improve a state of the art model by 1.1%. Xiangwei Shi, Seyran Khademi, Yunqiang Li, Jan C. van Gemert |
ICPR | 4 |
| 2020 | WeightAlign: Normalizing Activations by Weight AlignmentabstractBatch normalization (BN) allows training very deep networks by normalizing activations by mini-batch sample statistics which renders BN unstable for small batch sizes. Current small-batch solutions such as Instance Norm, Layer Norm, and Group Norm use channel statistics which can be computed even for a single sample. Such methods are less stable than BN as they critically depend on the statistics of a single input sample. To address this problem, we propose a normalization of activation without sample statistics. We present WeightAlign: a method that normalizes the weights by the mean and scaled standard derivation computed within a filter, which normalizes activations without computing any sample statistics. Our proposed method is independent of batch size and stable over a wide range of batch sizes. Because weight statistics are orthogonal to sample statistics, we can directly combine WeightAlign with any method for activation normalization. We experimentally demonstrate these benefits for classification on CIFAR-10, CIFAR-100, ImageNet, for semantic segmentation on PASCAL VOC 2012 and for domain adaptation on Office-31. Xiangwei Shi, Yunqiang Li, Jan C. van Gemert |
ICPR | 4 |
| 2020 | Respecting Domain Relations: Hypothesis Invariance for Domain GeneralizationabstractIn domain generalization, multiple labeled nonindependent and non-identically distributed source domains are available during training while neither the data nor the labels of target domains are. Currently, learning so-called domain invariant representations (DIRs) is the prevalent approach to domain generalization. In this work, we define DIRs employed by existing works in probabilistic terms and show that by learning DIRs, overly strict requirements are imposed concerning the invariance. Particularly, DIRs aim to perfectly align representations of different domains, i.e. their input distributions. This is, however, not necessary for good generalization to a target domain and may even dispose of valuable classification information. We propose to learn so-called hypothesis invariant representations (HIRs), which relax the invariance assumptions by merely aligning posteriors, instead of aligning representations. We report experimental results on public domain generalization datasets to show that learning HIRs is more effective than learning DIRs. In fact, our approach can even compete with approaches using prior knowledge about domains. Ziqi Wang 0004, Marco Loog, Jan C. van Gemert |
ICPR | 3 |
| 2019 | Push for Quantization: Deep Fisher Hashing
Yunqiang Li, Wenjie Pei, Yufei Zha, Jan C. van Gemert |
BMVC | 4 |
| 2019 | Divide and Count: Generic Object Counting by Image DivisionsabstractWe propose a general object counting method that does not use any prior category information. We learn from local image divisions to predict global image-level counts without using any form of local annotations. Our method separates the input image into a sets of image divisions - each fully covering the image. Each image division is composed of a set of region proposals or uniform grid cells. Our approach learns in an endto- end deep learning architecture to predict global image-level counts from local image divisions. The method incorporates a counting layer which predicts object counts in the complete image, by enforcing consistency in counts when dealing with overlapping image regions. Our counting layer is based on the inclusion-exclusion principle from set theory. We analyze the individual building blocks of our proposed approach on Pascal- VOC2007 and evaluate our method on the MS-COCO large scale generic object dataset as well as on three class-specific counting datasets: UCSD pedestrian dataset, and CARPK and PUCPR+ car datasets. Tobias Stahl, Silvia L. Pintea, Jan C. van Gemert |
IEEE Trans. Image Process. | 3 |
| 2018 | Sight-Seeing in the Eyes of Deep Neural NetworksabstractWe address the interpretability of convolutional neural networks (CNNs) for predicting a geo-location from an image. In a pilot experiment we classify images of Pittsburgh vs Tokyo and visualize the learned CNN filters. We found that varying the CNN architecture leads to variating in the visualized filters. This calls for further investigation of the effective parameters on the interpretability of CNNs. Seyran Khademi, Xiangwei Shi, Tino Mager, Ronny Siebes, Carola Hein, Victor de Boer, Jan C. van Gemert |
eScience | 7 |
| 2018 | Recurrent Knowledge DistillationabstractKnowledge distillation compacts deep networks by letting a small student network learn from a large teacher network. The accuracy of knowledge distillation recently benefited from adding residual layers. We propose to reduce the size of the student network even further by recasting multiple residual layers in the teacher network into a single recurrent student layer. We propose three variants of adding recurrent connections into the student network, and show experimentally on CIFAR-10, Scenes and MiniPlaces, that we can reduce the number of parameters at little loss in accuracy. Silvia L. Pintea, Jan C. van Gemert |
ICIP | 3 |
| 2018 | Asymmetric kernel in Gaussian Processes for learning target variance
Silvia L. Pintea, Jan C. van Gemert, Arnold W. M. Smeulders |
Pattern Recognit. Lett. | 2 |
| 2018 | DeepEyes: Progressive Visual Analytics for Designing Deep Neural NetworksabstractDeep neural networks are now rivaling human accuracy in several pattern recognition problems. Compared to traditional classifiers, where features are handcrafted, neural networks learn increasingly complex features directly from the data. Instead of handcrafting the features, it is now the network architecture that is manually engineered. The network architecture parameters such as the number of layers or the number of filters per layer and their interconnections are essential for good performance. Even though basic design guidelines exist, designing a neural network is an iterative trial-and-error process that takes days or even weeks to perform due to the large datasets used for training. In this paper, we present DeepEyes, a Progressive Visual Analytics system that supports the design of neural networks during training. We present novel visualizations, supporting the identification of layers that learned a stable set of patterns and, therefore, are of interest for a detailed analysis. The system facilitates the identification of problems, such as superfluous filters or layers, and information that is not being captured by the network. We demonstrate the effectiveness of our system through multiple use cases, showing how a trained network can be compressed, reshaped and adapted to different problems. Nicola Pezzotti, Thomas Höllt, Jan C. van Gemert, Boudewijn P. F. Lelieveldt, Elmar Eisemann, Anna Vilanova |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2017 | Object-Extent Pooling for Weakly Supervised Single-Shot Localization
Amogh Gudi, Nicolai van Rosmalen, Marco Loog, Jan C. van Gemert |
BMVC | 4 |
| 2017 | Video Acceleration MagnificationabstractThe ability to amplify or reduce subtle image changes over time is useful in contexts such as video editing, medical video analysis, product quality control and sports. In these contexts there is often large motion present which severely distorts current video amplification methods that magnify change linearly. In this work we propose a method to cope with large motions while still magnifying small changes. We make the following two observations: i) large motions are linear on the temporal scale of the small changes, ii) small changes deviate from this linearity. We ignore linear motion and propose to magnify acceleration. Our method is pure Eulerian and does not require any optical flow, temporal alignment or region annotations. We link temporal second-order derivative filtering to spatial acceleration magnification. We apply our method to moving objects where we show motion magnification and color magnification. We provide quantitative as well as qualitative evidence for our method while comparing to the state-of-the-art. Silvia L. Pintea, Jan C. van Gemert |
CVPR | 3 |
| 2017 | Active Decision Boundary Annotation with Deep Generative ModelsabstractThis paper is on active learning where the goal is to reduce the data annotation burden by interacting with a (human) oracle during training. Standard active learning methods ask the oracle to annotate data samples. Instead, we take a profoundly different approach: we ask for annotations of the decision boundary. We achieve this using a deep generative model to create novel instances along a 1d line. A point on the decision boundary is revealed where the instances change class. Experimentally we show on three data sets that our method can be plugged into other active learning schemes, that human oracles can effectively annotate points on the decision boundary, that our method is robust to annotation noise, and that decision boundary annotations improve over annotating data samples. Miriam W. Huijser, Jan C. van Gemert |
ICCV | 2 |
| 2017 | Tubelets: Unsupervised Action Proposals from Spatiotemporal Super-VoxelsabstractThis paper considers the problem of localizing actions in videos as sequences of bounding boxes. The objective is to generate action proposals that are likely to include the action of interest, ideally achieving high recall with few proposals. Our contributions are threefold. First, inspired by selective search for object proposals, we introduce an approach to generate action proposals from spatiotemporal super-voxels in an unsupervised manner, we call them Tubelets . Second, along with the static features from individual frames our approach advantageously exploits motion. We introduce independent motion evidence as a feature to characterize how the action deviates from the background and explicitly incorporate such motion information in various stages of the proposal generation. Finally, we introduce spatiotemporal refinement of Tubelets, for more precise localization of actions, and pruning to keep the number of Tubelets limited. We demonstrate the suitability of our approach by extensive experiments for action proposal quality and action localization on three public datasets: UCF Sports, MSR-II and UCF101. For action proposal quality, our unsupervised proposals beat all other existing approaches on the three datasets. For action localization, we show top performance on both the trimmed videos of UCF Sports and UCF101 as well as the untrimmed videos of MSR-II. Mihir Jain, Jan C. van Gemert, Hervé Jégou, Patrick Bouthemy, Cees Snoek |
Int. J. Comput. Vis. | 2 |
| 2017 | Con-Text: Text Detection for Fine-Grained Object ClassificationabstractThis paper focuses on fine-grained object classification using recognized scene text in natural images. While the state-of-the-art relies on visual cues only, this paper is the first work which proposes to combine textual and visual cues. Another novelty is the textual cue extraction. Unlike the state-of-the-art text detection methods, we focus more on the background instead of text regions. Once text regions are detected, they are further processed by two methods to perform text recognition, i.e., ABBYY commercial OCR engine and a state-of-the-art character recognition algorithm. Then, to perform textual cue encoding, bi- and trigrams are formed between the recognized characters by considering the proposed spatial pairwise constraints. Finally, extracted visual and textual cues are combined for fine-grained classification. The proposed method is validated on four publicly available data sets: ICDAR03, ICDAR13, Con-Text, and Flickr-logo. We improve the state-of-the-art end-to-end character recognition by a large margin of 15% on ICDAR03. We show that textual cues are useful in addition to visual cues for fine-grained classification. We show that textual cues are also useful for logo retrieval. Adding textual cues outperforms visual- and textual-only in fine-grained classification (70.7% to 60.3%) and logo retrieval (57.4% to 54.8%). Sezer Karaoglu, Ran Tao 0004, Jan C. van Gemert, Theo Gevers |
IEEE Trans. Image Process. | 3 |
| 2016 | Structured Receptive Fields in CNNsabstractLearning powerful feature representations with CNNs is hard when training data are limited. Pre-training is one way to overcome this, but it requires large datasets sufficiently similar to the target domain. Another option is to design priors into the model, which can range from tuned hyperparameters to fully engineered representations like Scattering Networks. We combine these ideas into structured receptive field networks, a model which has a fixed filter basis and yet retains the flexibility of CNNs. This flexibility is achieved by expressing receptive fields in CNNs as a weighted sum over a fixed basis which is similar in spirit to Scattering Networks. The key difference is that we learn arbitrary effective filter sets from the basis rather than modeling the filters. This approach explicitly connects classical multiscale image analysis with general CNNs. With structured receptive field networks, we improve considerably over unstructured CNNs for small and medium dataset scenarios as well as over Scattering for large datasets. We validate our findings on ILSVRC2012, Cifar-10, Cifar-100 and MNIST. As a realistic small dataset example, we show state-of-the-art classification results on popular 3D MRI brain-disease datasets where pre-training is difficult due to a lack of large public datasets in a similar domain. Jörn-Henrik Jacobsen, Jan C. van Gemert, Zhongyu Lou, Arnold W. M. Smeulders |
CVPR | 2 |
| 2016 | Depth-Aware Motion Magnification
Julian F. P. Kooij, Jan C. van Gemert |
ECCV (8) | 2 |
| 2016 | Spot On: Action Localization from Pointly-Supervised Proposals
Pascal Mettes, Jan C. van Gemert, Cees Snoek |
ECCV (5) | 2 |
| 2016 | Featureless: Bypassing feature extraction in action categorizationabstractThis method introduces an efficient manner of learning action categories without the need of feature estimation. The approach starts from low-level values, in a similar style to the successful CNN methods. However, rather than extracting general image features, we learn to predict specific video representations from raw video data. The benefit of such an approach is that at the same computational expense it can predict 2D video representations as well as 3D ones, based on motion. The proposed model relies on discriminative Wald-boost, which we enhance to a multiclass formulation for the purpose of learning video representations. The suitability of the proposed approach as well as its time efficiency are tested on the UCF11 action recognition dataset. Silvia L. Pintea, Pascal Mettes, Jan C. van Gemert, Arnold W. M. Smeulders |
ICIP | 3 |
| 2016 | Exploring the Long Tail of Social Media Tags
Svetlana Kordumova, Jan C. van Gemert, Cees Snoek |
MMM (1) | 2 |
| 2016 | No spare parts: Sharing part detectors for image categorization
Pascal Mettes, Jan C. van Gemert, Cees Snoek |
Comput. Vis. Image Underst. | 2 |
| 2016 | Large scale Gaussian Process for overlap-based object proposal scoring
Silvia L. Pintea, Sezer Karaoglu, Jan C. van Gemert, Arnold W. M. Smeulders |
Comput. Vis. Image Underst. | 3 |
| 2015 | APT: Action localization proposals from dense trajectoriesabstractThis paper is on action localization in video with the aid of spatio-temporal proposals. To alleviate the computational expensive segmentation step of existing proposals, we propose bypassing the segmentations completely by generating proposals directly from the dense trajectories used to represent videos during classification. Our Action localization Proposals from dense Trajectories (APT) use an efficient proposal generation algorithm to handle the high number of trajectories in a video. Our spatio-temporal proposals are faster than current methods and outperform the localization and classification accuracy of current proposals on the UCF Sports, UCF 101, and MSR-II video datasets. Corrected version: we fixed a mistake in our UCF-101 ground truth. Numbers are different; conclusions are unchanged Jan C. van Gemert, Mihir Jain, Ella Gati, Cees Snoek |
BMVC | 1 |
| 2015 | What do 15, 000 object categories tell us about classifying and localizing actions?abstractThis paper contributes to automatic classification and localization of human actions in video. Whereas motion is the key ingredient in modern approaches, we assess the benefits of having objects in the video representation. Rather than considering a handful of carefully selected and localized objects, we conduct an empirical study on the benefit of encoding 15,000 object categories for action using 6 datasets totaling more than 200 hours of video and covering 180 action classes. Our key contributions are i) the first in-depth study of encoding objects for actions, ii) we show that objects matter for actions, and are often semantically relevant as well. iii) We establish that actions have object preferences. Rather than using all objects, selection is advantageous for action recognition. iv)We reveal that object-action relations are generic, which allows to transferring these relationships from the one domain to the other. And, v) objects, when combined with motion, improve the state-of-the-art for both action classification and localization. Mihir Jain, Jan C. van Gemert, Cees Snoek |
CVPR | 2 |
| 2015 | Objects2action: Classifying and Localizing Actions without Any Video ExampleabstractThe goal of this paper is to recognize actions in video without the need for examples. Different from traditional zero-shot approaches we do not demand the design and specification of attribute classifiers and class-to-attribute mappings to allow for transfer from seen classes to unseen classes. Our key contribution is objects2action, a semantic word embedding that is spanned by a skip-gram model of thousands of object categories. Action labels are assigned to an object encoding of unseen video based on a convex combination of action and object affinities. Our semantic embedding has three main characteristics to accommodate for the specifics of actions. First, we propose a mechanism to exploit multiple-word descriptions of actions and objects. Second, we incorporate the automated selection of the most responsive objects per action. And finally, we demonstrate how to extend our zero-shot approach to the spatio-temporal localization of actions in video. Experiments on four action datasets demonstrate the potential of our approach. Mihir Jain, Jan C. van Gemert, Thomas Mensink, Cees Snoek |
ICCV | 2 |
| 2015 | Per-patch metric learning for robust image matchingabstractWe propose a patch-specific metric learning method to improve matching performance of local descriptors. Existing methodologies typically focus on invariance, by completely considering, or completely disregarding all variations. We propose a metric learning method that is robust to only a range of variations. The ability to choose the level of robustness allows us to fine-tune the trade-off between invariance and discriminative power. We learn a distance metric for each patch independently by sampling from a set of relevant image transformations. These transformations give a-priori knowledge about the behavior of the query patch under the applied transformation in feature space. We learn the robust metric by either fully generating only the relevant range of transformations, or by a novel direct metric. The matching between query patch and data is performed with this new metric. Results on the ALOI dataset show that the proposed method improves performance of SIFT by 6.22% for geometric and 4.43% for photometric transformations. Sezer Karaoglu, Ivo Everts, Jan C. van Gemert, Theo Gevers |
ICIP | 3 |
| 2015 | Bag-of-Fragments: Selecting and Encoding Video Fragments for Event Detection and RecountingabstractThe goal of this paper is event detection and recounting using a representation of concept detector scores. Different from existing work, which encodes videos by averaging concept scores over all frames, we propose to encode videos using fragments that are discriminatively learned per event. Our bag-of-fragments split a video into semantically coherent fragment proposals. From training video proposals we show how to select the most discriminative fragment for an event. An encoding of a video is in turn generated by matching and pooling these discriminative fragments to the fragment proposals of the video. The bag-of-fragments forms an effective encoding for event detection and is able to provide a precise temporally localized event recounting. Furthermore, we show how bag-of-fragments can be extended to deal with irrelevant concepts in the event recounting. Experiments on challenging web videos show that i) our modest number of fragment proposals give a high sub-event recall, ii) bag-of-fragments is complementary to global averaging and provides better event detection, iii) bag-of-fragments with concept filtering yields a desirable event recounting. We conclude that fragments matter for video event detection and recounting. Pascal Mettes, Jan C. van Gemert, Spencer Cappallo, Thomas Mensink, Cees Snoek |
ICMR | 2 |
| 2014 | Action Localization with Tubelets from MotionabstractThis paper considers the problem of action localization, where the objective is to determine when and where certain actions appear. We introduce a sampling strategy to produce 2D+t sequences of bounding boxes, called tubelets. Compared to state-of-the-art alternatives, this drastically reduces the number of hypotheses that are likely to include the action of interest. Our method is inspired by a recent technique introduced in the context of image localization. Beyond considering this technique for the first time for videos, we revisit this strategy for 2D+t sequences obtained from super-voxels. Our sampling strategy advantageously exploits a criterion that reflects how action related motion deviates from background motion. We demonstrate the interest of our approach by extensive experiments on two public datasets: UCF Sports and MSR-II. Our approach significantly outperforms the state-of-the-art on both datasets, while restricting the search of actions to a fraction of possible bounding box sequences. Mihir Jain, Jan C. van Gemert, Hervé Jégou, Patrick Bouthemy, Cees Snoek |
CVPR | 2 |
| 2014 | Déjà Vu: - Motion Prediction in Static Images
Silvia L. Pintea, Jan C. van Gemert, Arnold W. M. Smeulders |
ECCV (3) | 2 |
| 2014 | The Rijksmuseum Challenge: Museum-Centered Visual RecognitionabstractThis paper offers a challenge for visual classification and content-based retrieval of artistic content. The challenge is posed from a museum-centric point of view offering a wide range of object types including paintings, photographs, ceramics, furniture, etc. The freely available dataset consists of 112,039 photographic reproductions of the artworks exhibited in the Rijksmuseum in Amsterdam, the Netherlands. We offer four automatic visual recognition challenges consisting of predicting the artist, type, material and creation year. We include a set of baseline results, and make available state-of-the-art image features encoded with the Fisher vector. Progress on this challenge improves the tools of a museum curator while improving content-based exploration by online visitors of the museum collection. Thomas Mensink, Jan C. van Gemert |
ICMR | 2 |
| 2014 | Evaluation of Color Spatio-Temporal Interest Points for Human Action RecognitionabstractThis paper considers the recognition of realistic human actions in videos based on spatio-temporal interest points (STIPs). Existing STIP-based action recognition approaches operate on intensity representations of the image data. Because of this, these approaches are sensitive to disturbing photometric phenomena, such as shadows and highlights. In addition, valuable information is neglected by discarding chromaticity from the photometric representation. These issues are addressed by color STIPs. Color STIPs are multichannel reformulations of STIP detectors and descriptors, for which we consider a number of chromatic and invariant representations derived from the opponent color space. Color STIPs are shown to outperform their intensity-based counterparts on the challenging UCF sports, UCF11 and UCF50 action recognition benchmarks by more than 5% on average, where most of the gain is due to the multichannel descriptors. In addition, the results show that color STIPs are currently the single best low-level feature choice for STIP-based approaches to human action recognition. Ivo Everts, Jan C. van Gemert, Theo Gevers |
IEEE Trans. Image Process. | 2 |
| 2014 | Robustifying Descriptor Instability Using Fisher VectorsabstractMany computer vision applications, including image classification, matching, and retrieval use global image representations, such as the Fisher vector, to encode a set of local image patches. To describe these patches, many local descriptors have been designed to be robust against lighting changes and noise. However, local image descriptors are unstable when the underlying image signal is low. Such low-signal patches are sensitive to small image perturbations, which might come e.g., from camera noise or lighting effects. In this paper, we first quantify the relation between the signal strength of a patch and the instability of that patch, and second, we extend the standard Fisher vector framework to explicitly take the descriptor instabilities into account. In comparison to common approaches to dealing with descriptor instabilities, our results show that modeling local descriptor instability is beneficial for object matching, image retrieval, and classification. Ivo Everts, Jan C. van Gemert, Thomas Mensink, Theo Gevers |
IEEE Trans. Image Process. | 2 |
| 2013 | Evaluation of Color STIPs for Human Action RecognitionabstractThis paper is concerned with recognizing realistic human actions in videos based on spatio-temporal interest points (STIPs). Existing STIP-based action recognition approaches operate on intensity representations of the image data. Because of this, these approaches are sensitive to disturbing photometric phenomena such as highlights and shadows. Moreover, valuable information is neglected by discarding chromaticity from the photometric representation. These issues are addressed by Color STIPs. Color STIPs are multi-channel reformulations of existing intensity-based STIP detectors and descriptors, for which we consider a number of chromatic representations derived from the opponent color space. This enhanced modeling of appearance improves the quality of subsequent STIP detection and description. Color STIPs are shown to substantially outperform their intensity-based counterparts on the challenging UCF~sports, UCF11 and UCF50 action recognition benchmarks. Moreover, the results show that color STIPs are currently the single best low-level feature choice for STIP-based approaches to human action recognition. Ivo Everts, Jan C. van Gemert, Theo Gevers |
CVPR | 2 |
| 2013 | Automatic Egyptian hieroglyph recognition by retrieving images as textsabstractIn this paper we propose an approach for automatically recognizing ancient Egyptian hieroglyph from photographs. To this end we first manually annotated and segmented a large collection of nearly 4,000 hieroglyphs. In our automatic approach we localize and segment each individual hieroglyph, determine the reading order and subsequently evaluate 5 visual descriptors in 3 different matching schemes to evaluate visual hieroglyph recognition. In addition to visual-only cues, we use a corpus of Egyptian texts to learn language models that help re-rank the visual output. Morris Franken, Jan C. van Gemert |
ACM Multimedia | 2 |
| 2013 | Con-text: text detection using background connectivity for fine-grained object classificationabstractThis paper focuses on fine-grained classification by detecting photographed text in images. We introduce a text detection method that does not try to detect all possible foreground text regions but instead aims to reconstruct the scene background to eliminate non-text regions. Object cues such as color, contrast, and objectiveness are used in corporation with a random forest classifier to detect background pixels in the scene. Results on two publicly available datasets ICDAR03 and a fine-grained Building subcategories of ImageNet shows the effectiveness of the proposed method. Sezer Karaoglu, Jan C. van Gemert, Theo Gevers |
ACM Multimedia | 2 |
| 2013 | Spot the differences: from a photograph burst to the single best pictureabstractWith the rise of the digital camera, people nowadays typically take several near-identical photos of the same scene to maximize the chances of a good shot. This paper proposes a user-friendly tool for exploring a personal photo gallery for selecting or even creating the best shot of a scene between its multiple alternatives. This functionality is realized through a graphical user interface where the best viewpoint can be selected from a generated panorama of the scene. Once the viewpoint is selected, the user is able to go explore possible alternatives coming from the other images. Using this tool, one can explore a photo gallery efficiently. Moreover, additional compositions from other images are also possible. With such additional compositions, one can go from a burst of photographs to the single best one. Even funny compositions of images, where you can duplicate a person in the same image, are possible with our proposed tool. H. Emrah Tasli, Jan C. van Gemert, Theo Gevers |
ACM Multimedia | 2 |
| 2012 | Per-patch Descriptor Selection Using Surface and Scene Properties
Ivo Everts, Jan C. van Gemert, Theo Gevers |
ECCV (6) | 2 |
| 2011 | Exploiting photographic style for category-level image classification by generalizing the spatial pyramidabstractThis paper investigates the use of photographic style for category-level image classification. Specifically, we exploit the assumption that images within a category share a similar style defined by attributes such as colorfulness, lighting, depth of field, viewpoint and saliency. For these style attributes we create correspondences across images by a generalized spatial pyramid matching scheme. Where the spatial pyramid groups features spatially, we allow more general feature grouping and in this paper we focus on grouping images on photographic style. We evaluate our approach in an object classification task and investigate style differences between professional and amateur photographs. We show that a generalized pyramid with style-based attributes improves performance on the professional Corel and amateur Pascal VOC 2009 image datasets. Jan C. van Gemert |
ICMR | 1 |
| 2010 | Comparing compact codebooks for visual categorization
Jan C. van Gemert, Cees Snoek, Cor J. Veenman, Arnold W. M. Smeulders, Jan-Mark Geusebroek |
Comput. Vis. Image Underst. | 1 |
| 2010 | Visual Word AmbiguityabstractThis paper studies automatic image classification by modeling soft assignment in the popular codebook model. The codebook model describes an image as a bag of discrete visual words selected from a vocabulary, where the frequency distributions of visual words in an image allow classification. One inherent component of the codebook model is the assignment of discrete visual words to continuous image features. Despite the clear mismatch of this hard assignment with the nature of continuous features, the approach has been successfully applied for some years. In this paper, we investigate four types of soft assignment of visual words to image features. We demonstrate that explicitly modeling visual word assignment ambiguity improves classification performance compared to the hard assignment of the traditional codebook model. The traditional codebook model is compared against our method for five well-known data sets: 15 natural scenes, Caltech-101, Caltech-256, and Pascal VOC 2007/2008. We demonstrate that large codebook vocabulary sizes completely deteriorate the performance of the traditional model, whereas the proposed model performs consistently. Moreover, we show that our method profits in high-dimensional feature spaces and reaps higher benefits when increasing the number of image categories. Jan C. van Gemert, Cor J. Veenman, Arnold W. M. Smeulders, Jan-Mark Geusebroek |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Episode-Constrained Cross-Validation in Video Concept RetrievalabstractWhereas video tells a narrative by a composition of shots, current video retrieval methods focus mainly on single shots. In retrieval performance estimation, similar shots in a narrative may result in performance overestimation. We propose an episode-based version of cross-validation leading up to 14% classification improvement over shot-based cross-validation. Jan C. van Gemert, Cor J. Veenman, Jan-Mark Geusebroek |
IEEE Trans. Multim. | 1 |
| 2008 | Kernel Codebooks for Scene Categorization
Jan C. van Gemert, Jan-Mark Geusebroek, Cor J. Veenman, Arnold W. M. Smeulders |
ECCV (3) | 1 |
| 2008 | Emotional valence categorization using holistic image featuresabstractCan a machine learn to perceive emotions as evoked by an artwork? Here we propose an emotion categorization system, trained by ground truth from psychology studies. The training data contains emotional valences scored by human subjects on the International Affective Picture System (IAPS), a standard emotion evoking image set in psychology. Our approach is based on the assessment of local image statistics which are learned per emotional category using support vector machines. We show results for our system on the I APS dataset, and for a collection of masterpieces. Although the results are preliminary, they demonstrate the potential of machines to elicit realistic emotions when considering masterpieces. Victoria Yanulevskaya, Jan C. van Gemert, Katharina Roth, Ann-Katrin Herbold, Nicu Sebe, Jan-Mark Geusebroek |
ICIP | 2 |
| 2006 | The influence of cross-validation on video classification performanceabstractDigital video is sequential in nature. When video data is used in a semantic concept classification task, the episodes are usually summarized with shots. The shots are annotated as containing, or not containing, a certain concept resulting in a labeled dataset. These labeled shots can subsequently be used by supervised learning methods (classifiers) where they are trained to predict the absence or presence of the concept in unseen shots and episodes. The performance of such automatic classification systems is usually estimated with cross-validation. By taking random samples from the dataset for training and testing as such, part of the shots from an episode are in the training set and another part from the same episode is in the test set. Accordingly, data dependence between training and test set is introduced, resulting in too optimistic performance estimates. In this paper, we experimentally show this bias, and propose how this bias can be prevented using episode-constrained crossvalidation. Moreover, we show that a 17% higher classifier performance can be achieved by using episode constrained cross-validation for classifier parameter tuning. Jan C. van Gemert, Cees Snoek, Cor J. Veenman, Arnold W. M. Smeulders |
ACM Multimedia | 1 |
| 2006 | The challenge problem for automated detection of 101 semantic concepts in multimediaabstractWe introduce the challenge problem for generic video indexing to gain insight in intermediate steps that affect performance of multimedia analysis methods, while at the same time fostering repeatability of experiments. To arrive at a challenge problem, we provide a general scheme for the systematic examination of automated concept detection methods, by decomposing the generic video indexing problem into 2 unimodal analysis experiments, 2 multimodal analysis experiments, and 1 combined analysis experiment. For each experiment, we evaluate generic video indexing performance on 85 hours of international broadcast news data, from the TRECVID 2005/2006 benchmark, using a lexicon of 101 semantic concepts. By establishing a minimum performance on each experiment, the challenge problem allows for component-based optimization of the generic indexing issue, while simultaneously offering other researchers a reference for comparison during indexing methodology development. To stimulate further investigations in intermediate analysis steps that inuence video indexing performance, the challenge offers to the research community a manually annotated concept lexicon, pre-computed low-level multimedia features, trained classifier models, and five experiments together with baseline performance, which are all available at http://www.mediamill.nl/challenge/. Cees Snoek, Marcel Worring, Jan C. van Gemert, Jan-Mark Geusebroek, Arnold W. M. Smeulders |
ACM Multimedia | 3 |
| 2006 | The mediamill large.lexicon concept suggestion engineabstractIn this technical demonstration we show the current version of the MediaMill system, a search engine that facilitates access to news video archives at a semantic level. The core of the system is a lexicon of 436 automatically detected semantic concepts. To handle such a large lexicon in retrieval, an engine is developed which automatically selects a set of relevant concepts based on the textual query and example images. The result set can be browsed easily to obtain the final result for the query. Marcel Worring, Cees Snoek, Bouke Huurnink, Jan C. van Gemert, Dennis C. Koelma, Ork de Rooij |
ACM Multimedia | 4 |
| 2005 | MediaMill: exploring news video archives based on learned semanticsabstractIn this technical demonstration we showcase the MediaMill system. A search engine that facilitates access to news video archives at a semantic level. The core of the system is an unprecedented lexicon of 100 automatically detected semantic concepts. Based on this lexicon we demonstrate how users can obtain highly relevant retrieval results using query-by-concept. In addition, we show how the lexicon of concepts can be exploited for novel applications using advanced semantic visualizations. Several aspects of the MediaMill system are evaluated as part of our TRECVID 2005 efforts. Cees Snoek, Marcel Worring, Jan C. van Gemert, Jan-Mark Geusebroek, Dennis C. Koelma, Giang P. Nguyen, Ork de Rooij, Frank J. Seinstra |
ACM Multimedia | 3 |
| 2004 | Accessing video archives using interactive searchabstractIn this presentation, we present a system for interactive search in video archives. In our view, interactive search is a four-step process composed of indexing, filtering, browsing, and ranking. We have experimentally verified, using 22 groups of two participants each, how users apply these steps in the interactive search and how well they perform. Marcel Worring, Giang P. Nguyen, Laura Hollink, Jan C. van Gemert, Dennis C. Koelma |
ICME | 4 |
| 2004 | Spokenquery: an alternate approach to chosing items with speechabstractA majority of spoken user interfaces deal with the task of retrieving an element from a list. Conventionally, spoken UIs deal with such tasks through hierarchies of menus or dialogs, that navigate users through a series of steps, each of which present them with a limited set of choices. In a recent paper [2] we presented an alternative approach to such UIs, termed SpokenQuery, that recasts the problem of selection from lists as one of retrieval, and demonstrated that it could result in significantly lowered cognitive load on the user. In this paper, we examine varius aspects of retrieval from spoken queries, and UIs based on such retrieval, and demonstrate that in addition to reducing the cognitive load on the user, the system is effective for searching large databases, is robust to environment noise, and is effective as a UI. Joseph Woelfel, Jan C. van Gemert, Bhiksha Raj, David Wong 0006 |
INTERSPEECH | 3 |
| 2002 | Interactive Indexing and Retrieval of Multimedia Content
Marcel Worring, Andrew D. Bagdanov, Jan C. van Gemert, Jan-Mark Geusebroek, Hoang Minh, Guus Schreiber, Cees Snoek, Jeroen Vendrig, Jan Wielemaker, Arnold W. M. Smeulders |
SOFSEM | 3 |