VLDB 2026 Research / reviewers in the wild / expert
Stefan Duffner
dblp:64/6849
· DBLP profile ↗
54ranked-venue papers
14as first author
20since 2021 · last 2026
0000-0003-0374-3814ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 10 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bijective graph learning architecture with multi-level attributes interaction
Ikenna Oluigbo, Stefan Duffner, Kajal Eybpoosh, Catherine Pothier |
Data Min. Knowl. Discov. | 2 |
| 2025 | DML Mask R-CNN: addressing Dependent Multi-Label defect detection with severity estimation in manufacturingabstractIn this paper, we present our DML (Dependant Multi-Label) Mask R-CNN, a novel architecture designed for defect detection and severity estimation in manufacturing processes. Our approach aims to enhance automated visual inspection systems, addressing critical challenges of multi-label defect classification. The DML Mask R-CNN incorporates a modified Mask R-CNN framework with a dedicated classification head that effectively handles dependent multi-labels. We also introduced a simple geometric-based feature extraction module for defects severity classification, alongside an innovative inference algorithm involving a multi-threshold inference strategy and a modified Non-Maximum Suppression (NMS). Our architecture demonstrates significant improvements, with an average F1-score increase of 6.4% compared to our baseline across two self-collated tire datasets and the public Severstal dataset. This work emphasizes the architecture’s versatility and potential applicability to various domains requiring detailed instance segmentation and classification with multiple labels. Thomas Mignot, François Ponchon, Alexandre Derville, Stefan Duffner, Christophe Garcia |
IJCNN | 4 |
| 2024 | GroCo: Ground Constraint for Metric Self-supervised Monocular Depth
Aurélien Cecille, Stefan Duffner, Franck Davoine, Thibault Neveu, Rémi Agier |
ECCV (87) | 2 |
| 2024 | Deep Domain Isolation and Sample Clustered Federated Learning for Semantic Segmentation
Matthis Manthe, Carole Lartizien, Stefan Duffner |
ECML/PKDD (4) | 3 |
| 2024 | On GNN explainability with activation rules
Luca Veyrin-Forrer, Ataollah Kamal, Stefan Duffner, Marc Plantevit, Céline Robardet |
Data Min. Knowl. Discov. | 3 |
| 2024 | Federated brain tumor segmentation: An extensive benchmark
Matthis Manthe, Stefan Duffner, Carole Lartizien |
Medical Image Anal. | 2 |
| 2023 | Is My Neural Net Driven by the MDL Principle?
Eduardo Brandao, Stefan Duffner, Rémi Emonet, Amaury Habrard, François Jacquenet, Marc Sebban |
ECML/PKDD (2) | 2 |
| 2022 | Robust Variational Autoencoders and Normalizing Flows for Unsupervised Network Anomaly Detection
Naji Najari, Samuel Berlemont, Grégoire Lefebvre, Stefan Duffner, Christophe Garcia |
AINA (2) | 4 |
| 2022 | Improving Information Extraction on Business Documents with Specific Pre-training Tasks
Thibault Douzon, Stefan Duffner, Christophe Garcia, Jérémy Espinas |
DAS | 2 |
| 2022 | Efficient One-Shot Sports Field Image Registration with Arbitrary Keypoint SegmentationabstractAutomatic sports field registration aims at projecting a given image taken with unknown camera parameters to a known 3D coordinate system in order to obtain higher-level information like the position and speed of players. Existing methods generally detect specific visual landmarks on the field and then use an iterative refinement to get closer to the desired calibration. They are usually only compared in terms of precision on a standard benchmark without considering other metrics. However, execution speed is also important, mainly in the context of live broadcast TV and sports analysis. This work introduces a new automatic field registration method achieving excellent performance on the WorldCup Soccer benchmark, while neither depending on specific visible landmarks nor any refinement, resulting in a very high execution speed one-shot model. Finally, to complement the usual Soccer benchmark, we introduce a new Swimming Pool registration benchmark which is more challenging for the task at hand. Code and dataset available at https://github.com/njacquelin/sportsfieldregistration. Nicolas Jacquelin, Romain Vuillemot, Stefan Duffner |
ICIP | 3 |
| 2022 | What Does My GNN Really Capture? On Exploring Internal GNN RepresentationsabstractGraph Neural Networks (GNNs) are very efficient at classifying graphs but their internal functioning is opaque which limits their field of application. Existing methods to explain GNN focus on disclosing the relationships between input graphs and model decision. In this article, we propose a method that goes further and isolates the internal features, hidden in the network layers, that are automatically identified by the GNN and used in the decision process. We show that this method makes possible to know the parts of the input graphs used by GNN with much less bias that SOTA methods and thus to bring confidence in the decision process. Luca Veyrin-Forrer, Ataollah Kamal, Stefan Duffner, Marc Plantevit, Céline Robardet |
IJCAI | 3 |
| 2022 | In pursuit of the hidden features of GNN's internal representations
Luca Veyrin-Forrer, Ataollah Kamal, Stefan Duffner, Marc Plantevit, Céline Robardet |
Data Knowl. Eng. | 3 |
| 2022 | Periodicity counting in videos with unsupervised learning of cyclic embeddings
Nicolas Jacquelin, Romain Vuillemot, Stefan Duffner |
Pattern Recognit. Lett. | 3 |
| 2021 | Exploiting Visual Context to Identify People in TV Programs
Thomas Petit, Pierre Letessier, Stefan Duffner, Christophe Garcia |
CAIP (2) | 3 |
| 2021 | Self-supervised Continual Learning for Object Recognition in Image Sequences
Ruiqi Dai, Mathieu Lefort, Frédéric Armetta, Mathieu Guillermin, Stefan Duffner |
ICONIP (5) | 5 |
| 2021 | Novelty detection for unsupervised continual learning in image sequencesabstractRecent works in the domain of deep learning for object recognition on common image classification benchmarks often address the representation learning problem under the assumption of i.i.d. input data. Although achieving satisfying results, this assumption seems not realistic when agents have to learn autonomously. An autonomous agent receives a continual visual flow of objects which is far from an i.i.d. distribution of objects. Moreover, agents have to construct their representations of the world and adapt to unknown environments, without relying on external sources of information such as labels that would be provided post-classification and are unavoidable when an over-segmentation is done. Then, in order to exploit the learned representation effectively for object recognition, a clear and meaningful relationship w.r.t. real object categories is required, which has been largely neglected in existing unsupervised algorithms.In this paper, we propose a novelty detection method for continual and unsupervised object recognition, as an extension for the recent CURL model, which allows to moderate over-segmentation while preserving accuracy, in order to meet the requirements for autonomy. We experimentally validated our approach on two standard image classification benchmarks, MNIST and Fashion-MNIST, in this unsupervised and continual learning setting and improve the state of the art in terms of cluster purity, which is crucial for subsequent object recognition, since it facilitates clustering when information on ground truth labels is not available for free. Ruiqi Dai, Mathieu Lefort, Frédéric Armetta, Mathieu Guillermin, Stefan Duffner |
ICTAI | 5 |
| 2021 | Sequence Metric Learning as Synchronization of Recurrent Neural NetworksabstractSequence metric learning is becoming a widely adopted approach for various applications dealing with sequential multi-variate data such as activity recognition or natural language processing. It is most of the time tackled with sequence alignment approaches or representation learning. In this paper, we propose to study this subject from the point of view of dynamical system theory by drawing the analogy between synchronized trajectories produced by dynamical systems and the distance between similar sequences processed by a siamese recurrent neural network. Indeed, a siamese recurrent network comprises two identical sub-networks, two identical dynamical systems which can theoretically achieve complete synchronization if a coupling is introduced between them. We therefore propose a new neural network model that implements this coupling with a new gate integrated into the classical Gated Recurrent Unit architecture. This model is thus able to simultaneously learn a similarity metric and the synchronization of unaligned multi-variate sequences in a weakly supervised way. Our experiments show that introducing such a coupling improves the performance of the siamese Gated Recurrent Unit architecture on two datasets: one dedicated to activity recognition and another to transportation recognition. Paul Compagnon, Grégoire Lefebvre, Stefan Duffner, Christophe Garcia |
IJCNN | 3 |
| 2021 | RADON: Robust Autoencoder for Unsupervised Anomaly DetectionabstractAnomaly detection is a critical element in the design of secure and reliable networks. Networks are vulnerable to diverse anomalies, including malicious attacks, traffic congestion, operational problems, and hardware failures. In addition, novel unknown anomalies are constantly emerging, which challenges existing knowledge-based anomaly detectors. In this paper, we propose RADON, a robust algorithm for unsupervised anomaly detection, which does not require anomaly signatures or any labeled training data. RADON consists in a novel robust training strategy applied to an autoencoder. First, a filtering rule is applied on the reconstruction scores of traffic metadata to identify potential anomalies contaminating the training data. Then, RADON simultaneously learns to minimize the reconstruction scores of nominal instances, and to maximize that of the filtered anomalies. Finally, the trained model is leveraged to detect anomalous observations. Extensive experiments on the MNIST, NSL-KDD and MedBIoT benchmark datasets illustrate the effectiveness of our approach for anomaly detection, both in image and network traffic analyses. Naji Najari, Samuel Berlemont, Grégoire Lefebvre, Stefan Duffner, Christophe Garcia |
SIN | 4 |
| 2021 | List-wise learning-to-rank with convolutional neural networks for person re-identification
Yiqiang Chen 0003, Stefan Duffner, Andrei Stoian, Jean-Yves Dufour, Atilla Baskurt |
Mach. Vis. Appl. | 2 |
| 2021 | 2D Wasserstein loss for robust facial landmark detection
Yongzhe Yan, Stefan Duffner, Priyanka Phutane, Anthony Berthelier, Christophe Blanc, Christophe Garcia, Thierry Chateau |
Pattern Recognit. | 2 |
| 2020 | Multiple Instance Learning for Training Neural Networks under Label NoiseabstractIn this paper, we present an extensive study of different neural network-based approaches and loss functions applied to the Multiple Instance Learning (MIL) problem and binary classification. In the MIL setting, training is performed on small sets of instances called bags, where each positive bag contains at least one positive instance and each negative bag contains only negative instances. We propose a new loss function based on the generalised mean and an effective training strategy particularly suited to this setting and to problems where the instances of one class contain a considerable amount of label noise. Furthermore, we present a probabilistic approach to dynamically estimate the label noise in this unbalanced binary classification setting and utilise it to automatically modulate the hyper-parameter of our proposed loss function. We experimentally evaluated our approach on a number of standard benchmarks for binary classification and showed that it outperforms standard neural network optimisation algorithms as well as most state-of-the-art MIL methods, both on numerical/categorical vector data with MLP architectures and images with Convolutional Neural Networks. Stefan Duffner, Christophe Garcia |
IJCNN | 1 |
| 2020 | Unsupervised learning of co-occurrences for face images retrievalabstractDespite a huge leap in performance of face recognition systems in recent years, some cases remain challenging for them while being trivial for humans. This is because a human brain is exploiting much more information than the face appearance to identify a person. In this work, we aim at capturing the social context of unlabeled observed faces in order to improve face retrieval. In particular, we propose a framework that substantially improves face retrieval by exploiting the faces occurring simultaneously in a query's context to infer a multi-dimensional social context descriptor. Combining this compact structural descriptor with the individual visual face features in a common feature vector considerably increases the correct face retrieval rate and allows to disambiguate a large proportion of query results of different persons that are barely distinguishable visually. Thomas Petit, Pierre Letessier, Stefan Duffner, Christophe Garcia |
MMAsia | 3 |
| 2020 | Learning personalized ADL recognition models from few raw data
Paul Compagnon, Grégoire Lefebvre, Stefan Duffner, Christophe Garcia |
Artif. Intell. Medicine | 3 |
| 2020 | Fine-grained facial landmark detection exploiting intermediate feature representations
Yongzhe Yan, Stefan Duffner, Priyanka Phutane, Anthony Berthelier, Xavier Naturel, Christophe Blanc, Christophe Garcia, Thierry Chateau |
Comput. Vis. Image Underst. | 2 |
| 2020 | Two-stage human hair segmentation in the wild using deep shape prior
Yongzhe Yan, Stefan Duffner, Xavier Naturel, Anthony Berthelier, Christophe Garcia, Christophe Blanc, Thierry Chateau |
Pattern Recognit. Lett. | 2 |
| 2019 | Personalized Posture and Fall Classification with Shallow Gated Recurrent UnitsabstractActivities of Daily Living (ADL) classification is a key part of assisted living systems as it can be used to assess a person autonomy. We present in this paper an activity classification pipeline using Gated Recurrent Units (GRU) and inertial sequences. We aim to take advantage of the feature extraction properties of neural networks to free ourselves from defining rules or manually choosing features. We also investigate the advantages of resampling input sequences and personalizing GRU models to improve the performances. We evaluate our models on two datasets: a dataset containing five common postures: sitting, lying, standing, walking and transfer and a dataset named MobiAct V2 providing ADL and falls. Results show that the proposed approach could benefit eHealth services and particularly activity monitoring. Paul Compagnon, Grégoire Lefebvre, Stefan Duffner, Christophe Garcia |
CBMS | 3 |
| 2019 | Routine Modeling with Time Series Metric Learning
Paul Compagnon, Grégoire Lefebvre, Stefan Duffner, Christophe Garcia |
ICANN (2) | 3 |
| 2018 | Person Re-Identification with a Body Orientation-Specific Convolutional Neural Network
Yiqiang Chen 0003, Stefan Duffner, Andrei Stoian, Jean-Yves Dufour, Atilla Baskurt |
ACIVS | 2 |
| 2018 | Person Re-identification Using Group Context
Yiqiang Chen 0003, Stefan Duffner, Andrei Stoian, Jean-Yves Dufour, Atilla Baskurt |
ACIVS | 2 |
| 2018 | Similarity Learning with Listwise Ranking for Person Re-IdentificationabstractPerson re- identification is an important task in video surveillance systems. It consists in matching an image of a probe person among a gallery image set of people detected from a network of surveillance cameras with non-overlapping fields of view. The main challenge of person re- identification is to find image representations that are discriminating the persons' identities and that are robust to the viewpoint, body pose, illumination changes and partial occlusions. In this paper, we proposed a metric learning approach based on a deep neural network using a novel loss function which we call the Rank- Triplet loss. This proposed loss function is based on the predicted and ground truth ranking of a list of instances instead of pairs or triplets and takes into account the improvement of evaluation measures during training. Through our experiments on two person re- identification datasets, we show that the new loss outperforms other common loss functions and that our approach achieves state-of-the-art results on these two datasets. Yiqiang Chen 0003, Stefan Duffner, Atilla Baskurt, Andrei Stoian, Jean-Yves Dufour |
ICIP | 2 |
| 2018 | Class-balanced siamese neural networks
Samuel Berlemont, Grégoire Lefebvre, Stefan Duffner, Christophe Garcia |
Neurocomputing | 3 |
| 2018 | Deep and low-level feature based attribute learning for person re-identification
Yiqiang Chen 0003, Stefan Duffner, Andrei Stoian, Jean-Yves Dufour, Atilla Baskurt |
Image Vis. Comput. | 2 |
| 2018 | Pairwise Identity Verification via Linear Concentrative Metric LearningabstractThis paper presents a study of metric learning systems on pairwise identity verification, including pairwise face verification and pairwise speaker verification, respectively. These problems are challenging because the individuals in training and testing are mutually exclusive, and also due to the probable setting of limited training data. For such pairwise verification problems, we present a general framework of metric learning systems and employ the stochastic gradient descent algorithm as the optimization solution. We have studied both similarity metric learning and distance metric learning systems, of either a linear or shallow nonlinear model under both restricted and unrestricted training settings. Extensive experiments demonstrate that with limited training pairs, learning a linear system on similar pairs only is preferable due to its simplicity and superiority, i.e., it generally achieves competitive performance on both the labeled faces in the wild face dataset and the NIST speaker dataset. It is also found that a pretrained deep nonlinear model helps to improve the face verification results significantly. Lilei Zheng, Stefan Duffner, Khalid Idrissi, Christophe Garcia, Atilla Baskurt |
IEEE Trans. Cybern. | 2 |
| 2018 | Low-Complexity Approximate Convolutional Neural NetworksabstractIn this paper, we present an approach for minimizing the computational complexity of the trained convolutional neural networks (ConvNets). The idea is to approximate all elements of a given ConvNet and replace the original convolutional filters and parameters (pooling and bias coefficients; and activation function) with an efficient approximations capable of extreme reductions in computational complexity. Low-complexity convolution filters are obtained through a binary (zero and one) linear programming scheme based on the Frobenius norm over sets of dyadic rationals. The resulting matrices allow for multiplication-free computations requiring only addition and bit-shifting operations. Such low-complexity structures pave the way for low power, efficient hardware designs. We applied our approach on three use cases of different complexities: 1) a "light" but efficient ConvNet for face detection (with around 1000 parameters); 2) another one for hand-written digit classification (with more than 180 000 parameters); and 3) a significantly larger ConvNet: AlexNet with million matrices. We evaluated the overall performance on the respective tasks for different levels of approximations. In all considered applications, very low-complexity approximations have been derived maintaining an almost equal classification performance. Renato J. Cintra, Stefan Duffner, Christophe Garcia, André Leite |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Triplet CNN and pedestrian attribute recognition for improved person re-identificationabstractIn this paper, we propose a pedestrian attribute recognition approach and a CNN-based person re-identification framework enhanced by pedestrian attributes. The knowledge of person attributes can help video surveillance tasks like person re-identification as well as person search, semantic video indexing and retrieval to overcome viewpoint changes with their robustness to the inherent visual appearance variations. Compared to previous approaches, our attribute recognition method using Local Maximal Occurrence (LOMO) features and a Multi-Label Multi-Layer Perceptron (MLMLP) classifier proves to be more robust to different view points and is computationally more efficient. The experiments on three public benchmarks show that the proposed method improves the state-of-the art on attribute recognition. Furthermore, we integrate our attribute recognition algorithm into a triplet CNN similarity learning framework for person re-identification fusing both learned CNN features and attributes. This fusion leads to an overall improvement, and we achieve state-of-the-art results on person re-identification. Yiqiang Chen 0003, Stefan Duffner, Andrei Stoian, Jean-Yves Dufour, Atilla Baskurt |
AVSS | 2 |
| 2017 | Fast Pixelwise Adaptive Visual Tracking of Non-Rigid ObjectsabstractIn this paper, we present a new algorithm for real-time single-object tracking in videos in unconstrained environments. The algorithm comprises two different components that are trained "in one shot" at the first video frame: a detector that makes use of the generalized Hough transform with color and gradient descriptors and a probabilistic segmentation method based on global models for foreground and background color distributions. Both components work at pixel level and are used for tracking in a combined way adapting each other in a co-training manner. Moreover, we propose an adaptive shape model as well as a new probabilistic method for updating the scale of the tracker. Through effective model adaptation and segmentation, the algorithm is able to track objects that undergo rigid and non-rigid deformations and considerable shape and appearance variations. The proposed tracking method has been thoroughly evaluated on challenging benchmarks, and outperforms the state-of-the-art tracking methods designed for the same task. Finally, a very efficient implementation of the proposed models allows for extremely fast tracking. Stefan Duffner, Christophe Garcia |
IEEE Trans. Image Process. | 1 |
| 2016 | Polar Sine Based Siamese Neural Network for Gesture Recognition
Samuel Berlemont, Grégoire Lefebvre, Stefan Duffner, Christophe Garcia |
ICANN (2) | 3 |
| 2016 | Siamese multi-layer perceptrons for dimensionality reduction and face identification
Lilei Zheng, Stefan Duffner, Khalid Idrissi, Christophe Garcia, Atilla Baskurt |
Multim. Tools Appl. | 2 |
| 2016 | Using Discriminative Motion Context for Online Visual Object TrackingabstractIn this paper, we propose an algorithm for online, real-time tracking of arbitrary objects in videos from unconstrained environments. The method is based on a particle filter framework using different visual features and motion prediction models. We effectively integrate a discriminative online learning classifier into the model and propose a new method to collect negative training examples for updating the classifier at each video frame. Instead of taking negative examples only from the surroundings of the object region, or from specific background regions, our algorithm samples the negatives from a contextual motion density function in order to learn to discriminate the target as early as possible from potential distracting image regions. We experimentally show that this learning scheme improves the overall performance of the tracking algorithm. Moreover, we present quantitative and qualitative results on four challenging public data sets that show the robustness of the tracking algorithm with respect to appearance and view changes, lighting variations, partial occlusions, as well as object deformations. Finally, we compare the results with more than 30 state-of-the-art methods using two public benchmarks, showing very competitive results. Stefan Duffner, Christophe Garcia |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Visual Focus of Attention Estimation With Unsupervised Incremental LearningabstractIn this paper, we propose a new method for estimating the visual focus of attention (VFOA) in a video stream captured by a single distant camera and showing several persons sitting around a table, like in formal meeting or video conferencing settings. The visual targets for a given person are automatically extracted online using an unsupervised algorithm that incrementally learns the different appearance clusters from low-level visual features computed from face patches provided by a face tracker without the need of an intermediate error-prone step of head pose estimation as in classical approaches. The clusters learned in that way can then be used to classify the different visual attention targets of the person during a tracking run, without any prior knowledge on the environment and the configuration of the room or the visible persons. The experiments on public datasets containing almost 2 h of annotated videos from meetings and video conferencing show that the proposed algorithm produces state-of-the-art results and even outperforms a traditional supervised method that is based on head orientation estimation and that classifies VFOA using Gaussian mixture models. Stefan Duffner, Christophe Garcia |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Classifying Global Scene Context for On-line Multiple Tracker SelectionabstractIn this paper, we present a novel framework for combining several independent on-line trackers using visual scene context. The aim of our method is to decide automatically at each point in time which specific tracking algorithm works best under the given scene or acquisition conditions. To this end, we define a set of generic global context features computed on each frame of a set of training videos. At the same time, we record the performance of each individual tracker on these videos in terms of object bounding box overlap with the ground truth. Then a classifier is trained to estimate which tracker gives the best result given the global scene context in a particular frame. We experimentally show that such a classifier can predict the best tracker with a precision of over 80% in unknown videos with unknown environments. The proposed tracking method further filters the classifier responses temporarily using a Hidden Markov Model in order to avoid rapid oscillations between different trackers. Finally, we evaluated the overall tracking system and showed that this scene context-based tracker selection considerably improves the overall robustness and compares favourably with the state-of-the-art. Salma Moujtahid, Stefan Duffner, Atilla Baskurt |
BMVC | 2 |
| 2015 | Logistic similarity metric learning for face verificationabstractThis paper presents a new method for similarity metric learning, called Logistic Similarity Metric Learning (LSML), where the cost is formulated as the logistic loss function, which gives a probability estimation of a pair of faces being similar. Especially, we propose to shift the similarity decision boundary gaining significant performance improvement. We test the proposed method on the face verification problem using four single face descriptors: LBP, OCLBP, SIFT and Gabor wavelets. Extensive experimental results on the LFW-a data set demonstrate that the proposed method achieves competitive state-of-the-art performance on the problem of face verification. Lilei Zheng, Khalid Idrissi, Christophe Garcia, Stefan Duffner, Atilla Baskurt |
ICASSP | 4 |
| 2014 | 3D gesture classification with convolutional neural networksabstractIn this paper, we present an approach that classifies 3D gestures using jointly accelerometer and gyroscope signals from a mobile device. The proposed method is based on a convolutional neural network with a specific structure involving a combination of 1D convolution, averaging, and max-pooling operations. It directly classifies the fixed-length input matrix, composed of the normalised sensor data, as one of the gestures to be recognises. Experimental results on different datasets with varying training/testing configurations show that our method outperforms or is on par with current state-of-the-art methods for almost all data configurations. Stefan Duffner, Samuel Berlemont, Grégoire Lefebvre, Christophe Garcia |
ICASSP | 1 |
| 2014 | Leveraging colour segmentation for upper-body detection
Stefan Duffner, Jean-Marc Odobez |
Pattern Recognit. | 1 |
| 2013 | Unsupervised online learning of visual focus of attentionabstractIn this paper, we propose a novel approach for estimating visual focus of attention in video streams. The method is based on an unsupervised algorithm that incrementally learns the different appearance clusters from low-level visual features extracted from face patches provided by a face tracker. The clusters learnt in that way can then be used to classify the different visual attention targets of a given person during a tracking run, without any prior knowledge on the environment and the configuration of the room or the visible persons. Experiments on public datasets containing almost two hours of annotated videos from meetings and video-conferencing show that the proposed algorithm produces state-of-the-art results and even outperforms a traditional supervised method that is based on head orientation estimation and that classifies visual focus of attention using Gaussian Mixture Models. Stefan Duffner, Christophe Garcia |
AVSS | 1 |
| 2013 | PixelTrack: A Fast Adaptive Algorithm for Tracking Non-rigid ObjectsabstractIn this paper, we present a novel algorithm for fast tracking of generic objects in videos. The algorithm uses two components: a detector that makes use of the generalised Hough transform with pixel-based descriptors, and a probabilistic segmentation method based on global models for foreground and background. These components are used for tracking in a combined way, and they adapt each other in a co-training manner. Through effective model adaptation and segmentation, the algorithm is able to track objects that undergo rigid and non-rigid deformations and considerable shape and appearance variations. The proposed tracking method has been thoroughly evaluated on challenging standard videos, and outperforms state-of-the-art tracking methods designed for the same task. Finally, the proposed models allow for an extremely efficient implementation, and thus tracking is very fast. Stefan Duffner, Christophe Garcia |
ICCV | 1 |
| 2013 | Track Creation and Deletion Framework for Long-Term Online Multiface TrackingabstractTo improve visual tracking, a large number of papers study more powerful features, or better cue fusion mechanisms, such as adaptation or contextual models. A complementary approach consists of improving the track management, that is, deciding when to add a target or stop its tracking, for example, in case of failure. This is an essential component for effective multiobject tracking applications, and is often not trivial. Deciding whether or not to stop a track is a compromise between avoiding erroneous early stopping while tracking is fine, and erroneous continuation of tracking when there is an actual failure. This decision process, very rarely addressed in the literature, is difficult due to object detector deficiencies or observation models that are insufficient to describe the full variability of tracked objects and deliver reliable likelihood (tracking) information. This paper addresses the track management issue and presents a real-time online multiface tracking algorithm that effectively deals with the above difficulties. The tracking itself is formulated in a multiobject state-space Bayesian filtering framework solved with Markov Chain Monte Carlo. Within this framework, an explicit probabilistic filtering step decides when to add or remove a target from the tracker, where decisions rely on multiple cues such as face detections, likelihood measures, long-term observations, and track state characteristics. The method has been applied to three challenging data sets of more than 9 h in total, and demonstrate a significant performance increase compared to more traditional approaches (Markov Chain Monte Carlo, reversible-jump Markov Chain Monte Carlo) only relying on head detection and likelihood for track management. Stefan Duffner, Jean-Marc Odobez |
IEEE Trans. Image Process. | 1 |
| 2012 | Multimodal Cue Detection Engine for Orchestrated Entertainment
Danil Korchagin, Stefan Duffner, Petr Motlícek, Carl Scheffler |
MMM | 2 |
| 2011 | Exploiting long-term observations for track creation and deletion in online multi-face trackingabstractIn many visual multi-object tracking applications, the question when to add or remove a target is not trivial due to, for example, erroneous outputs of object detectors or observation models that cannot describe the full variability of the objects to track. In this paper, we present a real-time, online multi-face tracking algorithm that effectively deals with missing or uncertain detections in a principled way. The tracking is formulated in a multi-object state-space Bayesian filtering framework solved with Markov Chain Monte Carlo. Within this framework, an explicit probabilistic filtering step relying on head detections, likelihood models, and long term observations as well as object track characteristics has been designed to take the decision on when to add or remove a target from the tracker. The proposed method applied on three challenging datasets of more than 9 hours shows a significant performance increase compared to a traditional approach relying on head detection and likelihood models only. Stefan Duffner, Jean-Marc Odobez |
FG | 1 |
| 2011 | Just-in-time multimodal association and fusion from home entertainmentabstractIn this paper, we describe a real-time multimodal analysis system with just-in-time multimodal association and fusion for a living room environment, where multiple people may enter, interact and leave the observable world with no constraints. It comprises detection and tracking of up to 4 faces, detection and localisation of verbal and paralinguistic events, their association and fusion. The system is designed to be used in open, unconstrained environments like in next generation video conferencing systems that automatically "orchestrate" the transmitted video streams to improve the overall experience of interaction between spatially separated families and friends. Performance levels achieved to date on hand-labelled dataset have shown sufficient reliability at the same time as fulfilling real-time processing requirements. Danil Korchagin, Petr Motlícek, Stefan Duffner, Hervé Bourlard |
ICME | 3 |
| 2009 | Dynamic Partitioned Sampling For Tracking With Discriminative FeaturesabstractWe present a multi-cue fusion method for tracking with particle filters which relies on a novel hierarchical sampling strategy. Similarly to previous works, it tackles the problem of tracking in a relatively high-dimensional state space by dividing such a space into partitions, each one corresponding to a single cue, and sampling from them in a hierarchical manner. However, unlike other approaches, the order of partitions is not fixed a priori but changes dynamically depending on the reliability of each cue, i.e. more reliable cues are sampled first. We call this approach Dynamic Partitioned Sampling (DPS). The reliability of each cue is measured in terms of its ability to discriminate the object with respect to the background, where the background is not described by a fixed model or by random patches but is represented by a set of informative "background particles" which are tracked in order to be as similar as possible to the object. The effectiveness of this general framework is demonstrated on the specific problem of head tracking with three different cues: colour, edge and contours. Experimental results prove the robustness of our algorithm in several challenging video sequences. Stefan Duffner, Jean-Marc Odobez, Elisa Ricci 0001 |
BMVC | 1 |
| 2007 | Face recognition using non-linear image reconstructionabstractWe present a face recognition technique based on a special type of convolutional neural network that is trained to extract characteristic features from face images and reconstruct the corresponding reference face images which are chosen beforehand for each individual to recognize. The reconstruction is realized by a so-called "bottle-neck" neural network that learns to project face images into a low-dimensional vector space and to reconstruct the respective reference images from the projected vectors. In contrast to methods based on the Principal Component Analysis (PCA), the Linear Discriminant Analysis (LDA) etc., the projection is non-linear and depends on the choice of the reference images. Moreover, local and global processing are closely interconnected and the respective parameters are conjointly learnt. Having trained the neural network, new face images can then be classified by comparing the respective projected vectors. We experimentally show that the choice of the reference images influences the final recognition performance and that this method outperforms linear projection methods in terms of precision and robustness. Stefan Duffner, Christophe Garcia |
AVSS | 1 |
| 2007 | An Online Backpropagation Algorithm with Validation Error-Based Adaptive Learning Rate
Stefan Duffner, Christophe Garcia |
ICANN (1) | 1 |
| 2006 | A Neural Scheme for Robust Detection of Transparent Logos in TV Programs
Stefan Duffner, Christophe Garcia |
ICANN (2) | 1 |