EDBT 2026 Demo / reviewers in the wild / expert
Oswald Lanz
dblp:02/1449
· DBLP profile ↗
50ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0003-4793-4276ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 29 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal predictive process monitoring and its application to explainable clinical pathwaysabstractThis paper presents one of the first contributions in the context of Multimodal Predictive Process Monitoring (MM-PPM) . In recent years, Predictive Process Monitoring (PPM) has evolved at the intersection of process mining, machine learning, and data science, as organizations seek to anticipate the future course of ongoing processes. Traditional PPM mainly relies on structured event log data, but many real-world scenarios generate richer information, including text, images, audio, and video. MM-PPM promises to start addressing this rich data scenario by integrating complementary knowledge from heterogeneous modalities through modality-specific representations and information fusion techniques. The growing digitization of healthcare systems, combined with advances in Artificial Intelligence (AI), has accelerated AI-based PPM for analyzing sequences of clinical events, supporting decision-making, enabling personalized care, and improving clinical facility management. Given these characteristics, clinical pathways represent an ideal domain for experimenting with MM-PPM, as they may naturally involve diverse modalities such as structured records, free-text notes, or medical images. To handle multimodal information available with clinical pathways, we introduce MEDUSA , an MM-PPM approach for outcome prediction, which jointly processes medical image information coupled with the storytelling of structural records and text notes collected during the clinical pathway of a patient until the acquisition of the considered image. The evaluation of MEDUSA is done in a COVID-19 case study, to assess the performance of the proposed approach and explain how specific information within each modality influences the decisions of the predictive model. Vincenzo Pasquadibisceglie, Ivan Donadello, Annalisa Appice, Oswald Lanz, Fabrizio Maria Maggi, Giuseppe Fiameni, Donato Malerba |
Inf. Syst. | 4 |
| 2025 | L-SWAG: Layer-Sample Wise Activation with Gradients Information for Zero-Shot NAS on Vision TransformersabstractTraining-free Neural Architecture Search (NAS) efficiently identifies high-performing neural networks using zero-cost (ZC) proxies. Unlike multi-shot and one-shot NAS approaches, ZC-NAS is both (i) time-efficient, eliminating the need for model training, and (ii) interpretable, with proxy designs often theoretically grounded. Despite rapid developments in the field, current SOTA ZC proxies are typically constrained to well-established convolutional search spaces. With the rise of Large Language Models shaping the future of deep learning, this work extends ZC proxy applicability to Vision Transformers (ViTs). We present a new benchmark using the Autoformer search space evaluated on 6 distinct tasks and propose Layer-Sample Wise Activation with Gradients information (L-SWAG), a novel, generalizable metric that characterizes both convolutional and transformer architectures across 14 tasks. Additionally, previous works highlighted how different proxies contain complementary information, motivating the need for a ML model to identify useful combinations. To further enhance ZC-NAS, we therefore introduce LIBRA-NAS (Low Information gain and Bias Re-Alignment), a method that strategically combines proxies to best represent a specific benchmark. Integrated into the NAS search, LIBRA-NAS outperforms evolution and gradient-based NAS techniques by identifying an architecture with a 17.0% test error on ImageNet1k in just 0.1 GPU days. Sofia Casarin, Sergio Escalera, Oswald Lanz |
CVPR | 3 |
| 2025 | ONFOODS: A Substitute Recommendation System in Food RecipesabstractAbstract Food waste is a serious problem in modern society. A specific aspect of food waste concerns meat consumption in gastronomy, where typically only prime cuts of meat are used in the kitchen. To facilitate the usage of all parts of animals and thereby reducing food waste, we present Onfoods , a system that recommends alternative meat cuts in recipes and integrates inventory data to help with the creation of menus. Onfoods uses an ontology and a knowledge graph to model recipes, meat cuts and the relationships between the two, similarity measures to find candidates for alternative meat cuts, and inventory data to track the availability of different meat cuts. An intuitive user interface allows the user on one hand to update the knowledge graph and inventory data, and on the other hand to navigate through recipes and choose alternative meat cuts. Maryam Mozaffari, Anton Dignös, Oswald Lanz, Dominik T. Matt, Gabriele Pasetti Monizza, Matthias Gauly, Johann Gamper |
DEXA (2) | 3 |
| 2024 | Your Image Is My Video: Reshaping the Receptive Field via Image-to-Video Differentiable AutoAugmentation and FusionabstractThe landscape of deep learning research is moving towards innovative strategies to harness the true potential of data. Traditionally, emphasis has been on scaling model architectures, resulting in large and complex neural networks, which can be difficult to train with limited computational resources. However, independently of the model size, data quality (i.e. amount and variability) is still a major factor that affects model generalization. In this work, we propose a novel technique to exploit available data through the use of automatic data augmentation for the tasks of image classification and semantic segmentation. We introduce the first Differentiable Augmentation Search method (DAS) to generate variations of images that can be processed as videos. Compared to previous approaches, DAS is extremely fast and flexible, allowing the search on very large search spaces in less than a GPU day. Our intuition is that the increased receptive field in the temporal dimension provided by DAS could lead to benefits also to the spatial receptive field. More specifically, we leverage DAS to guide the reshaping of the spatial receptive field by selecting task-dependant transformations. As a result, compared to standard augmentation alternatives, we improve in terms of accuracy on ImageNet, Cifar10, Cijar100, Tiny-ImageNet, Pascal-VOC-2012 and CityScapes datasets when plugging-in our DAS over different light-weight video backbones. Sofia Casarin, Cynthia Ifeyinwa Ugwu, Sergio Escalera, Oswald Lanz |
CVPR | 4 |
| 2024 | Improving semantic video retrieval models by training with a relevance-aware online mining strategyabstractTo retrieve a video via a multimedia search engine, a textual query is usually created by the user and then used to perform the search. Recent state-of-the-art cross-modal retrieval methods learn a joint text-video embedding space by using contrastive loss functions, which maximize the similarity of positive pairs while decreasing that of the negative pairs. Although the choice of these pairs is fundamental for the construction of the joint embedding space, the selection procedure is usually driven by the relationships found within the dataset: a positive pair is commonly formed by a video and its own caption, whereas unrelated video-caption pairs represent the negative ones. We hypothesize that this choice results in a retrieval system with limited semantics understanding, as the standard training procedure requires the system to discriminate between groundtruth and negative even though there is no difference in their semantics. Therefore, differently from the previous approaches, in this paper we propose a novel strategy for the selection of both positive and negative pairs which takes into account both the annotations and the semantic contents of the captions. By doing so, the selected negatives do not share semantic concepts with the positive pair anymore, and it is also possible to discover new positives within the dataset. Based on our hypothesis, we provide a novel design of two popular contrastive loss functions, and explore their effectiveness on three heterogeneous state-of-the-art approaches. The extensive experimental analysis conducted on two datasets, EPIC-Kitchens-100 and MSR-VTT, validates the effectiveness of the proposed strategy, observing, e.g., more than +20% nDCG on EPIC-Kitchens-100. Furthermore, these results are corroborated with qualitative evidence both supporting our hypothesis and explaining why the proposed strategy effectively overcomes it. Alex Falcon, Giuseppe Serra 0001, Oswald Lanz |
Comput. Vis. Image Underst. | 3 |
| 2023 | Video question answering supported by a multi-task learning objectiveabstractAbstract Video Question Answering (VideoQA) concerns the realization of models able to analyze a video, and produce a meaningful answer to visual content-related questions. To encode the given question, word embedding techniques are used to compute a representation of the tokens suitable for neural networks. Yet almost all the works in the literature use the same technique, although recent advancements in NLP brought better solutions. This lack of analysis is a major shortcoming. To address it, in this paper we present a twofold contribution about this inquiry and its relation with question encoding. First of all, we integrate four of the most popular word embedding techniques in three recent VideoQA architectures, and investigate how they influence the performance on two public datasets: EgoVQA and PororoQA. Thanks to the learning process, we show that embeddings carry question type-dependent characteristics. Secondly, to leverage this result, we propose a simple yet effective multi-task learning protocol which uses an auxiliary task defined on the question types. By using the proposed learning strategy, significant improvements are observed in most of the combinations of network architecture and embedding under analysis. Alex Falcon, Giuseppe Serra 0001, Oswald Lanz |
Multim. Tools Appl. | 3 |
| 2023 | Learning to Recognize Actions on Objects in Egocentric Video With Attention DictionariesabstractWe present EgoACO, a deep neural architecture for video action recognition that learns to pool action-context-object descriptors from frame level features by leveraging the verb-noun structure of action labels in egocentric video datasets. The core component is class activation pooling (CAP), a differentiable pooling layer that combines ideas from bilinear pooling for fine-grained recognition and from feature learning for discriminative localization. CAP uses self-attention with a dictionary of learnable weights to pool from the most relevant feature regions. Through CAP, EgoACO learns to decode object and scene context descriptors from video frame features. For temporal modeling we design a recurrent version of class activation pooling termed Long Short-Term Attention (LSTA). LSTA extends convolutional gated LSTM with built-in spatial attention and a re-designed output gate. Action, object and context descriptors are fused by a multi-head prediction that accounts for the inter-dependencies between noun-verb-action structured labels in egocentric video datasets. EgoACO features built-in visual explanations, helping learning and interpretation of discriminative information in video. Results on the two largest egocentric action recognition datasets currently available, EPIC-KITCHENS and EGTEA Gaze+, show that by decoding action-context-object descriptors, the model achieves state-of-the-art recognition performance. Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Gate-Shift-Fuse for Video Action RecognitionabstractConvolutional Neural Networks are the de facto models for image recognition. However 3D CNNs, the straight forward extension of 2D CNNs for video recognition, have not achieved the same success on standard action recognition benchmarks. One of the main reasons for this reduced performance of 3D CNNs is the increased computational complexity requiring large scale annotated datasets to train them in scale. 3D kernel factorization approaches have been proposed to reduce the complexity of 3D CNNs. Existing kernel factorization approaches follow hand-designed and hard-wired techniques. In this paper we propose Gate-Shift-Fuse (GSF), a novel spatio-temporal feature extraction module which controls interactions in spatio-temporal decomposition and learns to adaptively route features through time and combine them in a data dependent manner. GSF leverages grouped spatial gating to decompose input tensor and channel weighting to fuse the decomposed tensors. GSF can be inserted into existing 2D CNNs to convert them into an efficient and high performing spatio-temporal feature extractor, with negligible parameter and compute overhead. We perform an extensive analysis of GSF using two popular 2D CNN families and achieve state-of-the-art or competitive performance on five standard action recognition benchmarks. Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Implicit texture mapping for multi-view video synthesis
Mohamed Ilyes Lakhal, Oswald Lanz, Andrea Cavallaro |
BMVC | 2 |
| 2022 | Higher-Order Recurrent Network with Space-Time Attention for Video Early Action RecognitionabstractEndowing visual agents with predictive capability is a key step towards video intelligence at scale. Early action recognition aims to predict the action labels before fully observing the complete video frames. Unlike action recognition, the model is asked to forecast the future or the effects by only observing the initial few frames. The strong reasoning ability over the temporal dimension is the key to success. To this end, in this paper, we propose a novel recurrent network with decomposed space-time attention and higher-order design to capture the temporal dependency associated with the specific actions. Our method achieves state-of-the-art performance on Something-Something and EPIC-Kitchens datasets under the early action recognition setting, showing evidence of predictive capability that we attribute to our higher-order recurrent design with space-time attention. Tsung-Ming Tai, Giuseppe Fiameni, Cheng-Kuang Lee, Oswald Lanz |
ICIP | 4 |
| 2022 | Unified Recurrence Modeling for Video Action AnticipationabstractForecasting future events based on evidence of current conditions is an innate skill of human beings, and key for predicting the outcome of any decision making. In artificial vision for example, we would like to predict the next human action before it happens, without observing the future video frames associated to it. Computer vision models for action anticipation are expected to collect the subtle evidence in the preamble of the target actions. In prior studies recurrence modeling often leads to better performance, the strong temporal inference is assumed to be a key element for reasonable prediction. To this end, we propose a unified recurrence modeling for video action anticipation via message passing framework. The information flow in space-time can be described by the interaction between vertices and edges, and the changes of vertices for each incoming frame reflects the underlying dynamics. Our model leverages self-attention as the building blocks for each of the message passing functions. In addition, we introduce different edge learning strategies that can be end-to-end optimized to gain better flexibility for the connectivity between vertices. Our experimental results demonstrate that our proposed method outperforms previous works on the large-scale EPIC-Kitchen dataset. Tsung-Ming Tai, Giuseppe Fiameni, Cheng-Kuang Lee, Simon See, Oswald Lanz |
ICPR | 5 |
| 2022 | Relevance-based Margin for Contrastively-trained Video Retrieval ModelsabstractVideo retrieval using natural language queries has attracted increasing interest due to its relevance in real-world applications, from intelligent access in private media galleries to web-scale video search. Learning the cross-similarity of video and text in a joint embedding space is the dominant approach. To do so, a contrastive loss is usually employed because it organizes the embedding space by putting similar items close and dissimilar items far. This framework leads to competitive recall rates, as they solely focus on the rank of the groundtruth items. Yet, assessing the quality of the ranking list is of utmost importance when considering intelligent retrieval systems, since multiple items may share similar semantics, hence a high relevance. Moreover, the aforementioned framework uses a fixed margin to separate similar and dissimilar items, treating all non-groundtruth items as equally irrelevant. In this paper we propose to use a variable margin: we argue that varying the margin used during training based on how much relevant an item is to a given query, i.e. a relevance-based margin, easily improves the quality of the ranking lists measured through nDCG and mAP. We demonstrate the advantages of our technique using different models on EPIC-Kitchens-100 and YouCook2. We show that even if we carefully tuned the fixed margin, our technique (which does not have the margin as a hyper-parameter) would still achieve better performance. Finally, extensive ablation studies and qualitative analysis support the robustness of our approach. Code will be released at \urlhttps://github.com/aranciokov/RelevanceMargin-ICMR22. Alex Falcon, Swathikiran Sudhakaran, Giuseppe Serra 0001, Sergio Escalera, Oswald Lanz |
ICMR | 5 |
| 2022 | A Feature-space Multimodal Data Augmentation Technique for Text-video RetrievalabstractEvery hour, huge amounts of visual contents are posted on social media and user-generated content platforms. To find relevant videos by means of a natural language query, text-video retrieval methods have received increased attention over the past few years. Data augmentation techniques were introduced to increase the performance on unseen test examples by creating new training samples with the application of semantics-preserving techniques, such as color space or geometric transformations on images. Yet, these techniques are usually applied on raw data, leading to more resource-demanding solutions and also requiring the shareability of the raw data, which may not always be true, e.g. copyright issues with clips from movies or TV series. To address this shortcoming, we propose a multimodal data augmentation technique which works in the feature space and creates new videos and captions by mixing semantically similar samples. We experiment our solution on a large scale public dataset, EPIC-Kitchens-100, and achieve considerable improvements over a baseline method, improved state-of-the-art performance, while at the same time performing multiple ablation studies. We release code and pretrained models on Github at https://github.com/aranciokov/FSMMDA\_VideoRetrieval. Alex Falcon, Giuseppe Serra 0001, Oswald Lanz |
ACM Multimedia | 3 |
| 2022 | Audio-Visual Tracking of Concurrent SpeakersabstractAudio-visual tracking of an unknown number of concurrent speakers in 3D is a challenging task, especially when sound and video are collected with a compact sensing platform. In this paper, we propose a tracker that builds on generative and discriminative audio-visual likelihood models formulated in a particle filtering framework. We localize multiple concurrent speakers with a de-emphasized acoustic map assisted by the image detection-derived 3D video observations. The 3D multi-modal observations are either assigned to existing tracks for discriminative likelihood computation or used to initialize new tracks. The generative likelihoods rely on color distribution of the target and the de-emphasized acoustic map value. Experiments on AV16.3 and CAV3D datasets show that the proposed tracker outperforms the uni-modal trackers and the state-of-the-art approaches both in 3D and on the image plane. Xinyuan Qian 0001, Alessio Brutti, Oswald Lanz, Maurizio Omologo, Andrea Cavallaro |
IEEE Trans. Multim. | 3 |
| 2020 | Novel-View Human Action Synthesis
Mohamed Ilyes Lakhal, Davide Boscaini, Fabio Poiesi, Oswald Lanz, Andrea Cavallaro |
ACCV (4) | 4 |
| 2020 | Gate-Shift Networks for Video Action RecognitionabstractDeep 3D CNNs for video action recognition are designed to learn powerful representations in the joint spatio-temporal feature space. In practice however, because of the large number of parameters and computations involved, they may under-perform in the lack of sufficiently large datasets for training them at scale. In this paper we introduce spatial gating in spatial-temporal decomposition of 3D kernels. We implement this concept with Gate-Shift Module (GSM). GSM is lightweight and turns a 2D-CNN into a highly efficient spatio-temporal feature extractor. With GSM plugged in, a 2D-CNN learns to adaptively route features through time and combine them, at almost no additional parameters and computational overhead. We perform an extensive evaluation of the proposed module to study its effectiveness in video action recognition, achieving state-of-the-art results on Something Something-V1 and Diving48 datasets, and obtaining competitive results on EPIC-Kitchens with far less model complexity. Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz |
CVPR | 3 |
| 2020 | A Spatio-Temporal Multi-Scale Binary DescriptorabstractBinary descriptors are widely used for multi-view matching and robotic navigation. However, their matching performance decreases considerably under severe scale and viewpoint changes in non-planar scenes. To overcome this problem, we propose to encode the varying appearance of selected 3D scene points tracked by a moving camera with compact spatio-temporal descriptors. To this end, we first track interest points and capture their temporal variations at multiple scales. Then, we validate feature tracks through 3D reconstruction and compress the temporal sequence of descriptors by encoding the most frequent and stable binary values. Finally, we determine multiscale correspondences across views with a matching strategy that handles severe scale differences. The proposed spatio-temporal multi-scale approach is generic and can be used with a variety of binary descriptors. We show the effectiveness of the joint multiscale extraction and temporal reduction through comparisons of different temporal reduction strategies and the application to several binary descriptors. Alessio Xompero, Oswald Lanz, Andrea Cavallaro |
IEEE Trans. Image Process. | 2 |
| 2019 | LSTA: Long Short-Term Attention for Egocentric Action RecognitionabstractEgocentric activity recognition is one of the most challenging tasks in video analysis. It requires a fine-grained discrimination of small objects and their manipulation. While some methods base on strong supervision and attention mechanisms, they are either annotation consuming or do not take spatio-temporal patterns into account. In this paper we propose LSTA as a mechanism to focus on features from spatial relevant parts while attention is being tracked smoothly across the video sequence. We demonstrate the effectiveness of LSTA on egocentric activity recognition with an end-to-end trainable two-stream architecture, achieving state-of-the-art performance on four standard benchmarks. Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz |
CVPR | 3 |
| 2019 | Accurate Target Annotation in 3D from Multimodal StreamsabstractAccurate annotation is fundamental to quantify the performance of multi-sensor and multi-modal object detectors and trackers. However, invasive or expensive instrumentation is needed to automatically generate these annotations. To mitigate this problem, we present a multi-modal approach that leverages annotations from reference streams (e.g. individual camera views) and measurements from unannotated additional streams (e.g. audio) to infer 3D trajectories through an optimization. The core of our approach is a multi-modal extension of Bundle Adjustment with a cross-modal correspondence detection that selectively uses measurements in the optimization. We apply the proposed approach to fully annotate a new multi-modal and multi-view dataset for multi-speaker 3D tracking. Oswald Lanz, Alessio Brutti, Alessio Xompero, Xinyuan Qian 0001, Maurizio Omologo, Andrea Cavallaro |
ICASSP | 1 |
| 2019 | View-LSTM: Novel-View Video Synthesis Through View DecompositionabstractWe tackle the problem of synthesizing a video of multiple moving people as seen from a novel view, given only an input video and depth information or human poses of the novel view as prior. This problem requires a model that learns to transform input features into target features while maintaining temporal consistency. To this end, we learn an invariant feature from the input video that is shared across all viewpoints of the same scene and a view-dependent feature obtained using the target priors. The proposed approach, View-LSTM, is a recurrent neural network structure that accounts for the temporal consistency and target feature approximation constraints. We validate View-LSTM by designing an end-to-end generator for novel-view video synthesis. Experiments on a large multi-view action recognition dataset validate the proposed model. Mohamed Ilyes Lakhal, Oswald Lanz, Andrea Cavallaro |
ICCV | 2 |
| 2019 | Learnable Masks for Pose-Guided View SynthesisabstractPose-guided human view synthesis uses a target pose to generate the appearance of a new view of a person. The input view and the target pose can be processed separately with UNet architectures that combine the results in a late fusion stage. UNet architectures link their encoder and decoder with skip connections that preserve the location of spatial features by injecting input information in the decoding process. However, direct skip connections may transfer irrelevant information to the decoder. We overcome this limitation with learnable masks for skip connections that encourage the decoder to use only relevant information from the encoder. We show that adding the proposed mask to UNet architectures improves the performance of view synthesis with only a slight increase in inference time. Mohamed Ilyes Lakhal, Oswald Lanz, Andrea Cavallaro |
ICIP | 2 |
| 2019 | Multi-Speaker Tracking From an Audio-Visual Sensing DeviceabstractCompact multi-sensor platforms are portable and thus desirable for robotics and personal-assistance tasks. However, compared to physically distributed sensors, the size of these platforms makes person tracking more difficult. To address this challenge, we propose a novel 3-D audio-visual people tracker that exploits visual observations (object detections) to guide the acoustic processing by constraining the acoustic likelihood on the horizontal plane defined by the predicted height of a speaker. This solution allows the tracker to estimate, with a small microphone array, the distance of a sound. Moreover, we apply a color-based visual likelihood on the image plane to compensate for misdetections. Finally, we use a 3-D particle filter and greedy data association to combine visual observations, color-based, and acoustic likelihoods to track the position of multiple simultaneous speakers. We compare the proposed multimodal 3-D tracker against two state-of-the-art methods on the AV16.3 dataset and on a newly collected dataset with co-located sensors, which we make available to the research community. Experimental results show that our multimodal approach outperforms the other methods both in 3-D and on the image plane. Xinyuan Qian 0001, Alessio Brutti, Oswald Lanz, Maurizio Omologo, Andrea Cavallaro |
IEEE Trans. Multim. | 3 |
| 2018 | Attention is All We Need: Nailing Down Object-centric Attention for Egocentric Activity Recognition
Swathikiran Sudhakaran, Oswald Lanz |
BMVC | 2 |
| 2018 | Multi-Camera Matching of Spatio-Temporal Binary FeaturesabstractLocal image features are generally robust to different geometric and photometric transformations on planar surfaces or under narrow baseline views. However, the matching performance decreases considerably across cameras with unknown poses separated by a wide baseline. To address this problem, we accumulate temporal information within each view by tracking local binary features, which encode intensity comparisons of pixel pairs in an image patch. We then encode the spatio-temporal features into fixed-length binary descriptors by selecting temporally dominant binary values. We complement the descriptor with a binary vector that identifies intensity comparisons that are temporally unstable. Finally, we use this additional vector to ignore the corresponding binary values in the fixed-length binary descriptor when matching the features across cameras. We analyse the performance of the proposed approach and compare it with baselines. Alessio Xompero, Oswald Lanz, Andrea Cavallaro |
FUSION | 2 |
| 2018 | 3D Mouth Tracking from a Compact Microphone Array Co-Located with a cameraabstractWe address the 3D audio-visual mouth tracking problem when using a compact platform with co-located audio-visual sensors, without a depth camera. In particular, we propose a multi-modal particle filter that combines a face detector and 3D hypothesis mapping to the image plane. The audio likelihood computation is assisted by video, which relies on a GCC-PHAT based acoustic map. By combining audio and video inputs, the proposed approach can cope with a reverberant and noisy environment, and can deal with situations when the person is occluded, outside the Field of View (FoV), or not facing the sensors. Experimental results show that the proposed tracker is accurate both in 3D and on the image plane. Xinyuan Qian 0001, Alessio Xompero, Andrea Cavallaro, Alessio Brutti, Oswald Lanz, Maurizio Omologo |
ICASSP | 5 |
| 2018 | MORB: A Multi-Scale Binary DescriptorabstractLocal image features play an important role in matching images under different geometric and photometric transformations. However, as the scale difference across views increases, the matching performance may considerably decrease. To address this problem we propose MORB, a multi-scale binary descriptor that is based on ORB and that improves the accuracy of feature matching under scale changes. MORB describes an image patch at different scales using an oriented sampling pattern of intensity comparisons in a predefined set of pixel pairs. We also propose a matching strategy that estimates the cross-scale match between MORB descriptors across views. Experiments show that MORB outperforms state-of-the-art binary descriptors under several transformations. Alessio Xompero, Oswald Lanz, Andrea Cavallaro |
ICIP | 2 |
| 2018 | Joint Estimation of Human Pose and Conversational Groups from Social Scenes
Jagannadan Varadarajan, Subramanian Ramanathan, Samuel Rota Bulò, Narendra Ahuja, Oswald Lanz, Elisa Ricci 0001 |
Int. J. Comput. Vis. | 5 |
| 2017 | Learning to detect violent videos using convolutional long short-term memoryabstractDeveloping a technique for the automatic analysis of surveillance videos in order to identify the presence of violence is of broad interest. In this work, we propose a deep neural network for the purpose of recognizing violent videos. A convolutional neural network is used to extract frame level features from a video. The frame level features are then aggregated using a variant of the long short term memory that uses convolutional gates. The convolutional neural network along with the convolutional long short term memory is capable of capturing localized spatio-temporal features which enables the analysis of local motion taking place in the video. We also propose to use adjacent frame differences as the input to the model thereby forcing it to encode the changes occurring in the video. The performance of the proposed feature extraction pipeline is evaluated on three standard benchmark datasets in terms of recognition accuracy. Comparison of the results obtained with the state of the art techniques revealed the promising capability of the proposed method in recognizing violent videos. Swathikiran Sudhakaran, Oswald Lanz |
AVSS | 2 |
| 2017 | An automatic image-to-DEM alignment approach for annotating mountains pictures on a smartphone
Lorenzo Porzi, Samuel Rota Bulò, Oswald Lanz, Paolo Valigi, Elisa Ricci 0001 |
Mach. Vis. Appl. | 3 |
| 2016 | SALSA: A Novel Dataset for Multimodal Group Behavior AnalysisabstractStudying free-standing conversational groups (FCGs) in unstructured social settings (e.g., cocktail party ) is gratifying due to the wealth of information available at the group (mining social networks) and individual (recognizing native behavioral and personality traits) levels. However, analyzing social scenes involving FCGs is also highly challenging due to the difficulty in extracting behavioral cues such as target locations, their speaking activity and head/body pose due to crowdedness and presence of extreme occlusions. To this end, we propose SALSA, a novel dataset facilitating multimodal and Synergetic sociAL Scene Analysis, and make two main contributions to research on automated social interaction analysis: (1) SALSA records social interactions among 18 participants in a natural, indoor environment for over 60 minutes, under the poster presentation and cocktail party contexts presenting difficulties in the form of low-resolution images, lighting variations, numerous occlusions, reverberations and interfering sound sources; (2) To alleviate these problems we facilitate multimodal analysis by recording the social interplay using four static surveillance cameras and sociometric badges worn by each participant, comprising the microphone, accelerometer, bluetooth and infrared sensors. In addition to raw data, we also provide annotations concerning individuals' personality as well as their position, head, body orientation and F-formation information over the entire event duration. Through extensive experiments with state-of-the-art approaches, we show (a) the limitations of current methods and (b) how the recorded multiple cues synergetically aid automatic analysis of social interactions. SALSA is available at http://tev.fbk.eu/salsa. Xavier Alameda-Pineda, Jacopo Staiano, Subramanian Ramanathan, Ligia Maria Batrinca, Elisa Ricci 0001, Bruno Lepri, Oswald Lanz, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2016 | A Multi-Task Learning Framework for Head Pose Estimation under Target MotionabstractRecently, head pose estimation (HPE) from low-resolution surveillance data has gained in importance. However, monocular and multi-view HPE approaches still work poorly under target motion, as facial appearance distorts owing to camera perspective and scale changes when a person moves around. To this end, we propose FEGA-MTL, a novel framework based on Multi-Task Learning (MTL) for classifying the head pose of a person who moves freely in an environment monitored by multiple, large field-of-view surveillance cameras. Upon partitioning the monitored scene into a dense uniform spatial grid, FEGA-MTL simultaneously clusters grid partitions into regions with similar facial appearance, while learning region-specific head pose classifiers. In the learning phase, guided by two graphs which a-priori model the similarity among (1) grid partitions based on camera geometry and (2) head pose classes, FEGA-MTL derives the optimal scene partitioning and associated pose classifiers. Upon determining the target's position using a person tracker at test time, the corresponding region-specific classifier is invoked for HPE. The FEGA-MTL framework naturally extends to a weakly supervised setting where the target's walking direction is employed as a proxy in lieu of head orientation. Experiments confirm that FEGA-MTL significantly outperforms competing single-task and multi-task learning methods in multi-view settings. Yan Yan 0002, Elisa Ricci 0001, Subramanian Ramanathan, Gaowen Liu, Oswald Lanz, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2015 | Uncovering Interactions and Interactors: Joint Estimation of Head, Body Orientation and F-Formations from Surveillance VideosabstractWe present a novel approach for jointly estimating targets' head, body orientations and conversational groups called F-formations from a distant social scene (e.g., a cocktail party captured by surveillance cameras). Differing from related works that have (i) coupled head and body pose learning by exploiting the limited range of orientations that the two can jointly take, or (ii) determined F-formations based on the mutual head (but not body) orientations of interactors, we present a unified framework to jointly infer both (i) and (ii). Apart from exploiting spatial and orientation relationships, we also integrate cues pertaining to temporal consistency and occlusions, which are beneficial while handling low-resolution data under surveillance settings. Efficacy of the joint inference framework reflects via increased head, body pose and F-formation estimation accuracy over the state-of-the-art, as confirmed by extensive experiments on two social datasets. Elisa Ricci 0001, Jagannadan Varadarajan, Subramanian Ramanathan, Samuel Rota Bulò, Narendra Ahuja, Oswald Lanz |
ICCV | 6 |
| 2015 | Analyzing Free-standing Conversational Groups: A Multimodal ApproachabstractDuring natural social gatherings, humans tend to organize themselves in the so-called free-standing conversational groups. In this context, robust head and body pose estimates can facilitate the higher-level description of the ongoing interplay. Importantly, visual information typically obtained with a distributed camera network might not suffice to achieve the robustness sought. In this line of thought, recent advances in wearable sensing technology open the door to multimodal and richer information flows. In this paper we propose to cast the head and body pose estimation problem into a matrix completion task. We introduce a framework able to fuse multimodal data emanating from a combination of distributed and wearable sensors, taking into account the temporal consistency, the head/body coupling and the noise inherent to the scenario. We report results on the novel and challenging SALSA dataset, containing visual, auditory and infrared recordings of 18 people interacting in a regular indoor environment. We demonstrate the soundness of the proposed method and the usability for higher-level tasks such as the detection of F-formations and the discovery of social attention attractors. Xavier Alameda-Pineda, Yan Yan 0002, Elisa Ricci 0001, Oswald Lanz, Nicu Sebe |
ACM Multimedia | 4 |
| 2015 | Jointly Estimating Interactions and Head, Body Pose of Interactors from Distant Social ScenesabstractWe present joint estimation of F-formations and head, body pose of interactors in a social scene captured by surveillance cameras. Unlike prior works that have focused on (a) discovering F-formations based on head pose and position cues, or (b) jointly learned head and body pose of individuals based on anatomic constraints, we exploit positional and pose cues characterizing interactors and interactions to jointly infer both (a) and (b). We show how the joint inference framework benefits both F-formation and head, body pose estimation accuracy via experiments on two social datasets. Subramanian Ramanathan, Jagannadan Varadarajan, Elisa Ricci 0001, Oswald Lanz, Stefan Winkler 0001 |
ACM Multimedia | 4 |
| 2015 | Dynamic task decomposition for decentralized object tracking in complex scenes
Tao Hu 0007, Stefano Messelodi, Oswald Lanz |
Comput. Vis. Image Underst. | 3 |
| 2014 | Dynamic Task Decomposition for Probabilistic Tracking in Complex ScenesabstractThe employment of visual sensor networks in surveillance systems has brought in as many challenges as advantages. While the integration of multiple cameras into a network has the potential advantage of fusing complementary observations from sensors and enlarging visual coverage, it also increases the complexity of tracking tasks and poses challenges to system scalability. A key approach to tackling these challenges is the mapping of the demanding global task onto a distributed sensing and processing infrastructure. In this paper, we present an efficient and scalable multi-camera multi-people tracking system with a three-layer architecture, in which we formulate the overall task (i.e. tracking all people using all available cameras) as a vision based state estimation problem and aim to maximize utility and sharing of available sensing and processing resources. By exploiting the geometric relations between sensing geometry and people's positions, our method is able to dynamically and adaptively partition the overall task into a number of nearly independent subtasks, each of which tracks a subset of people with a subset of cameras. The method hereby reduces task complexity dramatically and helps to boost parallelization and maximize the real-time throughput and available resources of the system while accounting for intrinsic uncertainty induced, e.g., by visual clutter, occlusion, and illumination changes. We demonstrate the efficiency of our method by testing it with a challenging video sequence. Tao Hu 0007, Stefano Messelodi, Oswald Lanz |
ICPR | 3 |
| 2014 | Evaluating Multi-task Learning for Multi-view Head-Pose Classification in Interactive EnvironmentsabstractSocial attention behavior offers vital cues towards inferring one's personality traits from interactive settings such as round-table meetings and cocktail parties. Head orientation is typically employed as a proxy for determining the social attention direction when faces are captured at low-resolution. Recently, multi-task learning has been proposed to robustly compute head pose under perspective and scale-based facial appearance variations when multiple, distant and large field-of-view cameras are employed for visual analysis in smart-room applications. In this paper, we evaluate the effectiveness of an SVM-based MTL (SVM+MTL) framework with various facial descriptors (KL, HOG, LBP, etc.). The KL+HOG feature combination is found to produce the best classification performance, with SVM+MTL outperforming classical SVM irrespective of the feature used. Yan Yan 0002, Subramanian Ramanathan, Elisa Ricci 0001, Oswald Lanz, Nicu Sebe |
ICPR | 4 |
| 2014 | Exploring Transfer Learning Approaches for Head Pose Classification from Multi-view Surveillance Images
Anoop Kolar Rajagopal, Subramanian Ramanathan, Elisa Ricci 0001, Radu L. Vieriu, Oswald Lanz, Kalpathi Ramakrishnan, Nicu Sebe |
Int. J. Comput. Vis. | 5 |
| 2013 | No Matter Where You Are: Flexible Graph-Guided Multi-task Learning for Multi-view Head Pose Classification under Target MotionabstractWe propose a novel Multi-Task Learning framework (FEGA-MTL) for classifying the head pose of a person who moves freely in an environment monitored by multiple, large field-of-view surveillance cameras. As the target (person) moves, distortions in facial appearance owing to camera perspective and scale severely impede performance of traditional head pose classification methods. FEGA-MTL operates on a dense uniform spatial grid and learns appearance relationships across partitions as well as partition-specific appearance variations for a given head pose to build region-specific classifiers. Guided by two graphs which a-priori model appearance similarity among (i) grid partitions based on camera geometry and (ii) head pose classes, the learner efficiently clusters appearance wise related grid partitions to derive the optimal partitioning. For pose classification, upon determining the target's position using a person tracker, the appropriate region specific classifier is invoked. Experiments confirm that FEGA-MTL achieves state-of-the-art classification with few training data. Yan Yan 0002, Elisa Ricci 0001, Subramanian Ramanathan, Oswald Lanz, Nicu Sebe |
ICCV | 4 |
| 2013 | Multi-scale f-formation discovery for group detectionabstractWe present an unsupervised approach for the automatic detection of static interactive groups. The approach builds upon a novel multi-scale Hough voting policy, which incorporates in a flexible way the sociological notion of group as F-formation; the goal is to model at the same time small arrangements of close friends and aggregations of many individuals spread over a large area. Our technique is based on a competition of different voting sessions, each one specialized for a particular group cardinality; all the votes are then evaluated using information theoretic criteria, producing the final set of groups. The proposed technique has been applied on public benchmark sequences and a novel cocktail party dataset, evaluating new group detection metrics and obtaining state-of-the-art performances. Francesco Setti, Oswald Lanz, Roberta Ferrario, Vittorio Murino, Marco Cristani |
ICIP | 2 |
| 2013 | On the relationship between head pose, social attention and personality prediction for unstructured and dynamic group interactionsabstractCorrelates between social attention and personality traits have been widely acknowledged in social psychology studies. Head pose has commonly been employed as a proxy for determining the social attention direction in small group interactions. However, the impact of head pose estimation errors on personality estimates has not been studied to our knowledge. Subramanian Ramanathan, Yan Yan 0002, Jacopo Staiano, Oswald Lanz, Nicu Sebe |
ICMI | 4 |
| 2012 | An Adaptation Framework for Head-Pose Classification in Dynamic Multi-view Scenarios
Anoop Kolar Rajagopal, Subramanian Ramanathan, Radu L. Vieriu, Elisa Ricci 0001, Oswald Lanz, Kalpathi Ramakrishnan, Nicu Sebe |
ACCV (2) | 5 |
| 2012 | Active transfer learning for multi-view head-pose classification
Yan Yan 0002, Subramanian Ramanathan, Oswald Lanz, Nicu Sebe |
ICPR | 3 |
| 2011 | Dynamic resource allocation for probabilistic tracking via attentive sensing and samplingabstractIn the context of Ambient Intelligence a fundamental challenge is the design of monitoring technologies able to infer activities of people at-a-distance, employing non-intrusive sensors. Ideally, such solutions should operate in real time using minimal resources and scale to environments with complex topologies. These requirements naturally emerge in application domains such as Security & Surveillance, Ambient Assisted Living, Retail Monitoring, etc., and new research challenges are to be faced to push current state-of-the-art towards meeting them. In line with this trend, our recent efforts detailed in this paper focus on some of the limitations of traditional multi-camera based tracking methods arising in this context, which are characterized by passive sensing and limited adaptation. Oswald Lanz, Tao Hu 0007 |
AVSS | 1 |
| 2010 | Tracking Multiple People with Illumination MapsabstractWe address the problem of multiple people tracking under non-homogenous and time-varying illumination conditions. We propose a unified framework for jointly estimating the position of the targets and their illumination conditions. For each target multiple templates are considered to model appearance variations due to lighting changes. The template choice is driven by an illumination map which describes the light conditions in different areas of the scene. This map is computed with a novel algorithm for efficient inference in a hierarchical Markov Random Field (MRF) and is updated online to adapt to slow lighting changes. Experimental results demonstrate the effectiveness of our approach. Gloria Zen, Oswald Lanz, Stefano Messelodi, Elisa Ricci 0001 |
ICPR | 2 |
| 2010 | BabyExp: Constructing a Huge Multimodal Resource to Acquire Commonsense Knowledge Like Children Do
Massimo Poesio, Marco Baroni, Oswald Lanz, Alessandro Lenci, Alexandros Potamianos, Hinrich Schütze, Sabine Schulte im Walde, Luca Surian |
LREC | 3 |
| 2009 | A Sampling Algorithm for Occlusion Robust Multi Target DetectionabstractBayesian methods for visual tracking, with the particle filter as its most prominent instance, have proven to work effectively in the presence of clutter, occlusions, and dynamic background. When applied to track a variable number of targets, however, they become inefficient due to the absence of strong priors. In this paper we present an efficient sampling algorithm for target detection build upon an informed prior that is derived as the inverse of an occlusion robust image likelihood. It has the advantage of being fully integrated in the Bayesian tracking framework, and reactive as it uses sparse features not explained by tracked objects. Oswald Lanz, Stefano Messelodi |
AVSS | 1 |
| 2009 | A HJS filter to track visually interacting targetsabstractVisual tracking with explicit occlusion models is computationally hard, in the sense that the complexity explodes as the number of targets increases. Recently, the hybrid joint-separable (HJS) model has been proposed that enables tracking the local appearance of a number of bodies through occlusions with a quadratic, no more exponential, upper bound. In this paper we extend that method to account for a larger spectrum of visual interactions, captured by a full-image likelihood enabling true Bayesian inference, without compromising scalability. The resulting tracker then proves to be significantly more robust, and able to resolve long term occlusion among five people aligned on a single line-of-sight, observed from a single camera, at a manageable computational cost. Oswald Lanz |
ICASSP | 1 |
| 2006 | Approximate Bayesian Multibody TrackingabstractVisual tracking of multiple targets is a challenging problem, especially when efficiency is an issue. Occlusions, if not properly handled, are a major source of failure. Solutions supporting principled occlusion reasoning have been proposed but are yet unpractical for online applications. This paper presents a new solution which effectively manages the trade-off between reliable modeling and computational efficiency. The Hybrid Joint-Separable (HJS) filter is derived from a joint Bayesian formulation of the problem, and shown to be efficient while optimal in terms of compact belief representation. Computational efficiency is achieved by employing a Markov random field approximation to joint dynamics and an incremental algorithm for posterior update with an appearance likelihood that implements a physically-based model of the occlusion process. A particle filter implementation is proposed which achieves accurate tracking during partial occlusions, while in cases of complete occlusion, tracking hypotheses are bound to estimated occlusion volumes. Experiments show that the proposed algorithm is efficient, robust, and able to resolve long-term occlusions between targets with identical appearance. Oswald Lanz |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Hybrid Joint-Separable Multibody TrackingabstractStatistical models for tracking different moving bodies must be able to reason about occlusions in order to be effective. Representing the joint statistics across different bodies is computationally hard, since the size of the representation grows exponentially with the number of bodies being tracked. Separable tracking, with one tracker per body, cannot deal with occlusions effectively. We propose a new model, dubbed Hybrid Joint-Separable (HJS), that uses a representation size that grows linearly with the number of bodies, and a computational complexity that grows quadratically. This model can reason explicitly about occlusions. We describe a particle filter implementation of this model, and present promising experimental results. Oswald Lanz, Roberto Manduchi |
CVPR (1) | 1 |