Marco La Cascia

dblp:35/4922 · DBLP profile ↗
← Back
37ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0002-8766-6395ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 16 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021
YearPublicationVenuePosition
2025 Sphere-Depth: A Benchmark for Depth Estimation Methods with Varying Spherical Camera Orientations
Soulayma Gazzeh, Giuseppe Mazzola, Liliana Lo Presti, Marco La Cascia
CAIP (1)4
2025 ABBIE: Attention-Based BI-Encoders for Predicting Where to Split Compound Sanskrit Words
abstract
Sanskrit is a highly composite language, morphologically and phonetically complex. One of the major challenges in processing Sanskrit is the splitting of compound words that are merged phonetically. Recognizing the exact location of splits in a compound word is difficult since several possible splits can be found, but only a few of them are semantically meaningful. This paper proposes a novel deep learning method that uses two bi-encoders and a multi-head attention module to predict the valid split location in Sanskrit compound words. The two bi-encoders process the input sequence in direct and reverse order respectively. The model learns the character-level context in which the splitting occurs by exploiting the correlation between the direct and reverse dynamics of the characters sequence. The results of the proposed model are compared with a stateof-the-art technique that adopts a bidirectional recurrent network to solve the same task. Experimental results show that the proposed model correctly identifies where the compound word should be split into its components in 89.27% of cases, outperforming the state-of-the-art technique. The paper also proposes a dataset developed from the repository of the Digital Corpus of Sanskrit (DCS) and the University of Hyderabad (UoH) corpus.
Irfan Ali 0003, Liliana Lo Presti, Igor Spano, Marco La Cascia
ICAART (2)4
2025 A Unified Attention-Based Model for Segmenting Compound Words in Sanskrit
Irfan Ali 0003, Liliana Lo Presti, Igor Spano, Marco La Cascia
ICDAR (4)4
2025 Context-ped: multi-modal context fusion for pedestrian crossing intention prediction
abstract
Abstract Predicting pedestrian crossing intentions is crucial for enhancing autonomous vehicles’ safety and decision-making capabilities. Timely and accurate predictions made well before potential crossing events are essential for enabling vehicles to respond appropriately and prevent accidents. Unlike state-of-the-art models that rely on complex architectures and multiple input modalities, this article proposes Context-Ped, a minimalist effective model for pedestrian intention prediction. To capture critical contextual information, Context-Ped processes two key visual inputs: cropped pedestrian images and environmental road structure. By integrating residual networks with recurrent architectures, the model extracts robust spatiotemporal features through the aggregation of hidden states from ConvLSTM layers. This operation serves as a temporal integration strategy that accumulates the spatiotemporal patterns over the sequence, capturing both dynamic motion cues and static contextual information to encode both pedestrian behavior and environmental context. As a result, Context-Ped achieves accurate intention prediction while maintaining computational efficiency. Additionally, the proposed approach addresses the often-overlooked issue of class imbalance by employing the focal loss function, which significantly enhances performance for the minority class. For evaluation, AUC-ROC and recall metrics are prioritized over traditional accuracy, providing clearer insight into the model’s performance. Experimental results on the JAAD and PIE datasets demonstrate competitive performance, achieving an AUC of 79% and a crossing recall of 82% on the JAAD dataset, outperforming more complex state-of-the-art models and by only using visual input data.
Soulayma Gazzeh, Liliana Lo Presti, Ali Douik, Marco La Cascia
Mach. Vis. Appl.4
2024 Automatic Lemmatization of Old Church Slavonic Language Using A Novel Dictionary-Based Approach
Usman Nawaz, Liliana Lo Presti, Marianna Napolitano, Marco La Cascia
DAS4
2024 Is text preprocessing still worth the time? A comparative survey on the influence of popular preprocessing methods on Transformers and traditional classifiers
abstract
With the advent of the modern pre-trained Transformers, the text preprocessing has started to be neglected and not specificly addressed in recent NLP literature. However, both from a linguistic and from a computer science point of view, we believe that even when using modern Transformers, text preprocessing can significantly impact on the performance of a classification model. We want to investigate and compare, through this study, how preprocessing impacts on the Text Classification (TC) performance of modern and traditional classification models. We report and discuss the preprocessing techniques found in the literature and their most recent variants or applications to address TC tasks in different domains. In order to assess how much the preprocessing affects classification performance, we apply the three top referenced preprocessing techniques (alone or in combination) to four publicly available datasets from different domains. Then, nine machine learning models - including modern Transformers - get the preprocessed text as input. The results presented show that an educated choice on the text preprocessing strategy to employ should be based on the task as well as on the model considered. Outcomes in this survey show that choosing the best preprocessing technique - in place of the worst - can significantly improve accuracy on the classification (up to 25%, as in the case of an XLNet on the IMDB dataset). In some cases, by means of a suitable preprocessing strategy, even a simple Naïve Bayes classifier proved to outperform (i.e., by 2% in accuracy) the best performing Transformer. We found that Transformers and traditional models exhibit a higher impact of the preprocessing on the TC performance. Our main findings are: (1) also on modern pre-trained language models, preprocessing can affect performance, depending on the datasets and on the preprocessing technique or combination of techniques used, (2) in some cases, using a proper preprocessing strategy, simple models can outperform Transformers on TC tasks, (3) similar classes of models exhibit similar level of sensitivity to text preprocessing.
Marco Siino, Ilenia Tinnirello, Marco La Cascia
Inf. Syst.3
2024 Learn & drop: fast learning of cnns based on layer dropping
abstract
Abstract This paper proposes a new method to improve the training efficiency of deep convolutional neural networks. During training, the method evaluates scores to measure how much each layer’s parameters change and whether the layer will continue learning or not. Based on these scores, the network is scaled down such that the number of parameters to be learned is reduced, yielding a speed-up in training. Unlike state-of-the-art methods that try to compress the network to be used in the inference phase or to limit the number of operations performed in the back-propagation phase, the proposed method is novel in that it focuses on reducing the number of operations performed by the network in the forward propagation during training. The proposed training strategy has been validated on two widely used architecture families: VGG and ResNet. Experiments on MNIST, CIFAR-10 and Imagenette show that, with the proposed method, the training time of the models is more than halved without significantly impacting accuracy. The FLOPs reduction in the forward propagation during training ranges from 17.83% for VGG-11 to 83.74% for ResNet-152. As for the accuracy, the impact depends on the depth of the model and the decrease is between 0.26% and 2.38% for VGGs and between 0.4 and 3.2% for ResNets. These results demonstrate the effectiveness of the proposed technique in speeding up learning of CNNs. The technique will be especially useful in applications where fine-tuning or online training of convolutional models is required, for instance because data arrive sequentially.
Giorgio Cruciata, Luca Cruciata, Liliana Lo Presti, Jan C. van Gemert, Marco La Cascia
Neural Comput. Appl.5
2023 RLSTM: A Novel Residual and Recurrent Network for Pedestrian Action Classification
Soulayma Gazzeh, Liliana Lo Presti, Ali Douik, Marco La Cascia
CAIP (2)4
2018 Multi-modal Medical Image Registration by Local Affine Transformations
abstract
Image registration is the process of finding the geometric transformation that, applied to the floating image, gives the registered image with the highest similarity to the reference image. Registering a pair of images involves the definition of a similarity function in terms of the parameters of the geometric transformation that allows the registration. This paper proposes to register a pair of images by iteratively maximizing the empirical mutual information through coordinate gradient descent. Hence, the registered image is obtained by applying a sequence of local affine transformations. Rather than adopting a uniformly spaced grid to select image blocks to locally register, as done by state-of-the-art techniques, this paper proposes a method which is similar in spirit to boosting strategies used in classification. In this work, a probability distribution over the pixels of the registered image is maintained. At each pixel, this distribution represents the probability that a local affine transformation of a block centered on this pixel should be computed to improve the similarity between the registered and the reference images. The distribution is updated iteratively during the registration process to move probability mass towards pixels unaffected by the estimated local transformation. The paper presents preliminary results by a qualitative evaluation on several pairs of medical images acquired by different sources.
Liliana Lo Presti, Marco La Cascia
ICPRAM2
2017 Boosting Hankel matrices for face emotion recognition and pain detection
Liliana Lo Presti, Marco La Cascia
Comput. Vis. Image Underst.2
2016 A Novel Time Series Kernel for Sequences Generated by LTI Systems
Liliana Lo Presti, Marco La Cascia
ACCV (3)2
2016 3D skeleton-based human action classification: A survey
Liliana Lo Presti, Marco La Cascia
Pattern Recognit.2
2015 Keyword based Keyframe Extraction in Online Video Collections
Edoardo Ardizzone, Marco La Cascia, Giuseppe Mazzola
ICPRAM (2)2
2015 Hankelet-based dynamical systems modeling for 3D action recognition
abstract
This paper proposes to model an action as the output of a sequence of atomic Linear Time Invariant (LTI) systems. The sequence of LTI systems generating the action is modeled as a Markov chain , where a Hidden Markov Model (HMM) is used to model the transition from one atomic LTI system to another. In turn, the LTI systems are represented in terms of their Hankel matrices. For classification purposes, the parameters of a set of HMMs (one for each action class) are learned via a discriminative approach. This work proposes a novel method to learn the atomic LTI systems from training data , and analyzes in detail the action representation in terms of a sequence of Hankel matrices. Extensive evaluation of the proposed approach on two publicly available datasets demonstrates that the proposed method attains state-of-the-art accuracy in action classification from the 3D locations of body joints (skeleton).
Liliana Lo Presti, Marco La Cascia, Stan Sclaroff, Octavia I. Camps
Image Vis. Comput.2
2014 Gesture Modeling by Hanklet-Based Hidden Markov Model
Liliana Lo Presti, Marco La Cascia, Stan Sclaroff, Octavia I. Camps
ACCV (3)2
2014 Video Object Recognition and Modeling by SIFT Matching Optimization
abstract
In this paper we present a novel technique for object modeling and object recognition in video. Given a set of videos containing 360 degrees views of objects we compute a model for each object, then we analyze short videos to determine if the object depicted in the video is one of the modeled objects. The object model is built from a video spanning a 360 degree view of the object taken against a uniform background. In order to create the object model, the proposed techniques selects a few representative frames from each video and local features of such frames. The object recognition is performed selecting a few frames from the query video, extracting local features from each frame and looking for matches in all the representative frames constituting the models of all the objects. If the number of matches exceed a fixed threshold the corresponding object is considered the recognized objects .To evaluate our approach we acquired a dataset of 25 videos representing 25 different objects and used these videos to build the objects model. Then we took 25 test videos containing only one of the known objects and 5 videos containing only unknown objects. Experiments showed that, despite a significant compression in the model, recognition results are satisfactory.
Alessandro Bruno, Luca Greco 0002, Marco La Cascia
ICPRAM3
2014 Concurrent photo sequence organization
Liliana Lo Presti, Marco La Cascia
Multim. Tools Appl.2
2013 Object Recognition and Modeling Using SIFT Features
Alessandro Bruno, Luca Greco 0002, Marco La Cascia
ACIVS3
2013 An Automated Visual Inspection System for the Classification of the Phases of Ti-6Al-4V Titanium Alloy
Antonino Ducato, Livan Fratini, Marco La Cascia, Giuseppe Mazzola
CAIP (2)3
2012 User detection through multi-sensor fusion in an AmI scenario
Alessandra De Paola, Marco La Cascia, Giuseppe Lo Re, Marco Morana, Marco Ortolani
FUSION2
2012 An on-line learning method for face association in personal photo collection
Liliana Lo Presti, Marco La Cascia
Image Vis. Comput.2
2012 A data association approach to detect and organize people in personal photo collections
Liliana Lo Presti, Marco Morana, Marco La Cascia
Multim. Tools Appl.3
2012 Path Modeling and Retrieval in Distributed Video Surveillance Databases
abstract
We propose a framework for querying a distributed database of video surveillance data in order to retrieve a set of likely paths of a person moving in the area under surveillance. In our framework, each camera of the surveillance system locally processes the data and stores video sequences in a storage unit and the metadata for each detected person in the distributed database. A pedestrian's path is formulated as a dynamic Bayesian network (DBN) to model the dependencies between subsequent observations of the person as he makes his way through the camera network. We propose a tool by which the analyst can pose queries about where a certain person appeared while moving in the site during a specified temporal window. The DBN is used in an algorithm that finds potentially relevant metadata records from the distributed databases and then assembles these into probable paths that the person took in the camera network. Finally, the system presents the analyst with the retrieved set of likely paths in ranked order. The computational complexity for our method is quadratic in the number of camera nodes and linear in the number of moving persons. Experiments were carried out on simulated data to test the system with large distributed databases and in a real setting in which six databases store the data from six video cameras. The simulations confirm that our method provides good results with varying numbers of cameras and persons moving in the network. In a real setting, the method reconstructs paths across the camera network with approximatively 75% accuracy at rank 1.
Liliana Lo Presti, Stan Sclaroff, Marco La Cascia
IEEE Trans. Multim.3
2010 Mobile Interface for Content-Based Image Management
abstract
People make more and more use of digital image acquisition devices to capture screenshots of their everyday life. The growing number of personal pictures raise the problem of their classification. Some of the authors proposed an automatic technique for personal photo album management dealing with multiple aspects (i. e., people, time and background) in a homogenous way. In this paper we discuss a solution that allows mobile users to remotely access such technique by means of their mobile phones, almost from everywhere, in a pervasive fashion. This allows users to classify pictures they store on their devices. The whole solution is presented, with particular regard to the user interface implemented on the mobile phone, along with some experimental results.
Marco La Cascia, Marco Morana, Salvatore Sorce
CISIS1
2010 A Data Association Algorithm for People Re-identification in Photo Sequences
abstract
In this paper, a new system is presented to support the user in the face annotation task. Every time a photo sequence becomes available, the system analyses it to detect and cluster faces in set corresponding to the same person. We propose to model the problem of people re-identification in photos as a data association problem. In this way, the system takes advantage from the assumption that each person can appear at most once in each photo. We propose a fully automated method for grouping facial images, the method does not require any initialization neither a priori knowledge of the number of persons that are in the photo sequence. We compare the results obtained with our method and with standard clustering methods on three personal collections and on a publicly available dataset.
Liliana Lo Presti, Marco Morana, Marco La Cascia
ISM3
2008 Mean shift clustering for personal photo album organization
abstract
In this paper we propose a probabilistic approach for the automatic organization of pictures in personal photo album. Images are analyzed in term of faces and low-level visual features of the background. The description of the background is based on RGB color histogram and on Gabor filter energy accounting for texture information. The face descriptor is obtained by projection of detected and rectified faces on a common low dimensional eigenspace. Vectors representing faces and background are clustered in an unsupervised fashion exploiting a mean shift clustering technique. We observed that, given the peculiarity of the domain of personal photo libraries where most of the pictures contain faces of a relatively small number of different individuals, clusters tend to be not only visually but also semantically significant. Experimental results are reported
Edoardo Ardizzone, Marco La Cascia, Filippo Vella
ICIP2
2007 A P2P Architecture for Multimedia Content Retrieval
Edoardo Ardizzone, Luca Gatani, Marco La Cascia, Giuseppe Lo Re, Marco Ortolani
MMM (1)3
2004 Fully automatic, real-time detection of facial gestures from generic video
abstract
A technique for the detection of facial gestures from low resolution video sequences is presented. The technique builds upon the automatic 3D head tracker formulation of [M. La Cascia et al., 2000]. The tracker is based on the registration of a texture-mapped cylindrical model. Facial gesture analysis is performed in the texture map by assuming that the residual registration error can be modeled as a linear combination of facial motion templates. Two formulations are proposed and tested. In one formulation, the head and facial motion are estimated in a single, combined linear system. In the other formulation, head motion and then facial motion are estimated in a two-step process. The two-step approach significantly yields better accuracy in facial gesture analysis. The system is demonstrated in detecting two types of facial gestures: "mouth opening" and "eyebrows raising." On a dataset with lots of head motion, the two-step algorithm achieved a recognition accuracy of 70% for the "mouth opening" and an accuracy of 66% for "eyebrows raising" gestures. The algorithm can reliably track and classify facial gestures without any user intervention and runs in real-time.
Marco La Cascia, Lorenzo Valenti, Stan Sclaroff
MMSP1
2000 Fast, Reliable Head Tracking under Varying Illumination: An Approach Based on Registration of Texture-Mapped 3D Models
abstract
A technique for 3D head tracking under varying illumination is proposed. The head is modeled as a texture mapped cylinder. Tracking is formulated as an image registration problem in the cylinder's texture map image. The resulting dynamic texture map provides a stabilized view of the face that can be used as input to many existing 2D techniques for face recognition, facial expressions analysis, lip reading, and eye tracking. To solve the registration problem with lighting variation and head motion, the residual registration error is modeled as a linear combination of texture warping templates and orthogonal illumination templates. Fast stable online tracking is achieved via regularized weighted least-squares error minimization. The regularization tends to limit potential ambiguities that arise in the warping and illumination templates. It enables stable tracking over extended sequences. Tracking does not require a precise initial model fit; the system is initialized automatically using a simple 2D face detector. It is assumed that the target is facing the camera in the first frame. The formulation uses texture mapping hardware. The nonoptimized implementation runs at about 15 frames per second on a SGI O2 graphic workstation. Extensive experiments evaluating the effectiveness of the formulation are reported. The sensitivity of the technique to illumination, regularization parameters, errors in the initial positioning, and internal camera parameters are analyzed. Examples and applications of tracking are reported.
Marco La Cascia, Stan Sclaroff, Vassilis Athitsos
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 Fast, Reliable Head Tracking under Varying Illumination
abstract
An improved technique for 3D head tracking under varying illumination conditions is proposed. The head is modeled as a texture mapped cylinder. Tracking is formulated as an image registration problem in the cylinder's texture map image. To solve the registration problem in the presence of lighting variation and head motion, the residual error of registration is modeled as a linear combination of texture warping templates and orthogonal illumination templates. Fast and stable on-line tracking is then achieved via regularized weighted least squares minimization of the registration error. The regularization term tends to limit potential ambiguities that arise in the warping and illumination templates. Tracking does not require a precise initial of the model; the system is initialized automatically using a simple 2D face detector. The only assumption is that the target is facing the camera in the first frame of the sequence. Experiments in tracking are reported.
Marco La Cascia, Stan Sclaroff
CVPR1
1999 Unifying Textual and Visual Cues for Content-Based Image Retrieval on the World Wide Web
abstract
A system is proposed that combines textual and visual statistics in a single index vector for content-based search of a WWW image database. Textual statistics are captured in vector form using latent semantic indexing based on text in the containing HTML document. Visual statistics are captured in vector form using color and orientation histograms. By using an integrated approach, it becomes possible to take advantage of possible statistical couplings between the content of the document (latent semantic content) and the contents of images (visual statistics). The combined approach allows improved performance in conducting content-based search. Search performance experiments are reported for a database containing 350,000 images collected from the WWW.
Stan Sclaroff, Marco La Cascia, Saratendu Sethi, Leonid Taycher
Comput. Vis. Image Underst.2
1998 Head Tracking via Robust Registration in Texture Map Images
abstract
A novel method for 3D head tracking in the presence of large head rotations and facial expression changes is described. Tracking is formulated in terms of color image registration in the texture map of a 3D surface model. Model appearance is recursively updated via image mosaicking in the texture map as the head orientation varies. The resulting dynamic texture map provides a stabilized view of the face that can be used as input to many existing 2D techniques for face recognition, facial expressions analysis, lip reading, and eye tracking. Parameters are estimated via a robust minimization procedure; this provides robustness to occlusions, wrinkles, shadows and specular highlights. The system was tested on a variety of sequences taken with low quality, uncalibrated video cameras. Experimental results are reported.
Marco La Cascia, John Isidoro, Stan Sclaroff
CVPR1
1997 Automatic Video Database Indexing and Retrieval
Edoardo Ardizzone, Marco La Cascia
Multim. Tools Appl.2
1996 JACOB: just a content-based query system for video databases
abstract
The increasing development of advanced multimedia applications requires new technologies for organizing and retrieving by content databases of still digital images or digital video sequences. The authors describe JACOB, a prototypal system allowing content-based browsing and querying in video databases. The JACOB system automatically splits a video into a sequence of shots, extracts a few representative frames (said r-frames) from each shot and computes r-frame descriptors based on features like color and texture. No user action is required during the database population step. Queries exploit this image content description and may be direct or by example.
Marco La Cascia, Edoardo Ardizzone
ICASSP1
1996 Video indexing using optical flow field
abstract
The increasing development of advanced multimedia applications requires new technologies for organizing and retrieving by content databases of digital video. Several content based features (color, texture, motion, etc.) are needed to perform a reliable content based retrieval. We present a method for automatic motion based video indexing and retrieval. A prototypal system has been developed to prove the validity of our approach. Our system automatically splits a video into a sequence of shots, extracts a few representative frames (said r-frames) from each shot and computes some motion based features related to the optical flow field. Motion based queries are then performed either in a qualitative or quantitative way. The results obtained with our system proved that motion based query can play a central role in content based video retrieval.
Edoardo Ardizzone, Marco La Cascia
ICIP (3)2
1996 Content-based indexing of image and video databases by global and shape features
abstract
Indexing and retrieval methods based on the image content are required to effectively use information from the large repositories of digital images and videos currently available. Both global (colour, texture, motion, etc.) and local (object shape, etc.) features are needed to perform a reliable content based retrieval. We present a method for automatic extraction of global image features, like colour and motion parameters, and their use for data restriction in video database querying. Further retrieval is therefore accomplished, in a restricted set of images, by shape feature (skeleton, local symmetry moments, correlation, etc.) local search. The proposed indexing methodology has been developed and tested inside JACOB, a prototypal system for content-based video database querying.
Edoardo Ardizzone, Marco La Cascia, Vito Di Gesù, Cesare Valenti
ICPR2
1996 Motion and color-based video indexing and retrieval
abstract
In this paper we present a method for automatic motion and color based video indexing and retrieval. Our system automatically splits a video into a sequence of shots and extracts a few representative frames (r-frames) from each shot. For each r-frame we compute the optical flow field; motion features are then derived from the flow field. Color features are related to the three-dimensional RGB color histogram. Queries (direct or by example) are based on these features. Obtained results proved that motion and color based querying can play a central role in content based video retrieval.
Edoardo Ardizzone, Marco La Cascia, Davide Molinelli
ICPR2