EDBT 2026 Demo / reviewers in the wild / expert
Alessandro Bissacco
dblp:93/2043
· DBLP profile ↗
19ranked-venue papers
9as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 9 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Image recognition and object detection · 40% Segmentation and scene understanding · 17% Face, body and person analysis · 11% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 84% Knowledge graphs · 16% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 61% Computer animation and physical simulation · 30% Audio and music processing · 9% |
Topics — the 30 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
scene text detection |
1.0 | 2 | 2022 | Towards End-to-End Unified Scene Text Detection and Layout Analysis · CVPR 2022 Towards Unconstrained End-to-End Text Spotting · ICCV 2019 |
Computer vision › Image recognition and object detection
scene text recognition |
0.5 | 2 | 2019 | Towards Unconstrained End-to-End Text Spotting · ICCV 2019 PhotoOCR: Reading Text in Uncontrolled Conditions · ICCV 2013 |
Computer vision › Video understanding and tracking › motion analysis
human motion analysis |
0.2 | 4 | 2009 | Hybrid Dynamical Models of Human Motion for the Recognition of Human Gaits · Int. J. Comput. Vis. 2009 Classifying Human Dynamics Without Contact Forces · CVPR (2) 2006 Modeling and Learning Contact Dynamics in Human Motion · CVPR (1) 2005 |
Computer vision › Face, body and person analysis
human pose estimation |
0.1 | 2 | 2007 | Fast Human Pose Estimation using Appearance and Motion via Multi-Dimensional Boosting Regression · CVPR 2007 Detecting Humans via Their Pose · NIPS 2006 |
Computer vision › Face, body and person analysis › gait analysis
gait recognition |
0.1 | 2 | 2009 | Hybrid Dynamical Models of Human Motion for the Recognition of Human Gaits · Int. J. Comput. Vis. 2009 Recognition of Human Gaits · CVPR (2) 2001 |
Information retrieval › image retrieval › large-scale image retrieval
web-scale image retrieval |
0.1 | 2 | 2009 | Tour the world: Building a web-scale landmark recognition engine · CVPR 2009 Tour the world: a technical demonstration of a web-scale landmark recognition engine · ACM Multimedia 2009 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.1 | 1 | 2019 | Towards Unconstrained End-to-End Text Spotting · ICCV 2019 |
Machine learning › Generative modeling
autoregressive model |
0.1 | 1 | 2009 | Hybrid Dynamical Models of Human Motion for the Recognition of Human Gaits · Int. J. Comput. Vis. 2009 |
Computer vision › Face, body and person analysis
face detection |
0.1 | 1 | 2009 | Large-scale privacy protection in Google Street View · ICCV 2009 |
Robotics › Motion planning and robot control › hybrid systems
hybrid dynamics |
0.1 | 1 | 2009 | Hybrid Dynamical Models of Human Motion for the Recognition of Human Gaits · Int. J. Comput. Vis. 2009 |
Robotics › Robot navigation and mapping
landmark detection |
0.1 | 1 | 2009 | Tour the world: Building a web-scale landmark recognition engine · CVPR 2009 |
Multimedia analysis and retrieval › object recognition
landmark recognition |
0.1 | 1 | 2009 | Tour the world: a technical demonstration of a web-scale landmark recognition engine · ACM Multimedia 2009 |
Privacy and data protection
image privacy |
0.1 | 1 | 2009 | Large-scale privacy protection in Google Street View · ICCV 2009 |
Computer vision › 3D vision
3d human pose estimation |
0.1 | 1 | 2007 | Fast Human Pose Estimation using Appearance and Motion via Multi-Dimensional Boosting Regression · CVPR 2007 |
Computer vision › Face, body and person analysis › human pose estimation
regression-based pose estimation |
0.1 | 1 | 2007 | Fast Human Pose Estimation using Appearance and Motion via Multi-Dimensional Boosting Regression · CVPR 2007 |
Mathematical optimization
optimal transport |
0.1 | 1 | 2007 | Classification and Recognition of Dynamical Models: The Role of Phase, Independent Components, Kernels and Optimal Transport · IEEE Trans. Pattern Anal. Mach. Intell. 2007 |
Computer vision › Video understanding and tracking › action recognition
gait classification |
0.1 | 1 | 2006 | Classifying Human Dynamics Without Contact Forces · CVPR (2) 2006 |
Natural language and speech › Information extraction and text analysis › topic model
latent dirichlet allocation |
0.1 | 1 | 2006 | Detecting Humans via Their Pose · NIPS 2006 |
Computer vision › Image recognition and object detection › object detection › category-specific object detection
person detection |
0.1 | 1 | 2006 | Detecting Humans via Their Pose · NIPS 2006 |
Natural language and speech › Information extraction and text analysis
topic model |
0.1 | 1 | 2006 | Detecting Humans via Their Pose · NIPS 2006 |
Robotics › Motion planning and robot control › robot dynamics
contact dynamics |
0.1 | 1 | 2005 | Modeling and Learning Contact Dynamics in Human Motion · CVPR (1) 2005 |
Machine learning › Time series and sequential data › linear dynamical systems
switching linear dynamical system |
0.1 | 1 | 2005 | Modeling and Learning Contact Dynamics in Human Motion · CVPR (1) 2005 |
Computer animation and physical simulation
facial animation |
0.0 | 1 | 2004 | Modeling and Synthesis of Facial Motion Driven by Speech · ECCV (3) 2004 |
Machine learning › Representation and self-supervised learning
dynamical system representation |
0.0 | 1 | 2001 | Recognition of Human Gaits · CVPR (2) 2001 |
Computer vision › Image recognition and object detection › object detection
false positive reduction |
0.0 | 1 | 2009 | Large-scale privacy protection in Google Street View · ICCV 2009 |
Information retrieval
image retrieval |
0.0 | 1 | 2009 | Tour the world: a technical demonstration of a web-scale landmark recognition engine · ACM Multimedia 2009 |
Knowledge graphs
knowledge graph construction |
0.0 | 1 | 2009 | Tour the world: Building a web-scale landmark recognition engine · CVPR 2009 |
Computer vision › Video understanding and tracking › motion analysis
human motion recognition |
0.0 | 1 | 2007 | Classification and Recognition of Dynamical Models: The Role of Phase, Independent Components, Kernels and Optimal Transport · IEEE Trans. Pattern Anal. Mach. Intell. 2007 |
Computer vision › Face, body and person analysis › human pose estimation
real-time pose estimation |
0.0 | 1 | 2007 | Fast Human Pose Estimation using Appearance and Motion via Multi-Dimensional Boosting Regression · CVPR 2007 |
Audio and music processing
speech processing |
0.0 | 1 | 2004 | Modeling and Synthesis of Facial Motion Driven by Speech · ECCV (3) 2004 |
Methods — techniques the papers use, named apart from their topics
end-to-end unified model · 0.6roi masking · 0.4partially labeled training data · 0.4attention · 0.4visual model learning · 0.2unsupervised clustering · 0.2image matching · 0.2image clustering · 0.2GPS-tagged photo mining · 0.2distributed language modelling · 0.2deep learning classification · 0.2autoregressive model · 0.2sliding-window detection · 0.1system identification · 0.1kernel methods · 0.1independent component analysis · 0.1motion synthesis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Hierarchical Text Spotter for Joint Text Spotting and Layout AnalysisabstractWe propose Hierarchical Text Spotter (HTS), a novel method for the joint task of word-level text spotting and geometric layout analysis. HTS can recognize text in an image and identify its 4-level hierarchical structure: characters, words, lines, and paragraphs. The proposed HTS is characterized by two novel components: (1) a Unified-DetectorPolygon (UDP) that produces Bezier Curve polygons of text lines and an affinity matrix for paragraph grouping between detected lines; (2) a Line-to-Character-to-Word (L2C2W) recognizer that splits lines into characters and further merges them back into words. HTS achieves stateof-the-art results on multiple word-level text spotting benchmark datasets as well as geometric layout analysis tasks. Shangbang Long, Siyang Qin, Yasuhisa Fujii, Alessandro Bissacco, Michalis Raptis |
WACV | 4 |
| 2023 | ICDAR 2023 Competition on Hierarchical Text Detection and Recognition
Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, Michalis Raptis |
ICDAR (2) | 4 |
| 2023 | Text Reading Order in Uncontrolled Conditions by Sparse Graph Segmentation
Renshen Wang, Yasuhisa Fujii, Alessandro Bissacco |
ICDAR (6) | 3 |
| 2022 | Towards End-to-End Unified Scene Text Detection and Layout AnalysisabstractScene text detection and document layout analysis have long been treated as two separate tasks in different image domains. In this paper, we bring them together and introduce the task of unified scene text detection and layout analysis. The first hierarchical scene text dataset is introduced to enable this novel research task. We also propose a novel method that is able to simultaneously detect scene text and form text clusters in a unified way. Comprehensive experiments show that our unified model achieves better performance than multiple well-designed baseline methods. Additionally, this model achieves state-of-the-art results on multiple scene text detection datasets without the need of complex post-processing. Dataset and code: https://github.com/google-research-datasets/hiertext. Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, Michalis Raptis |
CVPR | 4 |
| 2019 | Towards Unconstrained End-to-End Text SpottingabstractWe propose an end-to-end trainable network that can simultaneously detect and recognize text of arbitrary shape, making substantial progress on the open problem of reading scene text of irregular shape. We formulate arbitrary shape text detection as an instance segmentation problem; an attention model is then used to decode the textual content of each irregularly shaped text region without rectification. To extract useful irregularly shaped text instance features from image scale features, we propose a simple yet effective RoI masking step. Additionally, we show that predictions from an existing multi-step OCR engine can be leveraged as partially labeled training data, which leads to significant improvements in both the detection and recognition accuracy of our model. Our method surpasses the state-of-the-art for end-to-end recognition tasks on the ICDAR15 (straight) benchmark by 4.6%, and on the Total-Text (curved) benchmark by more than 16%. Siyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii |
ICCV | 2 |
| 2013 | PhotoOCR: Reading Text in Uncontrolled ConditionsabstractWe describe Photo OCR, a system for text extraction from images. Our particular focus is reliable text extraction from smartphone imagery, with the goal of text recognition as a user input modality similar to speech recognition. Commercially available OCR performs poorly on this task. Recent progress in machine learning has substantially improved isolated character classification, we build on this progress by demonstrating a complete OCR system using these techniques. We also incorporate modern data center-scale distributed language modelling. Our approach is capable of recognizing text in a variety of challenging imaging conditions where traditional OCR systems fail, notably in the presence of substantial blur, low resolution, low contrast, high image noise and other distortions. It also operates with low latency, mean processing time is 600 ms per image. We evaluate our system on public benchmark datasets for text extraction and outperform all previously reported results, more than halving the error rate on multiple benchmarks. The system is currently in use in many applications at Google, and is available as a user input modality in Google Translate for Android. Alessandro Bissacco, Mark Joseph Cummins, Yuval Netzer, Hartmut Neven |
ICCV | 1 |
| 2009 | Tour the world: Building a web-scale landmark recognition engineabstractModeling and recognizing landmarks at world-scale is a useful yet challenging task. There exists no readily available list of worldwide landmarks. Obtaining reliable visual models for each landmark can also pose problems, and efficiency is another challenge for such a large scale system. This paper leverages the vast amount of multimedia data on the Web, the availability of an Internet image search engine, and advances in object recognition and clustering techniques, to address these issues. First, a comprehensive list of landmarks is mined from two sources: (1) ~20 million GPS-tagged photos and (2) online tour guide Web pages. Candidate images for each landmark are then obtained from photo sharing Websites or by querying an image search engine. Second, landmark visual models are built by pruning candidate images using efficient image matching and unsupervised clustering techniques. Finally, the landmarks and their visual models are validated by checking authorship of their member images. The resulting landmark recognition engine incorporates 5312 landmarks from 1259 cities in 144 countries. The experiments demonstrate that the engine can deliver satisfactory recognition performance with high efficiency. Yantao Zheng, Ming Zhao 0003, Yang Song 0009, Hartwig Adam, Ulrich Buddemeier, Alessandro Bissacco, Fernando Brucher, Tat-Seng Chua, Hartmut Neven |
CVPR | 6 |
| 2009 | Large-scale privacy protection in Google Street ViewabstractThe last two years have witnessed the introduction and rapid expansion of products based upon large, systematically-gathered, street-level image collections, such as Google Street View, EveryScape, and Mapjack. In the process of gathering images of public spaces, these projects also capture license plates, faces, and other information considered sensitive from a privacy standpoint. In this work, we present a system that addresses the challenge of automatically detecting and blurring faces and license plates for the purpose of privacy protection in Google Street View. Though some in the field would claim face detection is “solved”, we show that state-of-the-art face detectors alone are not sufficient to achieve the recall desired for large-scale privacy protection. In this paper we present a system that combines a standard sliding-window detector tuned for a high recall, low-precision operating point with a fast post-processing stage that is able to remove additional false positives by incorporating domain-specific information not available to the sliding-window detector. Using a completely automatic system, we are able to sufficiently blur more than 89% of faces and 94 - 96% of license plates in evaluation sets sampled from Google Street View imagery. Andrea Frome, German Cheung, Ahmad Abdulkader, Marco Zennaro, Bo Wu 0001, Alessandro Bissacco, Hartwig Adam, Hartmut Neven, Luc Vincent |
ICCV | 6 |
| 2009 | Tour the world: a technical demonstration of a web-scale landmark recognition engineabstractWe present a technical demonstration of a world-scale touristic landmark recognition engine. To build such an engine, we leverage ~21.4 million images, from photo sharing websites and Google Image Search, and around two thousand web articles to mine the landmark names and learn the visual models. The landmark recognition engine incorporates 5312 landmarks from 1259 cities in 144 countries. This demonstration gives three exhibits: (1) a live landmark recognition engine that can visually recognize landmarks in a given image; (2) an interactive navigation tool showing landmarks on Google Earth; and (3) sample visual clusters (landmark model images) and a list of 1000 randomly selected landmarks from our recognition engine with their iconic images. Yantao Zheng, Ming Zhao 0003, Yang Song 0009, Hartwig Adam, Ulrich Buddemeier, Alessandro Bissacco, Fernando Brucher, Tat-Seng Chua, Hartmut Neven, Jay Yagnik |
ACM Multimedia | 6 |
| 2009 | Hybrid Dynamical Models of Human Motion for the Recognition of Human GaitsabstractWe propose a hybrid dynamical model of human motion and develop a classification algorithm for the purpose of analysis and recognition. We assume that some temporal statistics are extracted from the images, and use them to infer a dynamical model that explicitly represents ground contact events. Such events correspond to “switches” between symmetric sets of hidden parameters in an auto-regressive model. We propose novel algorithms to estimate switches and model parameters, and develop a distance between such models that explicitly factors out exogenous inputs that are not unique to an individual or his/her gait. We show that such a distance is more discriminative than the distance between simple linear systems for the task of gait recognition. Alessandro Bissacco, Stefano Soatto |
Int. J. Comput. Vis. | 1 |
| 2007 | On the Blind Classification of Time SeriesabstractWe propose a cord distance in the space of dynamical models that takes into account their dynamics, including transients, output maps and input distributions. In data analysis applications, as opposed to control, the input is often not known and is inferred as part of the (blind) identification. So it is an integral part of the model that should be considered when comparing different time series. Previous work on kernel distances between dynamical models assumed either identical or independent inputs. We extend it to arbitrary distributions, highlighting connections with system identification, independent component analysis, and optimal transport. The increased modeling power is demonstrated empirically on gait classification from simple visual features. Alessandro Bissacco, Stefano Soatto |
CVPR | 1 |
| 2007 | Fast Human Pose Estimation using Appearance and Motion via Multi-Dimensional Boosting RegressionabstractWe address the problem of estimating human pose in video sequences, where rough location has been determined. We exploit both appearance and motion information by defining suitable features of an image and its temporal neighbors, and learning a regression map to the parameters of a model of the human body using boosting techniques. Our algorithm can be viewed as a fast initialization step for human body trackers, or as a tracker itself. We extend gradient boosting techniques to learn a multi-dimensional map from (rotated and scaled) Haar features to the entire set of joint angles representing the full body pose. We test our approach by learning a map from image patches to body joint angles from synchronized video and motion capture walking data. We show how our technique enables learning an efficient real-time pose estimator, validated on publicly available datasets. Alessandro Bissacco, Ming-Hsuan Yang 0001, Stefano Soatto |
CVPR | 1 |
| 2007 | Classification and Recognition of Dynamical Models: The Role of Phase, Independent Components, Kernels and Optimal TransportabstractWe address the problem of performing decision tasks, and in particular classification and recognition, in the space of dynamical models in order to compare time series of data. Motivated by the application of recognition of human motion in image sequences, we consider a class of models that include linear dynamics, both stable and marginally stable (periodic), both minimum and non-minimum phase, driven by non-Gaussian processes. This requires extending existing learning and system identification algorithms to handle periodic modes and nonminimum phase behavior, while taking into account higher-order statistics of the data. Once a model is identified, we define a kernel-based cord distance between models that includes their dynamics, their initial conditions as well as input distribution. This is made possible by a novel kernel defined between two arbitrary (non-Gaussian) distributions, which is computed by efficiently solving an optimal transport problem. We validate our choice of models, inference algorithm, and distance on the tasks of human motion synthesis (sample paths of the learned models), and recognition (nearest-neighbor classification in the computed distance). However, our work can be applied more broadly where one needs to compare historical data while taking into account periodic trends, non-minimum phase behavior, and non-Gaussian input distributions. Alessandro Bissacco, Alessandro Chiuso, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | Classifying Human Dynamics Without Contact ForcesabstractWe develop a classification algorithm for hybrid autoregressive models of human motion for the purpose of videobased analysis and recognition. We assume that some temporal statistics are extracted from the images, and we use them to infer a dynamical system that explicitly models contact forces. We then develop a distance between such models that explicitly factors out exogenous inputs that are not unique to an individual or her gait. We show that such a distance is more discriminative than the distance between simple linear systems, where most of the energy is devoted to modeling the dynamics of spurious nuisances such as contact forces. Alessandro Bissacco, Stefano Soatto |
CVPR (2) | 1 |
| 2006 | High Performance Feature Detection on a Reconfigurable Co-ProcessorabstractIn this paper, the authors propose a new design for feature detection used for tracking, which eliminates the need of a central computer to complete computations for the feature selection algorithm. Such a system constrains performance due to the delay in which data is transferred from camera to computer for processing. Our design suggests that feature detection computation can be done on a processor within the camera helping to reduce overall computation time for detection and increase performance for overall tracking system. However, these systems are often constrained by the processing power available to the camera. But with Benedetti and Perona's approach to Tomasi and Kanade's detection algorithm, such a design is possible to implement onto a camera system which would eliminate the delay and also improve performance over a tracking system designed on software Jia Ming Mar, Alessandro Bissacco, Stefano Soatto, Soheil Ghiasi |
FCCM | 2 |
| 2006 | Detecting Humans via Their PoseabstractWe consider the problem of detecting humans and classifying their pose from a single image. Specifically, our goal is to devise a statistical model that simultaneously answers two questions: 1) is there a human in the image? and, if so, 2) what is a low-dimensional representation of her pose? We investigate models that can be learned in an unsupervised manner on unlabeled images of human poses, and provide information that can be used to match the pose of a new image to the ones present in the training set. Starting from a set of descriptors recently proposed for human detection, we apply the Latent Dirichlet Allocation framework to model the statistics of these features, and use the resulting model to answer the above questions. We show how our model can efficiently describe the space of images of humans with their pose, by providing an effective representation of poses for tasks such as classification and matching, while performing remarkably well in human/non human decision problems, thus enabling its use for human detection. We validate the model with extensive quantitative experiments and comparisons with other approaches on human detection and pose matching. Alessandro Bissacco, Ming-Hsuan Yang 0001, Stefano Soatto |
NIPS | 1 |
| 2005 | Modeling and Learning Contact Dynamics in Human MotionabstractWe propose a simple model of human motion as a switching linear dynamical system where the switches correspond to contact forces with the ground. This significantly improves the modeling performance when compared to simpler linear systems, with only marginal increase in complexity. We introduce a novel closed-form (non-iterative) algorithm to estimate the switches and learn the model parameters in between switches. We validate our model qualitatively by running simulations, and quantitatively by computing prediction errors that show significant improvements over previous approaches using linear models. Alessandro Bissacco |
CVPR (1) | 1 |
| 2004 | Modeling and Synthesis of Facial Motion Driven by Speech
Payam Saisan, Alessandro Bissacco, Alessandro Chiuso, Stefano Soatto |
ECCV (3) | 2 |
| 2001 | Recognition of Human GaitsabstractWe pose the problem of recognizing different types of human gait in the space of dynamical systems where each gait is represented Established techniques are employed to track a kinematic model of a human body in motion, and the trajectories of the parameters are used to learn a representation of a dynamical system, which defines a gait. Various types of distance between models are then computed These computations are non trivial due to the fact that, even for the case of linear systems, the space of canonical realizations is not linear. Alessandro Bissacco, Alessandro Chiuso, Yi Ma 0001, Stefano Soatto |
CVPR (2) | 1 |