EDBT 2026 Demo / reviewers in the wild / expert
Carlo Tomasi
dblp:24/2162
· DBLP profile ↗
78ranked-venue papers
8as first author
5since 2021 · last 2025
0000-0001-6104-6641ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 8 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7Computer networks · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
37 papers |
3D vision · 64% Video understanding and tracking · 14% Efficient and distributed learning · 5% | |
| Computer graphics and multimedia
15 papers |
Image and video processing · 67% Geometric modeling and processing · 22% Multimedia analysis and retrieval · 8% | |
| Theoretical computer science
11 papers |
Graph algorithms and graph theory · 41% Computational geometry · 24% Algorithms and data structures · 18% |
Topics — the 30 heaviest of 102, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › motion estimation
optical flow |
1.4 | 4 | 2023 | SemARFlow: Injecting Semantics into Unsupervised Optical Flow Estimation for Autonomous Driving · ICCV 2023 Optical Flow Training Under Limited Label Budget via Active Learning · ECCV (22) 2022 Video Motion for Every Visible Point · ICCV 2013 |
Computer vision › 3D vision
3d reconstruction |
1.1 | 4 | 2025 | Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance · NeurIPS 2025 Tree Topology Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2015 Surfaces with Occlusions from Layered Stereo · CVPR (1) 2003 |
Computer vision › 3D vision
motion estimation |
1.0 | 3 | 2023 | SemARFlow: Injecting Semantics into Unsupervised Optical Flow Estimation for Autonomous Driving · ICCV 2023 Video Motion for Every Visible Point · ICCV 2013 Dense Lagrangian motion estimation with occlusions · CVPR 2012 |
Computer vision › 3D vision › neural rendering
3d gaussian splatting |
0.9 | 1 | 2025 | Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance · NeurIPS 2025 |
Computer vision › 3D vision
pose estimation |
0.9 | 1 | 2025 | Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance · NeurIPS 2025 |
Computer vision › 3D vision › motion estimation › optical flow
unsupervised optical flow |
0.7 | 1 | 2023 | SemARFlow: Injecting Semantics into Unsupervised Optical Flow Estimation for Autonomous Driving · ICCV 2023 |
Machine learning › Efficient and distributed learning
active learning |
0.6 | 1 | 2022 | Optical Flow Training Under Limited Label Budget via Active Learning · ECCV (22) 2022 |
Computer vision › Video understanding and tracking
object tracking |
0.5 | 4 | 2012 | Twisted window search for efficient shape localization · CVPR 2012 Detecting motion synchrony by video tubes · ACM Multimedia 2011 Linear time offline tracking and lower envelope algorithms · ICCV 2011 |
Computer vision › 3D vision
structure from motion |
0.4 | 8 | 2015 | Tree Topology Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2015 Simultaneous Compaction and Factorization of Sparse Image Motion Matrices · ECCV (6) 2012 Linear and Incremental Acquisition of Invariant Shape Models From Image Sequences · IEEE Trans. Pattern Anal. Mach. Intell. 1995 |
Computer vision › Video understanding and tracking › multi-camera tracking
multi-target multi-camera tracking |
0.3 | 1 | 2018 | Features for Multi-Target Multi-Camera Tracking and Re-Identification · CVPR 2018 |
Computer vision › Face, body and person analysis
person re-identification |
0.3 | 1 | 2018 | Features for Multi-Target Multi-Camera Tracking and Re-Identification · CVPR 2018 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
image embedding |
0.3 | 1 | 2025 | Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance · NeurIPS 2025 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.2 | 1 | 2016 | Distance Minimization for Reward Learning from Scored Trajectories · AAAI 2016 |
Machine learning › Reinforcement learning
reward learning |
0.2 | 1 | 2016 | Distance Minimization for Reward Learning from Scored Trajectories · AAAI 2016 |
Robotics › Autonomous driving
perception |
0.2 | 1 | 2023 | SemARFlow: Injecting Semantics into Unsupervised Optical Flow Estimation for Autonomous Driving · ICCV 2023 |
Computer vision › Video understanding and tracking › feature tracking
dense point tracking |
0.1 | 1 | 2012 | Dense Lagrangian motion estimation with occlusions · CVPR 2012 |
Computer vision › Video understanding and tracking › object tracking
non-rigid object tracking |
0.1 | 1 | 2012 | Twisted window search for efficient shape localization · CVPR 2012 |
Computer vision › Image recognition and object detection
object detection |
0.1 | 1 | 2012 | Nested Pictorial Structures · ECCV (2) 2012 |
Computer vision › Image recognition and object detection
object localization |
0.1 | 1 | 2012 | Twisted window search for efficient shape localization · CVPR 2012 |
Computer vision › Video understanding and tracking › object tracking
occlusion handling |
0.1 | 1 | 2012 | Dense Lagrangian motion estimation with occlusions · CVPR 2012 |
Computer vision › Face, body and person analysis › human pose estimation
pictorial structures |
0.1 | 1 | 2012 | Nested Pictorial Structures · ECCV (2) 2012 |
Computer vision › 3D vision › 3d shape analysis
shape localization |
0.1 | 1 | 2012 | Twisted window search for efficient shape localization · CVPR 2012 |
Graph algorithms and graph theory
graph algorithms |
0.1 | 1 | 2012 | Fast Tiered Labeling with Topological Priors · ECCV (4) 2012 |
Geometric modeling and processing
3d reconstruction |
0.1 | 2 | 2011 | Detailed reconstruction of 3D plant root shape · ICCV 2011 Shape and motion from image streams under orthography: a factorization method · Int. J. Comput. Vis. 1992 |
Computer vision › Video understanding and tracking
motion analysis |
0.1 | 1 | 2011 | Detecting motion synchrony by video tubes · ACM Multimedia 2011 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.1 | 1 | 2011 | Branch and track · CVPR 2011 |
Computational geometry › distance computation
distance transform |
0.1 | 1 | 2011 | Linear time offline tracking and lower envelope algorithms · ICCV 2011 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.1 | 5 | 2003 | Surfaces with Occlusions from Layered Stereo · CVPR (1) 2003 Multiway Cut for Stereo and Motion with Slanted Surfaces · ICCV 1999 Stereo Matching as a Nearest-Neighbor Problem · IEEE Trans. Pattern Anal. Mach. Intell. 1998 |
Computer vision › 3D vision
feature matching |
0.1 | 1 | 2010 | Critical Nets and Beta-Stable Features for Image Matching · ECCV (3) 2010 |
Graph algorithms and graph theory
graph matching |
0.1 | 1 | 2010 | Critical Nets and Beta-Stable Features for Image Matching · ECCV (3) 2010 |
Methods — techniques the papers use, named apart from their topics
shape carving · 0.9rotation-invariant embedding · 0.93d gaussian splatting · 0.9semantic segmentation · 0.7self-supervision · 0.7active learning · 0.6dynamic programming · 0.5triplet loss · 0.3hard-identity mining · 0.3convolutional neural network · 0.3regularized visual hull · 0.2harmonic background modeling · 0.2global error minimization · 0.2heuristic search · 0.2generative tree-growth model · 0.2topological priors · 0.1sparse matrix compaction · 0.1matrix factorization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and AppearanceabstractAccurate and scalable quantification of animal pose and appearance is crucial for studying behavior. Current 3D pose estimation techniques, such as keypoint- and mesh-based techniques, often face challenges including limited representational detail, labor-intensive annotation requirements, and expensive per-frame optimization. These limitations hinder the study of subtle movements and can make large-scale analyses impractical. We propose *Pose Splatter*, a novel framework leveraging shape carving and 3D Gaussian splatting to model the complete pose and appearance of laboratory animals without prior knowledge of animal geometry, per-frame optimization, or manual annotations.
We also propose a rotation-invariant visual embedding technique for encoding pose and appearance, designed to be a plug-in replacement for 3D keypoint data in downstream behavioral analyses.
Experiments on datasets of mice, rats, and zebra finches show *Pose Splatter* learns accurate 3D animal geometries. Notably, *Pose Splatter* represents subtle variations in pose, provides better low-dimensional pose embeddings over state-of-the-art as evaluated by humans, and generalizes to unseen data.
By eliminating annotation and per-frame optimization bottlenecks, *Pose Splatter* enables analysis of large-scale, longitudinal behavior needed to map genotype, neural activity, and behavior at high resolutions. Jack Goffinet, Youngjo Min, Carlo Tomasi, David E. Carlson |
NeurIPS | 3 |
| 2023 | SemARFlow: Injecting Semantics into Unsupervised Optical Flow Estimation for Autonomous DrivingabstractUnsupervised optical flow estimation is especially hard near occlusions and motion boundaries and in low-texture regions. We show that additional information such as semantics and domain knowledge can help better constrain this problem. We introduce SemARFlow, an unsupervised optical flow network designed for autonomous driving data that takes estimated semantic segmentation masks as additional inputs. This additional information is injected into the encoder and into a learned upsampler that refines the flow output. In addition, a simple yet effective semantic augmentation module provides self-supervision when learning flow and its boundaries for vehicles, poles, and sky. Together, these injections of semantic information improve the KITTI-2015 optical flow test error rate from 11.80% to 8.38%. We also show visible improvements around object boundaries as well as a greater ability to generalize across datasets. Code is available at https://github.com/duke-vision/semantic-unsup-flow-release. Shuai Yuan 0015, Shuzhi Yu, Hannah Kim 0002, Carlo Tomasi |
ICCV | 4 |
| 2022 | Unsupervised Flow Refinement near Motion Boundaries
Shuzhi Yu, Hannah Kim 0002, Shuai Yuan 0015, Carlo Tomasi |
BMVC | 4 |
| 2022 | Optical Flow Training Under Limited Label Budget via Active Learning
Shuai Yuan 0015, Hannah Kim 0002, Shuzhi Yu, Carlo Tomasi |
ECCV (22) | 5 |
| 2021 | Joint Detection of Motion Boundaries and Occlusions
Hannah Kim 0002, Shuzhi Yu, Carlo Tomasi |
BMVC | 3 |
| 2018 | Features for Multi-Target Multi-Camera Tracking and Re-IdentificationabstractMulti-Target Multi-Camera Tracking (MTMCT) tracks many people through video taken from several cameras. Person Re-Identification (Re-ID) retrieves from a gallery images of people similar to a person query image. We learn good features for both MTMCT and Re-ID with a convolutional neural network. Our contributions include an adaptive weighted triplet loss for training and a new technique for hard-identity mining. Our method outperforms the state of the art both on the DukeMTMC benchmarks for tracking, and on the Market-1501 and DukeMTMC-ReID benchmarks for Re-ID. We examine the correlation between good Re-ID and good MTMCT scores, and perform ablation studies to elucidate the contributions of the main components of our system. Code is available1. Ergys Ristani, Carlo Tomasi |
CVPR | 2 |
| 2017 | Tracking Social Groups Within and Across CamerasabstractWe propose a method for tracking groups from single and multiple cameras with disjointed fields of view. Our formulation follows the tracking-by-detection paradigm in which groups are the atomic entities and are linked over time to form long and consistent trajectories. To this end, we formulate the problem as a supervised clustering problem in which a structural SVM classifier learns a similarity measure appropriate for group entities. Multicamera group tracking is handled inside the framework by adopting an orthogonal feature encoding that allows the classifier to learn inter- and intra-camera feature weights differently. Experiments were carried out on a novel annotated group tracking data set, the DukeMTMC-Groups data set. Since this is the first data set on the problem, it comes with the proposal of a suitable evaluation measure. Results of adopting learning for the task are encouraging, scoring a +15% improvement in F1measure over a nonlearning-based clustering baseline. To the best of our knowledge, this is the first proposal of its kind dealing with multicamera group tracking. Francesco Solera, Simone Calderara, Ergys Ristani, Carlo Tomasi, Rita Cucchiara |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Distance Minimization for Reward Learning from Scored TrajectoriesabstractMany planning methods rely on the use of an immediate reward function as a portable and succinct representation of desired behavior. Rewards are often inferred from demonstrated behavior that is assumed to be near-optimal. We examine a framework, Distance Minimization IRL (DM-IRL), for learning reward functions from scores an expert assigns to possibly suboptimal demonstrations. By changing the expert’s role from a demonstrator to a judge, DM-IRL relaxes some of the assumptions present in IRL, enabling learning from the scoring of arbitrary demonstration trajectories with unknown transition functions. DM-IRL complements existing IRL approaches by addressing different assumptions about the expert. We show that DM-IRL is robust to expert scoring error and prove that finding a policy that produces maximally informative trajectories for an expert to score is strongly NP-hard. Experimentally, we demonstrate that the reward function DM-IRL learns from an MDP with an unknown transition model can transfer to an agent with known characteristics in a novel environment, and we achieve successful learning with limited available training data. Benjamin Burchfiel, Carlo Tomasi, Ronald Parr |
AAAI | 2 |
| 2016 | Deformable Graph Model for Tracking Epithelial Cell Sheets in Fluorescence MicroscopyabstractWe propose a novel method for tracking cells that are connected through a visible network of membrane junctions. Tissues of this form are common in epithelial cell sheets and resemble planar graphs where each face corresponds to a cell. We leverage this structure and develop a method to track the entire tissue as a deformable graph. This coupled model in which vertices inform the optimal placement of edges and vice versa captures global relationships between tissue components and leads to accurate and robust cell tracking. We compare the performance of our method with that of four reference tracking algorithms on four data sets that present unique tracking challenges. Our method exhibits consistently superior performance in tracking all cells accurately over all image frames, and is robust over a wide range of image intensity and cell shape profiles. This may be an important tool for characterizing tissues of this type especially in the field of developmental biology where automated cell analysis can help elucidate the mechanisms behind controlled cell-shape changes. Roger S. Zou, Carlo Tomasi |
IEEE Trans. Medical Imaging | 2 |
| 2015 | Tree Topology EstimationabstractTree-like structures are fundamental in nature, and it is often useful to reconstruct the topology of a tree - what connects to what - from a two-dimensional image of it. However, the projected branches often cross in the image: the tree projects to a planar graph, and the inverse problem of reconstructing the topology of the tree from that of the graph is ill-posed. We regularize this problem with a generative, parametric tree-growth model. Under this model, reconstruction is possible in linear time if one knows the direction of each edge in the graph - which edge endpoint is closer to the root of the tree - but becomes NP-hard if the directions are not known. For the latter case, we present a heuristic search algorithm to estimate the most likely topology of a rooted, three-dimensional tree from a single two-dimensional image. Experimental results on retinal vessel, plant root, and synthetic tree data sets show that our methodology is both accurate and efficient. Rolando Estrada, Carlo Tomasi, Scott C. Schmidler, Sina Farsiu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Retinal Artery-Vein Classification via Topology EstimationabstractWe propose a novel, graph-theoretic framework for distinguishing arteries from veins in a fundus image. We make use of the underlying vessel topology to better classify small and midsized vessels. We extend our previously proposed tree topology estimation framework by incorporating expert, domain-specific features to construct a simple, yet powerful global likelihood model. We efficiently maximize this model by iteratively exploring the space of possible solutions consistent with the projected vessels. We tested our method on four retinal datasets and achieved classification accuracies of 91.0%, 93.5%, 91.7%, and 90.9%, outperforming existing methods. Our results show the effectiveness of our approach, which is capable of analyzing the entire vasculature, including peripheral vessels, in wide field-of-view fundus photographs. This topology-based method is a potentially important tool for diagnosing diseases with retinal vascular manifestation. Rolando Estrada, Michael J. Allingham, Priyatham S. Mettu, Scott W. Cousins, Carlo Tomasi, Sina Farsiu |
IEEE Trans. Medical Imaging | 5 |
| 2014 | Tracking Multiple People Online and in Real Time
Ergys Ristani, Carlo Tomasi |
ACCV (5) | 2 |
| 2013 | Video Motion for Every Visible PointabstractDense motion of image points over many video frames can provide important information about the world. However, occlusions and drift make it impossible to compute long motion paths by merely concatenating optical flow vectors between consecutive frames. Instead, we solve for entire paths directly, and flag the frames in which each is visible. As in previous work, we anchor each path to a unique pixel which guarantees an even spatial distribution of paths. Unlike earlier methods, we allow paths to be anchored in any frame. By explicitly requiring that at least one visible path passes within a small neighborhood of every pixel, we guarantee complete coverage of all visible points in all frames. We achieve state-of-the-art results on real sequences including both rigid and non-rigid motions with significant occlusions. Susanna Ricco, Carlo Tomasi |
ICCV | 2 |
| 2013 | A linear system form solution to compute the local space average color
Joaquín Salas, Carlo Tomasi |
Mach. Vis. Appl. | 2 |
| 2012 | Twisted window search for efficient shape localizationabstractMany computer vision systems approximate targets' shape with rectangular bounding boxes. This choice trades localization accuracy for efficient computation. We propose twisted window search, a strict generalization over rectangular window search, for the globally optimal localization of a target's shape. Despite its generality, we show that the new algorithm runs in O(n3), an asymptotic time complexity that is no greater than that of rectangular window search on an image of resolution n × n. We demonstrate improved results of twisted window search for localizing and tracking non-rigid objects with significant orientation, scale and shape change. Twisted window search runs at nearly 10 frames per second in our MATLAB/C++ implementation on images of resolution 240 × 320 on a quad-core laptop. Steve Gu, Carlo Tomasi |
CVPR | 3 |
| 2012 | Dense Lagrangian motion estimation with occlusionsabstractWe couple occlusion modeling and multi-frame motion estimation to compute dense, temporally extended point trajectories in video with significant occlusions. Our approach combines robust spatial regularization with spatially and temporally global occlusion labeling in a variational, Lagrangian framework with subspace constraints. We track points even through ephemeral occlusions. Experiments demonstrate accuracy superior to the state of the art while tracking more points through more frames. Susanna Ricco, Carlo Tomasi |
CVPR | 2 |
| 2012 | Nested Pictorial Structures
Steve Gu, Carlo Tomasi |
ECCV (2) | 3 |
| 2012 | Simultaneous Compaction and Factorization of Sparse Image Motion Matrices
Susanna Ricco, Carlo Tomasi |
ECCV (6) | 2 |
| 2012 | Fast Tiered Labeling with Topological Priors
Steve Gu, Carlo Tomasi |
ECCV (4) | 3 |
| 2012 | Shape from point featuresabstractWe present a nonparametric and efficient method for shape localization that improves on the traditional sub-window search in capturing the fine geometry of an object from a small number of feature points. Our method implies that the discrete set of features capture more appearance and shape information than is commonly exploited. We use the a-complex by Edelsbrunner et al. to build a filtration of simplicial complexes from a user-provided set of features. The optimal value of a is determined automatically by a search for the densest complex connected component, resulting in a parameter-free algorithm. Given K features, localization occurs in O(K log K) time. For VGA-resolution images, computation takes typically less than 10 milliseconds. We use our method for interactive object cut, with promising results. Steve Gu, Carlo Tomasi |
ICASSP | 3 |
| 2012 | Oscillation regularizationabstractWe measure the degree of oscillation of a sampled function f by the number of its local extrema. The greater this number, the more oscillatory and complex f becomes. In signal denoising, we want a restored function g that is simple and fits the data f well. We propose to model this by a global optimization, coined oscillation regularization, that reduces both the data fitting error and the number of local extrema of g: equation where err(f, g) measures the discrepancy between f and g and λ is a regularization parameter. To the best of our knowledge, the number of local extrema of g is a topological prior that is rarely exploited in the literature of regularization. Steve Gu, Carlo Tomasi |
ICASSP | 3 |
| 2012 | Topological persistence on a Jordan curveabstractTopological persistence measures the resilience of extrema of a function to perturbations, and has received increasing attention in computer graphics, visualization and computer vision. While the notion of topological persistence for piece-wise linear functions defined on a simplicial complex has been well studied, the time complexity of all the known algorithms are super-linear (e.g. O(n log n)) in the size n of the complex. We give an O(n) algorithm to compute topological persistence for a function defined on a Jordan curve. To the best of our knowledge, our algorithm is the first to attain linear asymptotic complexity, and is asymptotically optimal. We demonstrate the usefulness of persistence in shape abstraction and compression. Steve Gu, Carlo Tomasi |
ICASSP | 3 |
| 2011 | Branch and trackabstractWe present a new paradigm for tracking objects in video in the presence of other similar objects. This branch-and-track paradigm is also useful in the absence of motion, for the discovery of repetitive patterns in images. The object of interest is the lead object and the distracters are extras. The lead tracker branches out trackers for extras when they are detected, and all trackers share a common set of features. Sometimes, extras are tracked because they are of interest in their own right. In other cases, and perhaps more importantly, tracking extras makes tracking the lead nimbler and more robust, both because shared features provide a richer object model, and because tracking extras accounts for sources of confusion explicitly. Sharing features also makes joint tracking less expensive, and coordinating tracking across lead and extras allows optimizing window positions jointly rather than separately, for better results. The joint tracking of both lead and extras can be solved optimally by dynamic programming and branching is quickly determined by efficient subwindow search. Matlab experiments show near real time performance at 5-30 frames per second on a single-core laptop for 240 by 320 images. Steve Gu, Carlo Tomasi |
CVPR | 2 |
| 2011 | Linear time offline tracking and lower envelope algorithmsabstractOffline tracking of visual objects is particularly helpful in the presence of significant occlusions, when a frame-by-frame, causal tracker is likely to lose sight of the target. In addition, the trajectories found by offline tracking are typically smoother and more stable because of the global optimization this approach entails. In contrast with previous work, we show that this global optimization can be performed in O(MNT) time for T frames of video at M × N resolution, with the help of the generalized distance transform developed by Felzenszwalb and Huttenlocher [13]. Recognizing the importance of this distance transform, we extend the computation to a more general lower envelope algorithm in certain heterogeneous l1-distance metric spaces. The generalized lower envelope algorithm is of complexity O(MN(M+N)) and is useful for a more challenging offline tracking problem. Experiments show that trajectories found by offline tracking are superior to those computed by online tracking methods, and are computed at 100 frames per second. Steve Gu, Carlo Tomasi |
ICCV | 3 |
| 2011 | Detailed reconstruction of 3D plant root shapeabstractWe study the 3D reconstruction of plant roots from multiple 2D images. To meet the challenge caused by the delicate nature of thin branches, we make three innovations to cope with the sensitivity to image quality and calibration. First, we model the background as a harmonic function to improve the segmentation of the root in each 2D image. Second, we develop the concept of the regularized visual hull which reduces the effect of jittering and refraction by ensuring consistency with one 2D image. Third, we guarantee connectedness through adjustments to the 3D reconstruction that minimize global error. Our software is part of a biological phenotype/genotype study of agricultural root systems. It has been tested on more than 40 plant roots and results are promising in terms of reconstruction quality and efficiency. Steve Gu, Herbert Edelsbrunner, Carlo Tomasi, Philip Benfey |
ICCV | 4 |
| 2011 | Detecting motion synchrony by video tubesabstractMotion synchrony, i.e., the coordinated motion of a group of individuals, is an interesting phenomenon in nature or daily life. Fish swim in schools, birds fly in flocks, soldiers march in platoons, etc. Our goal is to detect motion synchrony that may be present in the video data, and to track the group of moving objects as a whole. This opens the door to novel algorithms and applications. To this end, we model individual motions as video tubes in space-time, define motion synchrony by the geometric relation among video tubes, and track a whole set of tubes by dynamic programming. The resulting algorithm is highly efficient in practice. Given a video clip of T frames of resolution XxY, we show that finding the K spatially correlated video tubes and determining the presence of synchrony can be solved optimally in O(XYTK) time. Preliminary experiments show that our method is both effective and efficient. Typical running times are 30 - 100 VGA-resolution frames per second after feature extraction, and the accuracy for the detection of synchrony is more than 90% as evaluated in our annotated data set. Steve Gu, Carlo Tomasi |
ACM Multimedia | 3 |
| 2010 | Efficient Visual Object Tracking with Online Nearest Neighbor Classifier
Steve Gu, Carlo Tomasi |
ACCV (1) | 3 |
| 2010 | Critical Nets and Beta-Stable Features for Image Matching
Steve Gu, Carlo Tomasi |
ECCV (3) | 3 |
| 2010 | Semi-Supervised Fisher Linear Discriminant (SFLD)abstractSupervised learning uses a training set of labeled examples to compute a classifier which is a mapping from feature vectors to class labels. The success of a learning algorithm is evaluated by its ability to generalize, i.e., to extend this mapping accurately to new data that is commonly referred to as the test data. Good generalization depends crucially on the quality of the training set. Because collecting labeled data is laborious, training sets are typically small. Furthermore, it is often difficult to represent all possible observation scenarios during training, so that the statistics of the training set end up differing from those of the test data, a problem known as the sample selection bias. To address sample selection bias, we introduce a Semi-Supervised Fisher Linear Discriminant (SFLD) that utilizes additional, unlabeled data to improve generalization for both small and biased training sets. We characterize the conditions under which SFLD helps, and illustrate its benefits through experiments on digit and car recognition applications. Seda Remus, Carlo Tomasi |
ICASSP | 2 |
| 2009 | Fingerspelling Recognition through Classification of Letter-to-Letter Transitions
Susanna Ricco, Carlo Tomasi |
ACCV (3) | 2 |
| 2009 | Phase diffusion for the synchronization of heterogenous sensor streamsabstractThe analysis of complex human activity typically requires multiple sensors: cameras that take videos from different directions and in different areas, microphones, proximity sensors, range finders, and more. Scenarios where it is not possible to associate reliable clocks to each of the sensors pose a synchronization problem between heterogeneous data streams. In this paper, we propose a new theoretical framework for measuring the synchrony between heterogenous sensor streams. The main idea is to model the phase disparity between two data streams explicitly as an Ornstein-Uhlenbeck random process. Based on this model, we derive a simple method for synchronizing of underlying sources. We illustrate the ideas with experiments on audio-visual synchronization and human motion categorization, and report promising results. Steve Gu, Carlo Tomasi |
ICASSP | 2 |
| 2009 | Manuscript Bleed-through Removal via Hysteresis ThresholdingabstractMany types of degradation can render ancient manuscripts very hard to read. In bleed-through, the text from the reverse, or verso, side of a page seeps through into the front, or recto. In this paper, we propose hysteresis thresholding to greatly reduce bleed-through. Thresholding alone cannot properly separate ink and bleed-through because the ranges of intensities for the two classes overlap. Hysteresis thresholding overcomes this limitation via the two steps of thresholding and ink regrowth. In order to provide quantitative measures of the effectiveness of this approach, we constructed a novel dataset which features bleed-through and has available ground truth. We evaluated our method and a number of previously proposed approaches on ink pixel precision and recall. Hysteresis thresholding significantly improves over existing methods. Rolando Estrada, Carlo Tomasi |
ICDAR | 2 |
| 2008 | Robust shape normalization based on implicit representationsabstractWe introduce a new shape normalization method based on implicit shape representations. The proposed method is robust with respect to deformations and invariant to similarity transformations (translation, isotropic scaling and rotation). The new method has been tested and compared to the classical shape normalization method and previous work in terms of aligning groups of shapes with deformations. Tingting Jiang 0001, Carlo Tomasi |
ICPR | 2 |
| 2008 | Editorial
Cordelia Schmid, Stefano Soatto, Carlo Tomasi |
Int. J. Comput. Vis. | 3 |
| 2007 | Finite-Element Level-Set Curve ParticlesabstractParticle filters encode a time-evolving probability density by maintaining a random sample from it. Level sets represent closed curves as zero crossings of functions of two variables. The combination of level sets and particle filters presents many conceptual advantages when tracking uncertain, evolving boundaries over time, but the cost of combining these two ideas seems prima facie prohibitive. A previous publication showed that a large number of virtual level set particles can be tracked with a logarithmic amount of work for propagation and update. We now make level- set curve particles more efficient by borrowing ideas from the Finite Element Method (FEM). This improves level-set curve particles in both running time (by a constant factor) and accuracy of the results. Tingting Jiang 0001, Carlo Tomasi |
ICCV | 2 |
| 2007 | Correspondence as energy-based segmentation
Stanley T. Birchfield, Braga Natarajan, Carlo Tomasi |
Image Vis. Comput. | 3 |
| 2006 | Level-Set Curve Particles
Tingting Jiang 0001, Carlo Tomasi |
ECCV (3) | 2 |
| 2005 | Mean Shift Is a Bound OptimizationabstractWe build on the current understanding of mean shift as an optimization procedure. We demonstrate that, in the case of piecewise constant kernels, mean shift is equivalent to Newton's method. Further, we prove that, for all kernels, the mean shift procedure is a quadratic bound maximization. Mark Fashing, Carlo Tomasi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | 3D Head Tracking Based on Recognition and Interpolation Using a Time-of-Flight Depth Sensor
Salih Burak Göktürk, Carlo Tomasi |
CVPR (2) | 2 |
| 2004 | Image Similarity Using Mutual Information of Regions
Daniel B. Russakoff, Carlo Tomasi, Torsten Rohlfing, Calvin R. Maurer Jr. |
ECCV (3) | 2 |
| 2004 | Surfaces with Occlusions from Layered StereoabstractWe propose a new binocular stereo algorithm that estimates scene structure as a collection of smooth surface patches. The disparities within each patch are modeled by a continuous-valued spline, while the extent of each patch is represented via a pixelwise partitioning of the images. Disparities and extents are alternately estimated in an iterative, energy minimization framework. Experimental results demonstrate that, for scenes consisting of smooth surfaces, the proposed algorithm significantly improves upon the state of the art. Michael H. Lin, Carlo Tomasi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Surfaces with Occlusions from Layered StereoabstractAlthough steady progress has been made in recent stereo algorithms, producing accurate results in the neighborhood of depth discontinuities remains a challenge. Moreover, among the techniques that best localize depth discontinuities, it is common to work only with a discrete set of disparity values, hindering the modeling of smooth, non-fronto-parallel surfaces. We propose to estimate scene structure as a set of smooth surface patches. The disparities within each patch are modeled by a spline, while the extent of each patch is represented by a pixelwise labeling of the source images. Disparities and extents are alternately estimated in an iterative, energy minimization framework. Segmentation is via graph cuts, aided by image gradients. Input images are treated symmetrically, and occlusions are addressed explicitly. Promising experimental results are presented. Michael H. Lin, Carlo Tomasi |
CVPR (1) | 2 |
| 2003 | 3D Tracking = Classification + InterpolationabstractHand gestures are examples of fast and complex motions. Computers fail to track these in fast video, but sleight of hand fools humans as well: what happens too quickly we just cannot see. We show a 3D tracker for these types of motions that relies on the recognition of familiar configurations in 2D images (classification), and fills the gaps in-between (interpolation). We illustrate this idea with experiments on hand motions similar to finger spelling. The penalty for a recognition failure is often small: if two configurations are confused, they are often similar to each other, and the illusion works well enough, for instance, to drive a graphics animation of the moving hand. We contribute advances in both feature design and classifier training: our image features are invariant to image scale, translation, and rotation, and we propose a classification method that combines VQPCA with discrimination trees. Carlo Tomasi, Slav Petrov, Arvind Sastry |
ICCV | 1 |
| 2002 | On the Consistency of Instantaneous Rigid Motion Estimation
Tong Zhang 0001, Carlo Tomasi |
Int. J. Comput. Vis. | 2 |
| 2002 | Edge Displacement Field-Based Classification for Improved Detection of Polyps in CT ColonographyabstractColorectal cancer can easily be prevented provided that the precursors to tumors, small colonic polyps, are detected and removed. Currently, the only definitive examination of the colon is fiber-optic colonoscopy, which is invasive and expensive. Computed tomographic colonography (CTC) is potentially a less costly and less invasive alternative to FOC. It would be desirable to have computer-aided detection (CAD) algorithms to examine the large amount of data CTC provides. Most current CAD algorithms have high false positive rates at the required sensitivity levels. We developed and evaluated a postprocessing algorithm to decrease the false positive rate of such a CAD method without sacrificing sensitivity. Our method attempts to model the way a radiologist recognizes a polyp while scrolling a cross-sectional plane through three-dimensional computed tomography data by classification of the changes in the location of the edges in the two-dimensional plane. We performed a tenfold cross-validation study to assess its performance using sensitivity/specificity analysis on data from 48 patients. The mean specificity over all experiments increased from 0.19 (0.35) to 0.47 (0.56) for a sensitivity of 1.00 (0.95). Burak Acar, Christopher F. Beaulieu, Salih Burak Göktürk, Carlo Tomasi, David S. Paik, R. Brooke Jeffrey Jr., Judy Yee, Sandy Napel |
IEEE Trans. Medical Imaging | 4 |
| 2001 | A New 3-D Pattern Recognition Technique With Application to Computer Aided ColonoscopyabstractTo utilize CT or MRI images for computer aided diagnosis applications, robust features that represent 3D image data need to be constructed and subsequently used by a classification method. We present a computer aided diagnosis system for early diagnosis of colon cancer. The system extracts features via a new 3D pattern processing method and processes them using a support vector machine classifier. Our 3D pattern processing method, called Random Orthogonal Shape Section (ROSS) mimics the radiologist's way of viewing these images and combines information from many random triples of mutually orthogonal sections going through the volume. Another contribution of the paper is a new feedback framework between the classification algorithm and the definition of the features. This framework, called Distinctive Component Analysis combines support vector samples with linear discriminant analysis to map the features of clustered support vectors to a lower dimensional space where the two classes of objects of interest are optimally separated to obtain better features. We show that the combination of these better features with support vector machine classification provides a good recognition rate. Salih Burak Göktürk, Carlo Tomasi |
CVPR (1) | 2 |
| 2001 | Using Optical Flow Fields for Polyp Detection in Virtual Colonoscopy
Burak Acar, Sandy Napel, David S. Paik, Salih Burak Göktürk, Carlo Tomasi, Christopher F. Beaulieu |
MICCAI | 5 |
| 2001 | A Learning Method for Automated Polyp Detection
Salih Burak Göktürk, Carlo Tomasi, Burak Acar, David S. Paik, Christopher F. Beaulieu, Sandy Napel |
MICCAI | 2 |
| 2001 | Empirical Evaluation of Dissimilarity Measures for Color and Texture
Yossi Rubner, Jan Puzicha, Carlo Tomasi, Joachim M. Buhmann |
Comput. Vis. Image Underst. | 3 |
| 2001 | Edge, Junction, and Corner Detection Using Color DistributionsabstractFor over 30 years (1970-2000) researchers in computer vision have been proposing new methods for performing low-level vision tasks such as detecting edges and corners. One key element shared by most methods is that they represent local image neighborhoods as constant in color or intensity with deviations modeled as noise. Due to computational considerations that encourage the use of small neighborhoods where this assumption holds, these methods remain popular. The research presented models a neighborhood as a distribution of colors. The goal is to show that the increase in accuracy of this representation translates into higher-quality results for low-level vision tasks on difficult, natural images, especially as neighborhood size increases. We emphasize large neighborhoods because small ones often do not contain enough information. We emphasize color because it subsumes gray scale as an image range and because it is the dominant form of human perception. We discuss distributions in the context of detecting edges, corners, and junctions, and we show results for each. Mark A. Ruzon, Carlo Tomasi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | A Statistical 3D Pattern Processing Method for Computer Aided Detection of Polyps in CT ColonographyabstractAdenomatous polyps in the colon are believed to be the precursor to colorectal carcinoma, the second leading cause of cancer deaths in United States. In this paper, we propose a new method for computer-aided detection of polyps in computed tomography (CT) colonography (virtual colonoscopy), a technique in which polyps are imaged along the wall of the air-inflated, cleansed colon with X-ray CT. Initial work with computer aided detection has shown high sensitivity, but at a cost of too many false positives. We present a statistical approach that uses support vector machines to distinguish the differentiating characteristics of polyps and healthy tissue, and uses this information for the classification of the new cases. One of the main contributions of the paper is the new three-dimensional pattern processing approach, called random orthogonal shape sections method, which combines the information from many random images to generate reliable signatures of shape. The input to the proposed system is a collection of volume data from candidate polyps obtained by a high-sensitivity, low-specificity system that we developed previously. The results of our ten-fold cross-validation experiments show that, on the average, the system increases the specificity from 0.19 (0.35) to 0.69 (0.74) at a sensitivity level of 1.0 (0.95). Salih Burak Göktürk, Carlo Tomasi, Burak Acar, Christopher F. Beaulieu, David S. Paik, R. Brooke Jeffrey Jr., Judy Yee, Sandy Napel |
IEEE Trans. Medical Imaging | 2 |
| 2000 | Alpha Estimation in Natural ImagesabstractMany boundaries between objects in the world project onto curves in an image. However, boundaries involving natural objects (e.g., trees, hair, water, smoke) are often unworkable under this model because many pixels receive light from more than one object. We propose a technique for estimating alpha, the proportion in which two colors mix to produce a color at the boundary. The technique extends blue screen matting to backgrounds that have almost arbitrary color distributions, though coarse knowledge of the boundary's location is required. Results show a number of different objects moved from one image to another while maintaining naturalism. Mark A. Ruzon, Carlo Tomasi |
CVPR | 2 |
| 2000 | The Earth Mover's Distance as a Metric for Image Retrieval
Yossi Rubner, Carlo Tomasi, Leonidas J. Guibas |
Int. J. Comput. Vis. | 2 |
| 1999 | Color Edge Detection with the Compass OperatorabstractThe compass operator detects step edges without assuming that the regions on either side have constant color. Using distributions of pixel colors rather than the mean, the operator finds the orientation of a diameter that maximizes the difference between two halves of a circular window. Junctions can also be detected by exploiting their lack of bilateral symmetry. This approach is superior to a multi-dimensional gradient method in situations that often result in false negatives, and it localizes edges better as scale increases. Mark A. Ruzon, Carlo Tomasi |
CVPR | 2 |
| 1999 | Fast, Robust, and Consistent Camera Motion EstimationabstractPrevious algorithms that recover camera motion from image velocities suffer from both bias and excessive variance in the results. We propose a robust estimator of camera motion that is statistically consistent when image noise is isotropic. Consistency means that the estimated motion converges in probability, to the true value as the number of image points increases. An algorithm based on reweighted Gauss-Newton iterations handles 100 velocity measurements in about 50 milliseconds on a workstation. Tong Zhang 0001, Carlo Tomasi |
CVPR | 2 |
| 1999 | Multiway Cut for Stereo and Motion with Slanted SurfacesabstractSlanted surfaces pose a problem for correspondence algorithms utilizing search because of the greatly increased number of possibilities, when compared with fronto-parallel surfaces. In this paper we propose an algorithm to compute correspondence between stereo images or between frames of a motion sequence by minimizing an energy functional that accounts for slanted surfaces. The energy is minimized in a greedy strategy that alternates between segmenting the image into a number of non-overlapping regions (using the multiway-cut algorithm of Boykov, Veksler, and Zabih) and finding the affine parameters describing the displacement function of each region. A follow-up step enables the algorithm to escape local minima due to oversegmentation. Experiments on real images show the algorithm's ability to find an accurate segmentation and displacement map, as well as discontinuities and creases, from a wide variety of stereo and motion imagery. Stanley T. Birchfield, Carlo Tomasi |
ICCV | 2 |
| 1999 | Representation Issues in the ML Estimation of Camera MotionabstractThe computation of camera motion from image measurements is a parameter estimation problem. We show that for the analysis of the problem's sensitivity, the parameterization must enjoy the property of fairness, which makes sensitivity results invariant to changes of coordinates. We prove that Cartesian unit norm vectors and quaternions are fair parameterizations of rotations and translations, respectively, and that spherical coordinates and Euler angles are not. We extend the Gauss-Markov theorem to implicit formulations with constrained parameters, a necessary step in order to take advantage of fair parameterizations. We show how maximum likelihood (ML) estimation problems whose sensitivity depends on a large number of parameters, such as coordinates of points in the scene, can be partitioned into equivalence classes, with problems in the same class exhibiting the same sensitivity. Joachim Hornegger, Carlo Tomasi |
ICCV | 2 |
| 1999 | Empirical Evaluation of Dissimilarity Measures for Color and TextureabstractThis paper empirically compares nine image dissimilarity measures that are based on distributions of color and texture features summarizing over 1,000 CPU hours of computational experiments. Ground truth is collected via a novel random sampling scheme for color and via an image partitioning method for texture. Quantitative performance evaluations are given for classification, image retrieval, and segmentation tasks, and for a wide variety of dissimilarity measures. It is demonstrated how the selection of a measure, based on large scale evaluation, substantially improves the quality of classification, retrieval, and unsupervised segmentation of color and texture images. Jan Puzicha, Yossi Rubner, Carlo Tomasi, Joachim M. Buhmann |
ICCV | 3 |
| 1999 | Texture-based Image Retrieval without SegmentationabstractImage segmentation is not only hard and unnecessary for texture-based image retrieval, but can even be harmful. Images of either individual or multiple textures are best described by distributions of spatial frequency descriptors, rather than single descriptor vectors over presegmented regions. A retrieval method based on the earth movers distance with an appropriate ground distance is shown to handle both complete and partial multi-textured queries. As an illustration, different images of the same type of animal are easily retrieved together. At the same time, animals with subtly different coats, like cheetahs and leopards, are properly distinguished. Yossi Rubner, Carlo Tomasi |
ICCV | 2 |
| 1999 | Corner Detection in Textured Color ImagesabstractCorner models in the literature have lagged behind edge models with respect to color and shading. We use both a region model, based on distributions of pixel colors, and an edge model, which removes false positives, to perform corner detection on color images whose regions contain texture. We show results on a variety of natural images at different scales that highlight the problems that occur when boundaries between regions have curvature. Mark A. Ruzon, Carlo Tomasi |
ICCV | 2 |
| 1999 | Depth Discontinuities by Pixel-to-Pixel Stereo
Stanley T. Birchfield, Carlo Tomasi |
Int. J. Comput. Vis. | 2 |
| 1998 | Adaptive Color-Image Embeddings for Database Navigation
Yossi Rubner, Carlo Tomasi, Leonidas J. Guibas |
ACCV (1) | 2 |
| 1998 | Depth Discontinuities by Pixel-to-Pixel StereoabstractAn algorithm to detect depth discontinuities from a stereo pair of images is presented. The algorithm matches individual pixels in corresponding scanline pairs while allowing occluded pixels to remain unmatched, then propagates the information between scanlines by means of a fast postprocessor. The algorithm handles large untextured regions, uses a measure of pixel dissimilarity that is insensitive to image sampling, and prunes bad search nodes to increase the speed of dynamic programming. The computation is relatively fast, taking about 1.5 microseconds per pixel per disparity on a workstation. Approximate disparity maps and precise depth discontinuities (along both horizontal and vertical boundaries) are shown for five stereo images containing textured, untextured, fronto-parallel, and slanted objects. Stanley T. Birchfield, Carlo Tomasi |
ICCV | 2 |
| 1998 | A Metric for Distributions with Applications to Image DatabasesabstractWe introduce a new distance between two distributions that we call the Earth Mover's Distance (EMD), which reflects the minimal amount of work that must be performed to transform one distribution into the other by moving "distribution mass" around. This is a special case of the transportation problem from linear optimization, for which efficient algorithms are available. The EMD also allows for partial matching. When used to compare distributions that have the same overall mass, the EMD is a true metric, and has easy-to-compute lower bounds. In this paper we focus on applications to image databases, especially color and texture. We use the EMD to exhibit the structure of color-distribution and texture spaces by means of Multi-Dimensional Scaling displays. We also propose a novel approach to the problem of navigating through a collection of color images, which leads to a new paradigm for image database search. Yossi Rubner, Carlo Tomasi, Leonidas J. Guibas |
ICCV | 2 |
| 1998 | Bilateral Filtering for Gray and Color ImagesabstractBilateral filtering smooths images while preserving edges, by means of a nonlinear combination of nearby image values. The method is noniterative, local, and simple. It combines gray levels or colors based on both their geometric closeness and their photometric similarity, and prefers near values to distant values in both domain and range. In contrast with filters that operate on the three bands of a color image separately, a bilateral filter can enforce the perceptual metric underlying the CIE-Lab color space, and smooth colors and preserve edges in a way that is tuned to human perception. Also, in contrast with standard filtering, bilateral filtering produces no phantom colors along edges in color images, and reduces phantom colors where they appear in the original image. Carlo Tomasi, Roberto Manduchi |
ICCV | 1 |
| 1998 | Texture metricsabstractWe introduce a class of metric perceptual distances between textures. The first metric is sensitive to both rotation and scale differences, and provides a basis for two other metrics, one invariant to rotation, and the other invariant to both rotation and scale. Our metrics are based on Earth Mover's Distance computations on log-polar distributions of spatial frequency computed from Gabor filters. We show consistency of our metrics with psychophysical findings on texture discrimination and classification. Yossi Rubner, Carlo Tomasi |
SMC | 2 |
| 1998 | A Pixel Dissimilarity Measure That Is Insensitive to Image SamplingabstractBecause of image sampling, traditional measures of pixel dissimilarity can assign a large value to two corresponding pixels in a stereo pair, even in the absence of noise and other degrading effects. We propose a measure of dissimilarity that is provably insensitive to sampling because it uses the linearly interpolated intensity functions surrounding the pixels. Experiments on real images show that our measure alleviates the problem of sampling with little additional computational overhead. Stanley T. Birchfield, Carlo Tomasi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Stereo Matching as a Nearest-Neighbor ProblemabstractWe propose a representation of images, called intrinsic curves, that transforms stereo matching from a search problem into a nearest-neighbor problem. Intrinsic curves are the paths that a set of local image descriptors trace as an image scanline is traversed from left to right. Intrinsic curves are ideally invariant with respect to disparity. Stereo correspondence then becomes a trivial lookup problem in the ideal case. We also show how to use intrinsic curves to match real images in the presence of noise, brightness bias, contrast fluctuations, moderate geometric distortion, image ambiguity, and occlusions. In this case, matching becomes a nearest-neighbor problem, even for very large disparity values. Carlo Tomasi, Roberto Manduchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1996 | Comparison of Approaches to Egomotion ComputationabstractWe evaluated six algorithms for computing egomotion from image velocities. We established benchmarks for quantifying bias and sensitivity to noise, and for quantifying the convergence properties of those algorithms that require numerical search. Our simulation results reveal some interesting and surprising results. First, it is often written in the literature that the egomotion problem is difficult because translation (e.g., along the X-axis) and rotation (e.g., about the Y-axis) produce similar image velocities. We found, to the contrary, that the bias and sensitivity of our six algorithms are totally invariant with respect to the axis of rotation. Second, it is also believed by some that fixating helps to make the egomotion problem easier: We found, to the contrary, that fixating does not help when the noise is independent of the image velocities. Fixation does help if the noise is proportional to speed, but this is only for the trivial reason that the speeds are slower under fixation. Third, it is widely believed that increasing the field of view will yield better performance. We found, to the contrary, that this is not necessarily true. Tina Yu Tian, Carlo Tomasi, David J. Heeger |
CVPR | 2 |
| 1996 | Stereo Without Search
Carlo Tomasi, Roberto Manduchi |
ECCV (1) | 1 |
| 1995 | Linear and Incremental Acquisition of Invariant Shape Models From Image SequencesabstractWe show how to automatically acquire Euclidian shape representations of objects from noisy image sequences under weak perspective. The proposed method is linear and incremental, requiring no more than pseudoinverse. A nonlinear, but numerically sound preprocessing stage is added to improve the accuracy of the results even further. Experiments show that attention to noise and computational techniques improves the shape results substantially with respect to previous methods proposed for ideal images.> Daphna Weinshall, Carlo Tomasi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1994 | Good features to trackabstractNo feature-based vision system can work until good features can be identified and tracked from frame to frame. Although tracking itself is by and large a solved problem, selecting features that can be tracked well and correspond to physical points in the world is still an open problem. We propose a feature selection criterion that is optimal by construction because it is based on how the tracker works, as well as a feature monitoring method that can detect occlusions, disocclusions, and features that do not correspond to points in the world. These methods are based on a new tracking algorithm that extends previous Newton-Raphson style search methods to work under affine image transformations. We test performance with several simulations and experiments on real images. Jianbo Shi, Carlo Tomasi |
CVPR | 2 |
| 1994 | Pictures and trails: a new framework for the computation of shape and motion from perspective image sequencesabstractThis paper presents a new framework for the computation of shape and motion from a sequence of images taken under perspective projection. The framework is based on two abstractions, the picture and trail loci, that represent respectively the set of all pictures of the same scene and the set of all trails that a point in the world can leave on the image for a given camera trajectory. These abstractions lead to a remarkably clean relation between perspective and orthography. A shape and motion reconstruction method is developed for the case of a two-dimensional world but all concepts also hold in three dimensions. Experiments show that the method is rather immune to noise but critically dependent on camera calibration.> Carlo Tomasi |
CVPR | 1 |
| 1993 | Direction of heading from image deformationsabstractA method is proposed to compute the direction of heading from the differential changes in the angles between the projection rays of pairs of point features. These angles, the image deformations, do not depend on viewer rotation. The key problem of separating the effects of rotation from those of translation is solved at the input. Experiments show both the feasibility of the method on real images and the advantages of using deformation rather than optical flow.> Carlo Tomasi, Jianbo Shi |
CVPR | 1 |
| 1993 | Linear and incremental acquisition of invariant shape models from image sequencesabstractThe authors show how to automatically acquire similarity-invariant shape representations of objects from noisy image sequences under a weak perspective. The incremental nature of the method makes it possible to process images one at a time, moving away from the storage-intensive batch methods of the past. It is based on the observation that the trajectories that points on the object form in weak-perspective image sequences are linear combinations of three of the trajectories themselves, and that the coefficients of the linear combinations represent shape in an affine-invariant basis. A nonlinear but numerically sound preprocessing state is added to improve the accuracy of the results even further. Experiments showed that attention to noise and computational techniques improved the shape results substantially with respect to previous methods.> Daphna Weinshall, Carlo Tomasi |
ICCV | 2 |
| 1992 | Shape and motion from image streams under orthography: a factorization method
Carlo Tomasi, Takeo Kanade |
Int. J. Comput. Vis. | 1 |
| 1990 | Shape and motion without depthabstractInferring the depth and shape of remote objects and the camera motion from a sequence of images is possible in principle, but is an ill-conditioned problem when the objects are distant with respect to their size. This problem is overcome by inferring shape and motion without computing depth as an intermediate step. On a single epipolar plane, an image sequence can be represented by the F*P matrix of the image coordinates of P points tracked through F frames. It is shown that under orthographic projection this matrix is of rank three. Using this result, the authors develop a shape-and-motion algorithm based on singular value decomposition. The algorithm gives accurate results, without relying on any smoothness assumption for either shape or motion.> Carlo Tomasi, Takeo Kanade |
ICCV | 1 |
| 1984 | Spectral Analysis of Line Regenerator Time JitterabstractA closed form expression for the spectral density of the time jitter produced in a line regenerator in the presence of a general polynomial nonlinear circuit, band-limited baseband pulses, and an arbitrary tuned filter is developed. The method applies to a PAM digital signal with mutually independent symbols but with an arbitrary number of levels and symbol probabilities. As an example of the method, the spectrum of the time jitter produced in the presence of a fourth power law nonlinear circuit is reported. Silvano Pupolin, Carlo Tomasi |
IEEE Trans. Commun. | 2 |