Carlo Tomasi

dblp:24/2162 · DBLP profile ↗
← Back
78ranked-venue papers
8as first author
5since 2021 · last 2025
0000-0001-6104-6641ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 63 · 8 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7Computer networks · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
37 papers
3D vision · 64% Video understanding and tracking · 14% Efficient and distributed learning · 5%
Computer graphics and multimedia
15 papers
Image and video processing · 67% Geometric modeling and processing · 22% Multimedia analysis and retrieval · 8%
Theoretical computer science
11 papers
Graph algorithms and graph theory · 41% Computational geometry · 24% Algorithms and data structures · 18%

Topics — the 30 heaviest of 102, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › motion estimation
optical flow
1.442023
SemARFlow: Injecting Semantics into Unsupervised Optical Flow Estimation for Autonomous Driving · ICCV 2023
Optical Flow Training Under Limited Label Budget via Active Learning · ECCV (22) 2022
Video Motion for Every Visible Point · ICCV 2013
Computer vision › 3D vision
3d reconstruction
1.142025
Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance · NeurIPS 2025
Tree Topology Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Surfaces with Occlusions from Layered Stereo · CVPR (1) 2003
Computer vision › 3D vision
motion estimation
1.032023
SemARFlow: Injecting Semantics into Unsupervised Optical Flow Estimation for Autonomous Driving · ICCV 2023
Video Motion for Every Visible Point · ICCV 2013
Dense Lagrangian motion estimation with occlusions · CVPR 2012
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.912025
Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance · NeurIPS 2025
Computer vision › 3D vision
pose estimation
0.912025
Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance · NeurIPS 2025
Computer vision › 3D vision › motion estimation › optical flow
unsupervised optical flow
0.712023
SemARFlow: Injecting Semantics into Unsupervised Optical Flow Estimation for Autonomous Driving · ICCV 2023
Machine learning › Efficient and distributed learning
active learning
0.612022
Optical Flow Training Under Limited Label Budget via Active Learning · ECCV (22) 2022
Computer vision › Video understanding and tracking
object tracking
0.542012
Twisted window search for efficient shape localization · CVPR 2012
Detecting motion synchrony by video tubes · ACM Multimedia 2011
Linear time offline tracking and lower envelope algorithms · ICCV 2011
Computer vision › 3D vision
structure from motion
0.482015
Tree Topology Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Simultaneous Compaction and Factorization of Sparse Image Motion Matrices · ECCV (6) 2012
Linear and Incremental Acquisition of Invariant Shape Models From Image Sequences · IEEE Trans. Pattern Anal. Mach. Intell. 1995
Computer vision › Video understanding and tracking › multi-camera tracking
multi-target multi-camera tracking
0.312018
Features for Multi-Target Multi-Camera Tracking and Re-Identification · CVPR 2018
Computer vision › Face, body and person analysis
person re-identification
0.312018
Features for Multi-Target Multi-Camera Tracking and Re-Identification · CVPR 2018
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
image embedding
0.312025
Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance · NeurIPS 2025
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.212016
Distance Minimization for Reward Learning from Scored Trajectories · AAAI 2016
Machine learning › Reinforcement learning
reward learning
0.212016
Distance Minimization for Reward Learning from Scored Trajectories · AAAI 2016
Robotics › Autonomous driving
perception
0.212023
SemARFlow: Injecting Semantics into Unsupervised Optical Flow Estimation for Autonomous Driving · ICCV 2023
Computer vision › Video understanding and tracking › feature tracking
dense point tracking
0.112012
Dense Lagrangian motion estimation with occlusions · CVPR 2012
Computer vision › Video understanding and tracking › object tracking
non-rigid object tracking
0.112012
Twisted window search for efficient shape localization · CVPR 2012
Computer vision › Image recognition and object detection
object detection
0.112012
Nested Pictorial Structures · ECCV (2) 2012
Computer vision › Image recognition and object detection
object localization
0.112012
Twisted window search for efficient shape localization · CVPR 2012
Computer vision › Video understanding and tracking › object tracking
occlusion handling
0.112012
Dense Lagrangian motion estimation with occlusions · CVPR 2012
Computer vision › Face, body and person analysis › human pose estimation
pictorial structures
0.112012
Nested Pictorial Structures · ECCV (2) 2012
Computer vision › 3D vision › 3d shape analysis
shape localization
0.112012
Twisted window search for efficient shape localization · CVPR 2012
Graph algorithms and graph theory
graph algorithms
0.112012
Fast Tiered Labeling with Topological Priors · ECCV (4) 2012
Geometric modeling and processing
3d reconstruction
0.122011
Detailed reconstruction of 3D plant root shape · ICCV 2011
Shape and motion from image streams under orthography: a factorization method · Int. J. Comput. Vis. 1992
Computer vision › Video understanding and tracking
motion analysis
0.112011
Detecting motion synchrony by video tubes · ACM Multimedia 2011
Computer vision › Video understanding and tracking
multi-object tracking
0.112011
Branch and track · CVPR 2011
Computational geometry › distance computation
distance transform
0.112011
Linear time offline tracking and lower envelope algorithms · ICCV 2011
Computer vision › 3D vision › stereo vision
stereo matching
0.152003
Surfaces with Occlusions from Layered Stereo · CVPR (1) 2003
Multiway Cut for Stereo and Motion with Slanted Surfaces · ICCV 1999
Stereo Matching as a Nearest-Neighbor Problem · IEEE Trans. Pattern Anal. Mach. Intell. 1998
Computer vision › 3D vision
feature matching
0.112010
Critical Nets and Beta-Stable Features for Image Matching · ECCV (3) 2010
Graph algorithms and graph theory
graph matching
0.112010
Critical Nets and Beta-Stable Features for Image Matching · ECCV (3) 2010

Methods — techniques the papers use, named apart from their topics

shape carving · 0.9rotation-invariant embedding · 0.93d gaussian splatting · 0.9semantic segmentation · 0.7self-supervision · 0.7active learning · 0.6dynamic programming · 0.5triplet loss · 0.3hard-identity mining · 0.3convolutional neural network · 0.3regularized visual hull · 0.2harmonic background modeling · 0.2global error minimization · 0.2heuristic search · 0.2generative tree-growth model · 0.2topological priors · 0.1sparse matrix compaction · 0.1matrix factorization · 0.1
YearPublicationVenuePosition
2025 Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance
abstract
Accurate and scalable quantification of animal pose and appearance is crucial for studying behavior. Current 3D pose estimation techniques, such as keypoint- and mesh-based techniques, often face challenges including limited representational detail, labor-intensive annotation requirements, and expensive per-frame optimization. These limitations hinder the study of subtle movements and can make large-scale analyses impractical. We propose *Pose Splatter*, a novel framework leveraging shape carving and 3D Gaussian splatting to model the complete pose and appearance of laboratory animals without prior knowledge of animal geometry, per-frame optimization, or manual annotations. We also propose a rotation-invariant visual embedding technique for encoding pose and appearance, designed to be a plug-in replacement for 3D keypoint data in downstream behavioral analyses. Experiments on datasets of mice, rats, and zebra finches show *Pose Splatter* learns accurate 3D animal geometries. Notably, *Pose Splatter* represents subtle variations in pose, provides better low-dimensional pose embeddings over state-of-the-art as evaluated by humans, and generalizes to unseen data. By eliminating annotation and per-frame optimization bottlenecks, *Pose Splatter* enables analysis of large-scale, longitudinal behavior needed to map genotype, neural activity, and behavior at high resolutions.
Jack Goffinet, Youngjo Min, Carlo Tomasi, David E. Carlson
NeurIPS3
2023 SemARFlow: Injecting Semantics into Unsupervised Optical Flow Estimation for Autonomous Driving
abstract
Unsupervised optical flow estimation is especially hard near occlusions and motion boundaries and in low-texture regions. We show that additional information such as semantics and domain knowledge can help better constrain this problem. We introduce SemARFlow, an unsupervised optical flow network designed for autonomous driving data that takes estimated semantic segmentation masks as additional inputs. This additional information is injected into the encoder and into a learned upsampler that refines the flow output. In addition, a simple yet effective semantic augmentation module provides self-supervision when learning flow and its boundaries for vehicles, poles, and sky. Together, these injections of semantic information improve the KITTI-2015 optical flow test error rate from 11.80% to 8.38%. We also show visible improvements around object boundaries as well as a greater ability to generalize across datasets. Code is available at https://github.com/duke-vision/semantic-unsup-flow-release.
Shuai Yuan 0015, Shuzhi Yu, Hannah Kim 0002, Carlo Tomasi
ICCV4
2022 Unsupervised Flow Refinement near Motion Boundaries
Shuzhi Yu, Hannah Kim 0002, Shuai Yuan 0015, Carlo Tomasi
BMVC4
2022 Optical Flow Training Under Limited Label Budget via Active Learning
Shuai Yuan 0015, Hannah Kim 0002, Shuzhi Yu, Carlo Tomasi
ECCV (22)5
2021 Joint Detection of Motion Boundaries and Occlusions
Hannah Kim 0002, Shuzhi Yu, Carlo Tomasi
BMVC3
2018 Features for Multi-Target Multi-Camera Tracking and Re-Identification
abstract
Multi-Target Multi-Camera Tracking (MTMCT) tracks many people through video taken from several cameras. Person Re-Identification (Re-ID) retrieves from a gallery images of people similar to a person query image. We learn good features for both MTMCT and Re-ID with a convolutional neural network. Our contributions include an adaptive weighted triplet loss for training and a new technique for hard-identity mining. Our method outperforms the state of the art both on the DukeMTMC benchmarks for tracking, and on the Market-1501 and DukeMTMC-ReID benchmarks for Re-ID. We examine the correlation between good Re-ID and good MTMCT scores, and perform ablation studies to elucidate the contributions of the main components of our system. Code is available1.
Ergys Ristani, Carlo Tomasi
CVPR2
2017 Tracking Social Groups Within and Across Cameras
abstract
We propose a method for tracking groups from single and multiple cameras with disjointed fields of view. Our formulation follows the tracking-by-detection paradigm in which groups are the atomic entities and are linked over time to form long and consistent trajectories. To this end, we formulate the problem as a supervised clustering problem in which a structural SVM classifier learns a similarity measure appropriate for group entities. Multicamera group tracking is handled inside the framework by adopting an orthogonal feature encoding that allows the classifier to learn inter- and intra-camera feature weights differently. Experiments were carried out on a novel annotated group tracking data set, the DukeMTMC-Groups data set. Since this is the first data set on the problem, it comes with the proposal of a suitable evaluation measure. Results of adopting learning for the task are encouraging, scoring a +15% improvement in F1measure over a nonlearning-based clustering baseline. To the best of our knowledge, this is the first proposal of its kind dealing with multicamera group tracking.
Francesco Solera, Simone Calderara, Ergys Ristani, Carlo Tomasi, Rita Cucchiara
IEEE Trans. Circuits Syst. Video Technol.4
2016 Distance Minimization for Reward Learning from Scored Trajectories
abstract
Many planning methods rely on the use of an immediate reward function as a portable and succinct representation of desired behavior. Rewards are often inferred from demonstrated behavior that is assumed to be near-optimal. We examine a framework, Distance Minimization IRL (DM-IRL), for learning reward functions from scores an expert assigns to possibly suboptimal demonstrations. By changing the expert’s role from a demonstrator to a judge, DM-IRL relaxes some of the assumptions present in IRL, enabling learning from the scoring of arbitrary demonstration trajectories with unknown transition functions. DM-IRL complements existing IRL approaches by addressing different assumptions about the expert. We show that DM-IRL is robust to expert scoring error and prove that finding a policy that produces maximally informative trajectories for an expert to score is strongly NP-hard. Experimentally, we demonstrate that the reward function DM-IRL learns from an MDP with an unknown transition model can transfer to an agent with known characteristics in a novel environment, and we achieve successful learning with limited available training data.
Benjamin Burchfiel, Carlo Tomasi, Ronald Parr
AAAI2
2016 Deformable Graph Model for Tracking Epithelial Cell Sheets in Fluorescence Microscopy
abstract
We propose a novel method for tracking cells that are connected through a visible network of membrane junctions. Tissues of this form are common in epithelial cell sheets and resemble planar graphs where each face corresponds to a cell. We leverage this structure and develop a method to track the entire tissue as a deformable graph. This coupled model in which vertices inform the optimal placement of edges and vice versa captures global relationships between tissue components and leads to accurate and robust cell tracking. We compare the performance of our method with that of four reference tracking algorithms on four data sets that present unique tracking challenges. Our method exhibits consistently superior performance in tracking all cells accurately over all image frames, and is robust over a wide range of image intensity and cell shape profiles. This may be an important tool for characterizing tissues of this type especially in the field of developmental biology where automated cell analysis can help elucidate the mechanisms behind controlled cell-shape changes.
Roger S. Zou, Carlo Tomasi
IEEE Trans. Medical Imaging2
2015 Tree Topology Estimation
abstract
Tree-like structures are fundamental in nature, and it is often useful to reconstruct the topology of a tree - what connects to what - from a two-dimensional image of it. However, the projected branches often cross in the image: the tree projects to a planar graph, and the inverse problem of reconstructing the topology of the tree from that of the graph is ill-posed. We regularize this problem with a generative, parametric tree-growth model. Under this model, reconstruction is possible in linear time if one knows the direction of each edge in the graph - which edge endpoint is closer to the root of the tree - but becomes NP-hard if the directions are not known. For the latter case, we present a heuristic search algorithm to estimate the most likely topology of a rooted, three-dimensional tree from a single two-dimensional image. Experimental results on retinal vessel, plant root, and synthetic tree data sets show that our methodology is both accurate and efficient.
Rolando Estrada, Carlo Tomasi, Scott C. Schmidler, Sina Farsiu
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Retinal Artery-Vein Classification via Topology Estimation
abstract
We propose a novel, graph-theoretic framework for distinguishing arteries from veins in a fundus image. We make use of the underlying vessel topology to better classify small and midsized vessels. We extend our previously proposed tree topology estimation framework by incorporating expert, domain-specific features to construct a simple, yet powerful global likelihood model. We efficiently maximize this model by iteratively exploring the space of possible solutions consistent with the projected vessels. We tested our method on four retinal datasets and achieved classification accuracies of 91.0%, 93.5%, 91.7%, and 90.9%, outperforming existing methods. Our results show the effectiveness of our approach, which is capable of analyzing the entire vasculature, including peripheral vessels, in wide field-of-view fundus photographs. This topology-based method is a potentially important tool for diagnosing diseases with retinal vascular manifestation.
Rolando Estrada, Michael J. Allingham, Priyatham S. Mettu, Scott W. Cousins, Carlo Tomasi, Sina Farsiu
IEEE Trans. Medical Imaging5
2014 Tracking Multiple People Online and in Real Time
Ergys Ristani, Carlo Tomasi
ACCV (5)2
2013 Video Motion for Every Visible Point
abstract
Dense motion of image points over many video frames can provide important information about the world. However, occlusions and drift make it impossible to compute long motion paths by merely concatenating optical flow vectors between consecutive frames. Instead, we solve for entire paths directly, and flag the frames in which each is visible. As in previous work, we anchor each path to a unique pixel which guarantees an even spatial distribution of paths. Unlike earlier methods, we allow paths to be anchored in any frame. By explicitly requiring that at least one visible path passes within a small neighborhood of every pixel, we guarantee complete coverage of all visible points in all frames. We achieve state-of-the-art results on real sequences including both rigid and non-rigid motions with significant occlusions.
Susanna Ricco, Carlo Tomasi
ICCV2
2013 A linear system form solution to compute the local space average color
Joaquín Salas, Carlo Tomasi
Mach. Vis. Appl.2
2012 Twisted window search for efficient shape localization
abstract
Many computer vision systems approximate targets' shape with rectangular bounding boxes. This choice trades localization accuracy for efficient computation. We propose twisted window search, a strict generalization over rectangular window search, for the globally optimal localization of a target's shape. Despite its generality, we show that the new algorithm runs in O(n3), an asymptotic time complexity that is no greater than that of rectangular window search on an image of resolution n × n. We demonstrate improved results of twisted window search for localizing and tracking non-rigid objects with significant orientation, scale and shape change. Twisted window search runs at nearly 10 frames per second in our MATLAB/C++ implementation on images of resolution 240 × 320 on a quad-core laptop.
Steve Gu, Carlo Tomasi
CVPR3
2012 Dense Lagrangian motion estimation with occlusions
abstract
We couple occlusion modeling and multi-frame motion estimation to compute dense, temporally extended point trajectories in video with significant occlusions. Our approach combines robust spatial regularization with spatially and temporally global occlusion labeling in a variational, Lagrangian framework with subspace constraints. We track points even through ephemeral occlusions. Experiments demonstrate accuracy superior to the state of the art while tracking more points through more frames.
Susanna Ricco, Carlo Tomasi
CVPR2
2012 Nested Pictorial Structures
Steve Gu, Carlo Tomasi
ECCV (2)3
2012 Simultaneous Compaction and Factorization of Sparse Image Motion Matrices
Susanna Ricco, Carlo Tomasi
ECCV (6)2
2012 Fast Tiered Labeling with Topological Priors
Steve Gu, Carlo Tomasi
ECCV (4)3
2012 Shape from point features
abstract
We present a nonparametric and efficient method for shape localization that improves on the traditional sub-window search in capturing the fine geometry of an object from a small number of feature points. Our method implies that the discrete set of features capture more appearance and shape information than is commonly exploited. We use the a-complex by Edelsbrunner et al. to build a filtration of simplicial complexes from a user-provided set of features. The optimal value of a is determined automatically by a search for the densest complex connected component, resulting in a parameter-free algorithm. Given K features, localization occurs in O(K log K) time. For VGA-resolution images, computation takes typically less than 10 milliseconds. We use our method for interactive object cut, with promising results.
Steve Gu, Carlo Tomasi
ICASSP3
2012 Oscillation regularization
abstract
We measure the degree of oscillation of a sampled function f by the number of its local extrema. The greater this number, the more oscillatory and complex f becomes. In signal denoising, we want a restored function g that is simple and fits the data f well. We propose to model this by a global optimization, coined oscillation regularization, that reduces both the data fitting error and the number of local extrema of g: equation where err(f, g) measures the discrepancy between f and g and λ is a regularization parameter. To the best of our knowledge, the number of local extrema of g is a topological prior that is rarely exploited in the literature of regularization.
Steve Gu, Carlo Tomasi
ICASSP3
2012 Topological persistence on a Jordan curve
abstract
Topological persistence measures the resilience of extrema of a function to perturbations, and has received increasing attention in computer graphics, visualization and computer vision. While the notion of topological persistence for piece-wise linear functions defined on a simplicial complex has been well studied, the time complexity of all the known algorithms are super-linear (e.g. O(n log n)) in the size n of the complex. We give an O(n) algorithm to compute topological persistence for a function defined on a Jordan curve. To the best of our knowledge, our algorithm is the first to attain linear asymptotic complexity, and is asymptotically optimal. We demonstrate the usefulness of persistence in shape abstraction and compression.
Steve Gu, Carlo Tomasi
ICASSP3
2011 Branch and track
abstract
We present a new paradigm for tracking objects in video in the presence of other similar objects. This branch-and-track paradigm is also useful in the absence of motion, for the discovery of repetitive patterns in images. The object of interest is the lead object and the distracters are extras. The lead tracker branches out trackers for extras when they are detected, and all trackers share a common set of features. Sometimes, extras are tracked because they are of interest in their own right. In other cases, and perhaps more importantly, tracking extras makes tracking the lead nimbler and more robust, both because shared features provide a richer object model, and because tracking extras accounts for sources of confusion explicitly. Sharing features also makes joint tracking less expensive, and coordinating tracking across lead and extras allows optimizing window positions jointly rather than separately, for better results. The joint tracking of both lead and extras can be solved optimally by dynamic programming and branching is quickly determined by efficient subwindow search. Matlab experiments show near real time performance at 5-30 frames per second on a single-core laptop for 240 by 320 images.
Steve Gu, Carlo Tomasi
CVPR2
2011 Linear time offline tracking and lower envelope algorithms
abstract
Offline tracking of visual objects is particularly helpful in the presence of significant occlusions, when a frame-by-frame, causal tracker is likely to lose sight of the target. In addition, the trajectories found by offline tracking are typically smoother and more stable because of the global optimization this approach entails. In contrast with previous work, we show that this global optimization can be performed in O(MNT) time for T frames of video at M × N resolution, with the help of the generalized distance transform developed by Felzenszwalb and Huttenlocher [13]. Recognizing the importance of this distance transform, we extend the computation to a more general lower envelope algorithm in certain heterogeneous l1-distance metric spaces. The generalized lower envelope algorithm is of complexity O(MN(M+N)) and is useful for a more challenging offline tracking problem. Experiments show that trajectories found by offline tracking are superior to those computed by online tracking methods, and are computed at 100 frames per second.
Steve Gu, Carlo Tomasi
ICCV3
2011 Detailed reconstruction of 3D plant root shape
abstract
We study the 3D reconstruction of plant roots from multiple 2D images. To meet the challenge caused by the delicate nature of thin branches, we make three innovations to cope with the sensitivity to image quality and calibration. First, we model the background as a harmonic function to improve the segmentation of the root in each 2D image. Second, we develop the concept of the regularized visual hull which reduces the effect of jittering and refraction by ensuring consistency with one 2D image. Third, we guarantee connectedness through adjustments to the 3D reconstruction that minimize global error. Our software is part of a biological phenotype/genotype study of agricultural root systems. It has been tested on more than 40 plant roots and results are promising in terms of reconstruction quality and efficiency.
Steve Gu, Herbert Edelsbrunner, Carlo Tomasi, Philip Benfey
ICCV4
2011 Detecting motion synchrony by video tubes
abstract
Motion synchrony, i.e., the coordinated motion of a group of individuals, is an interesting phenomenon in nature or daily life. Fish swim in schools, birds fly in flocks, soldiers march in platoons, etc. Our goal is to detect motion synchrony that may be present in the video data, and to track the group of moving objects as a whole. This opens the door to novel algorithms and applications. To this end, we model individual motions as video tubes in space-time, define motion synchrony by the geometric relation among video tubes, and track a whole set of tubes by dynamic programming. The resulting algorithm is highly efficient in practice. Given a video clip of T frames of resolution XxY, we show that finding the K spatially correlated video tubes and determining the presence of synchrony can be solved optimally in O(XYTK) time. Preliminary experiments show that our method is both effective and efficient. Typical running times are 30 - 100 VGA-resolution frames per second after feature extraction, and the accuracy for the detection of synchrony is more than 90% as evaluated in our annotated data set.
Steve Gu, Carlo Tomasi
ACM Multimedia3
2010 Efficient Visual Object Tracking with Online Nearest Neighbor Classifier
Steve Gu, Carlo Tomasi
ACCV (1)3
2010 Critical Nets and Beta-Stable Features for Image Matching
Steve Gu, Carlo Tomasi
ECCV (3)3
2010 Semi-Supervised Fisher Linear Discriminant (SFLD)
abstract
Supervised learning uses a training set of labeled examples to compute a classifier which is a mapping from feature vectors to class labels. The success of a learning algorithm is evaluated by its ability to generalize, i.e., to extend this mapping accurately to new data that is commonly referred to as the test data. Good generalization depends crucially on the quality of the training set. Because collecting labeled data is laborious, training sets are typically small. Furthermore, it is often difficult to represent all possible observation scenarios during training, so that the statistics of the training set end up differing from those of the test data, a problem known as the sample selection bias. To address sample selection bias, we introduce a Semi-Supervised Fisher Linear Discriminant (SFLD) that utilizes additional, unlabeled data to improve generalization for both small and biased training sets. We characterize the conditions under which SFLD helps, and illustrate its benefits through experiments on digit and car recognition applications.
Seda Remus, Carlo Tomasi
ICASSP2
2009 Fingerspelling Recognition through Classification of Letter-to-Letter Transitions
Susanna Ricco, Carlo Tomasi
ACCV (3)2
2009 Phase diffusion for the synchronization of heterogenous sensor streams
abstract
The analysis of complex human activity typically requires multiple sensors: cameras that take videos from different directions and in different areas, microphones, proximity sensors, range finders, and more. Scenarios where it is not possible to associate reliable clocks to each of the sensors pose a synchronization problem between heterogeneous data streams. In this paper, we propose a new theoretical framework for measuring the synchrony between heterogenous sensor streams. The main idea is to model the phase disparity between two data streams explicitly as an Ornstein-Uhlenbeck random process. Based on this model, we derive a simple method for synchronizing of underlying sources. We illustrate the ideas with experiments on audio-visual synchronization and human motion categorization, and report promising results.
Steve Gu, Carlo Tomasi
ICASSP2
2009 Manuscript Bleed-through Removal via Hysteresis Thresholding
abstract
Many types of degradation can render ancient manuscripts very hard to read. In bleed-through, the text from the reverse, or verso, side of a page seeps through into the front, or recto. In this paper, we propose hysteresis thresholding to greatly reduce bleed-through. Thresholding alone cannot properly separate ink and bleed-through because the ranges of intensities for the two classes overlap. Hysteresis thresholding overcomes this limitation via the two steps of thresholding and ink regrowth. In order to provide quantitative measures of the effectiveness of this approach, we constructed a novel dataset which features bleed-through and has available ground truth. We evaluated our method and a number of previously proposed approaches on ink pixel precision and recall. Hysteresis thresholding significantly improves over existing methods.
Rolando Estrada, Carlo Tomasi
ICDAR2
2008 Robust shape normalization based on implicit representations
abstract
We introduce a new shape normalization method based on implicit shape representations. The proposed method is robust with respect to deformations and invariant to similarity transformations (translation, isotropic scaling and rotation). The new method has been tested and compared to the classical shape normalization method and previous work in terms of aligning groups of shapes with deformations.
Tingting Jiang 0001, Carlo Tomasi
ICPR2
2008 Editorial
Cordelia Schmid, Stefano Soatto, Carlo Tomasi
Int. J. Comput. Vis.3
2007 Finite-Element Level-Set Curve Particles
abstract
Particle filters encode a time-evolving probability density by maintaining a random sample from it. Level sets represent closed curves as zero crossings of functions of two variables. The combination of level sets and particle filters presents many conceptual advantages when tracking uncertain, evolving boundaries over time, but the cost of combining these two ideas seems prima facie prohibitive. A previous publication showed that a large number of virtual level set particles can be tracked with a logarithmic amount of work for propagation and update. We now make level- set curve particles more efficient by borrowing ideas from the Finite Element Method (FEM). This improves level-set curve particles in both running time (by a constant factor) and accuracy of the results.
Tingting Jiang 0001, Carlo Tomasi
ICCV2
2007 Correspondence as energy-based segmentation
Stanley T. Birchfield, Braga Natarajan, Carlo Tomasi
Image Vis. Comput.3
2006 Level-Set Curve Particles
Tingting Jiang 0001, Carlo Tomasi
ECCV (3)2
2005 Mean Shift Is a Bound Optimization
abstract
We build on the current understanding of mean shift as an optimization procedure. We demonstrate that, in the case of piecewise constant kernels, mean shift is equivalent to Newton's method. Further, we prove that, for all kernels, the mean shift procedure is a quadratic bound maximization.
Mark Fashing, Carlo Tomasi
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 3D Head Tracking Based on Recognition and Interpolation Using a Time-of-Flight Depth Sensor
Salih Burak Göktürk, Carlo Tomasi
CVPR (2)2
2004 Image Similarity Using Mutual Information of Regions
Daniel B. Russakoff, Carlo Tomasi, Torsten Rohlfing, Calvin R. Maurer Jr.
ECCV (3)2
2004 Surfaces with Occlusions from Layered Stereo
abstract
We propose a new binocular stereo algorithm that estimates scene structure as a collection of smooth surface patches. The disparities within each patch are modeled by a continuous-valued spline, while the extent of each patch is represented via a pixelwise partitioning of the images. Disparities and extents are alternately estimated in an iterative, energy minimization framework. Experimental results demonstrate that, for scenes consisting of smooth surfaces, the proposed algorithm significantly improves upon the state of the art.
Michael H. Lin, Carlo Tomasi
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Surfaces with Occlusions from Layered Stereo
abstract
Although steady progress has been made in recent stereo algorithms, producing accurate results in the neighborhood of depth discontinuities remains a challenge. Moreover, among the techniques that best localize depth discontinuities, it is common to work only with a discrete set of disparity values, hindering the modeling of smooth, non-fronto-parallel surfaces. We propose to estimate scene structure as a set of smooth surface patches. The disparities within each patch are modeled by a spline, while the extent of each patch is represented by a pixelwise labeling of the source images. Disparities and extents are alternately estimated in an iterative, energy minimization framework. Segmentation is via graph cuts, aided by image gradients. Input images are treated symmetrically, and occlusions are addressed explicitly. Promising experimental results are presented.
Michael H. Lin, Carlo Tomasi
CVPR (1)2
2003 3D Tracking = Classification + Interpolation
abstract
Hand gestures are examples of fast and complex motions. Computers fail to track these in fast video, but sleight of hand fools humans as well: what happens too quickly we just cannot see. We show a 3D tracker for these types of motions that relies on the recognition of familiar configurations in 2D images (classification), and fills the gaps in-between (interpolation). We illustrate this idea with experiments on hand motions similar to finger spelling. The penalty for a recognition failure is often small: if two configurations are confused, they are often similar to each other, and the illusion works well enough, for instance, to drive a graphics animation of the moving hand. We contribute advances in both feature design and classifier training: our image features are invariant to image scale, translation, and rotation, and we propose a classification method that combines VQPCA with discrimination trees.
Carlo Tomasi, Slav Petrov, Arvind Sastry
ICCV1
2002 On the Consistency of Instantaneous Rigid Motion Estimation
Tong Zhang 0001, Carlo Tomasi
Int. J. Comput. Vis.2
2002 Edge Displacement Field-Based Classification for Improved Detection of Polyps in CT Colonography
abstract
Colorectal cancer can easily be prevented provided that the precursors to tumors, small colonic polyps, are detected and removed. Currently, the only definitive examination of the colon is fiber-optic colonoscopy, which is invasive and expensive. Computed tomographic colonography (CTC) is potentially a less costly and less invasive alternative to FOC. It would be desirable to have computer-aided detection (CAD) algorithms to examine the large amount of data CTC provides. Most current CAD algorithms have high false positive rates at the required sensitivity levels. We developed and evaluated a postprocessing algorithm to decrease the false positive rate of such a CAD method without sacrificing sensitivity. Our method attempts to model the way a radiologist recognizes a polyp while scrolling a cross-sectional plane through three-dimensional computed tomography data by classification of the changes in the location of the edges in the two-dimensional plane. We performed a tenfold cross-validation study to assess its performance using sensitivity/specificity analysis on data from 48 patients. The mean specificity over all experiments increased from 0.19 (0.35) to 0.47 (0.56) for a sensitivity of 1.00 (0.95).
Burak Acar, Christopher F. Beaulieu, Salih Burak Göktürk, Carlo Tomasi, David S. Paik, R. Brooke Jeffrey Jr., Judy Yee, Sandy Napel
IEEE Trans. Medical Imaging4
2001 A New 3-D Pattern Recognition Technique With Application to Computer Aided Colonoscopy
abstract
To utilize CT or MRI images for computer aided diagnosis applications, robust features that represent 3D image data need to be constructed and subsequently used by a classification method. We present a computer aided diagnosis system for early diagnosis of colon cancer. The system extracts features via a new 3D pattern processing method and processes them using a support vector machine classifier. Our 3D pattern processing method, called Random Orthogonal Shape Section (ROSS) mimics the radiologist's way of viewing these images and combines information from many random triples of mutually orthogonal sections going through the volume. Another contribution of the paper is a new feedback framework between the classification algorithm and the definition of the features. This framework, called Distinctive Component Analysis combines support vector samples with linear discriminant analysis to map the features of clustered support vectors to a lower dimensional space where the two classes of objects of interest are optimally separated to obtain better features. We show that the combination of these better features with support vector machine classification provides a good recognition rate.
Salih Burak Göktürk, Carlo Tomasi
CVPR (1)2
2001 Using Optical Flow Fields for Polyp Detection in Virtual Colonoscopy
Burak Acar, Sandy Napel, David S. Paik, Salih Burak Göktürk, Carlo Tomasi, Christopher F. Beaulieu
MICCAI5
2001 A Learning Method for Automated Polyp Detection
Salih Burak Göktürk, Carlo Tomasi, Burak Acar, David S. Paik, Christopher F. Beaulieu, Sandy Napel
MICCAI2
2001 Empirical Evaluation of Dissimilarity Measures for Color and Texture
Yossi Rubner, Jan Puzicha, Carlo Tomasi, Joachim M. Buhmann
Comput. Vis. Image Underst.3
2001 Edge, Junction, and Corner Detection Using Color Distributions
abstract
For over 30 years (1970-2000) researchers in computer vision have been proposing new methods for performing low-level vision tasks such as detecting edges and corners. One key element shared by most methods is that they represent local image neighborhoods as constant in color or intensity with deviations modeled as noise. Due to computational considerations that encourage the use of small neighborhoods where this assumption holds, these methods remain popular. The research presented models a neighborhood as a distribution of colors. The goal is to show that the increase in accuracy of this representation translates into higher-quality results for low-level vision tasks on difficult, natural images, especially as neighborhood size increases. We emphasize large neighborhoods because small ones often do not contain enough information. We emphasize color because it subsumes gray scale as an image range and because it is the dominant form of human perception. We discuss distributions in the context of detecting edges, corners, and junctions, and we show results for each.
Mark A. Ruzon, Carlo Tomasi
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 A Statistical 3D Pattern Processing Method for Computer Aided Detection of Polyps in CT Colonography
abstract
Adenomatous polyps in the colon are believed to be the precursor to colorectal carcinoma, the second leading cause of cancer deaths in United States. In this paper, we propose a new method for computer-aided detection of polyps in computed tomography (CT) colonography (virtual colonoscopy), a technique in which polyps are imaged along the wall of the air-inflated, cleansed colon with X-ray CT. Initial work with computer aided detection has shown high sensitivity, but at a cost of too many false positives. We present a statistical approach that uses support vector machines to distinguish the differentiating characteristics of polyps and healthy tissue, and uses this information for the classification of the new cases. One of the main contributions of the paper is the new three-dimensional pattern processing approach, called random orthogonal shape sections method, which combines the information from many random images to generate reliable signatures of shape. The input to the proposed system is a collection of volume data from candidate polyps obtained by a high-sensitivity, low-specificity system that we developed previously. The results of our ten-fold cross-validation experiments show that, on the average, the system increases the specificity from 0.19 (0.35) to 0.69 (0.74) at a sensitivity level of 1.0 (0.95).
Salih Burak Göktürk, Carlo Tomasi, Burak Acar, Christopher F. Beaulieu, David S. Paik, R. Brooke Jeffrey Jr., Judy Yee, Sandy Napel
IEEE Trans. Medical Imaging2
2000 Alpha Estimation in Natural Images
abstract
Many boundaries between objects in the world project onto curves in an image. However, boundaries involving natural objects (e.g., trees, hair, water, smoke) are often unworkable under this model because many pixels receive light from more than one object. We propose a technique for estimating alpha, the proportion in which two colors mix to produce a color at the boundary. The technique extends blue screen matting to backgrounds that have almost arbitrary color distributions, though coarse knowledge of the boundary's location is required. Results show a number of different objects moved from one image to another while maintaining naturalism.
Mark A. Ruzon, Carlo Tomasi
CVPR2
2000 The Earth Mover's Distance as a Metric for Image Retrieval
Yossi Rubner, Carlo Tomasi, Leonidas J. Guibas
Int. J. Comput. Vis.2
1999 Color Edge Detection with the Compass Operator
abstract
The compass operator detects step edges without assuming that the regions on either side have constant color. Using distributions of pixel colors rather than the mean, the operator finds the orientation of a diameter that maximizes the difference between two halves of a circular window. Junctions can also be detected by exploiting their lack of bilateral symmetry. This approach is superior to a multi-dimensional gradient method in situations that often result in false negatives, and it localizes edges better as scale increases.
Mark A. Ruzon, Carlo Tomasi
CVPR2
1999 Fast, Robust, and Consistent Camera Motion Estimation
abstract
Previous algorithms that recover camera motion from image velocities suffer from both bias and excessive variance in the results. We propose a robust estimator of camera motion that is statistically consistent when image noise is isotropic. Consistency means that the estimated motion converges in probability, to the true value as the number of image points increases. An algorithm based on reweighted Gauss-Newton iterations handles 100 velocity measurements in about 50 milliseconds on a workstation.
Tong Zhang 0001, Carlo Tomasi
CVPR2
1999 Multiway Cut for Stereo and Motion with Slanted Surfaces
abstract
Slanted surfaces pose a problem for correspondence algorithms utilizing search because of the greatly increased number of possibilities, when compared with fronto-parallel surfaces. In this paper we propose an algorithm to compute correspondence between stereo images or between frames of a motion sequence by minimizing an energy functional that accounts for slanted surfaces. The energy is minimized in a greedy strategy that alternates between segmenting the image into a number of non-overlapping regions (using the multiway-cut algorithm of Boykov, Veksler, and Zabih) and finding the affine parameters describing the displacement function of each region. A follow-up step enables the algorithm to escape local minima due to oversegmentation. Experiments on real images show the algorithm's ability to find an accurate segmentation and displacement map, as well as discontinuities and creases, from a wide variety of stereo and motion imagery.
Stanley T. Birchfield, Carlo Tomasi
ICCV2
1999 Representation Issues in the ML Estimation of Camera Motion
abstract
The computation of camera motion from image measurements is a parameter estimation problem. We show that for the analysis of the problem's sensitivity, the parameterization must enjoy the property of fairness, which makes sensitivity results invariant to changes of coordinates. We prove that Cartesian unit norm vectors and quaternions are fair parameterizations of rotations and translations, respectively, and that spherical coordinates and Euler angles are not. We extend the Gauss-Markov theorem to implicit formulations with constrained parameters, a necessary step in order to take advantage of fair parameterizations. We show how maximum likelihood (ML) estimation problems whose sensitivity depends on a large number of parameters, such as coordinates of points in the scene, can be partitioned into equivalence classes, with problems in the same class exhibiting the same sensitivity.
Joachim Hornegger, Carlo Tomasi
ICCV2
1999 Empirical Evaluation of Dissimilarity Measures for Color and Texture
abstract
This paper empirically compares nine image dissimilarity measures that are based on distributions of color and texture features summarizing over 1,000 CPU hours of computational experiments. Ground truth is collected via a novel random sampling scheme for color and via an image partitioning method for texture. Quantitative performance evaluations are given for classification, image retrieval, and segmentation tasks, and for a wide variety of dissimilarity measures. It is demonstrated how the selection of a measure, based on large scale evaluation, substantially improves the quality of classification, retrieval, and unsupervised segmentation of color and texture images.
Jan Puzicha, Yossi Rubner, Carlo Tomasi, Joachim M. Buhmann
ICCV3
1999 Texture-based Image Retrieval without Segmentation
abstract
Image segmentation is not only hard and unnecessary for texture-based image retrieval, but can even be harmful. Images of either individual or multiple textures are best described by distributions of spatial frequency descriptors, rather than single descriptor vectors over presegmented regions. A retrieval method based on the earth movers distance with an appropriate ground distance is shown to handle both complete and partial multi-textured queries. As an illustration, different images of the same type of animal are easily retrieved together. At the same time, animals with subtly different coats, like cheetahs and leopards, are properly distinguished.
Yossi Rubner, Carlo Tomasi
ICCV2
1999 Corner Detection in Textured Color Images
abstract
Corner models in the literature have lagged behind edge models with respect to color and shading. We use both a region model, based on distributions of pixel colors, and an edge model, which removes false positives, to perform corner detection on color images whose regions contain texture. We show results on a variety of natural images at different scales that highlight the problems that occur when boundaries between regions have curvature.
Mark A. Ruzon, Carlo Tomasi
ICCV2
1999 Depth Discontinuities by Pixel-to-Pixel Stereo
Stanley T. Birchfield, Carlo Tomasi
Int. J. Comput. Vis.2
1998 Adaptive Color-Image Embeddings for Database Navigation
Yossi Rubner, Carlo Tomasi, Leonidas J. Guibas
ACCV (1)2
1998 Depth Discontinuities by Pixel-to-Pixel Stereo
abstract
An algorithm to detect depth discontinuities from a stereo pair of images is presented. The algorithm matches individual pixels in corresponding scanline pairs while allowing occluded pixels to remain unmatched, then propagates the information between scanlines by means of a fast postprocessor. The algorithm handles large untextured regions, uses a measure of pixel dissimilarity that is insensitive to image sampling, and prunes bad search nodes to increase the speed of dynamic programming. The computation is relatively fast, taking about 1.5 microseconds per pixel per disparity on a workstation. Approximate disparity maps and precise depth discontinuities (along both horizontal and vertical boundaries) are shown for five stereo images containing textured, untextured, fronto-parallel, and slanted objects.
Stanley T. Birchfield, Carlo Tomasi
ICCV2
1998 A Metric for Distributions with Applications to Image Databases
abstract
We introduce a new distance between two distributions that we call the Earth Mover's Distance (EMD), which reflects the minimal amount of work that must be performed to transform one distribution into the other by moving "distribution mass" around. This is a special case of the transportation problem from linear optimization, for which efficient algorithms are available. The EMD also allows for partial matching. When used to compare distributions that have the same overall mass, the EMD is a true metric, and has easy-to-compute lower bounds. In this paper we focus on applications to image databases, especially color and texture. We use the EMD to exhibit the structure of color-distribution and texture spaces by means of Multi-Dimensional Scaling displays. We also propose a novel approach to the problem of navigating through a collection of color images, which leads to a new paradigm for image database search.
Yossi Rubner, Carlo Tomasi, Leonidas J. Guibas
ICCV2
1998 Bilateral Filtering for Gray and Color Images
abstract
Bilateral filtering smooths images while preserving edges, by means of a nonlinear combination of nearby image values. The method is noniterative, local, and simple. It combines gray levels or colors based on both their geometric closeness and their photometric similarity, and prefers near values to distant values in both domain and range. In contrast with filters that operate on the three bands of a color image separately, a bilateral filter can enforce the perceptual metric underlying the CIE-Lab color space, and smooth colors and preserve edges in a way that is tuned to human perception. Also, in contrast with standard filtering, bilateral filtering produces no phantom colors along edges in color images, and reduces phantom colors where they appear in the original image.
Carlo Tomasi, Roberto Manduchi
ICCV1
1998 Texture metrics
abstract
We introduce a class of metric perceptual distances between textures. The first metric is sensitive to both rotation and scale differences, and provides a basis for two other metrics, one invariant to rotation, and the other invariant to both rotation and scale. Our metrics are based on Earth Mover's Distance computations on log-polar distributions of spatial frequency computed from Gabor filters. We show consistency of our metrics with psychophysical findings on texture discrimination and classification.
Yossi Rubner, Carlo Tomasi
SMC2
1998 A Pixel Dissimilarity Measure That Is Insensitive to Image Sampling
abstract
Because of image sampling, traditional measures of pixel dissimilarity can assign a large value to two corresponding pixels in a stereo pair, even in the absence of noise and other degrading effects. We propose a measure of dissimilarity that is provably insensitive to sampling because it uses the linearly interpolated intensity functions surrounding the pixels. Experiments on real images show that our measure alleviates the problem of sampling with little additional computational overhead.
Stanley T. Birchfield, Carlo Tomasi
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Stereo Matching as a Nearest-Neighbor Problem
abstract
We propose a representation of images, called intrinsic curves, that transforms stereo matching from a search problem into a nearest-neighbor problem. Intrinsic curves are the paths that a set of local image descriptors trace as an image scanline is traversed from left to right. Intrinsic curves are ideally invariant with respect to disparity. Stereo correspondence then becomes a trivial lookup problem in the ideal case. We also show how to use intrinsic curves to match real images in the presence of noise, brightness bias, contrast fluctuations, moderate geometric distortion, image ambiguity, and occlusions. In this case, matching becomes a nearest-neighbor problem, even for very large disparity values.
Carlo Tomasi, Roberto Manduchi
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Comparison of Approaches to Egomotion Computation
abstract
We evaluated six algorithms for computing egomotion from image velocities. We established benchmarks for quantifying bias and sensitivity to noise, and for quantifying the convergence properties of those algorithms that require numerical search. Our simulation results reveal some interesting and surprising results. First, it is often written in the literature that the egomotion problem is difficult because translation (e.g., along the X-axis) and rotation (e.g., about the Y-axis) produce similar image velocities. We found, to the contrary, that the bias and sensitivity of our six algorithms are totally invariant with respect to the axis of rotation. Second, it is also believed by some that fixating helps to make the egomotion problem easier: We found, to the contrary, that fixating does not help when the noise is independent of the image velocities. Fixation does help if the noise is proportional to speed, but this is only for the trivial reason that the speeds are slower under fixation. Third, it is widely believed that increasing the field of view will yield better performance. We found, to the contrary, that this is not necessarily true.
Tina Yu Tian, Carlo Tomasi, David J. Heeger
CVPR2
1996 Stereo Without Search
Carlo Tomasi, Roberto Manduchi
ECCV (1)1
1995 Linear and Incremental Acquisition of Invariant Shape Models From Image Sequences
abstract
We show how to automatically acquire Euclidian shape representations of objects from noisy image sequences under weak perspective. The proposed method is linear and incremental, requiring no more than pseudoinverse. A nonlinear, but numerically sound preprocessing stage is added to improve the accuracy of the results even further. Experiments show that attention to noise and computational techniques improves the shape results substantially with respect to previous methods proposed for ideal images.>
Daphna Weinshall, Carlo Tomasi
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 Good features to track
abstract
No feature-based vision system can work until good features can be identified and tracked from frame to frame. Although tracking itself is by and large a solved problem, selecting features that can be tracked well and correspond to physical points in the world is still an open problem. We propose a feature selection criterion that is optimal by construction because it is based on how the tracker works, as well as a feature monitoring method that can detect occlusions, disocclusions, and features that do not correspond to points in the world. These methods are based on a new tracking algorithm that extends previous Newton-Raphson style search methods to work under affine image transformations. We test performance with several simulations and experiments on real images.
Jianbo Shi, Carlo Tomasi
CVPR2
1994 Pictures and trails: a new framework for the computation of shape and motion from perspective image sequences
abstract
This paper presents a new framework for the computation of shape and motion from a sequence of images taken under perspective projection. The framework is based on two abstractions, the picture and trail loci, that represent respectively the set of all pictures of the same scene and the set of all trails that a point in the world can leave on the image for a given camera trajectory. These abstractions lead to a remarkably clean relation between perspective and orthography. A shape and motion reconstruction method is developed for the case of a two-dimensional world but all concepts also hold in three dimensions. Experiments show that the method is rather immune to noise but critically dependent on camera calibration.>
Carlo Tomasi
CVPR1
1993 Direction of heading from image deformations
abstract
A method is proposed to compute the direction of heading from the differential changes in the angles between the projection rays of pairs of point features. These angles, the image deformations, do not depend on viewer rotation. The key problem of separating the effects of rotation from those of translation is solved at the input. Experiments show both the feasibility of the method on real images and the advantages of using deformation rather than optical flow.>
Carlo Tomasi, Jianbo Shi
CVPR1
1993 Linear and incremental acquisition of invariant shape models from image sequences
abstract
The authors show how to automatically acquire similarity-invariant shape representations of objects from noisy image sequences under a weak perspective. The incremental nature of the method makes it possible to process images one at a time, moving away from the storage-intensive batch methods of the past. It is based on the observation that the trajectories that points on the object form in weak-perspective image sequences are linear combinations of three of the trajectories themselves, and that the coefficients of the linear combinations represent shape in an affine-invariant basis. A nonlinear but numerically sound preprocessing state is added to improve the accuracy of the results even further. Experiments showed that attention to noise and computational techniques improved the shape results substantially with respect to previous methods.>
Daphna Weinshall, Carlo Tomasi
ICCV2
1992 Shape and motion from image streams under orthography: a factorization method
Carlo Tomasi, Takeo Kanade
Int. J. Comput. Vis.1
1990 Shape and motion without depth
abstract
Inferring the depth and shape of remote objects and the camera motion from a sequence of images is possible in principle, but is an ill-conditioned problem when the objects are distant with respect to their size. This problem is overcome by inferring shape and motion without computing depth as an intermediate step. On a single epipolar plane, an image sequence can be represented by the F*P matrix of the image coordinates of P points tracked through F frames. It is shown that under orthographic projection this matrix is of rank three. Using this result, the authors develop a shape-and-motion algorithm based on singular value decomposition. The algorithm gives accurate results, without relying on any smoothness assumption for either shape or motion.>
Carlo Tomasi, Takeo Kanade
ICCV1
1984 Spectral Analysis of Line Regenerator Time Jitter
abstract
A closed form expression for the spectral density of the time jitter produced in a line regenerator in the presence of a general polynomial nonlinear circuit, band-limited baseband pulses, and an arbitrary tuned filter is developed. The method applies to a PAM digital signal with mutually independent symbols but with an arbitrary number of levels and symbol probabilities. As an example of the method, the spectrum of the time jitter produced in the presence of a fourth power law nonlinear circuit is reported.
Silvano Pupolin, Carlo Tomasi
IEEE Trans. Commun.2