EDBT 2026 Demo / reviewers in the wild / expert
Erik G. Learned-Miller
dblp:l/ErikGLearnedMiller · also Erik G. Miller, Erik Gundersen Miller
· DBLP profile ↗
100ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0002-3778-9135ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 78 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 50 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 12 · 1 since 2021Systems, architecture and hardware · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Computer networks · 2Security and privacy · 2Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human Pose Aggregation for Multi-View Temporal Video AlignmentabstractWhen multiple videos of a scene are taken from differing viewpoints without precise synchronization, it can be difficult to temporally align them after the fact. Often the metadata or audio needed to do so is missing or inaccurate. But human motion in such videos can provide a strong signal for identifying matching time points across videos, through analysis of pose and movement. In this work, we leverage view-invariant human pose features to synchronize videos. Unlike previous human pose-based alignment techniques, our method can align videos containing multiple people without performing tracking or re-identification across views. We achieve this by aggregating pose information from multiple people into a single frame descriptor. This also enables fast ${\mathcal{O}}\left({n\log n}\right)$ search for the optimal alignment. This simple but effective strategy leads to major and consistent improvements over existing human-based and visual feature temporal alignment techniques. Fabien Delattre, Tsung-Wei Huang, Guan-Ming Su, Erik G. Learned-Miller |
WACV | 4 |
| 2026 | Toward Unified Expertise: Learning a Single Vision Model From Diverse PerceptionabstractMulti-task learning (MTL) presents greater optimization challenges than single-task learning (STL) due to conflicting gradients across tasks. While parameter sharing promotes cooperation among related tasks, many tasks require specialized representations. To balance cooperation and specialization, we propose Mod-Squad (Chen et al. 2023), a modular transformer-based model composed of a "squad" of experts. Each task activates a sparse subset of experts through a differentiable matching process, guided by a novel mutual information-based loss. This modular structure avoids full backbone sharing and scales effectively with the number of tasks and dataset size. In this extended version, we generalize Mod-Squad to support multi-dataset pre-training, enabling joint learning across disjoint, single-task datasets (e.g., ImageNet, COCO, ADE20 K). This is achieved via a new formulation of the mutual information loss that unifies learning across heterogeneous sources. More importantly, while most prior work in large models has focused on efficiency, few have explored adjustable efficiency. In this study, we further evaluate the model's generalization to downstream tasks and introduce a set of efficient adaptation techniques that leverage Mod-Squad's modularity for flexible fine-tuning-enabling dynamic adjustment of model size, parameter count, and computational cost. Additionally, we present a hybrid adaptation scheme that combines these techniques to achieve favorable performance-efficiency trade-offs. In summary, Mod-Squad provides a robust foundation for sparse modular models that can learn from diverse supervision and datasets. Its emergent modularity enables strong generalization, decomposition into high-performing components, and rapid, resource-efficient adaptation for downstream applications. Zitian Chen, Mingyu Ding, Yikang Shen, Erik G. Learned-Miller, Chuang Gan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Improving Pre-trained Self-Supervised Embeddings Through Effective Entropy MaximizationabstractA number of different architectures and loss functions have been applied to the problem of self-supervised learning (SSL), with the goal of developing embeddings that provide the best possible pre-training for as-yet-unknown, lightly supervised downstream tasks. One of these SSL criteria is to maximize the entropy of a set of embeddings in some compact space. But the goal of maximizing the embedding entropy often depends—whether explicitly or implicitly—upon high dimensional entropy estimates, which typically perform poorly in more than a few dimensions. In this paper, we motivate an effective entropy maximization criterion (E2MC), defined in terms of easy-to-estimate, low-dimensional constraints. We demonstrate that using it to continue training an already-trained SSL model for only a handful of epochs leads to a consistent and, in some cases, significant improvement in downstream performance. We perform careful ablation studies to show that the improved performance is due to the proposed add-on criterion. We also show that continued pre-training with alternative criteria does not lead to notable improvements, and in some cases, even degrades performance. Deep Chakraborty, Yann LeCun, Tim G. J. Rudner, Erik G. Learned-Miller |
AISTATS | 4 |
| 2024 | The Right Spin: Learning Object Motion from Rotation-Compensated Flow FieldsabstractAbstract A good understanding of geometrical concepts as well as a broad familiarity with objects lead to excellent human perception of moving objects. The human ability to detect and segment moving objects works in the presence of multiple objects, complex background geometry, motion of the observer and even camouflage. How we perceive moving objects so reliably is a longstanding research question in computer vision and borrows findings from related areas such as psychology, cognitive science and physics. One approach to the problem is to teach a deep network to model all of these effects. This is in contrast with the strategy used by human vision, where cognitive processes and body design are tightly coupled and each is responsible for certain aspects of correctly identifying moving objects. Similarly, from the computer vision perspective there is evidence that classical, geometry-based techniques are better suited to the “motion-based” parts of the problem, while deep networks are more suitable for modeling appearance. In this work, we argue that the coupling of camera rotation and camera translation can create complex motion fields that are difficult for a deep network to untangle directly. We present a novel probabilistic model to estimate the camera’s rotation given the motion field. We then rectify the flow field to obtain a rotation-compensated motion field for subsequent segmentation. This strategy of first estimating camera motion, and then allowing a network to learn the remaining parts of the problem, yields improved results on the widely used DAVIS benchmark as well as the more recent motion segmentation data set MoCA (Moving Camouflaged Animals). Pia Bideau, Erik G. Learned-Miller, Cordelia Schmid, Karteek Alahari |
Int. J. Comput. Vis. | 2 |
| 2023 | Mod-Squad: Designing Mixtures of Experts As Modular Multi-Task LearnersabstractOptimization in multi-task learning (MTL) is more challenging than single-task learning (STL), as the gradient from different tasks can be contradictory. When tasks are related, it can be beneficial to share some parameters among them (cooperation). However, some tasks require additional parameters with expertise in a specific type of data or discrimination (specialization). To address the MTL challenge, we propose Mod-Squad, a new model that is Modularized into groups of experts (a ‘Squad’). This structure allows us to formalize cooperation and specialization as the process of matching experts and tasks. We optimize this matching process during the training of a single model. Specifically, we incorporate mixture of experts (MoE) layers into a transformer model, with a new loss that incorporates the mutual dependence between tasks and experts. As a result, only a small set of experts are activated for each task. This prevents the sharing of the entire backbone model between all tasks, which strengthens the model, especially when the training set size and the number of tasks scale up. More interestingly, for each task, we can extract the small set of experts as a standalone model that maintains the same performance as the large model. Extensive experiments on the Taskonomy dataset with 13 vision tasks and the PASCAL-Context dataset with 5 vision tasks show the superiority of our approach. The project page can be accessed at https://vis-www.cs.umass.edu/Mod-Squad. Zitian Chen, Yikang Shen, Mingyu Ding, Zhenfang Chen, Hengshuang Zhao, Erik G. Learned-Miller, Chuang Gan 0001 |
CVPR | 6 |
| 2023 | EVAL: Explainable Video Anomaly LocalizationabstractWe develop a novel framework for single-scene video anomaly localization that allows for human-understandable reasons for the decisions the system makes. We first learn general representations of objects and their motions (using deep networks) and then use these representations to build a high-level, location-dependent model of any particular scene. This model can be used to detect anomalies in new videos of the same scene. Importantly, our approach is explainable - our high-level appearance and motion features can provide human-understandable reasons for why any part of a video is classified as normal or anomalous. We conduct experiments on standard video anomaly detection datasets (Street Scene, CUHK Avenue, ShanghaiTech and UCSD Ped1, Ped2) and show significant improvements over the previous state-of-the-art. All of our code and extra datasets will be made publicly available. Michael J. Jones 0001, Erik G. Learned-Miller |
CVPR | 3 |
| 2023 | Inv-Senet: Invariant Self Expression Network for Clustering Under Biased DataabstractSubspace clustering algorithms are used for understanding the cluster structure that explains the patterns prevalent in the dataset well. These methods are extensively used for data-exploration tasks in various areas of Natural Sciences. However, most of these methods fail to handle confounding attributes in the dataset. For datasets where a data sample represent multiple attributes, naively applying any clustering approach can result in undesired output. To this end, we propose a novel framework for jointly removing confounding attributes while learning to cluster data points in individual subspaces. Assuming we have label information about these confounding attributes, we regularize the clustering method by adversarially learning to minimize the mutual information between the data representation and the confounding attribute labels. Our experimental result on synthetic and real-world datasets demonstrate the effectiveness of our approach. Aria Masoomi, Tales Imbiriba, Erik G. Learned-Miller, Deniz Erdogmus |
ICASSP | 5 |
| 2023 | Robust Frame-to-Frame Camera Rotation Estimation in Crowded ScenesabstractWe present an approach to estimating camera rotation in crowded, real-world scenes from handheld monocular video. While camera rotation estimation is a well-studied problem, no previous methods exhibit both high accuracy and acceptable speed in this setting. Because the setting is not addressed well by other datasets, we provide a new dataset and benchmark, with high-accuracy, rigorously verified ground truth, on 17 video sequences. Methods developed for wide baseline stereo (e.g., 5-point methods) perform poorly on monocular video. On the other hand, methods used in autonomous driving (e.g., SLAM) leverage specific sensor setups, specific motion models, or local optimization strategies (lagging batch processing) and do not generalize well to handheld video. Finally, for dynamic scenes, commonly used robustification techniques like RANSAC require large numbers of iterations, and become prohibitively slow. We introduce a novel generalization of the Hough transform on SO(3) to efficiently and robustly find the camera rotation most compatible with optical flow. Among comparably fast methods, ours reduces error by almost 50% over the next best, and is more accurate than any method, irrespective of speed. This represents a strong new performance point for crowded scenes, an important setting for computer vision. The code and the dataset are available at https://fabiendelattre.com/robustrotation-estimation. Fabien Delattre, David Dirnfeld, Phat Nguyen, Stephen Scarano, Michael J. Jones 0001, Pedro Miraldo, Erik G. Learned-Miller |
ICCV | 7 |
| 2023 | Event Camera-Based Visual Odometry for Dynamic Motion Tracking of a Legged Robot Using Adaptive Time SurfaceabstractOur paper proposes a direct sparse visual odometry method that combines event and RGBD data to estimate the pose of agile-legged robots during dynamic locomotion and acrobatic behaviors. Event cameras offer high temporal resolution and dynamic range, which can eliminate the issue of blurred RGB images during fast movements. This unique strength holds a potential for accurate pose estimation of agile- legged robots, which has been a challenging problem to tackle. Our framework leverages the benefits of both RGBD and event cameras to achieve robust and accurate pose estimation, even during dynamic maneuvers such as jumping and landing a quadruped robot, the Mini-Cheetah. Our major contributions are threefold: Firstly, we introduce an adaptive time surface (ATS) method that addresses the whiteout and blackout issue in conventional time surfaces by formulating pixel-wise decay rates based on scene complexity and motion speed. Secondly, we develop an effective pixel selection method that directly samples from event data and applies sample filtering through ATS, enabling us to pick pixels on distinct features. Lastly, we propose a nonlinear pose optimization formula that simultaneously performs 3D-2D alignment on both RGB-based and event-based maps and images, allowing the algorithm to fully exploit the benefits of both data streams. We extensively evaluate the performance of our framework on both the public dataset and our own quadruped robot dataset, demonstrating its effectiveness in accurately estimating the pose of agile robots during dynamic movements. Supplemental video: https://youtu.be/-5ieQShOg3M Shifan Zhu, Zhipeng Tang, Michael Yang, Erik G. Learned-Miller, Donghyun Kim 0002 |
IROS | 4 |
| 2023 | Machine Learning for Automated Mitral Regurgitation Detection from Cardiac Imaging
Erik G. Learned-Miller, Evangelos Kalogerakis, James Priest, Madalina Fiterau |
MICCAI (7) | 2 |
| 2023 | DCVNet: Dilated Cost Volume Networks for Fast Optical FlowabstractThe cost volume, capturing the similarity of possible correspondences across two input images, is a key ingredient in state-of-the-art optical flow approaches. When sampling correspondences to build the cost volume, a large neighborhood radius is required to deal with large displacements, introducing a significant computational burden. To address this, coarse-to-fine or recurrent processing of the cost volume is usually adopted, where correspondence sampling in a local neighborhood with a small radius suffices. In this paper, we propose an alternative by constructing cost volumes with different dilation factors to capture small and large displacements simultaneously. A U-Net with sikp connections is employed to convert the dilated cost volumes into interpolation weights between all possible captured displacements to get the optical flow. Our proposed model DCVNet only needs to process the cost volume once in a simple feedforward manner and does not rely on the sequential processing strategy. DCVNet obtains comparable accuracy to existing approaches and achieves real-time inference (30 fps on a mid-end 1080ti GPU). Huaizu Jiang, Erik G. Learned-Miller |
WACV | 2 |
| 2023 | A domain-agnostic approach for characterization of lifelong learning systems
Megan M. Baker, Alexander New, Mario Aguilar-Simon, Ziad Al-Halah, Sébastien M. R. Arnold, Eseoghene Benjamin, Andrew P. Brna, Ethan Brooks, Ryan C. Brown, Zachary A. Daniels, Anurag Reddy Daram, Fabien Delattre, Ryan Dellana, Eric Eaton, Haotian Fu, Kristen Grauman, Jesse Hostetler, Shariq Iqbal, Cassandra Kent, Nicholas Ketz, Soheil Kolouri, George Dimitri Konidaris, Dhireesha Kudithipudi, Erik G. Learned-Miller, Michael L. Littman, Sandeep Madireddy, Jorge A. Mendez, Eric Q. Nguyen, Christine D. Piatko, Praveen K. Pilly, Aswin Raghavan, Abrar Rahman, Santhosh K. Ramakrishnan, Neale Ratzlaff, Andrea Soltoggio, Peter Stone 0001, Indranil Sur, Zhipeng Tang, Saket Tiwari, Kyle Vedder, Felix Wang, Zifan Xu, Angel Yanguas-Gil, Harel Yedidsion, Shangqun Yu, Gautam K. Vallabha |
Neural Networks | 24 |
| 2022 | PriFit: Learning to Fit Primitives Improves Few Shot Point Cloud SegmentationabstractAbstract We present PriFit, a semi‐supervised approach for label‐efficient learning of 3D point cloud segmentation networks. PriFit combines geometric primitive fitting with point‐based representation learning. Its key idea is to learn point representations whose clustering reveals shape regions that can be approximated well by basic geometric primitives, such as cuboids and ellipsoids. The learned point representations can then be re‐used in existing network architectures for 3D point cloud segmentation, and improves their performance in the few‐shot setting. According to our experiments on the widely used ShapeNet and PartNet benchmarks, PriFit outperforms several state‐of‐the‐art methods in this setting, suggesting that decomposability into primitives is a useful prior for learning representations predictive of semantic parts. We present a number of ablative experiments varying the choice of geometric primitives and downstream tasks to demonstrate the effectiveness of the method. Gopal Sharma, Bidya Dash, Aruni Roy Chowdhury, Matheus Gadelha, Marios Loizou, Liangliang Cao, Rui Wang 0003, Erik G. Learned-Miller, Subhransu Maji, Evangelos Kalogerakis |
Comput. Graph. Forum | 8 |
| 2021 | The Spatio-Temporal Poisson Point Process: A Simple Model for the Alignment of Event Camera DataabstractEvent cameras, inspired by biological vision systems, provide a natural and data efficient representation of visual information. Visual information is acquired in the form of events that are triggered by local brightness changes. However, because most brightness changes are triggered by relative motion of the camera and the scene, the events recorded at a single sensor location seldom correspond to the same world point. To extract meaningful information from event cameras, it is helpful to register events that were triggered by the same underlying world point. In this work we propose a new model of event data that captures its natural spatio-temporal structure. We start by developing a model for aligned event data. That is, we develop a model for the data as though it has been perfectly registered already. In particular, we model the aligned data as a spatio-temporal Poisson point process. Based on this model, we develop a maximum likelihood approach to registering events that are not yet aligned. That is, we find transformations of the observed events that make them as likely as possible under our model. In particular we extract the camera rotation that leads to the best event alignment. We show new state of the art accuracy for rotational velocity estimation on the DAVIS 240C dataset [20]. In addition, our method is also faster and has lower computational complexity than several competing methods. Code: https://github.com/pbideau/Event-ST-PPP Erik G. Learned-Miller, Daniel Sheldon, Guillermo Gallego 0002, Pia Bideau |
ICCV | 2 |
| 2021 | Towards Practical Mean Bounds for Small SamplesabstractHistorically, to bound the mean for small sample sizes, practitioners have had to choose between using methods with unrealistic assumptions about the unknown distribution (e.g., Gaussianity) and methods like Hoeffding’s inequality that use weaker assumptions but produce much looser (wider) intervals. In 1969, \citet{Anderson1969} proposed a mean confidence interval strictly better than or equal to Hoeffding’s whose only assumption is that the distribution’s support is contained in an interval $[a,b]$. For the first time since then, we present a new family of bounds that compares favorably to Anderson’s. We prove that each bound in the family has {\em guaranteed coverage}, i.e., it holds with probability at least $1-\alpha$ for all distributions on an interval $[a,b]$. Furthermore, one of the bounds is tighter than or equal to Anderson’s for all samples. In simulations, we show that for many distributions, the gain over Anderson’s bound is substantial. My Phan, Philip S. Thomas, Erik G. Learned-Miller |
ICML | 3 |
| 2021 | Universal Off-Policy EvaluationabstractWhen faced with sequential decision-making problems, it is often useful to be able to predict what would happen if decisions were made using a new policy. Those predictions must often be based on data collected under some previously used decision-making rule. Many previous methods enable such off-policy (or counterfactual) estimation of the expected value of a performance measure called the return. In this paper, we take the first steps towards a 'universal off-policy estimator' (UnO)---one that provides off-policy estimates and high-confidence bounds for any parameter of the return distribution. We use UnO for estimating and simultaneously bounding the mean, variance, quantiles/median, inter-quantile range, CVaR, and the entire cumulative distribution of returns. Finally, we also discuss UnO's applicability in various settings, including fully observable, partially observable (i.e., with unobserved confounders), Markovian, non-Markovian, stationary, smoothly non-stationary, and discrete distribution shifts. Yash Chandak, Scott Niekum, Bruno C. da Silva 0001, Erik G. Learned-Miller, Emma Brunskill, Philip S. Thomas |
NeurIPS | 4 |
| 2021 | Passage Retrieval for Outside-Knowledge Visual Question AnsweringabstractIn this work, we address multi-modal information needs that contain text questions and images by focusing on passage retrieval for outside-knowledge visual question answering. This task requires access to outside knowledge, which in our case we define to be a large unstructured passage collection. We first conduct sparse retrieval with BM25 and study expanding the question with object names and image captions. We verify that visual clues play an important role and captions tend to be more informative than object names in sparse retrieval. We then construct a dual-encoder dense retriever, with the query encoder being LXMERT, a multi-modal pre-trained transformer. We further show that dense retrieval significantly outperforms sparse retrieval that uses object expansion. Moreover, dense retrieval matches the performance of sparse retrieval that leverages human-generated captions. Chen Qu 0001, Hamed Zamani, Liu Yang 0005, W. Bruce Croft, Erik G. Learned-Miller |
SIGIR | 5 |
| 2021 | Image registration: Maximum likelihood, minimum entropy and deep learning
Alireza Sedghi, Lauren O'Donnell, Tina Kapur, Erik G. Learned-Miller, Parvin Mousavi, William M. Wells III |
Medical Image Anal. | 4 |
| 2020 | In Defense of Grid Features for Visual Question AnsweringabstractPopularized as `bottom-up' attention, bounding box (or region) based visual features have recently surpassed vanilla grid-based convolutional features as the de facto standard for vision and language tasks like visual question answering (VQA). However, it is not clear whether the advantages of regions (e.g. better localization) are the key reasons for the success of bottom-up attention. In this paper, we revisit grid features for VQA, and find they can work surprisingly well -- running more than an order of magnitude faster with the same accuracy (e.g. if pre-trained in a similar fashion). Through extensive experiments, we verify that this observation holds true across different VQA models (reporting a state-of-the-art accuracy on VQA 2.0 test-std, 72.71), datasets, and generalizes well to other tasks like image captioning. As grid features make the model design and training process much simpler, this enables us to train them end-to-end and also use a more flexible network design. We learn VQA models end-to-end, from pixels directly to answers, and show that strong performance is achievable without using any region annotations in pre-training. We hope our findings help further improve the scientific understanding and the practical application of VQA. Code and features will be made available. Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik G. Learned-Miller, Xinlei Chen |
CVPR | 4 |
| 2020 | Label-Efficient Learning on Point Clouds Using Approximate Convex Decompositions
Matheus Gadelha, Aruni Roy Chowdhury, Gopal Sharma, Evangelos Kalogerakis, Liangliang Cao, Erik G. Learned-Miller, Rui Wang 0003, Subhransu Maji |
ECCV (10) | 6 |
| 2020 | Improving Face Recognition by Clustering Unlabeled Faces in the Wild
Aruni Roy Chowdhury, Xiang Yu 0002, Kihyuk Sohn, Erik G. Learned-Miller, Manmohan Krishna Chandraker |
ECCV (24) | 4 |
| 2020 | C-14: assured timestamps for drone videosabstractInexpensive and highly capable unmanned aerial vehicles (aka drones) have enabled people to contribute high-quality videos at a global scale. However, a key challenge exists for accepting videos from untrusted sources: establishing when a particular video was taken. Once a video has been received or posted publicly, it is evident that the video was created before that time, but there are no current methods for establishing how long it was made before that time. Zhipeng Tang, Fabien Delattre, Pia Bideau, Mark D. Corner, Erik G. Learned-Miller |
MobiCom | 5 |
| 2019 | Automatic Adaptation of Object Detectors to New Domains Using Self-TrainingabstractThis work addresses the unsupervised adaptation of an existing object detector to a new target domain. We assume that a large number of unlabeled videos from this domain are readily available. We automatically obtain labels on the target data by using high-confidence detections from the existing detector, augmented with hard (misclassified) examples acquired by exploiting temporal cues using a tracker. These automatically-obtained labels are then used for re-training the original model. A modified knowledge distillation loss is proposed, and we investigate several ways of assigning soft-labels to the training examples from the target domain. Our approach is empirically evaluated on challenging face and pedestrian detection tasks: a face detector trained on WIDER-Face, which consists of high-quality images crawled from the web, is adapted to a large-scale surveillance data set; a pedestrian detector trained on clear, daytime images from the BDD-100K driving data set is adapted to all other scenarios such as rainy, foggy, night-time. Our results demonstrate the usefulness of incorporating hard examples obtained from tracking, the advantage of using soft-labels via distillation loss versus hard-labels, and show promising performance as a simple method for unsupervised domain adaptation of object detectors, with minimal dependence on hyper-parameters. Aruni Roy Chowdhury, Prithvijit Chakrabarty, SouYoung Jin, Huaizu Jiang, Liangliang Cao, Erik G. Learned-Miller |
CVPR | 7 |
| 2019 | Pixel-Adaptive Convolutional Neural NetworksabstractConvolutions are the fundamental building blocks of CNNs. The fact that their weights are spatially shared is one of the main reasons for their widespread use, but it is also a major limitation, as it makes convolutions content-agnostic. We propose a pixel-adaptive convolution (PAC) operation, a simple yet effective modification of standard convolutions, in which the filter weights are multiplied with a spatially varying kernel that depends on learnable, local pixel features. PAC is a generalization of several popular filtering techniques and thus can be used for a wide range of use cases. Specifically, we demonstrate state-of-the-art performance when PAC is used for deep joint image upsampling. PAC also offers an effective alternative to fully-connected CRF (Full-CRF), called PAC-CRF, which performs competitively compared to Full-CRF, while being considerably faster. In addition, we also demonstrate that PAC can be used as a drop-in replacement for convolution layers in pre-trained networks, resulting in consistent performance improvements. Hang Su 0005, Varun Jampani, Deqing Sun, Orazio Gallo, Erik G. Learned-Miller, Jan Kautz |
CVPR | 5 |
| 2019 | SENSE: A Shared Encoder Network for Scene-Flow EstimationabstractWe introduce a compact network for holistic scene flow estimation, called SENSE, which shares common encoder features among four closely-related tasks: optical flow estimation, disparity estimation from stereo, occlusion estimation, and semantic segmentation. Our key insight is that sharing features makes the network more compact, induces better feature representations, and can better exploit interactions among these tasks to handle partially labeled data. With a shared encoder, we can flexibly add decoders for different tasks during training. This modular design leads to a compact and efficient model at inference time. Exploiting the interactions among these tasks allows us to introduce distillation and self-supervised losses in addition to supervised losses, which can better handle partially labeled real-world data. SENSE achieves state-of-the-art results on several optical flow benchmarks and runs as fast as networks specifically designed for optical flow. It also compares favorably against the state of the art on stereo and scene flow, while consuming much less memory. Huaizu Jiang, Deqing Sun, Varun Jampani, Zhaoyang Lv, Erik G. Learned-Miller, Jan Kautz |
ICCV | 5 |
| 2019 | Concentration Inequalities for Conditional Value at RiskabstractIn this paper we derive new concentration inequalities for the conditional value at risk (CVaR) of a random variable, and compare them to the previous state of the art (Brown, 2007). We show analytically that our lower bound is strictly tighter than Brown’s, and empirically that this difference is significant. While our upper bound may be looser than Brown’s in some cases, we show empirically that in most cases our bound is significantly tighter. After discussing when each upper bound is superior, we conclude with empirical results which suggest that both of our bounds will often be significantly tighter than Brown’s. Philip S. Thomas, Erik G. Learned-Miller |
ICML | 2 |
| 2019 | Server-Side Traffic Analysis Reveals Mobile Location Information over the InternetabstractUsers can attempt to thwart third-party services from discovering their location by disabling location services on their mobile device. In this paper, we show that web services can use throughput information to reveal the path taken by the phone and its owner among a set of possibilities. For example, a TCP-based music streaming service can compile a sequence of throughputs over several minutes. We collected hundreds of traces of music that we streamed to phones in two different scenarios: a user traveling to four different towns from campus (or the reverse direction); and a user traveling within our campus. We evaluate three classifiers: k-Nearest Neighbors (k-NN), which compares a test sequence with respective time points of training sequences; a Hidden Markov Model (HMM), which computes the transition and emission probabilities of different geographic areas and chooses the most likely sequence of a test trace; and a Naive Bayes Classifier with KDE-based throughput estimates (NB-KDE), which looks at the density of throughputs at each time point along a path. In our study, the k-NN, HMM, and NB-KDE approaches can distinguish between a small number of geographic routes taken by mobile users using only throughput measurements. The NB-KDE method performed best, using throughput alone to identify the path and direction among two roads within a University campus (four classes) with 77 percent accuracy, and the path and direction among four roads (eight classes) out of town with 83 percent accuracy. Furthermore, it was able to classify among eight paths with greater than 59 percent accuracy after one minute. We examine the limitations of these techniques. Keen Sung, Joydeep Biswas, Erik G. Learned-Miller, Brian Neil Levine, Marc Liberatore |
IEEE Trans. Mob. Comput. | 3 |
| 2018 | From Neural Re-Ranking to Neural Ranking: Learning a Sparse Representation for Inverted IndexingabstractThe availability of massive data and computing power allowing for effective data driven neural approaches is having a major impact on machine learning and information retrieval research, but these models have a basic problem with efficiency. Current neural ranking models are implemented as multistage rankers: for efficiency reasons, the neural model only re-ranks the top ranked documents retrieved by a first-stage efficient ranker in response to a given query. Neural ranking models learn dense representations causing essentially every query term to match every document term, making it highly inefficient or intractable to rank the whole collection. The reliance on a first stage ranker creates a dual problem: First, the interaction and combination effects are not well understood. Second, the first stage ranker serves as a "gate-keeper" or filter, effectively blocking the potential of neural models to uncover new relevant documents. In this work, we propose a standalone neural ranking model (SNRM) by introducing a sparsity property to learn a latent sparse representation for each query and document. This representation captures the semantic relationship between the query and documents, but is also sparse enough to enable constructing an inverted index for the whole collection. We parameterize the sparsity of the model to yield a retrieval model as efficient as conventional term based models. Our model gains in efficiency without loss of effectiveness: it not only outperforms the existing term matching baselines, but also performs similarly to the recent re-ranking based neural models with dense representations. Our model can also take advantage of pseudo-relevance feedback for further improvements. More generally, our results demonstrate the importance of sparsity in neural IR models and show that dense representations can be pruned effectively, giving new insights about essential semantic features and their distributions. Hamed Zamani, Mostafa Dehghani 0001, W. Bruce Croft, Erik G. Learned-Miller, Jaap Kamps |
CIKM | 4 |
| 2018 | The Best of Both Worlds: Combining CNNs and Geometric Constraints for Hierarchical Motion SegmentationabstractTraditional methods of motion segmentation use powerful geometric constraints to understand motion, but fail to leverage the semantics of high-level image understanding. Modern CNN methods of motion analysis, on the other hand, excel at identifying well-known structures, but may not precisely characterize well-known geometric constraints. In this work, we build a new statistical model of rigid motion flow based on classical perspective projection constraints. We then combine piecewise rigid motions into complex deformable and articulated objects, guided by semantic segmentation from CNNs and a second "object-level" statistical model. This combination of classical geometric knowledge combined with the pattern recognition abilities of CNNs yields excellent performance on a wide range of motion segmentation benchmarks, from complex geometric scenes to camouflaged animals. Pia Bideau, Aruni Roy Chowdhury, Rakesh R. Menon, Erik G. Learned-Miller |
CVPR | 4 |
| 2018 | Super SloMo: High Quality Estimation of Multiple Intermediate Frames for Video InterpolationabstractGiven two consecutive frames, video interpolation aims at generating intermediate frame(s) to form both spatially and temporally coherent video sequences. While most existing methods focus on single-frame interpolation, we propose an end-to-end convolutional neural network for variable-length multi-frame video interpolation, where the motion interpretation and occlusion reasoning are jointly modeled. We start by computing bi-directional optical flow between the input images using a U-Net architecture. These flows are then linearly combined at each time step to approximate the intermediate bi-directional optical flows. These approximate flows, however, only work well in locally smooth regions and produce artifacts around motion boundaries. To address this shortcoming, we employ another U-Net to refine the approximated flow and also predict soft visibility maps. Finally, the two input images are warped and linearly fused to form each intermediate frame. By applying the visibility maps to the warped images before fusion, we exclude the contribution of occluded pixels to the interpolated intermediate frame to avoid artifacts. Since none of our learned network parameters are time-dependent, our approach is able to produce as many intermediate frames as needed. To train our network, we use 1,132 240-fps video clips, containing 300K individual video frames. Experimental results on several datasets, predicting different numbers of interpolated frames, demonstrate that our approach performs consistently better than existing methods. Huaizu Jiang, Deqing Sun, Varun Jampani, Ming-Hsuan Yang 0001, Erik G. Learned-Miller, Jan Kautz |
CVPR | 5 |
| 2018 | Self-Supervised Relative Depth Learning for Urban Scene Understanding
Huaizu Jiang, Gustav Larsson, Michael Maire, Gregory Shakhnarovich, Erik G. Learned-Miller |
ECCV (11) | 5 |
| 2018 | Unsupervised Hard Example Mining from Videos for Improved Object Detection
SouYoung Jin, Aruni Roy Chowdhury, Huaizu Jiang, Aditya Prasad, Deep Chakraborty, Erik G. Learned-Miller |
ECCV (13) | 7 |
| 2018 | A Framework for Dexterous ManipulationabstractIn this work, we introduce a framework for performing dexterous manipulations on the humanoid robot Robonaut-2. This framework memorizes how actions change perceptions and can learn a sequence of actions based on demonstrations. With the anthropomorphic Robonaut-2 hand and arm, a variety of manipulation tasks such as grasping novel objects, rotating a drill for grasping, and tightening a bolt with a ratchet can be accomplished. This framework was also used to compete in the IROS2018 Fan Robotic Challenge that requires manipulating a hand fan and was a winner of the phase I modality A competition. Li Yang Ku, Jonathan Rogers, Philip Strawser, Julia Badger, Erik G. Learned-Miller, Roderic A. Grupen |
IROS | 5 |
| 2018 | Citation Worthiness of Sentences in Scientific ReportsabstractDoes this sentence need citation? In this paper, we introduce the task of citation worthiness for scientific texts at a sentence-level granularity. The task is to detect whether a sentence in a scientific article needs to be cited or not. It can be incorporated into citation recommendation systems to help automate the citation process by marking sentences where needed. It may also be useful for publishers to regularize the citation process. We construct a dataset using the ACL Anthology Reference Corpus; consisting of over 1.1M "not_cite" and 85K "cite" sentences. We study the performance of a set of state-of-the-art sentence classifiers for the citation worthiness task and show the practical challenges. We also explore section-wise difficulty of the task and analyze the performance of our best model on a published article. Hamed R. Bonab, Hamed Zamani, Erik G. Learned-Miller, James Allan 0001 |
SIGIR | 3 |
| 2017 | Face Detection with the Faster R-CNNabstractWhile deep learning based methods for generic object detection have improved rapidly in the last two years, most approaches to face detection are still based on the R-CNN framework [11], leading to limited accuracy and processing speed. In this paper, we investigate applying the Faster RCNN [26], which has recently demonstrated impressive results on various object detection benchmarks, to face detection. By training a Faster R-CNN model on the large scale WIDER face dataset [34], we report state-of-the-art results on the WIDER test set as well as two other widely used face detection benchmarks, FDDB and the recently released IJB-A. Huaizu Jiang, Erik G. Learned-Miller |
FG | 2 |
| 2017 | End-to-End Face Detection and Cast Grouping in Movies Using Erdös-Rényi ClusteringabstractWe present an end-to-end system for detecting and clustering faces by identity in full-length movies. Unlike works that start with a predefined set of detected faces, we consider the end-to-end problem of detection and clustering together. We make three separate contributions. First, we combine a state-of-the-art face detector with a generic tracker to extract high quality face tracklets. We then introduce a novel clustering method, motivated by the classic graph theory results of Erdös and Rényi. It is based on the observations that large clusters can be fully connected by joining just a small fraction of their point pairs, while just a single connection between two different people can lead to poor clustering results. This suggests clustering using a verification system with very few false positives but perhaps moderate recall. We introduce a novel verification method, rank-1 counts verification, that has this property, and use it in a link-based clustering scheme. Finally, we define a novel end-to-end detection and clustering evaluation metric allowing us to assess the accuracy of the entire end-to-end system. We present state-of-the-art results on multiple video data sets and also on standard face databases. SouYoung Jin, Hang Su 0005, Chris Stauffer, Erik G. Learned-Miller |
ICCV | 4 |
| 2017 | An aspect representation for object manipulation based on convolutional neural networksabstractWe propose an intelligent visuomotor system that interacts with the environment and memorizes the consequences of actions. As more memories are recorded and more interactions are observed, the agent becomes more capable of predicting the consequences of actions and is, thus, better at planning sequences of actions to solve tasks. In previous work, we introduced the aspect transition graph (ATG) which represents how actions lead from one observation to another using a directed multi-graph. In this work, we propose a novel aspect representation based on hierarchical CNN features, learned with convolutional neural networks, that supports manipulation and captures the essential affordances of an object based on RGB-D images. In a traditional planning system, robots are given a pre-defined set of actions that take the robot from one symbolic state to another. However symbolic states often lack the flexibility to generalize across similar situations. Our proposed representation is grounded in the robot's observations and lies in a continuous space that allows the robot to handle similar unseen situations. The hierarchical CNN features within a representation also allow the robot to act precisely with respect to the spatial location of individual features. We evaluate the robustness of this representation using the Washington RGB-D Objects Dataset and show that it achieves state of the art results for instance pose estimation. We then test this representation in conjunction with an ATG on a drill grasping task on Robonaut-2. We show that given grasp, drag, and turn demonstrations on the drill, the robot is capable of planning sequences of learned actions to compensate for reachability constraints. Li Yang Ku, Erik G. Learned-Miller, Roderic A. Grupen |
ICRA | 2 |
| 2017 | Associating grasp configurations with hierarchical features in convolutional neural networksabstractIn this work, we provide a solution for posturing the anthropomorphic Robonaut-2 hand and arm for grasping based on visual information. A mapping from visual features extracted from a convolutional neural network (CNN) to grasp points is learned. We demonstrate that a CNN pre-trained for image classification can be applied to a grasping task based on a small set of grasping examples. Our approach takes advantage of the hierarchical nature of the CNN by identifying features that capture the hierarchical support relations between filters in different CNN layers and locating their 3D positions by tracing activations backwards in the CNN. When this backward trace terminates in the RGB-D image, important manipulable structures are thereby localized. These features that reside in different layers of the CNN are then associated with controllers that engage different kinematic subchains in the hand/arm system for grasping. A grasping dataset is collected using demonstrated hand/object relationships for Robonaut-2 to evaluate the proposed approach in terms of the precision of the resulting preshape postures. We demonstrate that this approach outperforms baseline approaches in cluttered scenarios on the grasping dataset and a point cloud based approach on a grasping task using Robonaut-2. Li Yang Ku, Erik G. Learned-Miller, Roderic A. Grupen |
IROS | 2 |
| 2017 | Active Bias: Training More Accurate Neural Networks by Emphasizing High Variance SamplesabstractSelf-paced learning and hard example mining re-weight training instances to improve learning accuracy. This paper presents two improved alternatives based on lightweight estimates of sample uncertainty in stochastic gradient descent (SGD): the variance in predicted probability of the correct class across iterations of mini-batch SGD, and the proximity of the correct class probability to the decision threshold. Extensive experimental results on six datasets show that our methods reliably improve accuracy in various network architectures, including additional gains on top of other popular training techniques, such as residual learning, momentum, ADAM, batch normalization, dropout, and distillation. Haw-Shiuan Chang, Erik G. Learned-Miller, Andrew McCallum |
NIPS | 2 |
| 2017 | Guest Editorial: Best of CVPR 2015
Kristen Grauman, Erik G. Learned-Miller, Antonio Torralba 0001, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | It's Moving! A Probabilistic Model for Causal Motion Segmentation in Moving Camera Videos
Pia Bideau, Erik G. Learned-Miller |
ECCV (8) | 2 |
| 2016 | One-to-many face recognition with bilinear CNNsabstractThe recent explosive growth in convolutional neural network (CNN) research has produced a variety of new architectures for deep learning. One intriguing new architecture is the bilinear CNN (B-CNN), which has shown dramatic performance gains on certain fine-grained recognition problems [15]. We apply this new CNN to the challenging new face recognition benchmark, the IARPA Janus Benchmark A (IJB-A) [12]. It features faces from a large number of identities in challenging real-world conditions. Because the face images were not identified automatically using a computerized face detection system, it does not have the bias inherent in such a database. We demonstrate the performance of the B-CNN model beginning from an AlexNet-style network pre-trained on ImageNet. We then show results for fine-tuning using a moderate-sized and public external database, FaceScrub [17]. We also present results with additional fine-tuning on the limited training data provided by the protocol. In each case, the fine-tuned bilinear model shows substantial improvements over the standard CNN. Finally, we demonstrate how a standard CNN pre-trained on a large face database, the recently released VGG-Face model [20], can be converted into a B-CNN without any additional feature training. This B-CNN improves upon the CNN performance on the IJB-A benchmark, achieving 89.5% rank-1 recall. Aruni Roy Chowdhury, Tsung-Yu Lin, Subhransu Maji, Erik G. Learned-Miller |
WACV | 4 |
| 2015 | Multi-view Convolutional Neural Networks for 3D Shape RecognitionabstractA longstanding question in computer vision concerns the representation of 3D shapes for recognition: should 3D shapes be represented with descriptors operating on their native 3D formats, such as voxel grid or polygon mesh, or can they be effectively represented with view-based descriptors? We address this question in the context of learning to recognize 3D shapes from a collection of their rendered views on 2D images. We first present a standard CNN architecture trained to recognize the shapes' rendered views independently of each other, and show that a 3D shape can be recognized even from a single view at an accuracy far higher than using state-of-the-art 3D shape descriptors. Recognition rates further increase when multiple views of the shapes are provided. In addition, we present a novel CNN architecture that combines information from multiple views of a 3D shape into a single and compact shape descriptor offering even better recognition performance. The same architecture can be applied to accurately recognize human hand-drawn sketches of shapes. We conclude that a collection of 2D views can be highly informative for 3D shape recognition and is amenable to emerging CNN architectures and their derivatives. Hang Su 0005, Subhransu Maji, Evangelos Kalogerakis, Erik G. Learned-Miller |
ICCV | 4 |
| 2015 | Modeling Objects as Aspect Transition Graphs to Support Manipulation
Li Yang Ku, Erik G. Learned-Miller, Roderic A. Grupen |
ISRR (2) | 2 |
| 2014 | The Shape-Time Random Field for Semantic Video LabelingabstractWe propose a novel discriminative model for semantic labeling in videos by incorporating a prior to model both the shape and temporal dependencies of an object in video. A typical approach for this task is the conditional random field (CRF), which can model local interactions among adjacent regions in a video frame. Recent work has shown how to incorporate a shape prior into a CRF for improving labeling performance, but it may be difficult to model temporal dependencies present in video by using this prior. The conditional restricted Boltzmann machine (CRBM) can model both shape and temporal dependencies, and has been used to learn walking styles from motion- capture data. In this work, we incorporate a CRBM prior into a CRF framework and present a new state-of-the-art model for the task of semantic labeling in videos. In particular, we explore the task of labeling parts of complex face scenes from videos in the YouTube Faces Database (YFDB). Our combined model outperforms competitive baselines both qualitatively and quantitatively. Andrew Kae, Benjamin M. Marlin, Erik G. Learned-Miller |
CVPR | 3 |
| 2014 | Optical Flow Estimation with Channel Constancy
Laura Sevilla-Lara, Deqing Sun, Erik G. Learned-Miller, Michael J. Black |
ECCV (1) | 3 |
| 2014 | Background subtraction: separating the modeling and the inference
Manjunath Narayana, Allen R. Hanson, Erik G. Learned-Miller |
Mach. Vis. Appl. | 3 |
| 2013 | Distribution Fields with Adaptive Kernels for Large Displacement Image AlignmentabstractWhile region-based image alignment algorithms that use gradient descent can achieve sub-pixel accuracy when they converge, their convergence depends on the smoothness of the image intensity values.Image smoothness is often enforced through the use of multiscale approaches in which images are smoothed and downsampled.Yet, these approaches typically use fixed smoothing parameters which may be appropriate for some images but not for others.Even for a particular image, the optimal smoothing parameters may depend on the magnitude of the transformation.When the transformation is large, the image should be smoothed more than when the transformation is small.Further, with gradient-based approaches, the optimal smoothing parameters may change with each iteration as the algorithm proceeds towards convergence.We address convergence issues related to the choice of smoothing parameters by deriving a Gauss-Newton gradient descent algorithm based on distribution fields (DFs) and proposing a method to dynamically select smoothing parameters at each iteration.DF and DF-like representations have previously been used in the context of tracking.In this work we incorporate DFs into a full affine model for region-based alignment and simultaneously search over parameterized sets of geometric and photometric transforms.We use a probabilistic interpretation of DFs to select smoothing parameters at each step in the optimization and show that this results in improved convergence rates. Benjamin Mears, Laura Sevilla-Lara, Erik G. Learned-Miller |
BMVC | 3 |
| 2013 | Augmenting CRFs with Boltzmann Machine Shape Priors for Image LabelingabstractConditional random fields (CRFs) provide powerful tools for building models to label image segments. They are particularly well-suited to modeling local interactions among adjacent regions (e.g., super pixels). However, CRFs are limited in dealing with complex, global (long-range) interactions between regions. Complementary to this, restricted Boltzmann machines (RBMs) can be used to model global shapes produced by segmentation models. In this work, we present a new model that uses the combined power of these two network types to build a state-of-the-art labeler. Although the CRF is a good baseline labeler, we show how an RBM can be added to the architecture to provide a global shape bias that complements the local modeling provided by the CRF. We demonstrate its labeling performance for the parts of complex face images from the Labeled Faces in the Wild data set. This hybrid model produces results that are both quantitatively and qualitatively better than the CRF alone. In addition to high-quality labeling results, we demonstrate that the hidden units in the RBM portion of our model can be interpreted as face attributes that have been learned without any attribute-level supervision. Andrew Kae, Kihyuk Sohn, Honglak Lee, Erik G. Learned-Miller |
CVPR | 4 |
| 2013 | Coherent Motion Segmentation in Moving Camera Videos Using Optical Flow OrientationsabstractIn moving camera videos, motion segmentation is commonly performed using the image plane motion of pixels, or optical flow. However, objects that are at different depths from the camera can exhibit different optical flows even if they share the same real-world motion. This can cause a depth-dependent segmentation of the scene. Our goal is to develop a segmentation algorithm that clusters pixels that have similar real-world motion irrespective of their depth in the scene. Our solution uses optical flow orientations instead of the complete vectors and exploits the well-known property that under camera translation, optical flow orientations are independent of object depth. We introduce a probabilistic model that automatically estimates the number of observed independent motions and results in a labeling that is consistent with real-world motion in the scene. The result of our system is that static objects are correctly identified as one segment, even if they are at different depths. Color features and information from previous frames in the video sequence are used to correct occasional errors due to the orientation-based segmentation. We present results on more than thirty videos from different benchmarks. The system is particularly robust on complex background scenes containing objects at significantly different depths. Manjunath Narayana, Allen R. Hanson, Erik G. Learned-Miller |
ICCV | 3 |
| 2013 | Improving Open-Vocabulary Scene Text RecognitionabstractThis paper presents a system for open-vocabulary text recognition in images of natural scenes. First, we describe a novel technique for text segmentation that models smooth color changes across images. We combine this with a recognition component based on a conditional random field with histogram of oriented gradients descriptors and incorporate language information from a lexicon to improve recognition performance. Many existing techniques for this problem use language information from a standard lexicon, but these may not include many of the words found in images of the environment, such as storefront signs and street signs. We avoid this limitation by incorporating language information from a large web-based lexicon of around 13.5 million words. This lexicon contains words encountered during a crawl of the web, so it is likely to contain proper nouns, like business names and street names. We show that our text segmentation method allows for better recognition performance than the current state-of-the-art text segmentation method. We also evaluate this full system on two standard data sets, ICDAR 2003 and ICDAR 2011, and show an increase in word recognition performance compared to the current state-of-the-art methods. Jacqueline L. Feild, Erik G. Learned-Miller |
ICDAR | 2 |
| 2013 | Using a Probabilistic Syllable Model to Improve Scene Text RecognitionabstractThis paper presents a new language model for text recognition in natural images. Many existing techniques incorporate n-gram information as an additional source of information. One problem is that some n-grams are very uncommon, but will still appear in a word across a syllable boundary. These words are given a low probability under an n-gram model. To overcome this problem, we introduce a probabilistic syllable model that uses a probabilistic context-free grammar to generate recognized word labels that are consistent with syllables. In other words, labels generated by this model are pronounceable. This is important for scene text recognition where text often includes proper nouns and standard dictionary information cannot be a useful resource. We show that this language model leads to increased recognition accuracy over a big ram model and discuss the benefits over a dictionary model. Jacqueline L. Feild, Erik G. Learned-Miller, David A. Smith |
ICDAR | 2 |
| 2013 | Scene Text Segmentation via Inverse RenderingabstractRecognizing text in natural photographs that contain specular highlights and focal blur is a challenging problem. In this paper we describe a new text segmentation method based on inverse rendering, i.e. decomposing an input image into basic rendering elements. Our technique uses iterative optimization to solve the rendering parameters, including light source, material properties (e.g. diffuse/specular reflectance and shininess) as well as blur kernel size. We combine our segmentation method with a recognition component and show that by accounting for the rendering parameters, our approach achieves higher text recognition accuracy than previous work, particularly in the presence of color changes and image blur. In addition, the derived rendering parameters can be used to synthesize new text images that imitate the appearance of an existing image. Yahan Zhou, Jacqueline L. Feild, Erik G. Learned-Miller, Rui Wang 0003 |
ICDAR | 3 |
| 2013 | Turning Off GPS Is Not Enough: Cellular Location Leaks over the Internet
Hamed Soroush, Keen Sung, Erik G. Learned-Miller, Brian Neil Levine, Marc Liberatore |
Privacy Enhancing Technologies | 3 |
| 2013 | Half-Wits: Software Techniques for Low-Voltage Probabilistic Storage on Microcontrollers with NOR Flash MemoryabstractThis work analyzes the stochastic behavior of writing to embedded flash memory at voltages lower than recommended by a microcontroller’s specifications in order to reduce energy consumption. Flash memory integrated within a microcontroller typically requires the entire chip to operate on a common supply voltage almost twice as much as what the CPU portion requires. Our software approach allows the flash memory to tolerate a lower supply voltage so that the CPU may operate in a more energy-efficient manner. Energy-efficient coding algorithms then cope with flash memory writes that behave unpredictably. Our software-only coding algorithms ( in-place writes, multiple-place writes, RS-Berger codes , and slow writes ) enable reliable storage at low voltages on unmodified hardware by exploiting the electrically cumulative nature of half-written data in write-once bits. For a sensor monitoring application using the MSP430, coding with in-place writes reduces the overall energy consumption by 34%. In-place writes are competitive when the time spent on low-voltage operations such as computation are at least four times greater than the time spent on writes to flash memory. Our evaluation shows that tightly maintaining the digital abstraction for storage in embedded flash memory comes at a significant cost to energy consumption with minimal gain in reliability. We find our techniques most effective for embedded workloads that have significant duty cycling, rare writes, or energy harvesting. Mastooreh Salajegheh, Anxiao Jiang, Erik G. Learned-Miller, Kevin Fu |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2012 | Improvements in Joint Domain-Range Modeling for Background SubtractionabstractIn many algorithms for background modeling, a distribution over feature values is modeled at each pixel. These models, however, do not account for the dependencies that may exist among nearby pixels. The joint domain-range kernel density estimate (KDE) model by Sheikh and Shah [7], which is not a pixel-wise model, represents the background and foreground processes by combining the three color dimensions and two spatial dimensions into a five-dimensional joint space. The Sheikh and Shah model, as we will show, has a peculiar dependence on the size of the image. In contrast, we build three-dimensional color distributions at each pixel and allow neighboring pixels to influence each other’s distributions. Our model is easy to interpret, does not exhibit the dependency on image size, and results in higher accuracy. Also, unlike Sheikh and Shah, we build an explicit model of the prior probability of the background and the foreground at each pixel. Finally, we use the adaptive kernel variance method of Narayana et al. [5] to adapt the KDE covariance at each pixel. With a simpler and more intuitive model, we can better interpret and visualize the effects of the adaptive kernel variance method, while achieving accuracy comparable to state-of-the-art on a standard backgrounding benchmark. 1 Manjunath Narayana, Allen R. Hanson, Erik G. Learned-Miller |
BMVC | 3 |
| 2012 | Learning hierarchical representations for face verification with convolutional deep belief networksabstractMost modern face recognition systems rely on a feature representation given by a hand-crafted image descriptor, such as Local Binary Patterns (LBP), and achieve improved performance by combining several such representations. In this paper, we propose deep learning as a natural source for obtaining additional, complementary representations. To learn features in high-resolution images, we make use of convolutional deep belief networks. Moreover, to take advantage of global structure in an object class, we develop local convolutional restricted Boltzmann machines, a novel convolutional learning model that exploits the global structure by not assuming stationarity of features across the image, while maintaining scalability and robustness to small misalignments. We also present a novel application of deep learning to descriptors other than pixel intensity values, such as LBP. In addition, we compare performance of networks trained using unsupervised learning against networks with random filters, and empirically show that learning weights not only is necessary for obtaining good multilayer representations, but also provides robustness to the choice of the network architecture parameters. Finally, we show that a recognition system using only representations obtained from deep learning can achieve comparable accuracy with a system using a combination of hand-crafted image descriptors. Moreover, by combining these representations, we achieve state-of-the-art results on a real-world face verification database. Gary B. Huang, Honglak Lee, Erik G. Learned-Miller |
CVPR | 3 |
| 2012 | Background modeling using adaptive pixelwise kernel variances in a hybrid feature spaceabstractRecent work on background subtraction has shown developments on two major fronts. In one, there has been increasing sophistication of probabilistic models, from mixtures of Gaussians at each pixel [7], to kernel density estimates at each pixel [1], and more recently to joint domainrange density estimates that incorporate spatial information [6]. Another line of work has shown the benefits of increasingly complex feature representations, including the use of texture information, local binary patterns, and recently scale-invariant local ternary patterns [4]. In this work, we use joint domain-range based estimates for background and foreground scores and show that dynamically choosing kernel variances in our kernel estimates at each individual pixel can significantly improve results. We give a heuristic method for selectively applying the adaptive kernel calculations which is nearly as accurate as the full procedure but runs much faster. We combine these modeling improvements with recently developed complex features [4] and show significant improvements on a standard backgrounding benchmark. Manjunath Narayana, Allen R. Hanson, Erik G. Learned-Miller |
CVPR | 3 |
| 2012 | Distribution fields for trackingabstractVisual tracking of general objects often relies on the assumption that gradient descent of the alignment function will reach the global optimum. A common technique to smooth the objective function is to blur the image. However, blurring the image destroys image information, which can cause the target to be lost. To address this problem we introduce a method for building an image descriptor using distribution fields (DFs), a representation that allows smoothing the objective function without destroying information about pixel values. We present experimental evidence on the superiority of the width of the basin of attraction around the global optimum of DFs over other descriptors. DFs also allow the representation of uncertainty about the tracked object. This helps in disregarding outliers during tracking (like occlusions or small misalignments) without modeling them explicitly. Finally, this provides a convenient way to aggregate the observations of the object through time and maintain an updated model. We present a simple tracking algorithm that uses DFs and obtains state-of-the-art results on standard benchmarks. Laura Sevilla-Lara, Erik G. Learned-Miller |
CVPR | 2 |
| 2012 | Learning to Align from ScratchabstractUnsupervised joint alignment of images has been demonstrated to improve performance on recognition tasks such as face verification. Such alignment reduces undesired variability due to factors such as pose, while only requiring weak supervision in the form of poorly aligned examples. However, prior work on unsupervised alignment of complex, real world images has required the careful selection of feature representation based on hand-crafted image descriptors, in order to achieve an appropriate, smooth optimization landscape. In this paper, we instead propose a novel combination of unsupervised joint alignment with unsupervised feature learning. Specifically, we incorporate deep learning into the {\em congealing} alignment framework. Through deep learning, we obtain features that can represent the image at differing resolutions based on network depth, and that are tuned to the statistics of the specific data being aligned. In addition, we modify the learning algorithm for the restricted Boltzmann machine by incorporating a group sparsity penalty, leading to a topographic organization on the learned filters and improving subsequent alignment results. We apply our method to the Labeled Faces in the Wild database (LFW). Using the aligned images produced by our proposed unsupervised algorithm, we achieve a significantly higher accuracy in face verification than obtained using the original face images, prior work in unsupervised alignment, and prior work in supervised alignment. We also match the accuracy for the best available, but unpublished method. Gary B. Huang, Marwan A. Mattar, Honglak Lee, Erik G. Learned-Miller |
NIPS | 4 |
| 2012 | Unsupervised Joint Alignment and Clustering using Bayesian Nonparametrics
Marwan A. Mattar, Allen R. Hanson, Erik G. Learned-Miller |
UAI | 3 |
| 2012 | Bounding the Probability of Error for High Precision Optical Character Recognition
Gary B. Huang, Andrew Kae, Carl Doersch, Erik G. Learned-Miller |
J. Mach. Learn. Res. | 4 |
| 2011 | Online domain adaptation of a pre-trained cascade of classifiersabstractMany classifiers are trained with massive training sets only to be applied at test time on data from a different distribution. How can we rapidly and simply adapt a classifier to a new test distribution, even when we do not have access to the original training data? We present an on-line approach for rapidly adapting a “black box” classifier to a new test data set without retraining the classifier or examining the original optimization criterion. Assuming the original classifier outputs a continuous number for which a threshold gives the class, we reclassify points near the original boundary using a Gaussian process regression scheme. We show how this general procedure can be used in the context of a classifier cascade, demonstrating performance that far exceeds state-of-the-art results in face detection on a standard data set. We also draw connections to work in semi-supervised learning, domain adaptation, and information regularization. Vidit Jain, Erik G. Learned-Miller |
CVPR | 2 |
| 2011 | Enforcing similarity constraints with integer programming for better scene text recognitionabstractThe recognition of text in everyday scenes is made difficult by viewing conditions, unusual fonts, and lack of linguistic context. Most methods integrate a priori appearance information and some sort of hard or soft constraint on the allowable strings. Weinman and Learned-Miller [14] showed that the similarity among characters, as a supplement to the appearance of the characters with respect to a model, could be used to improve scene text recognition. In this work, we make further improvements to scene text recognition by taking a novel approach to the incorporation of similarity. In particular, we train a similarity expert that learns to classify each pair of characters as equivalent or not. After removing logical inconsistencies in an equivalence graph, we formulate the search for the maximum likelihood interpretation of a sign as an integer program. We incorporate the equivalence information as constraints in the integer program and build an optimization criterion out of appearance features and character bigrams. Finally, we take the optimal solution from the integer program, and compare all “nearby” solutions using a probability model for strings derived from search engine queries. We demonstrate word error reductions of more than 30% relative to previous methods on the same data set. David L. Smith, Jacqueline L. Feild, Erik G. Learned-Miller |
CVPR | 3 |
| 2011 | Exploiting Half-Wits: Smarter Storage for Low-Power Devices
Mastooreh Salajegheh, Kevin Fu, Anxiao Jiang, Erik G. Learned-Miller |
FAST | 5 |
| 2011 | Forensic Triage for Mobile Phones with DEC0DE
Robert J. Walls, Erik G. Learned-Miller, Brian Neil Levine |
USENIX Security Symposium | 2 |
| 2011 | Learning on the fly: a font-free approach toward multilingual OCR
Andrew Kae, David A. Smith, Erik G. Learned-Miller |
Int. J. Document Anal. Recognit. | 3 |
| 2011 | Introduction to the Special Section on Real-World Face RecognitionabstractThe motivations for organizing this special section were to better address the challenges of face recognition in real-world scenarios, to promote systematic research and evaluation of promising methods and systems, to provide a snapshot of where we are in this domain, and to stimulate discussion about future directions. We solicited original contributions of research on all aspects of real-world face recognition, including: the design of robust face similarity features and metrics; robust face clustering and sorting algorithms; novel user interaction models and face recognition algorithms for face tagging; novel applications of web face recognition; novel computational paradigms for face recognition; challenges in large scale face recognition tasks, e.g., on the Internet; face recognition with contextual information; face recognition benchmarks and evaluation methodology for moderately controlled or uncontrolled environments; and video face recognition. We received 42 original submissions, four of which were rejected without review; the other 38 papers entered the normal review process. Each paper was reviewed by three reviewers who are experts in their respective topics. More than 100 expert reviewers have been involved in the review process. The papers were equally distributed among the guest editors. A final decision for each paper was made by at least two guest editors assigned to it. To avoid conflict of interest, no guest editor submitted any papers to this special section. Gang Hua 0001, Ming-Hsuan Yang 0001, Erik G. Learned-Miller, Yi Ma 0001, Matthew Turk 0001, David J. Kriegman, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | Improving state-of-the-art OCR through high-precision document-specific modelingabstractOptical character recognition (OCR) remains a difficult problem for noisy documents or documents not scanned at high resolution. Many current approaches rely on stored font models that are vulnerable to cases in which the document is noisy or is written in a font dissimilar to the stored fonts. We address these problems by learning character models directly from the document itself, rather than using pre-stored font models. This method has had some success in the past, but we are able to achieve substantial improvement in error reduction through a novel method for creating nearly error-free document-specific training data and building character appearance models from this data. In particular, we first use the state-of-the-art OCR system Tesseract to produce an initial translation. Then, our method identifies a subset of words that we have high confidence have been recognized correctly and uses this subset to bootstrap document-specific character models. We present theoretical justification that a word in the selected subset is very unlikely to be incorrectly recognized, and empirical results on a data set of difficult historical newspaper scans demonstrating that we make only two errors in 56 documents. We then relax the theoretical constraint in order to create a larger training set, and using document-specific character models generated from this data, we are able to reduce the error over properly segmented characters by 34.1% overall from the initial Tesseract translation. Andrew Kae, Gary B. Huang, Carl Doersch, Erik G. Learned-Miller |
CVPR | 4 |
| 2009 | Nonparametric curve alignmentabstractCongealing is a flexible nonparametric data-driven framework for the joint alignment of data. It has been successfully applied to the joint alignment of binary images of digits, binary images of object silhouettes, grayscale MRI images, color images of cars and faces, and 3D brain volumes. This research enhances congealing to practically and effectively apply it to curve data. We develop a parameterized set of nonlinear transformations that allow us to apply congealing to this type of data. We present positive results on aligning synthetic and real curve data sets and conclude with a discussion on extending this work to simultaneous alignment and clustering. Marwan A. Mattar, Michael G. Ross, Erik G. Learned-Miller |
ICASSP | 3 |
| 2009 | Learning on the Fly: Font-Free Approaches to Difficult OCR ProblemsabstractDespite ubiquitous claims that optical character recognition (OCR) is a "solved problem,'' many categories of documents continue to break modern OCR software such as documents with moderate degradation or unusual fonts. Many approaches rely on pre-computed or stored character models, but these are vulnerable to cases when the font of a particular document was not part of the training set, or when there is so much noise in a document that the font model becomes weak. To address these difficult cases, we present a form of iterative contextual modeling that learns character models directly from the document it is trying to recognize. We use these learned models both to segment the characters and to recognize them in an incremental, iterative process. We present results comparable to those of a commercial OCR system on a subset of characters from a difficult test document. Andrew Kae, Erik G. Learned-Miller |
ICDAR | 2 |
| 2009 | Scene Text Recognition Using Similarity and a Lexicon with Sparse Belief PropagationabstractScene text recognition (STR) is the recognition of text anywhere in the environment, such as signs and storefronts. Relative to document recognition, it is challenging because of font variability, minimal language context, and uncontrolled conditions. Much information available to solve this problem is frequently ignored or used sequentially. Similarity between character images is often overlooked as useful information. Because of language priors, a recognizer may assign different labels to identical characters. Directly comparing characters to each other, rather than only a model, helps ensure that similar instances receive the same label. Lexicons improve recognition accuracy but are used post hoc. We introduce a probabilistic model for STR that integrates similarity, language properties, and lexical decision. Inference is accelerated with sparse belief propagation, a bottom-up method for shortening messages by reducing the dependency between weakly supported hypotheses. By fusing information sources in one model, we eliminate unrecoverable errors that result from sequential processing, improving accuracy. In experimental results recognizing text from images of signs in outdoor scenes, incorporating similarity reduces character recognition error by 19 percent, the lexicon reduces word recognition error by 35 percent, and sparse belief propagation reduces the lexicon words considered by 99.9 percent with a 12X speedup and no loss in accuracy. Jerod J. Weinman, Erik G. Learned-Miller, Allen R. Hanson |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | A discriminative semi-Markov model for robust scene text recognitionabstractWe present a semi-Markov model for recognizing scene text that integrates character and word segmentation with recognition. Using wavelet features, it requires only approximate location of the text baseline and font size; no binarization or prior word segmentation is necessary. Our system is aided by a lexicon, yet it also allows non-lexicon words. To facilitate inference with a large lexicon, we use an approximate Viterbi beam search. Our system performs robustly on low-resolution images of signs containing text in fonts atypical of documents. Jerod J. Weinman, Erik G. Learned-Miller, Allen R. Hanson |
ICPR | 2 |
| 2008 | Learning to Locate Informative Features for Visual Identification
Andras Ferencz, Erik G. Learned-Miller, Jitendra Malik |
Int. J. Comput. Vis. | 2 |
| 2008 | A Probabilistic Upper Bound on Differential EntropyabstractA novel probabilistic upper bound on the entropy of an unknown one-dimensional distribution, given the support of the distribution and a sample from that distribution, is presented. No knowledge beyond the support of the unknown distribution is required. Previous distribution-free bounds on the cumulative distribution function of a random variable given a sample of that variable are used to construct the bound. A simple, fast, and intuitive algorithm for computing the entropy bound from a sample is provided. Erik G. Learned-Miller, Joseph DeStefano |
IEEE Trans. Inf. Theory | 1 |
| 2007 | Unsupervised Joint Alignment of Complex ImagesabstractMany recognition algorithms depend on careful positioning of an object into a canonical pose, so the position of features relative to a fixed coordinate system can be examined. Currently, this positioning is done either manually or by training a class-specialized learning algorithm with samples of the class that have been hand-labeled with parts or poses. In this paper, we describe a novel method to achieve this positioning using poorly aligned examples of a class with no additional labeling. Given a set of unaligned examplars of a class, such as faces, we automatically build an alignment mechanism, without any additional labeling of parts or poses in the data set. Using this alignment mechanism, new members of the class, such as faces resulting from a face detector, can be precisely aligned for the recognition process. Our alignment method improves performance on a face recognition task, both over unaligned images and over images aligned with a face alignment algorithm specifically developed for and trained on hand-labeled face images. We also demonstrate its use on an entirely different class of objects (cars), again without providing any information about parts or pose to the learning algorithm. Gary B. Huang, Vidit Jain, Erik G. Learned-Miller |
ICCV | 3 |
| 2007 | People-LDA: Anchoring Topics to People using Face RecognitionabstractTopic models have recently emerged as powerful tools for modeling topical trends in documents. Often the resulting topics are broad and generic, associating large groups of people and issues that are loosely related. In many cases, it may be desirable to influence the direction in which topic models develop. In this paper, we explore the idea of centering topics around people. In particular, given a large corpus of images featuring collections of people and associated captions, it seems natural to extract topics specifically focussed on each person. What words are most associated with George Bush? Which with Condoleezza Rice? Since people play such an important role in life, it is natural to anchor one topic to each person. In this paper, we present People-LDA, which uses the coherence efface images in news captions to guide the development of topics. In particular, we show how topics can be refined to be more closely related to a single person (like George Bush) rather than describing groups of people in a related area (like politics). To do this we introduce a new graphical model that tightly couples images and captions through a modern face recognizer. In addition to producing topics that are people specific (using images as a guiding force), the model also performs excellent soft clustering efface images, using the language model to boost performance. We present a variety of experiments comparing our method to recent developments in topic modeling and joint image-language modeling, showing that our model has lower perplexity for face identification than competing models and produces more refined topics. Vidit Jain, Erik G. Learned-Miller, Andrew McCallum |
ICCV | 2 |
| 2007 | Cryptogram Decoding for OCR Using Numerization StringsabstractOCR systems for printed documents typically require large numbers of font styles and character models to work well. When given an unseen font, performance degrades even in the absence of noise. In this paper, we perform OCR in an unsupervised fashion without using any character models by using a cryptogram decoding algorithm. We present results on real and artificial OCR data. Gary B. Huang, Erik G. Learned-Miller, Andrew McCallum |
ICDAR | 2 |
| 2007 | Fast Lexicon-Based Scene Text Recognition with Sparse Belief PropagationabstractUsing a lexicon can often improve character recognition under challenging conditions, such as poor image quality or unusual fonts. We propose a flexible probabilistic model for character recognition that integrates local language properties, such as bigrams, with lexical decision, having open and closed vocabulary modes that operate simultaneously. Lexical processing is accelerated by performing inference with sparse belief propagation, a bottom-up method for hypothesis pruning. We give experimental results on recognizing text from images of signs in outdoor scenes. Incorporating the lexicon reduces word recognition error by 42% and sparse belief propagation reduces the number of lexicon words considered by 97%. Jerod J. Weinman, Erik G. Learned-Miller, Allen R. Hanson |
ICDAR | 2 |
| 2007 | Context-Sensitive Error Correction: Using Topic Models to Improve OCRabstractModern optical, character recognition software relies on human interaction to correct mis recognized characters. Even though the software often reliably identifies low-confidence output, the simple language and vocabulary models employed are insufficient to automatically correct mistakes. This paper demonstrates that topic models, which automatically detect and represent an article's semantic context, reduces error by 7% over a global word distribution in a simulated OCR correction task. Detecting and leveraging context in this manner is an important step towards improving OCR. Michael L. Wick, Michael G. Ross, Erik G. Learned-Miller |
ICDAR | 3 |
| 2007 | Techniques and Applications for Persistent Backgrounding in a Humanoid Torso RobotabstractOne of the most basic capabilities for an agent with a vision system is to recognize its own surroundings. Yet surprisingly, despite the ease of doing so, many robots store little or no record of their own visual surroundings. This paper explores the utility of keeping the simplest possible persistent record of the environment of a stationary torso robot, in the form of a collection of images captured from various pan-tilt angles around the robot. We demonstrate that this particularly simple process of storing background images can be useful for a variety of tasks, and can relieve the system designer of certain requirements as well. We explore three uses for such a record: auto-calibration, novel object detection with a moving camera, and developing attentional saliency maps. David Walker Duhon, Jerod J. Weinman, Erik G. Learned-Miller |
ICRA | 3 |
| 2007 | Analyzing in situ gene expression in the mouse brain with image registration, feature extraction and block clusteringabstractBACKGROUND: Many important high throughput projects use in situ hybridization and may require the analysis of images of spatial cross sections of organisms taken with cellular level resolution. Projects creating gene expression atlases at unprecedented scales for the embryonic fruit fly as well as the embryonic and adult mouse already involve the analysis of hundreds of thousands of high resolution experimental images mapping mRNA expression patterns. Challenges include accurate registration of highly deformed tissues, associating cells with known anatomical regions, and identifying groups of genes whose expression is coordinately regulated with respect to both concentration and spatial location. Solutions to these and other challenges will lead to a richer understanding of the complex system aspects of gene regulation in heterogeneous tissue. RESULTS: We present an end-to-end approach for processing raw in situ expression imagery and performing subsequent analysis. We use a non-linear, information theoretic based image registration technique specifically adapted for mapping expression images to anatomical annotations and a method for extracting expression information within an anatomical region. Our method consists of coarse registration, fine registration, and expression feature extraction steps. From this we obtain a matrix for expression characteristics with rows corresponding to genes and columns corresponding to anatomical sub-structures. We perform matrix block cluster analysis using a novel row-column mixture model and we relate clustered patterns to Gene Ontology (GO) annotations. CONCLUSION: Resulting registrations suggest that our method is robust over intensity levels and shape variations in ISH imagery. Functional enrichment studies from both simple analysis and block clustering indicate that gene relationships consistent with biological knowledge of neuronal gene functions can be extracted from large ISH image databases such as the Allen Brain Atlas 1 and the Max-Planck Institute 2 using our method. While we focus here on imagery and experiments of the mouse brain our approach should be applicable to a variety of in situ experiments. Manjunatha Jagalur, Christopher Joseph Pal, Erik G. Learned-Miller, R. Thomas Zoeller, David Kulp |
BMC Bioinform. | 3 |
| 2006 | Discriminative Training of Hyper-feature Models for Object IdentificationabstractObject identification is the task of identifying specific objects belonging to the same class such as cars. We often need to recognize an object that we have only seen a few times. In fact, we often observe only one example of a particular object before we need to recognize it again. Thus we are interested in building a system which can learn to extract distinctive markers from a single example and which can then be used to identify the object in another image as “same ” or “different”. Previous work by Ferencz et al. introduced the notion of hyper-features, which are properties of an image patch that can be used to estimate the utility of the patch in subsequent matching tasks. In this work, we show that hyper-feature based models can be more efficiently estimated using discriminative training techniques. In particular, we describe a new hyper-feature model based upon logistic regression that shows improved performance over previously published techniques. Our approach significantly outperforms Bayesian face recognition that is considered as a standard benchmark for face recognition. 1 Vidit Jain, Andras Ferencz, Erik G. Learned-Miller |
BMVC | 3 |
| 2006 | Improving Recognition of Novel Input with SimilarityabstractMany sources of information relevant to computer vision and machine learning tasks are often underused. One example is the similarity between the elements from a novel source, such as a speaker, writer, or printed font. By comparing instances emitted by a source, we help ensure that similar instances are given the same label. Previous approaches have clustered instances prior to recognition. We propose a probabilistic framework that unifies similarity with prior identity and contextual information. By fusing information sources in a single model, we eliminate unrecoverable errors that result from processing the information in separate stages and improve overall accuracy. The framework also naturally integrates dissimilarity information, which has previously been ignored. We demonstrate with an application in printed character recognition from images of signs in natural scenes. Jerod J. Weinman, Erik G. Learned-Miller |
CVPR (1) | 2 |
| 2006 | Combinatorial Markov Random Fields
Ron Bekkerman, Mehran Sahami, Erik G. Learned-Miller |
ECML | 3 |
| 2006 | Detecting Acromegaly: Screening for Disease with a Morphable Model
Erik G. Learned-Miller, Qifeng Lu, Angela Paisley, Peter Trainer, Volker Blanz, Katrin Dedden, Ralph Miller |
MICCAI (2) | 1 |
| 2006 | Data Driven Image Models through Continuous Joint AlignmentabstractThis paper presents a family of techniques that we call congealing for modeling image classes from data. The idea is to start with a set of images and make them appear as similar as possible by removing variability along the known axes of variation. This technique can be used to eliminate "nuisance" variables such as affine deformations from handwritten digits or unwanted bias fields from magnetic resonance images. In addition to separating and modeling the latent images-i.e., the images without the nuisance variables-we can model the nuisance variables themselves, leading to factorized generative image models. When nuisance variable distributions are shared between classes, one can share the knowledge learned in one task with another task, leading to efficient learning. We demonstrate this process by building a handwritten digit classifier from just a single example of each class. In addition to applications in handwritten character recognition, we describe in detail the application of bias removal from magnetic resonance images. Unlike previous methods, we use a separate, nonparametric model for the intensity values at each pixel. This allows us to leverage the data from the MR images of different patients to remove bias from each other. Only very weak assumptions are made about the distributions of intensity values in the images. In addition to the digit and MR applications, we discuss a number of other uses of congealing and describe experiments about the robustness and consistency of the method. Erik G. Learned-Miller |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Building a Classification Cascade for Visual Identification from One ExampleabstractObject identification (OID) is specialized recognition where the category is known (e.g. cars) and the algorithm recognizes an object's exact identity (e.g. Bob's BMW). Two special challenges characterize OID. (1) Interclass variation is often small (many cars look alike) and may be dwarfed by illumination or pose changes. (2) There may be many classes but few or just one positive "training" examples per class. Due to (1), a solution must locate possibly subtle object-specific salient features (a door handle) while avoiding distracting ones (a specular highlight). However, (2) rules out direct techniques of feature selection. We describe an online algorithm that takes one model image from a known category and builds an efficient "same" vs. "different" classification cascade by predicting the most discriminative feature set for that object. Our method not only estimates the saliency and scoring function for each candidate feature, but also models the dependency between features, building an ordered feature sequence unique to a specific model image, maximizing cumulative information content. Learned stopping thresholds make the classifier very efficient. To make this possible, category-specific characteristics are learned automatically in an off-line training procedure from labeled image pairs of the category, without prior knowledge about the category. Our method, using the same algorithm for both cars and faces, outperforms a wide variety of other methods. Andras Ferencz, Erik G. Learned-Miller, Jitendra Malik |
ICCV | 2 |
| 2004 | Names and Faces in the News
Tamara L. Berg, Alexander C. Berg, Jaety Edwards, Michael Maire, Ryan White, Yee Whye Teh, Erik G. Learned-Miller, David A. Forsyth |
CVPR (2) | 7 |
| 2004 | Learning Hyper-Features for Visual IdentificationabstractWe address the problem of identifying specific instances of a class (cars) from a set of images all belonging to that class. Although we cannot build a model for any particular instance (as we may be provided with only one "training" example of it), we can use information extracted from observ- ing other members of the class. We pose this task as a learning problem, in which the learner is given image pairs, labeled as matching or not, and must discover which image features are most consistent for matching in- stances and discriminative for mismatches. We explore a patch based representation, where we model the distributions of similarity measure- ments defined on the patches. Finally, we describe an algorithm that selects the most salient patches based on a mutual information criterion. This algorithm performs identification well for our challenging dataset of car images, after matching only a few, well chosen patches. Andras Ferencz, Erik G. Learned-Miller, Jitendra Malik |
NIPS | 2 |
| 2004 | Joint MRI Bias Removal Using Entropy Minimization Across ImagesabstractThe correction of bias in magnetic resonance images is an important problem in medical image processing. Most previous approaches have used a maximum likelihood method to increase the likelihood of the pix- els in a single image by adaptively estimating a correction to the unknown image bias field. The pixel likelihoods are defined either in terms of a pre-existing tissue model, or non-parametrically in terms of the image's own pixel values. In both cases, the specific location of a pixel in the im- age is not used to calculate the likelihoods. We suggest a new approach in which we simultaneously eliminate the bias from a set of images of the same anatomy, but from different patients. We use the statistics from the same location across different images, rather than within an image, to eliminate bias fields from all of the images simultaneously. The method builds a "multi-resolution" non-parametric tissue model conditioned on image location while eliminating the bias fields associated with the orig- inal image set. We present experiments on both synthetic and real MR data sets, and present comparisons with other methods. Erik G. Learned-Miller, Parvez Ahammad |
NIPS | 1 |
| 2003 | Practical Non-parametric Density Estimation on a Transformation Group for VisionabstractIt is now common practice in machine vision to define the variability in an object's appearance in a factored manner, as a combination of shape and texture transformations. In this context, we present a simple and practical method for estimating non-parametric probability densities over a group of linear shape deformations. Samples drawn from such a distribution do not lie in a Euclidean space, and standard kernel density estimates may perform poorly. While variable kernel estimators may mitigate this problem to some extent, the geometry of the underlying configuration space ultimately demands a kernel, which accommodates its group structure. In this perspective, we propose a suitable invariant estimator on the linear group of non-singular matrices with positive determinant. We illustrate this approach by modeling image transformations in digit recognition problems, and present results showing the superiority of our estimator to comparable Euclidean estimators in this domain. Erik G. Learned-Miller, Christophe Chefd'Hotel |
CVPR (2) | 1 |
| 2003 | A new class of entropy estimators for multi-dimensional densitiesabstractWe present a new class of estimators for approximating the entropy of multi-dimensional probability densities based on a sample of the density. These estimators extend the classic "m-spacing" estimators of Vasicek (1976) and others for estimating entropies of one-dimensional probability densities. Unlike plug-in estimators of entropy, which first estimate a probability density and then compute its entropy. our estimators avoid the difficult intermediate step of density estimation. For fixed dimension. the estimators an polynomial in the sample size. Similarities to consistent and asymptotically efficient one-dimensional estimators of entropy suggest that our estimators may sham these properties. Erik G. Learned-Miller |
ICASSP (3) | 1 |
| 2003 | ICA Using Spacings Estimates of Entropy
Erik G. Learned-Miller, John W. Fisher III |
J. Mach. Learn. Res. | 1 |
| 2002 | Unsupervised Color ConstancyabstractIn [1] we introduced a linear statistical model of joint color changes in images due to variation in lighting and certain non-geometric camera pa- rameters. We did this by measuring the mappings of colors in one image of a scene to colors in another image of the same scene under different lighting conditions. Here we increase the flexibility of this color flow model by allowing flow coefficients to vary according to a low order polynomial over the image. This allows us to better fit smoothly vary- ing lighting conditions as well as curved surfaces without endowing our model with too much capacity. We show results on image matching and shadow removal and detection. Kinh Tieu, Erik G. Learned-Miller |
NIPS | 2 |
| 2001 | Color Eigenflows: Statistical Modeling of Joint Color Changes
Erik G. Learned-Miller, Kinh Tieu |
ICCV | 1 |
| 2001 | A Binary Entropy Measure to Assess Nonrigid Registration Algorithms
Simon K. Warfield, Jan Rexilius, Petra S. Huppi, Terrie E. Inder, Erik G. Learned-Miller, William M. Wells III, Gary P. Zientara, Ferenc A. Jolesz, Ron Kikinis |
MICCAI | 5 |
| 2001 | Transform-invariant Image Decomposition with Similarity TemplatesabstractRecent work has shown impressive transform-invariant modeling and clustering for sets of images of objects with similar appearance. We seek to expand these capabilities to sets of images of an object class that show considerable variation across individual instances (e.g. pedestrian images) using a representation based on pixel-wise similarities, similarity templates. Because of its invariance to the colors of particular components of an object, this representation en- ables detection of instances of an object class and enables alignment of those instances. Further, this model implicitly represents the re- gions of color regularity in the class-speci(cid:12)c image set enabling a decomposition of that object class into component regions. Chris Stauffer, Erik G. Learned-Miller, Kinh Tieu |
NIPS | 2 |
| 2000 | Learning from One Example through Shared Densities on TransformsabstractWe define a process called congealing in which elements of a dataset (images) are brought into correspondence with each other jointly, producing a data-defined model. It is based upon minimizing the summed component-wise (pixel-wise) entropies over a continuous set of transforms on the data. One of the biproducts of this minimization is a set of transform, one associated with each original training sample. We then demonstrate a procedure for effectively bringing test data into correspondence with the data-defined model produced in the congealing process. Subsequently; we develop a probability density over the set of transforms that arose from the congealing process. We suggest that this density over transforms may be shared by many classes, and demonstrate how using this density as "prior knowledge" can be used to develop a classifier based on only a single training example for each class. Erik G. Learned-Miller, Nicholas E. Matsakis, Paul A. Viola |
CVPR | 1 |
| 1999 | Alternative Tilings for Improved Surface Area Estimates by Local Counting Algorithms
Erik G. Learned-Miller |
Comput. Vis. Image Underst. | 1 |