Robinson Piramuthu

dblp:29/1333 · DBLP profile ↗
← Back
35ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-1767-8382ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 19 · 8 since 2021Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 3 · 3 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 START: Spatial and Textual Learning for Chart Understanding
Zhuoming Liu 0001, Xiaofeng Gao 0002, Feiyang Niu, Qiaozi Gao, Liu Liu 0018, Robinson Piramuthu
WACV6
2025 T2V-Turbo-v2: Enhancing Video Model Post-Training through Data, Reward, and Conditional Guidance Design
abstract
In this paper, we focus on enhancing a diffusion-based text-to-video (T2V) model during the post-training phase by distilling a highly capable consistency model from a pretrained T2V model. Our proposed method, T2V-Turbo-v2, introduces a significant advancement by integrating various supervision signals, including high-quality training data, reward model feedback, and conditional guidance, into the consistency distillation process. Through comprehensive ablation studies, we highlight the crucial importance of tailoring datasets to specific learning objectives and the effectiveness of learning from diverse reward models for enhancing both the visual quality and text-video alignment. Additionally, we highlight the vast design space of conditional guidance strategies, which centers on designing an effective energy function to augment the teacher ODE solver. We demonstrate the potential of this approach by extracting motion guidance from the training datasets and incorporating it into the ODE solver, showcasing its effectiveness in improving the motion quality of the generated videos with the improved motion-related metrics from VBench and T2V-CompBench. Empirically, our T2V-Turbo-v2 establishes a new state-of-the-art result on VBench, **with a Total score of 85.13**, surpassing proprietary systems such as Gen-3 and Kling.
Qian Long, Xiaofeng Gao 0002, Robinson Piramuthu, Wenhu Chen, William Yang Wang
ICLR5
2025 Is the House Ready For Sleeptime? Generating and Evaluating Situational Queries for Embodied Question Answering
abstract
We present and tackle the problem of Embodied Question Answering (EQA) with Situational Queries (S-EQA) in a household environment. Unlike prior EQA work tackling simple queries that directly reference target objects and properties ("What is the color of the car?"), situational queries (such as "Is the house ready for sleeptime?") are challenging as they require the agent to correctly identify multiple object-states (Doors: Closed, Lights: Off, etc.) and reach a consensus on their states for an answer. Towards this objective, we first introduce a novel Prompt-Generate-Evaluate (PGE) scheme that wraps around an LLM’s output to generate unique situational queries and corresponding consensus object information. PGE is used to generate 2K datapoints in the VirtualHome simulator, which is then annotated for ground truth answers via a large scale user-study conducted on M-Turk. With a high rate of answerability (97.26%) on this study, we establish that LLMs are good at generating situational data. However, in evaluating the data using an LLM, we observe a low correlation of 46.2% with the ground truth human annotations; indicating that while LLMs are good at generating situational data, they struggle to answer them according to consensus. When asked for reasoning, we observe the LLM often goes against commonsense in justifying its answer. Finally, we utilize PGE to generate situational data in a real-world environment, exposing LLM hallucination in generating reliable object-states when a structured scene graph is unavailable. To the best of our knowledge, this is the first work to introduce EQA in the context of situational queries and also the first to present a generative approach for query creation. We aim to foster research on improving the real-world usability of embodied agents through this work.
Vishnu Sashank Dorbala, Prasoon Goyal, Robinson Piramuthu, Michael Johnston, Reza Ghanadan, Dinesh Manocha
IROS3
2024 Decision Making for Human-in-the-loop Robotic Agents via Uncertainty-Aware Reinforcement Learning
abstract
In a Human-in-the-Loop paradigm, a robotic agent is able to act mostly autonomously in solving a task, but can request help from an external expert when needed. However, knowing when to request such assistance is critical: too few requests can lead to the robot making mistakes, but too many requests can overload the expert. In this paper, we present a Reinforcement Learning based approach to this problem, where a semi-autonomous agent asks for external assistance when it has low confidence in the eventual success of the task. The confidence level is computed by estimating the variance of the return from the current state. We show that this estimate can be iteratively improved during training using a Bellman-like recursion. On discrete navigation problems with both fully-and partially-observable state information, we show that our method makes effective use of a limited budget of expert calls at run-time, despite having no access to the expert at training time.
Siddharth Singi, Zhanpeng He, Alvin Pan, Sandip Patel, Gunnar A. Sigurdsson, Robinson Piramuthu, Shuran Song, Matei T. Ciocarlie
ICRA6
2024 Zero-Shot Controllable Image-to-Video Animation via Motion Decomposition
abstract
In this paper, we introduce a new challenging task called Zero-Shot Controllable Image-to-Video Animation, where the goal is to animate an image based on motion trajectories defined by the user, without fine-tuning the base model. Primary challenges include maintaining consistency of background, consistency of object in motion, faithfulness to the user-defined trajectory, and quality of motion animation. We also introduce a novel approach for this task, leveraging diffusion models called Img2VidAnim-Zero (IVA0). IVA0 tackles our controllable Image-to-Video (I2V) task by decomposing it into two subtasks: 'out-of-place' and 'in-place' motion animation. Due to this decomposition, IVA0 can leverage existing work on layout-conditioned image generation for out-of-place motion generation, and existing text-conditioned video generation methods for in-place motion animation, thus facilitating zero-shot generation. Our model also addresses key challenges for controllable animation, such as Layout Conditioning via Spatio-Temporal Masking to incorporate user guidance and Motion Afterimage Suppression (MAS) scheme to reduce object ghosting during out-of-place animation. Finally, we design a novel controllable I2V benchmark featuring diverse local- and global-level metrics. Results show IVA0 as a new state-of-the-art, establishing a new standard for the zero-shot controllable I2V task. Our method highlights the simplicity and effectiveness of task decomposition and modularization for this novel task for future studies. Our code and visualizations are available at https://img2vidanim-0.github.io/
Shoubin Yu, Jacob Zhiyuan Fang, Gunnar A. Sigurdsson, Vicente Ordonez, Robinson Piramuthu, Mohit Bansal
ACM Multimedia6
2023 A Simple Approach for Visual Room Rearrangement: 3D Mapping and Semantic Search
Brandon Trabucco, Gunnar A. Sigurdsson, Robinson Piramuthu, Gaurav S. Sukhatme, Ruslan Salakhutdinov
ICLR3
2023 RREx-BoT: Remote Referring Expressions with a Bag of Tricks
abstract
Household robots operate in the same space for years. Such robots incrementally build dynamic maps that can be used for tasks requiring remote object localization. However, benchmarks in robot learning often test generalization through inference on tasks in unobserved environments. In an observed environment, locating an object is reduced to choosing from among all object proposals in the environment, which may number in the 100,000s. Armed with this intuition, using only a generic vision-language scoring model with minor modifications for 3d encoding and operating in an embodied environment, we demonstrate an absolute performance gain of 9.84% on remote object grounding above state of the art models for REVERIE and of 5.04% on FAO. When allowed to pre-explore an environment, we also exceed the previous state of the art pre-exploration method on REVERIE. Additionally, we demonstrate our model on a real-world TurtleBot platform, highlighting the simplicity and usefulness of the approach. Our analysis outlines a “bag of tricks” essential for accomplishing this task, from utilizing 3d coordinates and context, to gener-alizing vision-language models to large 3d search spaces.
Gunnar A. Sigurdsson, Jesse Thomason, Gaurav S. Sukhatme, Robinson Piramuthu
IROS4
2022 TEACh: Task-Driven Embodied Agents That Chat
abstract
Robots operating in human spaces must be able to engage in natural language interaction, both understanding and executing instructions, and using conversation to resolve ambiguity and correct mistakes. To study this, we introduce TEACh, a dataset of over 3,000 human-human, interactive dialogues to complete household tasks in simulation. A Commander with access to oracle information about a task communicates in natural language with a Follower. The Follower navigates through and interacts with the environment to complete tasks varying in complexity from "Make Coffee" to "Prepare Breakfast", asking questions and getting additional information from the Commander. We propose three benchmarks using TEACh to study embodied intelligence challenges, and we evaluate initial models' abilities in dialogue understanding, language grounding, and task execution.
Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gökhan Tür, Dilek Hakkani-Tür
AAAI7
2022 SSL Enables Learning from Sparse Rewards in Image-Goal Navigation
Arjun Majumdar, Gunnar A. Sigurdsson, Robinson Piramuthu, Jesse Thomason, Dhruv Batra, Gaurav S. Sukhatme
ICML3
2021 Self-attentive 3D human pose and shape estimation from videos
Yun-Chun Chen, Marco Piccirilli, Robinson Piramuthu, Ming-Hsuan Yang 0001
Comput. Vis. Image Underst.3
2020 Mixup-CAM: Weakly-supervised Semantic Segmentation via Uncertainty Regularization
Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001
BMVC4
2020 Weakly-Supervised Semantic Segmentation via Sub-Category Exploration
abstract
Existing weakly-supervised semantic segmentation methods using image-level annotations typically rely on initial responses to locate object regions. However, such response maps generated by the classification network usually focus on discriminative object parts, due to the fact that the network does not need the entire object for optimizing the objective function. To enforce the network to pay attention to other parts of an object, we propose a simple yet effective approach that introduces a self-supervised task by exploiting the sub-category information. Specifically, we perform clustering on image features to generate pseudo sub-categories labels within each annotated parent class, and construct a sub-category objective to assign the network to a more challenging task. By iteratively clustering image features, the training process does not limit itself to the most discriminative object parts, hence improving the quality of the response maps. We conduct extensive analysis to validate the proposed method and show that our approach performs favorably against the state-of-the-art approaches.
Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001
CVPR4
2019 Adversarial Learning for Fine-Grained Image Search
abstract
Fine-grained image search is still a challenging problem due to the difficulty in capturing subtle differences regardless of pose variations of objects from fine-grained categories. In practice, a dynamic inventory with new fine-grained categories adds another dimension to this challenge. In this work, we propose an end-to-end network, called FGGAN, that learns discriminative representations by implicitly learning a geometric transformation from multi-view images for fine-grained rigid object retrieval. We integrate a generative adversarial network (GAN) that can automatically handle complex view and pose variations by converting them to a canonical view without any predefined transformations. Moreover, in an open-set scenario, our network is able to better match rigid objects from unseen and unknown fine-grained categories. Extensive experiments on the public CompCars dataset and a newly collected dataset have demonstrated the effectiveness of the proposed method in both closed-set and open-set scenarios.
Fan Yang 0076, Qiaosong Wang, Robinson Piramuthu
ICME4
2019 Understanding Image Quality and Trust in Peer-to-Peer Marketplaces
abstract
As any savvy online shopper knows, second-hand peer-to-peer marketplaces are filled with images of mixed quality. How does image quality impact marketplace outcomes, and can quality be automatically predicted? In this work, we conducted a large-scale study on the quality of user-generated images in peer-to-peer marketplaces. By gathering a dataset of common second-hand products (≈75,000 images) and annotating a subset with human-labeled quality judgments, we were able to model and predict image quality with decent accuracy (≈87%). We then conducted two studies focused on understanding the relationship between these image quality scores and two marketplace outcomes: sales and perceived trustworthiness. We show that image quality is associated with higher likelihood that an item will be sold, though other factors such as view count were better predictors of sales. Nonetheless, we show that high quality user-generated images selected by our models outperform stock imagery in eliciting perceptions of trust from users. Our findings can inform the design of future marketplaces and guide potential sellers to take better product images.
Xiao Ma 0010, Lina Mezghani, Kimberly Wilber, Hui Hong, Robinson Piramuthu, Mor Naaman, Serge J. Belongie
WACV5
2019 Give Me a Hint! Navigating Image Databases Using Human-in-the-Loop Feedback
abstract
In this paper, we introduce an attribute-based interactive image search which can leverage human-in-the-loop feedback to iteratively refine image search results. We study active image search where human feedback is solicited exclusively in visual form, without using relative attribute annotations used by prior work which are not typically found in many datasets. In order to optimize the image selection strategy, a deep reinforcement model is trained to learn what images are informative rather than rely on hand-crafted measures typically leveraged in prior work. Additionally, we extend the recently introduced Conditional Similarity Network to incorporate global similarity in training visual embeddings, which results in more natural transitions as the user explores the learned similarity embeddings. Our experiments demonstrate the effectiveness of our approach, producing compelling results on both active image search and image attribute representation tasks.
Bryan A. Plummer, M. Hadi Kiapour, Shuai Zheng 0004, Robinson Piramuthu
WACV4
2018 Conditional Image-Text Embedding Networks
Bryan A. Plummer, Paige Kordas, M. Hadi Kiapour, Shuai Zheng 0004, Robinson Piramuthu, Svetlana Lazebnik
ECCV (12)5
2018 ModaNet: A Large-scale Street Fashion Dataset with Polygon Annotations
abstract
Understanding clothes from a single image would have huge commercial and cultural impacts on modern societies. However, this task remains a challenging computer vision problem due to wide variations in the appearance, style, brand and layering of clothing items. We present a new database called ModaNet, a large-scale collection of images based on Paperdoll dataset. Our dataset provides 55,176 street images, fully annotated with polygons on top of the 1 million weakly annotated street images in Paperdoll. ModaNet aims to provide a technical benchmark to fairly evaluate the progress of applying the latest computer vision techniques that rely on large data for fashion understanding. The rich annotation of the dataset allows to measure the performance of state-of-the-art algorithms for object detection, semantic segmentation and polygon prediction on street fashion images in detail.
Shuai Zheng 0004, Fan Yang 0076, M. Hadi Kiapour, Robinson Piramuthu
ACM Multimedia4
2018 Towards the Success Rate of One: Real-Time Unconstrained Salient Object Detection
abstract
In this work, we propose an efficient and effective approach for unconstrained salient object detection in images using deep convolutional neural networks. Instead of generating thousands of candidate bounding boxes and refining them, our network directly learns to generate the saliency map containing the exact number of salient objects. During training, we convert the ground-truth rectangular boxes to Gaussian distributions that better capture the ROI regarding individual salient objects. During inference, the network predicts Gaussian distributions centered at salient objects with an appropriate covariance, from which bounding boxes are easily inferred. Notably, our network performs saliency map prediction without pixel-level annotations, salient object detection without object proposals, and salient object subitizing simultaneously, all in a single pass within a unified framework. Extensive experiments show that our approach outperforms existing methods on various datasets by a large margin, and achieves more than 100 fps with VGG16 network on a single GPU during inference.
Mahyar Najibi, Fan Yang 0076, Qiaosong Wang, Robinson Piramuthu
WACV4
2017 Visual Search at eBay
abstract
In this paper, we propose a novel end-to-end approach for scalable visual search infrastructure. We discuss the challenges we faced for a massive volatile inventory like at eBay and present our solution to overcome those. We harness the availability of large image collection of eBay listings and state-of-the-art deep learning techniques to perform visual search at scale. Supervised approach for optimized search limited to top predicted categories and also for compact binary signature are key to scale up without compromising accuracy and precision. Both use a common deep neural network requiring only a single forward inference. The system architecture is presented with in-depth discussions of its basic components and optimizations for a trade-off between search relevance and latency. This solution is currently deployed in a distributed cloud infrastructure and fuels visual search in eBay ShopBot and Close5. We show benchmark on ImageNet dataset on which our approach is faster and more accurate than several unsupervised baselines. We share our learnings with the hope that visual search becomes a first class citizen for all large scale search engines rather than an afterthought.
Fan Yang 0076, Ajinkya Kale, Yury Bubnov, Leon Stein, Qiaosong Wang, M. Hadi Kiapour, Robinson Piramuthu
KDD7
2016 GraB: Visual Saliency via Novel Graph Model and Background Priors
abstract
We propose an unsupervised bottom-up saliency detection approach by exploiting novel graph structure and background priors. The input image is represented as an undirected graph with superpixels as nodes. Feature vectors are extracted from each node to cover regional color, contrast and texture information. A novel graph model is proposed to effectively capture local and global saliency cues. To obtain more accurate saliency estimations, we optimize the saliency map by using a robust background measure. Comprehensive evaluations on benchmark datasets indicate that our algorithm universally surpasses state-of-the-art unsupervised solutions and performs favorably against supervised approaches.
Qiaosong Wang, Robinson Piramuthu
CVPR3
2016 Fashion apparel detection: The role of deep convolutional neural network and pose-dependent priors
abstract
In this work, we propose and address a new computer vision task, which we call fashion item detection, where the aim is to detect various fashion items a person in the image is wearing or carrying. The types of fashion items we consider in this work include hat, glasses, bag, pants, shoes and so on. The detection of fashion items can be an important first step of various e-commerce applications for fashion industry. Our method is based on state-of-the-art object detection method pipeline which combines object proposal methods with a Deep Convolutional Neural Network. Since the locations of fashion items are in strong correlation with the locations of body joints positions, we incorporate contextual information from body poses in order to improve the detection performance. Through the experiments, we demonstrate the effectiveness of the proposed method.
Kota Hara, Vignesh Jagadeesh, Robinson Piramuthu
WACV3
2015 ConceptLearner: Discovering visual concepts from weakly labeled image collections
abstract
Discovering visual knowledge from weakly labeled data is crucial to scale up computer vision recognition systems, since it is expensive to obtain fully labeled data for a large number of concept categories. In this paper, we propose ConceptLearner, which is a scalable approach to discover visual concepts from weakly labeled image collections. Thousands of visual concept detectors are learned automatically, without human in the loop for additional annotation. We show that these learned detectors could be applied to recognize concepts at image-level and to detect concepts at image region-level accurately. Under domain-specific supervision, we further evaluate the learned concepts for scene recognition on SUN database and for object detection on Pascal VOC 2007. ConceptLearner shows promising performance compared to fully supervised and weakly supervised methods.
Bolei Zhou, Vignesh Jagadeesh, Robinson Piramuthu
CVPR3
2015 HD-CNN: Hierarchical Deep Convolutional Neural Networks for Large Scale Visual Recognition
abstract
In image classification, visual separability between different object categories is highly uneven, and some categories are more difficult to distinguish than others. Such difficult categories demand more dedicated classifiers. However, existing deep convolutional neural networks (CNN) are trained as flat N-way classifiers, and few efforts have been made to leverage the hierarchical structure of categories. In this paper, we introduce hierarchical deep CNNs (HD-CNNs) by embedding deep CNNs into a two-level category hierarchy. An HD-CNN separates easy classes using a coarse category classifier while distinguishing difficult classes using fine category classifiers. During HDCNN training, component-wise pretraining is followed by global fine-tuning with a multinomial logistic loss regularized by a coarse category consistency term. In addition, conditional executions of fine category classifiers and layer parameter compression make HD-CNNs scalable for largescale visual recognition. We achieve state-of-the-art results on both CIFAR100 and large-scale ImageNet 1000-class benchmark datasets. In our experiments, we build up three different two-level HD-CNNs, and they lower the top-1 error of the standard CNNs by 2:65%, 3:1%, and 1:1%.
Zhicheng Yan 0001, Hao Zhang 0025, Robinson Piramuthu, Vignesh Jagadeesh, Dennis DeCoste, Wei Di, Yizhou Yu
ICCV3
2015 Mine the fine: Fine-grained fragment discovery
abstract
While discriminative visual element mining has been introduced before, in this paper we present an approach that requires minimal annotation in both training and test time. Given only a bounding box localization of the foreground objects, our approach automatically transforms the input images into a roughly-aligned pose space and discovers the most discriminative visual fragments for each category. These fragments are then used to learn robust classifiers that discriminate between very similar categories under challenging conditions such as large variations in pose or habitats. The minimal required input, is a critical characteristic that enables our approach to generalize over visual domains where expert knowledge is not readily available. Moreover, our approach takes advantage of deep networks that are targeted towards fine-grained classification. It learns mid-level representations that are specific to a category and generalize well across the category instances at the same time. Our evaluations demonstrate that the automatically learned representation based on discriminative fragments, significantly outperforms globally extracted deep features in classification accuracy.
M. Hadi Kiapour, Wei Di, Vignesh Jagadeesh, Robinson Piramuthu
ICIP4
2015 Efficient Media Retrieval from Non-Cooperative Queries
Kevin Shih, Wei Di, Vignesh Jagadeesh, Robinson Piramuthu
ICVS4
2014 Region-Based Discriminative Feature Pooling for Scene Text Recognition
abstract
We present a new feature representation method for scene text recognition problem, particularly focusing on improving scene character recognition. Many existing methods rely on Histogram of Oriented Gradient (HOG) or part-based models, which do not span the feature space well for characters in natural scene images, especially given large variation in fonts with cluttered backgrounds. In this work, we propose a discriminative feature pooling method that automatically learns the most informative sub-regions of each scene character within a multi-class classification framework, whereas each sub-region seamlessly integrates a set of low-level image features through integral images. The proposed feature representation is compact, computationally efficient, and able to effectively model distinctive spatial structures of each individual character class. Extensive experiments conducted on challenging datasets (Chars74K, ICDAR'03, ICDAR'11, SVT) show that our method significantly outperforms existing methods on scene character classification and scene text recognition tasks.
Chen-Yu Lee, Anurag Bhardwaj, Wei Di, Vignesh Jagadeesh, Robinson Piramuthu
CVPR5
2014 Cascaded sparse color-localized matching for logo retrieval
abstract
In this paper we present a framework for logo retrieval in natural images. Color-localized spatial masks are used as an alternative to computationally expensive spatial verification techniques like RANSAC. First, keypoints are detected using traditional techniques such as the SIFT detector. Local masks are defined around each keypoint that take its scale and orientation information into account. To exploit inherent color information presented in brand logos, ordered color histograms are extracted from masked regions. A separate vocabulary is constructed for both SIFT descriptors (visual word) and color histograms (color word). For faster matching during runtime, a two-stage cascaded index is designed, which maps the visual word and color word tuple to a list of relevant images. This list is finally re-ranked with BoW cosine similarity to generate relevant matches for the input query. To demonstrate the efficacy of our method, we conduct experiments on two popular logo datasets: Flickr27 and Flickr32. Our experimental results illustrate State-of-the-art retrieval performance on these datasets with potential for added speed and a lower memory footprint as indicated by the low response ratio.
Rohit Pandey, Wei Di, Vignesh Jagadeesh, Robinson Piramuthu, Anurag Bhardwaj
ICIP4
2014 Large scale visual recommendations from street fashion images
abstract
We describe a completely automated large scale visual recommendation system for fashion. Our focus is to efficiently harness the availability of large quantities of online fashion images and their rich meta-data. Specifically, we propose two classes of data driven models in the Deterministic Fashion Recommenders (DFR) and Stochastic Fashion Recommenders (SFR) for solving this problem. We analyze relative merits and pitfalls of these algorithms through extensive experimentation on a large-scale data set and baseline them against existing ideas from color science. We also illustrate key fashion insights learned through these experiments and show how they can be employed to design better recommendation systems. The industrial applicability of proposed models is in the context of mobile fashion shopping. Finally, we also outline a large-scale annotated data set of fashion images Fashion-136K) that can be exploited for future research in data driven visual fashion.
Vignesh Jagadeesh, Robinson Piramuthu, Anurag Bhardwaj, Wei Di, Neel Sundaresan
KDD2
2014 Im2depth: Scalable exemplar based depth transfer
abstract
The rapid increase in number of high quality mobile cameras have opened up an array of new problems in mobile vision. Mobile cameras are predominantly monocular and are devoid of any sense of depth, making them heavily reliant on 2D image processing. Understanding 3D structure of scenes being imaged can greatly improve the performance of existing vision/graphics techniques. In this regard, recent availability of large scale RGB-D datasets beg for more effective data driven strategies to leverage the scale of data. We propose a depth recovery mechanism “im2depth”, that is lightweight enough to run on mobile platforms, while leveraging the large scale nature of modern RGB-D datasets. Our key observation is to form a basis (dictionary) over the RGB and depth spaces, and represent depth maps by a sparse linear combination of weights over dictionary elements. Subsequently, a prediction function is estimated between weight vectors in RGB to depth space to recover depth maps from query images. A final superpixel post processor aligns depth maps with occlusion boundaries, creating physically plausible results. We conclude with thorough experimentation with four state of the art depth recovery algorithms, and observe an improvement of over 6.5 percent in shape recovery, and over 10cm reduction in average L1 error.
Mohammad Haris Baig, Vignesh Jagadeesh, Robinson Piramuthu, Anurag Bhardwaj, Wei Di, Neel Sundaresan
WACV3
2014 Furniture-geek: Understanding fine-grained furniture attributes from freely associated text and tags
abstract
As the amount of user generated content on the internet grows, it becomes ever more important to come up with vision systems that learn directly from weakly annotated and noisy data. We leverage a large scale collection of user generated content comprising of images, tags and title/captions of furniture inventory from an e-commerce website to discover and categorize learnable visual attributes. Furniture categories have long been the quintessential example of why computer vision is hard, and we make one of the first attempts to understand them through a large scale weakly annotated dataset. We focus on a handful of furniture categories that are associated with a large number of fine-grained attributes. We propose a set of localized feature representations built on top of state-of-the-art computer vision representations originally designed for fine-grained object categorization. We report a thorough empirical characterization on the visual identifiability of various fine-grained attributes using these representations and show encouraging results on finding iconic images and on multi-attribute prediction.
Vicente Ordonez, Vignesh Jagadeesh, Wei Di, Anurag Bhardwaj, Robinson Piramuthu
WACV5
2014 Is a picture really worth a thousand words?: - on the role of images in e-commerce
abstract
In online peer-to-peer commerce places where physical examination of the goods is infeasible, textual descriptions, images of the products, reputation of the participants, play key roles. Visual image is a powerful channel to convey crucial information towards e-shoppers and influence their choice. In this paper, we investigate a well-known online marketplace where over millions of products change hands and most are described with the help of one or more images. We present a systematic data mining and knowledge discovery approach that aims to quantitatively dissect the role of images in e-commerce in great detail. Our goal is two-fold. First, we aim to get a thorough understanding of impact of images across various dimensions: product categories, user segments, conversion rate. We present quantitative evaluation of the influence of images and show how to leverage different image aspects, such as quantity and quality, to effectively raise sale. Second, we study interaction of image data with other selling dimensions by jointly modeling them with user behavior data. Results suggest that "watch" behavior encodes complex signals combining both attention and hesitation from buyer, in which image still holds an important role when compared to other selling variables, especially for products for which appearance is important. We conclude on how these findings can benefit sellers in a high competitive online e-commerce market.
Wei Di, Neel Sundaresan, Robinson Piramuthu, Anurag Bhardwaj
WSDM3
2013 Palette power: enabling visual search through colors
abstract
With the explosion of mobile devices with cameras, online search has moved beyond text to other modalities like images, voice, and writing. For many applications like Fashion, image-based search offers a compelling interface as compared to text forms by better capturing the visual attributes. In this paper we present a simple and fast search algorithm that uses color as the main feature for building visual search. We show that low level cues such as color can be used to quantify image similarity and also to discriminate among products with different visual appearances. We demonstrate the effectiveness of our approach through a mobile shopping application\footnote{eBay Fashion App available at https://itunes.apple.com/us/app/ebay-fashion/id378358380?mt=8 and eBay image swatch is the feature indexing millions of real world fashion images}. Our approach outperforms several other state-of-the-art image retrieval algorithms for large scale image data.
Anurag Bhardwaj, Atish Das Sarma, Wei Di, Raffay Hamid, Robinson Piramuthu, Neel Sundaresan
KDD5
1999 Minimax Emission Computed Tomography using High-Resolution Anatomical Side Information and B-Spline Models
abstract
In this paper a minimax methodology is presented for combining information from two imaging modalities having different intrinsic spatial resolutions. The focus application is emission computed tomography (ECT), a low-resolution modality for reconstruction of radionuclide tracer density, when supplemented by high-resolution anatomical boundary information extracted from a magnetic resonance image (MRI) of the same imaging volume. The MRI boundary within the two-dimensional (2-D) slice of interest is parameterized by a closed planar curve. The Cramer-Rao (CR) lower bound is used to analyze estimation errors for different boundary shapes. Under a spatially inhomogeneous Gibbs field model for the tracer density a representation for the minimax MRI-enhanced tracer density estimator is obtained. It is shown that the estimator is asymptotically equivalent to a penalized maximum likelihood (PML) estimator with resolution-selective Gibbs penalty. Quantitative comparisons are presented using the iterative space alternating generalized expectation maximization (SAGE-FM) algorithm to implement the PML estimator with and without minimax weight averaging.
Alfred O. Hero III, Robinson Piramuthu, Jeffrey A. Fessler, Stephen R. Titus
IEEE Trans. Inf. Theory2
1998 Penalized maximum likelihood image reconstruction with min-max incorporation of noisy side information
abstract
A method for incorporating anatomical MRI boundary side information into penalized maximum likelihood (PML) emission computed tomography (ECT) image reconstructions using a set of averaged Gibbs weights was proposed by Hero and Piramuthu (see Proc. of IEEE/EURASIP Workshop on Nonlinear Signal and Image Processing, 1997). A quadratic penalty based on Gibbs weights was used to enforce smoothness constraints everywhere in the image except across the estimated boundary of the ROI. In this methodology, a limiting form of the posterior distribution of the MRI boundary parameters was used to average the Gibbs weights obtained by Titus, Hero and Fessler (see IEEE Int. Conf. on Image Processing, vol.2, Laussane, 1996). There is an improvement in performance over the method proposed by Titus et al., when the variance of boundary estimates from the MRI data becomes significant. Here, we present the empirical performance analysis of the proposed method of averaged Gibbs weights.
Robinson Piramuthu, Alfred O. Hero III
ICASSP1
1998 Side Information Averaging Method for PML Emission Tomography
abstract
The authors previously presented a methodology for incorporating perfect extracted MRI anatomical boundary estimates to improve the performance of penalized likelihood (PL) emission computed tomography (ECT) image reconstruction and ECT tracer uptake estimation. This technique used a spatially variant quadratic Gibbs penalty which enforced smoothness everywhere in the ECT image except across the MRI-extracted boundary of the ROI. When high quality estimates of the anatomical boundary are available and MRI and ECT images are perfectly registered, the performance of this Gibbs penalty method is very close to that attainable using perfect side information, i.e., an errorless anatomical boundary estimate. However when the variance of the MRI-extracted boundary estimate becomes significant this method performs poorly. Here we present a modified Gibbs penalty function which accounts for errors in side information based on an asymptotic min-max robustness approach. The resulting penalty is implemented with a set of averaged Gibbs weights where the averaging is performed with respect to a limiting form of the min-max induced posterior distribution of the MRI boundary parameters. Examples are presented for tracer uptake estimation using the SAGE version of the EM algorithm and various parameterizations of the anatomical boundaries.
Robinson Piramuthu, Alfred O. Hero III
ICIP (2)1