Hailin Jin

dblp:78/102 · DBLP profile ↗
← Back
113ranked-venue papers
14as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 99 · 12 first-author · 18 since 2021Artificial intelligence and machine learning · 84 · 11 first-author · 19 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 The Orlicz Gaussian Minkowski Problem for General Measures
Hailin Jin, Dandan Lai, Suwei Li
Discret. Comput. Geom.1
2025 Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval
abstract
Video moment retrieval (VMR) aims to locate the most likely video moment(s) corresponding to a text query in untrimmed videos. Training of existing methods is limited by the lack of diverse and generalisable VMR datasets, hindering their ability to generalise moment-text associations to queries containing novel semantic concepts (unseen both visually and textually in a training source domain). For model generalisation to novel semantics, existing methods rely heavily on assuming to have access to both video and text sentence pairs from a target domain in addition to the source domain pair-wise training data. This is neither practical nor scalable. In this work, we introduce a more generalisable approach by assuming only text sentences describing new semantics are available in model training without having seen any videos from a target domain. To that end, we propose a Fine-grained Video Editing framework, termed FVE, that explores generative video diffusion to facilitate fine-grained video editing from the seen source concepts to the unseen target sentences consisting of new concepts. This enables generative hypotheses of unseen video moments corresponding to the novel concepts in the target domain. This fine-grained generative video diffusion retains the original video structure and subject specifics from the source domain while introducing semantic distinctions of unseen novel vocabularies in the target domain. A critical challenge is how to enable this generative fine-grained diffusion process to be meaningful in optimising VMR, more than just synthesising visually pleasing videos. We solve this problem by introducing a hybrid selection mechanism that integrates three quantitative metrics to selectively incorporate synthetic video moments (novel video hypotheses) as enlarged additions to the original source training data, whilst minimising potential detrimental noise or unnecessary repetitions in the novel synthetic videos harmful to VMR learning. Experiments on three datasets demonstrate the effectiveness of FVE to unseen novel semantic video moment retrieval tasks
Dezhao Luo, Shaogang Gong, Jiabo Huang, Hailin Jin, Yang Liu 0105
AAAI4
2025 TeachText: CrossModal text-video retrieval through generalized distillation
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Yang Liu 0105, Samuel Albanie
Artif. Intell.4
2025 MLLM as video narrator: Mitigating modality imbalance in video moment retrieval
Weitong Cai, Jiabo Huang, Shaogang Gong, Hailin Jin, Yang Liu 0105
Pattern Recognit.4
2024 Zero-Shot Video Moment Retrieval from Frozen Vision-Language Models
abstract
Accurate video moment retrieval (VMR) requires universal visual-textual correlations that can handle unknown vocabulary and unseen scenes. However, the learned correlations are likely either biased when derived from a limited amount of moment-text data which is hard to scale up because of the prohibitive annotation cost (fully-supervised), or unreliable when only the video-text pairwise relationships are available without fine-grained temporal annotations (weakly-supervised). Recently, the vision-language models (VLM) demonstrate a new transfer learning paradigm to benefit different vision tasks through the universal visual-textual correlations derived from large-scale vision-language pairwise web data, which has also shown benefits to VMR by fine-tuning in the target domains.In this work, we propose a zero-shot method for adapting generalisable visual-textual priors from arbitrary VLM to facilitate moment-text alignment, without the need for accessing the VMR data. To this end, we devise a conditional feature refinement module to generate boundary-aware visual features conditioned on text queries to enable better moment boundary understanding. Additionally, we design a bottom-up proposal generation strategy that mitigates the impact of domain discrepancies and breaks down complex-query retrieval tasks into individual action retrievals, thereby maximizing the benefits of VLM. Extensive experiments conducted on three VMR benchmark datasets demonstrate the notable performance advantages of our zero-shot algorithm, especially in the novel-word and novel-location out-of-distribution setups.
Dezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin, Yang Liu 0105
WACV4
2023 Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence Localization
abstract
Video sentence localization aims to locate moments in an unstructured video according to a given natural language query.A main challenge is the expensive annotation costs and the annotation bias.In this work, we study video sentence localization in a zero-shot setting, which learns with only video data without any annotation.Existing zero-shot pipelines usually generate event proposals and then generate a pseudo query for each event proposal.However, their event proposals are obtained via visual feature clustering, which is query-independent and inaccurate; and the pseudo-queries are short or less interpretable.Moreover, existing approaches ignores the risk of pseudo-label noise when leveraging them in training.To address the above problems, we propose a Structurebased Pseudo Label generation (SPL), which first generate free-form interpretable pseudo queries before constructing query-dependent event proposals by modeling the event temporal structure.To mitigate the effect of pseudolabel noise, we propose a noise-resistant iterative method that repeatedly re-weight the training sample based on noise estimation to train a grounding model and correct pseudo labels.Experiments on the ActivityNet Captions and Charades-STA datasets demonstrate the advantages of our approach.Code can be found at https://github.com/minghangz/SPL.
Minghang Zheng, Shaogang Gong, Hailin Jin, Yuxin Peng 0001, Yang Liu 0105
ACL (1)3
2023 Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-Training
abstract
The correlation between the vision and text is essential for video moment retrieval (VMR), however, existing methods heavily rely on separate pre-training feature extractors for visual and textual understanding. Without sufficient temporal boundary annotations, it is non-trivial to learn universal video-text alignments. In this work, we explore multi-modal correlations derived from large-scale image-text data to facilitate generalisable VMR. To address the limitations of image-text pre-training models on capturing the video changes, we propose a generic method, referred to as Visual-Dynamic Injection (VDI), to empower the model's understanding of video moments. Whilst existing VMR methods are focusing on building temporalaware video features, being aware of the text descriptions about the temporal changes is also critical but originally overlooked in pre-training by matching static images with sentences. Therefore, we extract visual context and spatial dynamic information from video frames and explicitly enforce their alignments with the phrases describing video changes (e.g. verb). By doing so, the potentially relevant visual and motion patterns in videos are encoded in the corresponding text embeddings (injected) so to enable more accurate video-text alignments. We conduct extensive experiments on two VMR benchmark datasets (Charades-STA and ActivityNet-Captions) and achieve state-of-the-art performances. Especially, VDI yields notable advantages when being tested on the out-of-distribution splits where the testing samples involve novel scenes and vocabulary.
Dezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin, Yang Liu 0105
CVPR4
2023 Moment Detection in Long Tutorial Videos
abstract
Tutorial videos play an increasingly important role in professional development and self-directed education. For users to realise the full benefits of this medium, tutorial videos must be efficiently searchable. In this work, we focus on the task of moment detection, in which the goal is to localise the temporal window where a given event occurs within a given tutorial video. Prior work on moment detection has focused primarily on short videos (typically on videos shorter than three minutes). However, many tutorial videos are substantially longer (stretching to hours in duration), presenting significant challenges for existing moment detection approaches.To study this problem, we propose the first dataset of untrimmed, long-form tutorial videos for the task of Moment Detection called the Behance Moment Detection (BMD) dataset. BMD videos have an average duration of over one hour and are characterised by slowly evolving visual content and wide-ranging dialogue. To meet the unique challenges of this dataset, we propose a new framework, LongMoment-Detr, and demonstrate that it outperforms strong baselines. Additionally, we introduce a variation of the dataset that contains YouTube Chapter annotations and show that the features obtained by our framework can be successfully used to boost the performance on the task of chapter detection. Code and data can be found at https://github.com/ioanacroi/longmoment-detr.
Ioana Croitoru, Simion-Vlad Bogolin, Samuel Albanie, Yang Liu 0105, Seunghyun Yoon 0002, Franck Dernoncourt, Hailin Jin, Trung Bui
ICCV8
2023 Efficient Adaptive Human-Object Interaction Detection with Concept-guided Memory
abstract
Human Object Interaction (HOI) detection aims to localize and infer the relationships between a human and an object. Arguably, training supervised models for this task from scratch presents challenges due to the performance drop over rare classes and the high computational cost and time required to handle long-tailed distributions of HOIs in complex HOI scenes in realistic settings. This observation motivates us to design an HOI detector that can be trained even with long-tailed labeled data and can leverage existing knowledge from pre-trained models. Inspired by the powerful generalization ability of the large Vision-Language Models (VLM) on classification and retrieval tasks, we propose an efficient Adaptive HOI Detector with Concept-guided Memory (ADA-CM). ADA-CM has two operating modes. The first mode makes it tunable without learning new parameters in a training-free paradigm. Its second mode incorporates an instance-aware adapter mechanism that can further efficiently boost performance if updating a lightweight set of parameters can be afforded. Our proposed method achieves competitive results with state-of-the-art on the HICO-DET and V-COCO datasets with much less training time. Code can be found at https://github.com/ltttpku/ADA-CM.
Ting Lei 0001, Fabian Caba Heilbron, Qingchao Chen, Hailin Jin, Yuxin Peng 0001, Yang Liu 0105
ICCV4
2023 LiveSeg: Unsupervised Multimodal Temporal Segmentation of Long Livestream Videos
abstract
Livestream videos have become a significant part of online learning, where design, digital marketing, creative painting, and other skills are taught by experienced experts in the sessions, making them valuable materials. However, Livestream tutorial videos are usually hours long, recorded, and uploaded to the Internet directly after the live sessions, making it hard for other people to catch up quickly. An outline will be a beneficial solution, which requires the video to be temporally segmented according to topics. In this work, we introduced a large Livestream video dataset named MultiLive, and formulated the temporal segmentation of the long Livestream videos (TSLLV) task. We propose LiveSeg, an unsupervised Livestream video temporal Segmentation solution, which takes advantage of multimodal features from different domains. Our method achieved a 16.8% F1-score performance improvement compared with the state-of-the-art method.
Jielin Qiu, Franck Dernoncourt, Trung Bui, Ding Zhao, Hailin Jin
WACV6
2022 Cross Modal Retrieval with Querybank Normalisation
abstract
Profiting from large-scale training datasets, advances in neural architecture design and efficient inference, joint embeddings have become the dominant approach for tackling cross-modal retrieval. In this work we first show that, despite their effectiveness, state-of-the-art joint embeddings suffer significantly from the longstanding “hubness problem” in which a small number of gallery embeddings form the nearest neighbours of many queries. Drawing inspiration from the NLP literature, we formulate a simple but effective framework called Querybank Normalisation (QB-NORM) that re-normalises query similarities to account for hubs in the embedding space. QB-NORM improves retrieval performance without requiring retraining. Differently from prior work, we show that QB-NORM works effectively without concurrent access to any test set queries. Within the QB-NORM framework, we also propose a novel similarity normalisation method, the Dynamic Inverted Softmax, that is significantly more robust than existing approaches. We showcase QB-NORM across a range of cross modal retrieval models and benchmarks where it consistently enhances strong baselines beyond the state of the art. Code is available at https://vladbogo.github.io/QB-Norm/.
Simion-Vlad Bogolin, Ioana Croitoru, Hailin Jin, Yang Liu 0105, Samuel Albanie
CVPR3
2022 Video Activity Localisation with Uncertainties in Temporal Boundary
Jiabo Huang, Hailin Jin, Shaogang Gong, Yang Liu 0105
ECCV (34)2
2022 StyleBabel: Artistic Style Tagging and Captioning
abstract
We present StyleBabel, a unique open access dataset of natural language captions and free-form tags describing the artistic style of over 135K digital artworks, collected via a novel participatory method from experts studying at specialist art and design schools. StyleBabel was collected via an iterative method, inspired by ‘Grounded Theory’: a qualitative approach that enables annotation while co-evolving a shared language for fine-grained artistic style attribute description. We demonstrate several downstream tasks for StyleBabel, adapting the recent ALADIN architecture for fine-grained style similarity, to train cross-modal embeddings for: 1) free-form tag generation; 2) natural language description of artistic style; 3) fine-grained text search of style. To do so, we extend ALADIN with recent advances in Visual Transformer (ViT) and cross-modal representation learning, achieving a state of the art accuracy in fine-grained style retrieval.
Dan Ruta, Andrew Gilbert, Pranav Aggarwal, Naveen Marri, Ajinkya Kale, Jo Briggs, Chris Speed, Hailin Jin, Baldo Faieta, Alex Filipkowski, Zhe Lin 0001, John P. Collomosse
ECCV (8)8
2022 Privacy-Preserving Deep Action Recognition: An Adversarial Learning Framework and A New Dataset
abstract
We investigate privacy-preserving, video-based action recognition in deep learning, a problem with growing importance in smart camera applications. A novel adversarial training framework is formulated to learn an anonymization transform for input videos such that the trade-off between target utility task performance and the associated privacy budgets is explicitly optimized on the anonymized videos. Notably, the privacy budget, often defined and measured in task-driven contexts, cannot be reliably indicated using any single model performance because strong protection of privacy should sustain against any malicious model that tries to steal private information. To tackle this problem, we propose two new optimization strategies of model restarting and model ensemble to achieve stronger universal privacy protection against any attacker models. Extensive experiments have been carried out and analyzed. On the other hand, given few public datasets available with both utility and privacy labels, the data-driven (supervised) learning cannot exert its full power on this task. We first discuss an innovative heuristic of cross-dataset training and evaluation, enabling the use of multiple single-task datasets (one with target task labels and the other with privacy labels) in our problem. To further address this dataset challenge, we have constructed a new dataset, termed PA-HMDB51, with both target task labels (action) and selected privacy attributes (skin color, face, gender, nudity, and relationship) annotated on a per-frame basis. This first-of-its-kind video dataset and evaluation protocol can greatly facilitate visual privacy research and open up other opportunities. Our codes, models, and the PA-HMDB51 dataset are available at: https://github.com/VITA-Group/PA-HMDB51.
Zhenyu Wu 0002, Haotao Wang, Hailin Jin, Zhangyang Wang
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Magic Layouts: Structural Prior for Component Detection in User Interface Designs
Dipu Manandhar, Hailin Jin, John P. Collomosse
CVPR2
2021 StreamHover: Livestream Transcript Summarization and Annotation
abstract
Sangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui, Nedim Lipka, Walter Chang, Hailin Jin, Jonathan Brandt, Hassan Foroosh, Fei Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Sangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui, Nedim Lipka, Walter Chang, Hailin Jin, Jonathan Brandt, Hassan Foroosh, Fei Liu 0004
EMNLP (1)7
2021 TeachText: CrossModal Generalized Distillation for Text-Video Retrieval
abstract
In recent years, considerable progress on the task of text-video retrieval has been achieved by leveraging large-scale pretraining on visual and audio datasets to construct powerful video encoders. By contrast, despite the natural symmetry, the design of effective algorithms for exploiting large-scale language pretraining remains under-explored. In this work, we are the first to investigate the design of such algorithms and propose a novel generalized distillation method, TeachText, which leverages complementary cues from multiple text encoders to provide an enhanced supervisory signal to the retrieval model. Moreover, we extend our method to video side modalities and show that we can effectively reduce the number of used modalities at test time without compromising performance. Our approach advances the state of the art on several video retrieval benchmarks by a significant margin and adds no computational overhead at test time. Last but not least, we show an effective application of our method for eliminating noise from retrieval datasets. Code and data can be found at https://www.robots.ox.ac.uk/˜vgg/research/teachtext/.
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Samuel Albanie, Yang Liu 0105
ICCV4
2021 Cross-Sentence Temporal and Semantic Relations in Video Activity Localisation
abstract
Video activity localisation has recently attained increasing attention due to its practical values in automatically localising the most salient visual segments corresponding to their language descriptions (sentences) from untrimmed and unstructured videos. For supervised model training, a temporal annotation of both the start and end time index of each video segment for a sentence (a video moment) must be given. This is not only very expensive but also sensitive to ambiguity and subjective annotation bias, a much harder task than image labelling. In this work, we develop a more accurate weakly-supervised solution by introducing Cross-Sentence Relations Mining (CRM) in video moment proposal generation and matching when only a paragraph description of activities without per-sentence temporal annotation is available. Specifically, we explore two cross-sentence relational constraints: (1) Temporal ordering and (2) semantic consistency among sentences in a paragraph description of video activities. Existing weakly-supervised techniques only consider within-sentence video segment correlations in training without considering cross-sentence paragraph context. This can mislead due to ambiguous expressions of individual sentences with visually indiscriminate video moment proposals in isolation. Experiments on two publicly available activity localisation datasets show the advantages of our approach over the state-of-the-art weakly supervised methods, especially so when the video activity descriptions become more complex.
Jiabo Huang, Yang Liu 0105, Shaogang Gong, Hailin Jin
ICCV4
2021 Time-Equivariant Contrastive Video Representation Learning
abstract
We introduce a novel self-supervised contrastive learning method to learn representations from unlabelled videos. Existing approaches ignore the specifics of input distortions, e.g., by learning invariance to temporal transformations. Instead, we argue that video representation should preserve video dynamics and reflect temporal manipulations of the input. Therefore, we exploit novel constraints to build representations that are equivariant to temporal transformations and better capture video dynamics. In our method, relative temporal transformations between augmented clips of a video are encoded in a vector and contrasted with other transformation vectors. To support temporal equivariance learning, we additionally propose the self-supervised classification of two clips of a video into 1. overlapping 2. ordered, or 3. unordered. Our experiments show that time-equivariant representations achieve state-of-the-art results in video retrieval and action recognition benchmarks on UCF101, HMDB51, and Diving48.
Simon Jenni, Hailin Jin
ICCV2
2021 ALADIN: All Layer Adaptive Instance Normalization for Fine-grained Style Similarity
abstract
We present ALADIN (All Layer AdaIN); a novel architecture for searching images based on the similarity of their artistic style. Representation learning is critical to visual search, where distance in the learned search embedding reflects image similarity. Learning an embedding that discriminates fine-grained variations in style is hard, due to the difficulty of defining and labelling style. ALADIN takes a weakly supervised approach to learning a representation for fine-grained style similarity of digital artworks, leveraging BAM-FG, a novel large-scale dataset of user generated content groupings gathered from the web. ALADIN sets a new state of the art accuracy for style-based visual search over both coarse labelled style data (BAM) and BAM-FG; a new 2.62 million image dataset of 310,000 fine-grained style groupings also contributed by this work.
Dan Ruta, Saeid Motiian, Baldo Faieta, Zhe Lin 0001, Hailin Jin, Alex Filipkowski, Andrew Gilbert, John P. Collomosse
ICCV5
2021 Compositional Sketch Search
abstract
We present an algorithm for searching image collections using free-hand sketches that describe the appearance and relative positions of multiple objects1Sketch based image retrieval (SBIR) methods predominantly match queries containing a single, dominant object invariant to its position within an image. Our work exploits drawings as a concise and intuitive representation for specifying entire scene compositions. We train a convolutional neural network (CNN) to encode masked visual features from sketched objects, pooling these into a spatial descriptor encoding the spatial relationships and appearances of objects in the composition. Training the CNN backbone as a Siamese network under triplet loss yields a metric search embedding for measuring compositional similarity which may be efficiently leveraged for visual search by applying product quantization.
Alexander Black 0001, Tu Bui, Long Mai, Hailin Jin, John P. Collomosse
ICIP4
2021 A Multi-Implicit Neural Representation for Fonts
abstract
Fonts are ubiquitous across documents and come in a variety of styles. They are either represented in a native vector format or rasterized to produce fixed resolution images. In the first case, the non-standard representation prevents benefiting from latest network architectures for neural representations; while, in the latter case, the rasterized representation, when encoded via networks, results in loss of data fidelity, as font-specific discontinuities like edges and corners are difficult to represent using neural networks. Based on the observation that complex fonts can be represented by a superposition of a set of simpler occupancy functions, we introduce multi-implicits to represent fonts as a permutation-invariant set of learned implict functions, without losing features (e.g., edges and corners). However, while multi-implicits locally preserve font features, obtaining supervision in the form of ground truth multi-channel signals is a problem in itself. Instead, we propose how to train such a representation with only local supervision, while the proposed neural architecture directly finds globally consistent multi-implicits for font families. We extensively evaluate the proposed representation for various tasks including reconstruction, interpolation, and synthesis to demonstrate clear advantages with existing alternatives. Additionally, the representation naturally enables glyph completion, wherein a single characteristic font is used to synthesize a whole font family in the target style.
Pradyumna Reddy, Matthew Fisher, Hailin Jin, Niloy J. Mitra
NeurIPS5
2021 Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos
abstract
We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision on a large corpus of videos with accompanying transcribed narrations. To achieve this goal, we propose a multilayer cross-modal attention network that enables effective optimization of a contrastive loss during training. We introduce a divided strategy that alternates between computing inter- and intra-modal attention across the visual and natural language modalities, which allows effective training via directly contrasting the two modalities' representations. We demonstrate the effectiveness of our approach by self-training on the HowTo100M instructional video dataset and evaluating on a newly collected dataset of localized described interactions in the YouCook2 dataset. We show that our approach outperforms alternative baselines, including shallow co-attention and full cross-modal attention. We also apply our approach to grounding phrases in images with weak supervision on Flickr30K and show that stacking multiple attention layers is effective and, when combined with a word-to-region loss, achieves state of the art on recall-at-one and pointing hand accuracies.
Reuben Tan, Bryan A. Plummer, Kate Saenko, Hailin Jin, Bryan C. Russell
NeurIPS4
2021 Neural architecture search for deep image prior
Kary Ho, Andrew Gilbert, Hailin Jin, John P. Collomosse
Comput. Graph.3
2021 Black-Box Diagnosis and Calibration on GAN Intra-Mode Collapse: A Pilot Study
abstract
Generative adversarial networks (GANs) nowadays are capable of producing images of incredible realism. Two concerns raised are whether the state-of-the-art GAN’s learned distribution still suffers from mode collapse and what to do if so. Existing diversity tests of samples from GANs are usually conducted qualitatively on a small scale and/or depend on the access to original training data as well as the trained model parameters. This article explores GAN intra-mode collapse and calibrates that in a novel black-box setting: access to neither training data nor the trained model parameters is assumed. The new setting is practically demanded yet rarely explored and significantly more challenging. As a first stab, we devise a set of statistical tools based on sampling that can visualize, quantify, and rectify intra-mode collapse . We demonstrate the effectiveness of our proposed diagnosis and calibration techniques, via extensive simulations and experiments, on unconditional GAN image generation (e.g., face and vehicle). Our study reveals that the intra-mode collapse is still a prevailing problem in state-of-the-art GANs and the mode collapse is diagnosable and calibratable in black-box settings. Our codes are available at https://github.com/VITA-Group/BlackBoxGANCollapse .
Zhenyu Wu 0002, Ye Yuan 0012, Jianming Zhang 0001, Zhangyang Wang, Hailin Jin
ACM Trans. Multim. Comput. Commun. Appl.6
2020 Spatial Class Distribution Shift in Unsupervised Domain Adaptation: Local Alignment Comes to Rescue
Safa Cicek, Ning Xu 0007, Hailin Jin, Stefano Soatto
ACCV (3)4
2020 DeepVoxels++: Enhancing the Fidelity of Novel View Synthesis from 3D Voxel Embeddings
Tong He 0002, John P. Collomosse, Hailin Jin, Stefano Soatto
ACCV (1)3
2020 Steering Self-Supervised Feature Learning Beyond Local Pixel Statistics
abstract
We introduce a novel principle for self-supervised feature learning based on the discrimination of specific transformations of an image. We argue that the generalization capability of learned features depends on what image neighborhood size is sufficient to discriminate different image transformations: The larger the required neighborhood size and the more global the image statistics that the feature can describe. An accurate description of global image statistics allows to better represent the shape and configuration of objects and their context, which ultimately generalizes better to new tasks such as object classification and detection. This suggests a criterion to choose and design image transformations. Based on this criterion, we introduce a novel image transformation that we call limited context inpainting (LCI). This transformation inpaints an image patch conditioned only on a small rectangular pixel boundary (the limited context). Because of the limited boundary information, the inpainter can learn to match local pixel statistics, but is unlikely to match the global statistics of the image. We claim that the same principle can be used to justify the performance of transformations such as image rotations and warping. Indeed, we demonstrate experimentally that learning to discriminate transformations such as LCI, image warping and rotations, yields features with state of the art generalization capabilities on several datasets such as Pascal VOC, STL-10, CelebA, and ImageNet. Remarkably, our trained features achieve a performance on Places on par with features trained through supervised learning with ImageNet labels.
Simon Jenni, Hailin Jin, Paolo Favaro
CVPR2
2020 Screencast Tutorial Video Understanding
abstract
Screencast tutorials are videos created by people to teach how to use software applications or demonstrate procedures for accomplishing tasks. It is very popular for both novice and experienced users to learn new skills, compared to other tutorial media such as text, because of the visual guidance and the ease of understanding. In this paper, we propose visual understanding of screencast tutorials as a new research problem to the computer vision community. We collect a new dataset of Adobe Photoshop video tutorials and annotate it with both low-level and high-level semantic labels. We introduce a bottom-up pipeline to understand Photoshop video tutorials. We leverage state-of-the-art object detection algorithms with domain specific visual cues to detect important events in a video tutorial and segment it into clips according to the detected events. We propose a visual cue reasoning algorithm for two high-level tasks: video retrieval and video captioning. We conduct extensive evaluations of the proposed pipeline. Experimental results show that it is effective in terms of understanding video tutorials. We believe our work will serves as a starting point for future research on this important application domain of video understanding.
Seokhwan Kim, Hailin Jin, Yun Fu 0001
CVPR5
2020 Superpixel Segmentation With Fully Convolutional Networks
abstract
In computer vision, superpixels have been widely used as an effective way to reduce the number of image primitives for subsequent processing. But only a few attempts have been made to incorporate them into deep neural networks. One main reason is that the standard convolution operation is defined on regular grids and becomes inefficient when applied to superpixels. Inspired by an initialization strategy commonly adopted by traditional superpixel algorithms, we present a novel method that employs a simple fully convolutional network to predict superpixels on a regular image grid. Experimental results on benchmark datasets show that our method achieves state-of-the-art superpixel segmentation performance while running at about 50fps. Based on the predicted superpixels, we further develop a downsampling/upsampling scheme for deep networks with the goal of generating high-resolution outputs for dense prediction tasks. Specifically, we modify a popular network architecture for stereo matching to simultaneously predict superpixels and disparities. We show that improved disparity estimation accuracy can be obtained on public datasets.
Fengting Yang, Hailin Jin, Zihan Zhou 0001
CVPR3
2020 Video Question Answering on Screencast Tutorials
abstract
This paper presents a new video question answering task on screencast tutorials. We introduce a dataset including question, answer and context triples from the tutorial videos for a software. Unlike other video question answering works, all the answers in our dataset are grounded to the domain knowledge base. An one-shot recognition algorithm is designed to extract the visual cues, which helps enhance the performance of video question answering. We also propose several baseline neural network architectures based on various aspects of video contexts from the dataset. The experimental results demonstrate that our proposed models significantly improve the question answering performances by incorporating multi-modal contexts and domain knowledge.
Wentian Zhao, Seokhwan Kim, Hailin Jin
IJCAI4
2020 Geo-PIFu: Geometry and Pixel Aligned Implicit Functions for Single-view Human Reconstruction
abstract
We propose Geo-PIFu, a method to recover a 3D mesh from a monocular color image of a clothed person. Our method is based on a deep implicit function-based representation to learn latent voxel features using a structure-aware 3D U-Net, to constrain the model in two ways: first, to resolve feature ambiguities in query point encoding, second, to serve as a coarse human shape proxy to regularize the high-resolution mesh and encourage global shape regularity. We show that, by both encoding query points and constraining global shape using latent voxel features, the reconstruction we obtain for clothed human meshes exhibits less shape distortion and improved surface details compared to competing methods. We evaluate Geo-PIFu on a recent human mesh public dataset that is 10x larger than the private commercial dataset used in PIFu and previous derivative work. On average, we exceed the state of the art by 42.7% reduction in Chamfer and Point-to-Surface Distances, and 19.4% reduction in normal estimation errors.
Tong He 0002, John P. Collomosse, Hailin Jin, Stefano Soatto
NeurIPS3
2020 Adversarial training for fast arbitrary style transfer
Zheng Xu 0002, Kimberly Wilber, Aaron Hertzmann, Hailin Jin
Comput. Graph.5
2020 Product Quantization Network for Fast Visual Search
Jingjing Meng, Hailin Jin, Junsong Yuan 0001
Int. J. Comput. Vis.4
2020 Visual Font Pairing
abstract
This paper introduces the problem of automatic font pairing. Font pairing is an important design task that is difficult for novices. Given a font selection for one part of a document (e.g., header), our goal is to recommend a font to be used in another part (e.g., body) such that the two fonts used together look visually pleasing. There are three main challenges in font pairing. First, this is a fine-grained problem, in which the subtle distinctions between fonts may be important. Second, rules and conventions of font pairing given by human experts are difficult to formalize. Third, font pairing is an asymmetric problem in that the roles played by header and body fonts are not interchangeable. To address these challenges, we propose automatic font pairing through learning visual relationships from large-scale human-generated font pairs. We introduce a new database for font pairing constructed from millions of PDF documents available on the Internet. We propose two font pairing algorithms: dual-space k-NN and asymmetric similarity metric learning (ASML). These two methods automatically learn fine-grained relationships from large-scale data. We also investigate several baseline methods based on the rules from professional designers. Experiments and user studies demonstrate the effectiveness of our proposed dataset and methods.
Shuhui Jiang, Aaron Hertzmann, Hailin Jin, Yun Fu 0001
IEEE Trans. Multim.4
2019 LiveSketch: Query Perturbations for Guided Sketch-Based Visual Search
abstract
LiveSketch is a novel algorithm for searching large image collections using hand-sketched queries. LiveSketch tackles the inherent ambiguity of sketch search by creating visual suggestions that augment the query as it is drawn, making query specification an iterative rather than one-shot process that helps disambiguate users' search intent. Our technical contributions are: a triplet convnet architecture that incorporates an RNN based variational autoencoder to search for images using vector (stroke-based) queries; real-time clustering to identify likely search intents (and so, targets within the search embedding); and the use of backpropagation from those targets to perturb the input stroke sequence, so suggesting alterations to the query in order to guide the search. We show improvements in accuracy and time-to-task over contemporary baselines using a 67M image corpus.
John P. Collomosse, Tu Bui, Hailin Jin
CVPR3
2019 Learning Video Representations From Correspondence Proposals
abstract
Correspondences between frames encode rich information about dynamic content in videos. However, it is challenging to effectively capture and learn those due to their irregular structure and complex dynamics. In this paper, we propose a novel neural network that learns video representations by aggregating information from potential correspondences. This network, named CPNet, can learn evolving 2D fields with temporal consistency. In particular, it can effectively learn representations for videos by mixing appearance and long-range motion with an RGB-only input. We provide extensive ablation experiments to validate our model. CPNet shows stronger performance than existing methods on Kinetics and achieves the state-of-the-art performance on Something-Something and Jester. We provide analysis towards the behavior of our model and show its robustness to errors in proposals.
Xingyu Liu 0001, Joon-Young Lee, Hailin Jin
CVPR3
2019 Large-Scale Tag-Based Font Retrieval With Generative Feature Learning
abstract
Font selection is one of the most important steps in a design workflow. Traditional methods rely on ordered lists which require significant domain knowledge and are often difficult to use even for trained professionals. In this paper, we address the problem of large-scale tag-based font retrieval which aims to bring semantics to the font selection process and enable people without expert knowledge to use fonts effectively. We collect a large-scale font tagging dataset of high-quality professional fonts. The dataset contains nearly 20,000 fonts, 2,000 tags, and hundreds of thousands of font-tag relations. We propose a novel generative feature learning algorithm that leverages the unique characteristics of fonts. The key idea is that font images are synthetic and can therefore be controlled by the learning algorithm. We design an integrated rendering and learning process so that the visual feature from one image can be used to reconstruct another image with different text. The resulting feature captures important font design details while is robust to nuisance factors such as text. We propose a novel attention mechanism to re-weight the visual feature for joint visual-text modeling. We combine the feature and the attention mechanism in a novel recognition-retrieval model. Experimental results show that our method significantly outperforms the state-of-the-art for the important problem of large-scale tag-based font retrieval.
Ning Xu 0007, Hailin Jin, Jiebo Luo 0001
ICCV4
2019 An Internal Learning Approach to Video Inpainting
abstract
We propose a novel video inpainting algorithm that simultaneously hallucinates missing appearance and motion (optical flow) information, building upon the recent 'Deep Image Prior' (DIP) that exploits convolutional network architectures to enforce plausible texture in static images. In extending DIP to video we make two important contributions. First, we show that coherent video inpainting is possible without a priori training. We take a generative approach to inpainting based on internal (within-video) learning without reliance upon an external corpus of visual data to train a one-size-fits-all model for the large space of general videos. Second, we show that such a framework can jointly generate both appearance and flow, whilst exploiting these complementary modalities to ensure mutual consistency. We show that leveraging appearance statistics specific to each video achieves visually plausible results whilst handling the challenging problem of long-term consistency.
Long Mai, Hailin Jin, Ning Xu 0007, John P. Collomosse
ICCV3
2018 Learning from PhotoShop Operation Videos: The PSOV Dataset
Jingchun Cheng, Han-Kai Hsu, Hailin Jin, Shengjin Wang, Ming-Hsuan Yang 0001
ACCV (4)4
2018 Disentangling Structure and Aesthetics for Style-Aware Image Completion
abstract
Content-aware image completion or in-painting is a fundamental tool for the correction of defects or removal of objects in images. We propose a non-parametric in-painting algorithm that enforces both structural and aesthetic (style) consistency within the resulting image. Our contributions are two-fold: (1) we explicitly disentangle image structure and style during patch search and selection to ensure a visually consistent look and feel within the target image. (2) we perform adaptive stylization of patches to conform the aesthetics of selected patches to the target image, so harmonizing the integration of selected patches into the final composition. We show that explicit consideration of visual style during in-painting delivers excellent qualitative and quantitative results across the varied image styles and content, over the Places2 scene photographic dataset and a challenging new in-painting dataset of artwork derived from BAM!
Andrew Gilbert, John P. Collomosse, Hailin Jin, Brian L. Price
CVPR3
2018 Multi-Task Adversarial Network for Disentangled Feature Learning
abstract
We address the problem of image feature learning for the applications where multiple factors exist in the image generation process and only some factors are of our interest. We present a novel multi-task adversarial network based on an encoder-discriminator-generator architecture. The encoder extracts a disentangled feature representation for the factors of interest. The discriminators classify each of the factors as individual tasks. The encoder and the discriminators are trained cooperatively on factors of interest, but in an adversarial way on factors of distraction. The generator provides further regularization on the learned feature by reconstructing images with shared factors as the input image. We design a new optimization scheme to stabilize the adversarial optimization process when multiple distributions need to be aligned. The experiments on face recognition and font recognition tasks show that our method outperforms the state-of-the-art methods in terms of both recognizing the factors of interest and generalization to images with unseen variations.
Yang Liu 0105, Hailin Jin, Ian J. Wassell
CVPR3
2018 "Factual" or "Emotional": Stylized Image Captioning with Adaptive Learning and Attention
Zhongping Zhang, Quanzeng You, Hailin Jin, Jiebo Luo 0001
ECCV (10)6
2018 What Do I Annotate Next? An Empirical Study of Active Learning for Action Localization
Fabian Caba Heilbron, Joon-Young Lee, Hailin Jin, Bernard Ghanem
ECCV (11)3
2018 Interactive Boundary Prediction for Object Selection
Hoang Le, Long Mai, Brian L. Price, Scott Cohen, Hailin Jin, Feng Liu 0015
ECCV (14)5
2018 Synthetically Supervised Feature Learning for Scene Text Recognition
Yang Liu 0105, Hailin Jin, Ian J. Wassell
ECCV (5)3
2018 Towards Privacy-Preserving Visual Recognition via Adversarial Training: A Pilot Study
Zhenyu Wu 0002, Zhangyang Wang, Hailin Jin
ECCV (16)4
2018 Product Quantization Network for Fast Image Retrieval
Junsong Yuan 0001, Hailin Jin
ECCV (1)4
2018 Characterizing User Skills from Application Usage Traces with Hierarchical Attention Recurrent Networks
abstract
Predicting users’ proficiencies is a critical component of AI-powered personal assistants. This article introduces a novel approach for the prediction based on users’ diverse, noisy, and passively generated application usage histories. We propose a novel bi-directional recurrent neural network with hierarchical attention mechanism to extract sequential patterns and distinguish informative traces from noise. Our model is able to attend to the most discriminative actions and sessions to make more accurate and directly interpretable predictions while requiring 50× less training data than the state-of-the-art sequential learning approach. We evaluate our model with two large scale datasets collected from 68K Photoshop users: a digital design skill dataset where the user skill is determined by the quality of the end products and a software skill dataset where users self-disclose their software usage skill levels. The empirical results demonstrate our model’s superior performance compared to existing user representation learning techniques that leverage action frequencies and sequential patterns. In addition, we qualitatively illustrate the model’s significant interpretative power. The proposed approach is broadly relevant to applications that generate user time-series analytics.
Longqi Yang 0001, Hailin Jin, Matthew Hoffman 0001, Deborah Estrin
ACM Trans. Intell. Syst. Technol.3
2017 Visual Sentiment Analysis by Attending on Local Image Regions
abstract
Visual sentiment analysis, which studies the emotional response of humans on visual stimuli such as images and videos, has been an interesting and challenging problem. It tries to understand the high-level content of visual data. The success of current models can be attributed to the development of robust algorithms from computer vision. Most of the existing models try to solve the problem by proposing either robust features or more complex models. In particular, visual features from the whole image or video are the main proposed inputs. Little attention has been paid to local areas, which we believe is pretty relevant to human's emotional response to the whole image. In this work, we study the impact of local image regions on visual sentiment analysis. Our proposed model utilizes the recent studied attention mechanism to jointly discover the relevant local regions and build a sentiment classifier on top of these local regions. The experimental results suggest that 1) our model is capable of automatically discovering sentimental local regions of given images and 2) it outperforms existing state-of-the-art algorithms to visual sentiment analysis.
Quanzeng You, Hailin Jin, Jiebo Luo 0001
AAAI2
2017 Multiple Instance Visual-Semantic Embedding
Zhou Ren, Hailin Jin, Zhe Lin 0001, Alan L. Yuille
BMVC2
2017 Spatial-Semantic Image Search by Visual Feature Synthesis
abstract
The performance of image retrieval has been improved tremendously in recent years through the use of deep feature representations. Most existing methods, however, aim to retrieve images that are visually similar or semantically relevant to the query, irrespective of spatial configuration. In this paper, we develop a spatial-semantic image search technology that enables users to search for images with both semantic and spatial constraints by manipulating concept text-boxes on a 2D query canvas. We train a convolutional neural network to synthesize appropriate visual features that captures the spatial-semantic constraints from the user canvas query. We directly optimize the retrieval performance of the visual features when training our deep neural network. These visual features then are used to retrieve images that are both spatially and semantically relevant to the user query. The experiments on large-scale datasets such as MS-COCO and Visual Genome show that our method outperforms other baseline and state-of-the-art methods in spatial-semantic image search.
Long Mai, Hailin Jin, Zhe Lin 0001, Jonathan Brandt, Feng Liu 0015
CVPR2
2017 Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks
abstract
Indoor scene understanding is central to applications such as robot navigation and human companion assistance. Over the last years, data-driven deep neural networks have outperformed many traditional approaches thanks to their representation learning capabilities. One of the bottlenecks in training for better representations is the amount of available per-pixel ground truth data that is required for core scene understanding tasks such as semantic segmentation, normal prediction, and object boundary detection. To address this problem, a number of works proposed using synthetic data. However, a systematic study of how such synthetic data is generated is missing. In this work, we introduce a large-scale synthetic dataset with 500K physically-based rendered images from 45K realistic 3D indoor scenes. We study the effects of rendering methods and scene lighting on training for three computer vision tasks: surface normal prediction, semantic segmentation, and object boundary detection. This study provides insights into the best practices for training with synthetic data (more realistic rendering is worth it) and shows that pretraining with our new synthetic dataset can improve results beyond the current state of the art on all three tasks.
Yinda Zhang 0001, Shuran Song, Ersin Yumer, Manolis Savva, Joon-Young Lee, Hailin Jin, Thomas A. Funkhouser
CVPR6
2017 Sketching with Style: Visual Search with Sketches and Aesthetic Context
abstract
We propose a novel measure of visual similarity for image retrieval that incorporates both structural and aesthetic (style) constraints. Our algorithm accepts a query as sketched shape, and a set of one or more contextual images specifying the desired visual aesthetic. A triplet network is used to learn a feature embedding capable of measuring style similarity independent of structure, delivering significant gains over previous networks for style discrimination. We incorporate this model within a hierarchical triplet network to unify and learn a joint space from two discriminatively trained streams for style and structure. We demonstrate that this space enables, for the first time, style-constrained sketch search over a diverse domain of digital artwork comprising graphics, paintings and drawings. We also briefly explore alternative query modalities.
John P. Collomosse, Tu Bui, Kimberly Wilber, Hailin Jin
ICCV5
2017 BAM! The Behance Artistic Media Dataset for Recognition Beyond Photography
abstract
Computer vision systems are designed to work well within the context of everyday photography. However, artists often render the world around them in ways that do not resemble photographs. Artwork produced by people is not constrained to mimic the physical world, making it more challenging for machines to recognize.,,This work is a step toward teaching machines how to categorize images in ways that are valuable to humans. First, we collect a large-scale dataset of contemporary artwork from Behance, a website containing millions of portfolios from professional and commercial artists. We annotate Behance imagery with rich attribute labels for content, emotions, and artistic media. Furthermore, we carry out baseline experiments to show the value of this dataset for artistic style prediction, for improving the generality of existing object classifiers, and for the study of visual domain adaptation. We believe our Behance Artistic Media dataset will be a good starting point for researchers wishing to study artistic imagery and relevant problems. This dataset can be found at https://bam-dataset.org/.
Kimberly Wilber, Hailin Jin, Aaron Hertzmann, John P. Collomosse, Serge J. Belongie
ICCV3
2017 6-DOF VR videos with a single 360-camera
abstract
Recent breakthroughs in consumer level virtual reality (VR) headsets are creating a growing user-base in demand for immersive, full 3D VR experiences. While monoscopic 360-videos are perhaps the most prevalent type of content for VR headsets, they lack 3D information and thus cannot be viewed with full 6 degree-of-freedom (DOF). We present an approach that addresses this limitation via a novel warping algorithm that can synthesize new views both with rotational and translational motion of the viewpoint. This enables the ability to perform VR playback of input monoscopic 360-videos files in full stereo with full 6-DOF of head motion. Our method synthesizes novel views for each eye in accordance with the 6-DOF motion of the headset. Our solution tailors standard structure-from-motion and dense reconstruction algorithms to work accurately for 360-videos and is optimized for GPUs to achieve VR frame rates (>120 fps). We demonstrate the effectiveness our approach on a variety of videos with interesting content.
Jingwei Huang 0001, Duygu Ceylan, Hailin Jin
VR4
2016 Building a Large Scale Dataset for Image Emotion Recognition: The Fine Print and The Benchmark
abstract
Psychological research results have confirmed that people can have different emotional reactions to different visual stimuli. Several papers have been published on the problem of visual emotion analysis. In particular, attempts have been made to analyze and predict people's emotional reaction towards images. To this end, different kinds of hand-tuned features are proposed. The results reported on several carefully selected and labeled small image data sets have confirmed the promise of such features. While the recent successes of many computer vision related tasks are due to the adoption of Convolutional Neural Networks (CNNs), visual emotion analysis has not achieved the same level of success. This may be primarily due to the unavailability of confidently labeled and relatively large image data sets for visual emotion analysis. In this work, we introduce a new data set, which started from 3+ million weakly labeled images of different emotions and ended up 30 times as large as the current largest publicly available visual emotion data set. We hope that this data set encourages further research on visual emotion analysis. We also perform extensive benchmarking analyses on this large data set using the state of the art methods including CNNs.
Quanzeng You, Jiebo Luo 0001, Hailin Jin, Jianchao Yang
AAAI3
2016 Composition-Preserving Deep Photo Aesthetics Assessment
abstract
Photo aesthetics assessment is challenging. Deep convolutional neural network (ConvNet) methods have recently shown promising results for aesthetics assessment. The performance of these deep ConvNet methods, however, is often compromised by the constraint that the neural network only takes the fixed-size input. To accommodate this requirement, input images need to be transformed via cropping, scaling, or padding, which often damages image composition, reduces image resolution, or causes image distortion, thus compromising the aesthetics of the original images. In this paper, we present a composition-preserving deep Con-vNet method that directly learns aesthetics features from the original input images without any image transformations. Specifically, our method adds an adaptive spatial pooling layer upon the regular convolution and pooling layers to directly handle input images with original sizes and aspect ratios. To allow for multi-scale feature extraction, we develop the Multi-Net Adaptive Spatial Pooling ConvNet architecture which consists of multiple sub-networks with different adaptive spatial pooling sizes and leverage a scene-based aggregation layer to effectively combine the predictions from multiple sub-networks. Our experiments on the large-scale aesthetics assessment benchmark (AVA [29]) demonstrate that our method can significantly improve the state-of-the-art results in photo aesthetics assessment.
Long Mai, Hailin Jin, Feng Liu 0015
CVPR2
2016 Image Captioning with Semantic Attention
abstract
Automatically generating a natural language description of an image has attracted interests recently both because of its importance in practical applications and because it connects two major artificial intelligence fields: computer vision and natural language processing. Existing approaches are either top-down, which start from a gist of an image and convert it into words, or bottom-up, which come up with words describing various aspects of an image and then combine them. In this paper, we propose a new algorithm that combines both approaches through a model of semantic attention. Our algorithm learns to selectively attend to semantic concept proposals and fuse them into hidden states and outputs of recurrent neural networks. The selection and fusion form a feedback connecting the top-down and bottom-up computation. We evaluate our algorithm on two public benchmarks: Microsoft COCO and Flickr30K. Experimental results show that our algorithm significantly outperforms the state-of-the-art approaches consistently across different evaluation metrics.
Quanzeng You, Hailin Jin, Jiebo Luo 0001
CVPR2
2016 Joint Image-Text Representation by Gaussian Visual-Semantic Embedding
abstract
How to jointly represent images and texts is important for tasks involving both modalities. Visual-semantic embedding models have been recently proposed and shown to be effective. The key idea is that by learning a mapping from images into a semantic text space, the algorithm is able to learn a compact and effective joint representation. However, existing approaches simply map each text concept to a single point in the semantic space. Mapping instead to a density distribution provides many interesting advantages, including better capturing uncertainty about each text concept, and enabling better geometric interpretation of concepts such as inclusion, intersection, etc. In this work, we present a novel Gaussian Visual-Semantic Embedding (GVSE) model, which leverages the visual information to model text concepts as Gaussian distributions in semantic space. Experiments in two tasks, image classification and text-based image retrieval on the large scale MIT Places205 dataset, have demonstrated the superiority of our method over existing approaches, with higher accuracy and better robustness.
Zhou Ren, Hailin Jin, Zhe Lin 0001, Alan L. Yuille
ACM Multimedia2
2016 Robust Visual-Textual Sentiment Analysis: When Attention meets Tree-structured Recursive Neural Networks
abstract
Sentiment analysis is crucial for extracting social signals from social media content. Due to huge variation in social media, the performance of sentiment classifiers using single modality (visual or textual) still lags behind satisfaction. In this paper, we propose a new framework that integrates textual and visual information for robust sentiment analysis. Different from previous work, we believe visual and textual information should be treated jointly in a structural fashion. Our system first builds a semantic tree structure based on sentence parsing, aimed at aligning textual words and image regions for accurate analysis. Next, our system learns a robust joint visual-textual semantic representation by incorporating 1) an attention mechanism with LSTM (long short term memory) and 2) an auxiliary semantic learning task. Extensive experimental results on several known data sets show that our method outperforms existing the state-of-the-art joint models in sentiment analysis. We also investigate different tree-structured LSTM (T-LSTM) variants and analyze the effect of the attention mechanism in order to provide deeper insight on how the attention mechanism helps the learning of the joint visual-textual sentiment classifier.
Quanzeng You, Liangliang Cao, Hailin Jin, Jiebo Luo 0001
ACM Multimedia3
2016 Cross-modality Consistent Regression for Joint Visual-Textual Sentiment Analysis of Social Multimedia
abstract
Sentiment analysis of online user generated content is important for many social media analytics tasks. Researchers have largely relied on textual sentiment analysis to develop systems to predict political elections, measure economic indicators, and so on. Recently, social media users are increasingly using additional images and videos to express their opinions and share their experiences. Sentiment analysis of such large-scale textual and visual content can help better extract user sentiments toward events or topics. Motivated by the needs to leverage large-scale social multimedia content for sentiment analysis, we propose a cross-modality consistent regression (CCR) model, which is able to utilize both the state-of-the-art visual and textual sentiment analysis techniques. We first fine-tune a convolutional neural network (CNN) for image sentiment analysis and train a paragraph vector model for textual sentiment analysis. On top of them, we train our multi-modality regression model. We use sentimental queries to obtain half a million training samples from Getty Images. We have conducted extensive experiments on both machine weakly labeled and manually labeled image tweets. The results show that the proposed model can achieve better performance than the state-of-the-art textual and visual sentiment analysis algorithms alone.
Quanzeng You, Jiebo Luo 0001, Hailin Jin, Jianchao Yang
WSDM3
2015 Robust Image Sentiment Analysis Using Progressively Trained and Domain Transferred Deep Networks
abstract
Sentiment analysis of online user generated content is important for many social media analytics tasks. Researchers have largely relied on textual sentiment analysis to develop systems to predict political elections, measure economic indicators, and so on. Recently, social media users are increasingly using images and videos to express their opinions and share their experiences. Sentiment analysis of such large scale visual content can help better extract user sentiments toward events or topics, such as those in image tweets, so that prediction of sentiment from visual content is complementary to textual sentiment analysis. Motivated by the needs in leveraging large scale yet noisy training data to solve the extremely challenging problem of image sentiment analysis, we employ Convolutional Neural Networks (CNN). We first design a suitable CNN architecture for image sentiment analysis. We obtain half a million training samples by using a baseline sentiment algorithm to label Flickr images. To make use of such noisy machine labeled data, we employ a progressive strategy to fine-tune the deep network. Furthermore, we improve the performance on Twitter images by inducing domain transfer with a small number of manually labeled Twitter images. We have conducted extensive experiments on manually labeled Twitter images. The results show that the proposed CNN can achieve better performance in image sentiment analysis than competing algorithms.
Quanzeng You, Jiebo Luo 0001, Hailin Jin, Jianchao Yang
AAAI3
2015 Collaborative feature learning from social media
abstract
Image feature representation plays an essential role in image recognition and related tasks. The current state-of-the-art feature learning paradigm is supervised learning from labeled data. However, this paradigm requires large-scale category labels, which limits its applicability to domains where labels are hard to obtain. In this paper, we propose a new data-driven feature learning paradigm which does not rely on category labels. Instead, we learn from user behavior data collected on social media. Concretely, we use the image relationship discovered in the latent space from the user behavior data to guide the image feature learning. We collect a large-scale image and user behavior dataset from Behance.net. The dataset consists of 1.9 million images and over 300 million view records from 1.9 million users. We validate our feature learning paradigm on this dataset and find that the learned feature significantly outperforms the state-of-the-art image features in learning better image similarities. We also show that the learned feature performs competitively on various recognition benchmarks.
Hailin Jin, Jianchao Yang, Zhe Lin 0001
CVPR2
2015 Fine-grained recognition without part annotations
abstract
Scaling up fine-grained recognition to all domains of fine-grained objects is a challenge the computer vision community will need to face in order to realize its goal of recognizing all object categories. Current state-of-the-art techniques rely heavily upon the use of keypoint or part annotations, but scaling up to hundreds or thousands of domains renders this annotation cost-prohibitive for all but the most important categories. In this work we propose a method for fine-grained recognition that uses no part annotations. Our method is based on generating parts using co-segmentation and alignment, which we combine in a discriminative mixture. Experimental results show its efficacy, demonstrating state-of-the-art results even when compared to methods that use part annotations during training.
Jonathan Krause, Hailin Jin, Jianchao Yang, Li Fei-Fei 0001
CVPR2
2015 DeepFont: A System for Font Recognition and Similarity
abstract
We develop the DeepFont system, a large-scale learning-based solution for automatic font identification, organization and selection. In this proposed technical demonstration, we will give our audience a tour to the DeepFont system, with the focus on its impacts on real consumer products, including but not limited to: 1) a cloud-based iOS App for font recognition; 2) a web-based tool for font similarity evaluation and discovery.
Zhangyang Wang, Jianchao Yang, Hailin Jin, Jonathan Brandt, Eli Shechtman, Aseem Agarwala, Yuyan Song, Joseph Hsieh, Sarah Kong, Thomas S. Huang
ACM Multimedia3
2015 DeepFont: Identify Your Font from An Image
abstract
As font is one of the core design concepts, automatic font identification and similar font suggestion from an image or photo has been on the wish list of many designers. We study the Visual Font Recognition (VFR) problem [4] LFE, and advance the state-of-the-art remarkably by developing the DeepFont system. First of all, we build up the first available large-scale VFR dataset, named AdobeVFR, consisting of both labeled synthetic data and partially labeled real-world data. Next, to combat the domain mismatch between available training and testing data, we introduce a Convolutional Neural Network (CNN) decomposition approach, using a domain adaptation technique based on a Stacked Convolutional Auto-Encoder (SCAE) that exploits a large corpus of unlabeled real-world text images combined with synthetic data preprocessed in a specific way. Moreover, we study a novel learning-based model compression approach, in order to reduce the DeepFont model size without sacrificing its performance. The DeepFont system achieves an accuracy of higher than 80% (top-5) on our collected dataset, and also produces a good font similarity measure for font selection and suggestion. We also achieve around 6 times compression of the model without any visible loss of recognition accuracy.
Zhangyang Wang, Jianchao Yang, Hailin Jin, Eli Shechtman, Aseem Agarwala, Jonathan Brandt, Thomas S. Huang
ACM Multimedia3
2015 Joint Visual-Textual Sentiment Analysis with Deep Neural Networks
abstract
Sentiment analysis of online user generated content is important for many social media analytics tasks. Researchers have largely relied on textual sentiment analysis to develop systems to predict political elections, measure economic indicators, and so on. Recently, social media users are increasingly using additional images and videos to express their opinions and share their experiences. Sentiment analysis of such large-scale textual and visual content can help better extract user sentiments toward events or topics. Motivated by the needs to leverage large-scale social multimedia content for sentiment analysis, we utilize both the state-of-the-art visual and textual sentiment analysis techniques for joint visual-textual sentiment analysis. We first fine-tune a convolutional neural network (CNN) for image sentiment analysis and train a paragraph vector model for textual sentiment analysis. We have conducted extensive experiments on both machine weakly labeled and manually labeled image tweets. The results show that joint visual-textual features can achieve the state-of-the-art performance than textual and visual sentiment analysis algorithms alone.
Quanzeng You, Jiebo Luo 0001, Hailin Jin, Jianchao Yang
ACM Multimedia3
2015 Selective Pooling Vector for Fine-Grained Recognition
abstract
We propose a new framework for image recognition by selectively pooling local visual descriptors, and show its superior discriminative power on fine-grained image classification tasks. The representation is based on selecting the most confident local descriptors for nonlinear function learning using a linear approximation in an embedded higher dimensional space. The advantage of our Selective Pooling Vector over the previous state-of-the-art Super Vector and Fisher Vector representations, is that it ensures a more accurate learning function, which proves to be important for classifying details in fine-grained image recognition. Our experimental results corroborate this claim: with a simple linear SVM as the classifier, the selective pooling vector achieves significant performance gains on standard benchmark datasets for various fine-grained tasks such as the CMU Multi-PIE dataset for face recognition, the Caltech-UCSD Bird dataset and the Stanford Dogs dataset for fine-grained object categorization. On all datasets we outperform the state of the arts and boost the recognition rates to 96.4%, 48.9%, 52.0% respectively.
Jianchao Yang, Hailin Jin, Eli Shechtman, Jonathan Brandt, Tony X. Han
WACV3
2015 Tree-Based Locally Linear Regression for Image Denoising
abstract
We present a new patch-based approach for image denoising that combines similar patches in the same image and from a set of training images. The key idea of our method is that we can partition the training samples according to the clean patches and efficiently learn a denoising operator for each partition. Given a noisy patch, we use self-similarity to compute an initial denoising result which is used to locate the relevant partitions. We apply the corresponding learned denoising operator to the original noisy patch. Our method does not suffer either from the blurring effect that commonly exists in self-similarity based methods or from the training size problem that is associated with training-based methods. We evaluate our method on three benchmark datasets as well as real mobile images. Experimental results show that our approach consistently outperforms BM3D in terms of both peak signal-to-noise ratio and visual quality.
Xin Lu 0006, Zhe Lin 0001, Hailin Jin
WACV3
2015 Image-Specific Prior Adaptation for Denoising
abstract
Image priors are essential to many image restoration applications, including denoising, deblurring, and inpainting. Existing methods use either priors from the given image (internal) or priors from a separate collection of images (external). We find through statistical analysis that unifying the internal and external patch priors may yield a better patch prior. We propose a novel prior learning algorithm that combines the strength of both internal and external priors. In particular, we first learn a generic Gaussian mixture model from a collection of training images and then adapt the model to the given image by simultaneously adding additional components and refining the component parameters. We apply this image-specific prior to image denoising. The experimental results show that our approach yields better or competitive denoising results in terms of both the peak signal-to-noise ratio and structural similarity.
Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001
IEEE Trans. Image Process.3
2015 Rating Image Aesthetics Using Deep Learning
abstract
This paper investigates unified feature learning and classifier training approaches for image aesthetics assessment . Existing methods built upon handcrafted or generic image features and developed machine learning and statistical modeling techniques utilizing training examples. We adopt a novel deep neural network approach to allow unified feature learning and classifier training to estimate image aesthetics. In particular, we develop a double-column deep convolutional neural network to support heterogeneous inputs, i.e., global and local views, in order to capture both global and local characteristics of images . In addition, we employ the style and semantic attributes of images to further boost the aesthetics categorization performance . Experimental results show that our approach produces significantly better results than the earlier reported results on the AVA dataset for both the generic image aesthetics and content -based image aesthetics. Moreover, we introduce a 1.5-million image dataset (IAD) for image aesthetics assessment and we further boost the performance on the AVA test set by training the proposed deep neural networks on the IAD dataset.
Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001
IEEE Trans. Multim.3
2014 Fast Edge-Preserving PatchMatch for Large Displacement Optical Flow
abstract
We present a fast optical flow algorithm that can handle large displacement motions. Our algorithm is inspired by recent successes of local methods in visual correspondence searching as well as approximate nearest neighbor field algorithms. The main novelty is a fast randomized edge-preserving approximate nearest neighbor field algorithm which propagates self-similarity patterns in addition to offsets. Experimental results on public optical flow benchmarks show that our method is significantly faster than state-of-the-art methods without compromising on quality, especially when scenes contain large motions.
Linchao Bao, Qingxiong Yang, Hailin Jin
CVPR3
2014 Large-Scale Visual Font Recognition
abstract
This paper addresses the large-scale visual font recognition (VFR) problem, which aims at automatic identification of the typeface, weight, and slope of the text in an image or photo without any knowledge of content. Although visual font recognition has many practical applications, it has largely been neglected by the vision community. To address the VFR problem, we construct a large-scale dataset containing 2,420 font classes, which easily exceeds the scale of most image categorization datasets in computer vision. As font recognition is inherently dynamic and open-ended, i.e., new classes and data for existing categories are constantly added to the database over time, we propose a scalable solution based on the nearest class mean classifier (NCM). The core algorithm is built on local feature embedding, local feature metric learning and max-margin template selection, which is naturally amenable to NCM and thus to such open-ended classification problems. The new algorithm can generalize to new classes and new data at little added cost. Extensive experiments demonstrate that our approach is very effective on our synthetic test images, and achieves promising results on real world test images.
Jianchao Yang, Hailin Jin, Jonathan Brandt, Eli Shechtman, Aseem Agarwala, Tony X. Han
CVPR3
2014 Epitomic image colorization
abstract
Image colorization adds color to grayscale images. It not only increases the visual appeal of grayscale images, but also enriches the information conveyed by scientific images that lack color information. We develop a new image colorization method, epitomic image colorization, which automatically transfers color from the reference color image to the target grayscale image by a robust feature matching scheme using a new feature representation, namely the heterogeneous feature epitome. As a generative model, heterogeneous feature epitome is a condensed representation of image appearance which is employed for measuring the dissimilarity between reference patches and target patches in a way robust to noise in the reference image. We build a Markov Random Field (MRF) model with the learned heterogeneous feature epitome from the reference image, and inference in the MRF model achieves robust feature matching for transferring color. Our method renders better colorization results than the current state-of-the-art automatic colorization methods in our experiments.
Yingzhen Yang, Xinqi Chu, Tian-Tsong Ng, Alex Yong Sang Chia, Jianchao Yang, Hailin Jin, Thomas S. Huang
ICASSP6
2014 RAPID: Rating Pictorial Aesthetics using Deep Learning
abstract
Effective visual features are essential for computational aesthetic quality rating systems. Existing methods used machine learning and statistical modeling techniques on handcrafted features or generic image descriptors. A recently-published large-scale dataset, the AVA dataset, has further empowered machine learning based approaches. We present the RAPID (RAting PIctorial aesthetics using Deep learning) system, which adopts a novel deep neural network approach to enable automatic feature learning. The central idea is to incorporate heterogeneous inputs generated from the image, which include a global view and a local view, and to unify the feature learning and classifier training using a double-column deep convolutional neural network. In addition, we utilize the style attributes of images to help improve the aesthetic quality categorization accuracy. Experimental results show that our approach significantly outperforms the state of the art on the AVA dataset.
Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001
ACM Multimedia3
2014 Fast Edge-Preserving PatchMatch for Large Displacement Optical Flow
abstract
The speed of optical flow algorithm is crucial for many video editing tasks such as slow motion synthesis, selection propagation, tone adjustment propagation, and so on. Variational coarse-to-fine optical flow algorithms can generally produce high-quality results but cannot fulfil the speed requirement of many practical applications. Besides, large motions in real-world videos also pose a difficult problem to coarse-to-fine variational approaches. We, in this paper, present a fast optical flow algorithm that can handle large displacement motions. Our algorithm is inspired by recent successes of local methods in visual correspondence searching as well as approximate nearest neighbor field algorithms. The main novelty is a fast randomized edge-preserving approximate nearest neighbor field algorithm, which propagates self-similarity patterns in addition to offsets. Experimental results on public optical flow benchmarks show that our method is significantly faster than state-of-the-art methods without compromising on quality, especially when scenes contain large motions. Finally, we show some demo applications by applying our technique into real-world video editing tasks.
Linchao Bao, Qingxiong Yang, Hailin Jin
IEEE Trans. Image Process.3
2014 Automatic Scene Inference for 3D Object Compositing
abstract
We present a user-friendly image editing system that supports a drag-and-drop object insertion (where the user merely drags objects into the image, and the system automatically places them in 3D and relights them appropriately), postprocess illumination editing, and depth-of-field manipulation. Underlying our system is a fully automatic technique for recovering a comprehensive 3D scene model (geometry, illumination, diffuse albedo, and camera parameters) from a single, low dynamic range photograph. This is made possible by two novel contributions: an illumination inference algorithm that recovers a full lighting model of the scene (including light sources that are not directly visible in the photograph), and a depth estimation algorithm that combines data-driven depth transfer with geometric reasoning about the scene layout. A user study shows that our system produces perceptually convincing results, and achieves the same level of realism as techniques that require significant user interaction.
Kevin Karsch, Kalyan Sunkavalli, Sunil Hadap, Nathan Carr 0001, Hailin Jin, Rafael Fonte, Michael Sittig, David A. Forsyth
ACM Trans. Graph.5
2013 Large Displacement Optical Flow from Nearest Neighbor Fields
abstract
We present an optical flow algorithm for large displacement motions. Most existing optical flow methods use the standard coarse-to-fine framework to deal with large displacement motions which has intrinsic limitations. Instead, we formulate the motion estimation problem as a motion segmentation problem. We use approximate nearest neighbor fields to compute an initial motion field and use a robust algorithm to compute a set of similarity transformations as the motion candidates for segmentation. To account for deviations from similarity transformations, we add local deformations in the segmentation process. We also observe that small objects can be better recovered using translations as the motion candidates. We fuse the motion results obtained under similarity transformations and under translations together before a final refinement. Experimental validation shows that our method can successfully handle large displacement motions. Although we particularly focus on large displacement motions in this work, we make no sacrifice in terms of overall performance. In particular, our method ranks at the top of the Middlebury benchmark.
Zhuoyuan Chen, Hailin Jin, Zhe Lin 0001, Scott Cohen, Ying Wu 0001
CVPR2
2013 Specular Reflection Separation Using Dark Channel Prior
abstract
We present a novel method to separate specular reflection from a single image. Separating an image into diffuse and specular components is an ill-posed problem due to lack of observations. Existing methods rely on a specular-free image to detect and estimate specularity, which however may confuse diffuse pixels with the same hue but a different saturation value as specular pixels. Our method is based on a novel observation that for most natural images the dark channel can provide an approximate specular-free image. We also propose a maximum a posteriori formulation which robustly recovers the specular reflection and chromaticity despite of the hue-saturation ambiguity. We demonstrate the effectiveness of the proposed algorithm on real and synthetic examples. Experimental results show that our method significantly outperforms the state-of-the-art methods in separating specular reflection.
Hyeongwoo Kim, Hailin Jin, Sunil Hadap, In-So Kweon
CVPR2
2013 Plane-Based Content Preserving Warps for Video Stabilization
abstract
Recently, a new image deformation technique called content-preserving warping (CPW) has been successfully employed to produce the state-of-the-art video stabilization results in many challenging cases. The key insight of CPW is that the true image deformation due to viewpoint change can be well approximated by a carefully constructed warp using a set of sparsely constructed 3D points only. However, since CPW solely relies on the tracked feature points to guide the warping, it works poorly in large texture less regions, such as ground and building interiors. To overcome this limitation, in this paper we present a hybrid approach for novel view synthesis, observing that the texture less regions often correspond to large planar surfaces in the scene. Particularly, given a jittery video, we first segment each frame into piecewise planar regions as well as regions labeled as non-planar using Markov random fields. Then, a new warp is computed by estimating a single homography for regions belong to the same plane, while inheriting results from CPW in the non-planar regions. We demonstrate how the segmentation information can be efficiently obtained and seamlessly integrated into the stabilization framework. Experimental results on a variety of real video sequences verify the effectiveness of our method.
Zihan Zhou 0001, Hailin Jin, Yi Ma 0001
CVPR2
2013 Joint Subspace Stabilization for Stereoscopic Video
abstract
Shaky stereoscopic video is not only unpleasant to watch but may also cause 3D fatigue. Stabilizing the left and right view of a stereoscopic video separately using a monocular stabilization method tends to both introduce undesirable vertical disparities and damage horizontal disparities, which may destroy the stereoscopic viewing experience. In this paper, we present a joint subspace stabilization method for stereoscopic video. We prove that the low-rank subspace constraint for monocular video [10] also holds for stereoscopic video. Particularly, the feature trajectories from the left and right video share the same subspace. Based on this proof, we develop a stereo subspace stabilization method that jointly computes a common subspace from the left and right video and uses it to stabilize the two videos simultaneously. Our method meets the stereoscopic constraints without 3D reconstruction or explicit left-right correspondence. We test our method on a variety of stereoscopic videos with different scene content and camera motion. The experiments show that our method achieves high-quality stabilization for stereoscopic video in a robust and efficient way.
Feng Liu 0015, Yuzhen Niu, Hailin Jin
ICCV3
2013 Casual Stereoscopic Photo Authoring
abstract
Stereoscopic 3D displays become more and more popular these years. However, authoring high-quality stereoscopic 3D content remains challenging. In this paper, we present a method for easy stereoscopic photo authoring with a regular (monocular) camera. Our method takes two images or video frames using a monocular camera as input and transforms them into a stereoscopic image pair that provides a pleasant viewing experience. The key technique of our method is a perceptual-plausible image rectification algorithm that warps the input image pairs to meet the stereoscopic geometric constraint while avoiding noticeable visual distortion. Our method uses spatially-varying mesh-based image warps. Our warping method encodes a variety of constraints to best meet the stereoscopic geometric constraint and minimize visual distortion. Since each energy term is quadratic, our method eventually formulates the warping problem as a quadratic energy minimization which is solved efficiently using a sparse linear solver. Our method also allows both local and global adjustments of the disparities, an important property for adapting resulting stereoscopic images to different viewing conditions. Our experiments demonstrate that our spatially-varying warping technique can better support casual stereoscopic photo authoring than existing methods and our results and user study show that our method can effectively use casually-taken photos to create high-quality stereoscopic photos that deliver a pleasant 3D viewing experience.
Feng Liu 0015, Yuzhen Niu, Hailin Jin
IEEE Trans. Multim.3
2012 Robust plane-based structure from motion
abstract
We introduce a new approach to structure and motion recovery directly from one or more large planes in the scene. When such a plane exists, we demonstrate how to automatically detect and track it robustly and consistently over a long video sequence, and how to efficiently self-calibrate the camera using the homographies induced by this plane. We build a complete structure from motion system which does not use any additional off-the-plane information about the scene, and show its advantage over conventional systems in handling two important issues which often occur in real world videos, namely, the plane degeneracy and the dynamic foreground problems. Experimental results on a variety of real video sequences verify the effectiveness and efficiency of our system.
Zihan Zhou 0001, Hailin Jin, Yi Ma 0001
CVPR2
2012 A Locally Linear Regression Model for Boundary Preserving Regularization in Stereo Matching
Shengqi Zhu 0003, Li Zhang 0003, Hailin Jin
ECCV (5)3
2012 Video upscaling via spatio-temporal self-similarity
Alper Ayvaci, Hailin Jin, Zhe Lin 0001, Scott Cohen, Stefano Soatto
ICPR2
2012 Asymmetry of Convex Bodies of Constant Width
Hailin Jin
Discret. Comput. Geom.1
2012 Aesthetics-Based Stereoscopic Photo Cropping for Heterogeneous Displays
abstract
Stereoscopic displays are becoming ubiquitous, ranging from large 3-D TVs to small mobile phones. Stereoscopic photos need to be carefully adapted to be effectively viewed on the displays other than originally intended. In this paper, we present a method that can automatically crop and scale an existing stereoscopic photo to a variety of displays while preserving its aesthetic value. We formulate stereoscopic photo adaptation as an optimization problem that aims to preserve the aesthetic value of the input photo. We define a wide range of energy terms to preserve the stereoscopic photo aesthetics by borrowing rules from stereoscopic photography. Our experiments on a wide variety of stereoscopic photos demonstrate that our method can robustly produce display-dependent stereoscopic photos that deliver pleasant viewing experiences.
Yuzhen Niu, Feng Liu 0015, Wu-chi Feng, Hailin Jin
IEEE Trans. Multim.4
2011 Subspace video stabilization
abstract
We present a robust and efficient approach to video stabilization that achieves high-quality camera motion for a wide range of videos. In this article, we focus on the problem of transforming a set of input 2D motion trajectories so that they are both smooth and resemble visually plausible views of the imaged scene; our key insight is that we can achieve this goal by enforcing subspace constraints on feature trajectories while smoothing them. Our approach assembles tracked features in the video into a trajectory matrix, factors it into two low-rank matrices, and performs filtering or curve fitting in a low-dimensional linear space. In order to process long videos, we propose a moving factorization that is both efficient and streamable. Our experiments confirm that our approach can efficiently provide stabilization results comparable with prior 3D methods in cases where those methods succeed, but also provides smooth camera motions in cases where such approaches often fail, such as videos that lack parallax. The presented approach offers the first method that both achieves high-quality video stabilization and is practical enough for consumer applications.
Feng Liu 0015, Michael Gleicher, Jue Wang 0001, Hailin Jin, Aseem Agarwala
ACM Trans. Graph.4
2009 Stereo matching with nonparametric smoothness priors in feature space
abstract
We propose a novel formulation of stereo matching that considers each pixel as a feature vector. Under this view, matching two or more images can be cast as matching point clouds in feature space. We build a nonparametric depth smoothness model in this space that correlates the image features and depth values. This model induces a sparse graph that links pixels with similar features, thereby converting each point cloud into a connected network. This network defines a neighborhood system that captures pixel grouping hierarchies without resorting to image segmentation. We formulate global stereo matching over this neighborhood system and use graph cuts to match pixels between two or more such networks. We show that our stereo formulation is able to recover surfaces with different orders of smoothness, such as those with high-curvature details and sharp discontinuities. Furthermore, compared to other single-frame stereo methods, our method produces more temporally stable results from videos of dynamic scenes, even when applied to each frame independently.
Brandon M. Smith 0001, Li Zhang 0003, Hailin Jin
CVPR3
2009 Multiple view image denoising
abstract
We present a novel multi-view denoising algorithm. Our algorithm takes noisy images taken from different viewpoints as input and groups similar patches in the input images using depth estimation. We model intensity-dependent noise in low-light conditions and use the principal component analysis and tensor analysis to remove such noise. The dimensionalities for both PCA and tensor analysis are automatically computed in a way that is adaptive to the complexity of image structures in the patches. Our method is based on a probabilistic formulation that marginalizes depth maps as hidden variables and therefore does not require perfect depth estimation. We validate our algorithm on both synthetic and real images with different content. Our algorithm compares favorably against several state-of-the-art denoising algorithms.
Li Zhang 0003, Sundeep Vaddadi, Hailin Jin, Shree K. Nayar
CVPR3
2009 GroupSAC: Efficient consensus in the presence of groupings
abstract
We present a novel variant of the RANSAC algorithm that is much more efficient, in particular when dealing with problems with low inlier ratios. Our algorithm assumes that there exists some grouping in the data, based on which we introduce a new binomial mixture model rather than the simple binomial model as used in RANSAC. We prove that in the new model it is more efficient to sample data from a smaller numbers of groups and groups with more tentative correspondences, which leads to a new sampling procedure that uses progressive numbers of groups. We demonstrate our algorithm on two classical geometric vision problems: wide-baseline matching and camera resectioning. The experiments show that the algorithm serves as a general framework that works well with three possible grouping strategies investigated in this paper, including a novel optical flow based clustering approach. The results show that our algorithm is able to achieve a significant performance gain compared to the standard RANSAC and PROSAC.
Kai Ni 0001, Hailin Jin, Frank Dellaert
ICCV2
2009 Light field video stabilization
abstract
We describe a method for producing a smooth, stabilized video from the shaky input of a hand-held light field video camera—specifically, a small camera array. Traditional stabilization techniques dampen shake with 2D warps, and thus have limited ability to stabilize a significantly shaky camera motion through a 3D scene. Other recent stabilization techniques synthesize novel views as they would have been seen along a virtual, smooth 3D camera path, but are limited to static scenes. We show that video camera arrays enable much more powerful video stabilization, since they allow changes in viewpoint for a single time instant. Furthermore, we point out that the straightforward approach to light field video stabilization requires computing structure-from-motion, which can be brittle for typical consumer-level video of general dynamic scenes. We present a more robust approach that avoids input camera path reconstruction. Instead, we employ a spacetime optimization that directly computes a sequence of relative poses between the virtual camera and the camera array, while minimizing acceleration of salient visual features in the virtual image plane. We validate our novel method by comparing it to state-of-the-art stabilization software, such as Apple iMovie and 2d3 SteadyMove Pro, on a number of challenging scenes.
Brandon M. Smith 0001, Li Zhang 0003, Hailin Jin, Aseem Agarwala
ICCV3
2009 Content-preserving warps for 3D video stabilization
abstract
We describe a technique that transforms a video from a hand-held video camera so that it appears as if it were taken with a directed camera motion. Our method adjusts the video to appear as if it were taken from nearby viewpoints, allowing 3D camera movements to be simulated. By aiming only for perceptual plausibility, rather than accurate reconstruction, we are able to develop algorithms that can effectively recreate dynamic scenes from a single source video. Our technique first recovers the original 3D camera motion and a sparse set of 3D, static scene points using an off-the-shelf structure-from-motion system. Then, a desired camera path is computed either automatically (e.g., by fitting a linear or quadratic path) or interactively. Finally, our technique performs a least-squares optimization that computes a spatially-varying warp from each input video frame into an output frame. The warp is computed to both follow the sparse displacements suggested by the recovered 3D structure, and avoid deforming the content in the video frame. Our experiments on stabilizing challenging videos of dynamic scenes demonstrate the effectiveness of our technique.
Feng Liu 0015, Michael Gleicher, Hailin Jin, Aseem Agarwala
ACM Trans. Graph.3
2008 A three-point minimal solution for panoramic stitching with lens distortion
abstract
We present a minimal solution for aligning two images taken by a rotating camera from point correspondences. The solution particularly addresses the case where there is lens distortion in the images. We assume to know the two camera centers but not the focal lengths and allow the latter to vary. Our solution uses a minimal number (three) of point correspondences and is well suited to be used in a hypothesis testing framework. It does not suffer from numerical instabilities observed in other algebraic minimal solvers and is also efficient. We validate our solution in multi-image panoramic stitching on real images with lens distortion.
Hailin Jin
CVPR1
2008 Stereoscopic inpainting: Joint color and depth completion from stereo images
abstract
We present a novel algorithm for simultaneous color and depth inpainting. The algorithm takes stereo images and estimated disparity maps as input and fills in missing color and depth information introduced by occlusions or object removal. We first complete the disparities for the occlusion regions using a segmentation-based approach. The completed disparities can be used to facilitate the user in labeling objects to be removed. Since part of the removed regions in one image is visible in the other, we mutually complete the two images through 3D warping. Finally, we complete the remaining unknown regions using a depth-assisted texture synthesis technique, which simultaneously fills in both color and depth. We demonstrate the effectiveness of the proposed algorithm on several challenging data sets.
Liang Wang 0002, Hailin Jin, Ruigang Yang, Minglun Gong
CVPR2
2008 Search Space Reduction for MRF Stereo
Liang Wang 0002, Hailin Jin, Ruigang Yang
ECCV (1)2
2008 3-D Reconstruction of Shaded Objects from Multiple Images Under Unknown Illumination
Hailin Jin, Daniel Cremers, Emmanuel Prados, Anthony J. Yezzi, Stefano Soatto
Int. J. Comput. Vis.1
2005 Visual Tracking in the Presence of Motion Blur
abstract
We consider the problem of visual tracking of regions of interest in a sequence of motion blurred images. Traditional methods couple tracking with deblurring in order to correctly account for the effects of motion blur. Such coupling is usually appropriate, but computationally wasteful when visual tracking is the lone objective. Instead of deblurring images, we propose to match regions by blurring them. The matching score for two image regions is governed by a cost function that only involves the region deformation parameters and two motion blur vectors. We present an efficient algorithm to minimize the proposed cost function and demonstrate it on sequences of real blurred images.
Hailin Jin, Paolo Favaro, Roberto Cipolla
CVPR (2)1
2005 KALMANSAC: Robust Filtering by Consensus
abstract
We propose an algorithm to perform causal inference of the state of a dynamical model when the measurements are corrupted by outliers. While the optimal (maximum-likelihood) solution has doubly exponential complexity due to the combinatorial explosion of possible choices of inliers, we exploit the structure of the problem to design a sampling-based algorithm that has constant complexity. We derive our algorithm from the equations of the optimal filter, which makes our approximation explicit. Our work is motivated by real-time tracking and the estimation of structure from motion (SFM). We test our algorithm for on-line outlier rejection both for tracking and for SFM. We show that our approach can tolerate a large proportion of outliers, whereas previous causal robust statistical inference methods failed with less than half as many. Our work can be thought of as the extension of random sample consensus algorithms to dynamic data, or as the implementation of pseudo-Bayesian filtering algorithms in a sampling framework.
Andrea Vedaldi, Hailin Jin, Paolo Favaro, Stefano Soatto
ICCV2
2005 Multi-View Stereo Reconstruction of Dense Shape and Complex Appearance
Hailin Jin, Stefano Soatto, Anthony J. Yezzi
Int. J. Comput. Vis.1
2004 Shedding Light on Stereoscopic Segmentation
Hailin Jin, Daniel Cremers, Anthony J. Yezzi, Stefano Soatto
CVPR (1)1
2004 Region-Based Segmentation on Evolving Surfaces with Application to 3D Reconstruction of Shape and Piecewise Constant Radiance
Hailin Jin, Anthony J. Yezzi, Stefano Soatto
ECCV (2)1
2003 Multi-view Stereo Beyond Lambert
abstract
We consider the problem of estimating the shape and radiance of an object from a calibrated set of views under the assumption that the reflectance of the object is non-Lambertian. Unlike traditional stereo, we do not solve the correspondence problem by comparing image-to-image. Instead, we exploit a rank constraint on the radiance tensor field of the surface in space, and use it to define a discrepancy measure between each image and the underlying model. Our approach automatically returns an estimate of the radiance of the scene, along with its shape, represented by a dense surface. The former can be used to generate novel views that capture the non-Lambertian appearance of the scene.
Hailin Jin, Stefano Soatto, Anthony J. Yezzi
CVPR (1)1
2003 Tales of Shape and Radiance in Multi-view Stereo
abstract
To what extent can three-dimensional shape and radiance be inferred from a collection of images? Can the two be estimated separately while retaining optimality? How should the optimality criterion be computed? When is it necessary to employ an explicit model of the reflectance properties of a scene? In this paper we introduce a separation principle for shape and radiance estimation that applies to Lambertian scenes and holds for any choice of norm. When the scene is not Lambertian, however, shape cannot be decoupled from radiance, and therefore matching image-to-image is not possible directly. We employ a rank constraint on the radiance tensor, which is commonly used in computer graphics, and construct a novel cost functional whose minimization leads to an estimate of both shape and radiance for nonLambertian objects, which we validate experimentally.
Stefano Soatto, Anthony J. Yezzi, Hailin Jin
ICCV3
2003 A semi-direct approach to structure from motion
Hailin Jin, Paolo Favaro, Stefano Soatto
Vis. Comput.1
2002 A Variational Approach to Shape from Defocus
Hailin Jin, Paolo Favaro
ECCV (2)1
2002 Structure from Motion Causally Integrated Over Time
abstract
We describe an algorithm for reconstructing three-dimensional structure and motion causally, in real time from monocular sequences of images. We prove that the algorithm is minimal and stable, in the sense that the estimation error remains bounded with probability one throughout a sequence of arbitrary length. We discuss a scheme for handling occlusions (point features appearing and disappearing) and drift in the scale factor. These issues are crucial for the algorithm to operate in real time on real scenes. We describe in detail the implementation of the algorithm, which runs on a personal computer and has been made available to the community. We report the performance of our implementation on a few representative long sequences of real and synthetic images. The algorithm, which has been tested extensively over the course of the past few years, exhibits honest performance when the scene contains at least 20-40 points with high contrast, when the relative motion is "slow" compared to the sampling frequency of the frame grabber (30 Hz), and the lens aperture is "large enough" (typically more than 30/spl deg/ of visual field).
Alessandro Chiuso, Paolo Favaro, Hailin Jin, Stefano Soatto
IEEE Trans. Pattern Anal. Mach. Intell.3
2001 Real-time Virtual Object Insertion
abstract
We present a system to insert virtual objects into real image sequences in real time. The system consists of offthe- shelf hardware (a camera connected to a Pentium PC) and software to (a) automatically select and track region features despite changes in illumination, (b) estimate threedimensional position and orientation of surface patches relative to an inertial reference frame despite individual pointfeatures appearing and disappearing, (c) insert a texturemapped virtual object into the scene so as to make it appear to be part of the scene and moving with it. This is all done in real time. The multi-thread C++ code, which is readily interfaced with a frame grabber as well as Matlab for development, will be made available to the public at the demonstration.
Paolo Favaro, Hailin Jin, Stefano Soatto
ICCV2
2001 Real-Time Feature Tracking and Outlier Rejection with Changes in Illumination
Hailin Jin, Paolo Favaro, Stefano Soatto
ICCV1
2000 Real-Time 3-D Motion and Structure of Point-Features: A Front-End for Vision-Based Control and Interaction
abstract
We present a system that consists of one camera connected to a personal computer that can (a) select and track a number of high-contrast point features on a sequence of images, (b) estimate their three-dimensional motion and position relative to an inertial reference frame, assuming rigidity, (c) handle occlusions that cause point-features to disappear as well as new features to appear. The system can also (d) perform partial self-calibration and (e) check for consistency of the rigidity assumption, although these features are not implemented in the current release. All of this is done automatically and in real-time (30 Hz) for 40-50 point features using commercial off-the-shelf hardware. The system is based on an algorithm presented by Chiuso et al. (2000), the properties of which have been analyzed by Chiuso and Soatto (2000). In particular, the algorithm is provably observable, provably minimal and provably stable- under suitable conditions. The core of the system, consisting of C++ code ready to interface with a frame grabber as well as Matlab code for development, is available at http://ee.wustl.edu/-soatto/research.html. We demonstrate the system by showing its use as (1) an ego-motion estimator, (2) an object tracker, and (3) an interactive input device, all without any modification of the system settings.
Hailin Jin, Paolo Favaro, Stefano Soatto
CVPR1
2000 Stereoscopic Shading: Integrating Mult1Frame Shape Cues in a Variational Framework
abstract
We address the problem of integrating multi-frame stereo and shading cues within the framework of optimization in the infinite-dimensional space of piecewise smooth surfaces. Cue integration then reduces to the determination of regions where prior assumptions on the reflectance of the surfaces can be enforced. By combining cues, our formulation allows defining a well-posed problem even when reconstruction from stereo or shading in isolation would be ill-posed. For a simplified model we prove the necessary conditions for optimality, and propose an iterative optimization algorithm, which we implement using ultra-narrowband level set methods.
Hailin Jin, Stefano Soatto, Anthony J. Yezzi
CVPR1
2000 3-D Motion and Structure from 2-D Motion Causally Integrated over Time: Implementation
Alessandro Chiuso, Paolo Favaro, Hailin Jin, Stefano Soatto
ECCV (2)3