Liangliang Cao

dblp:95/6915 · DBLP profile ↗
← Back
112ranked-venue papers
26as first author
17since 2021 · last 2025
0000-0003-0900-1512ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 82 · 22 first-author · 11 since 2021Artificial intelligence and machine learning · 49 · 10 first-author · 10 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-authorSystems, architecture and hardware · 2Computer networks · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling Laws
abstract
This paper formalizes an emerging learning paradigm that uses a trained model as a reference to guide and enhance the training of a target model through strategic data selection or weighting, named **model steering**. While ad-hoc methods have been used in various contexts, including the training of large foundation models, its underlying principles remain insufficiently understood, leading to sub-optimal performance. In this work, we propose a theory-driven framework for model steering called **DRRho risk minimization**, which is rooted in Distributionally Robust Optimization (DRO). Through a generalization analysis, we provide theoretical insights into why this approach improves generalization and data efficiency compared to training without a reference model. To the best of our knowledge, this is the first time such theoretical insights are provided for the new learning paradigm, which significantly enhance our understanding and practice of model steering. Building on these insights and the connection between contrastive learning and DRO, we introduce a novel method for Contrastive Language-Image Pretraining (CLIP) with a reference model, termed DRRho-CLIP. Extensive experiments validate the theoretical insights, reveal a superior scaling law compared to CLIP without a reference model, and demonstrate its strength over existing heuristic approaches. Code is released at [github.com/Optimization-AI/DRRho-CLIP](https://github.com/Optimization-AI/DRRho-CLIP)
Xiyuan Wei, Ming Lin 0002, Fanjiang Ye, Fengguang Song, Liangliang Cao, My T. Thai, Tianbao Yang
ICML5
2025 Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention
abstract
In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera control into the generation process, but their results are often limited to simple trajectories or lack the ability to generate consistent videos from multiple distinct camera paths for the same scene. To address these limitations, we introduce Cavia, a novel framework for camera-controllable, multi-view video generation, capable of converting an input image into multiple spatiotemporally consistent videos. Our framework extends the spatial and temporal attention modules into view-integrated attention modules, improving both viewpoint and temporal consistency. This flexible design allows for joint training with diverse curated data sources, including scene-level static videos, object-level synthetic multi-view dynamic videos, and real-world monocular dynamic videos. To the best of our knowledge, Cavia is the first framework that enables users to generate multiple videos of the same scene with precise control over camera motion, while simultaneously preserving object motion. Extensive experiments demonstrate that Cavia surpasses state-of-the-art methods in terms of geometric consistency and perceptual quality.
Dejia Xu, Yifan Jiang 0001, Liangchen Song, Thorsten Gernoth, Liangliang Cao, Zhangyang Wang, Hao Tang 0001
ICML6
2025 Diffusion Model-Based Image Editing: A Survey
abstract
Denoising diffusion models have emerged as a powerful tool for various image generation and editing tasks, facilitating the synthesis of visual content in an unconditional or input-conditional manner. The core idea behind them is learning to reverse the process of gradually adding noise to images, allowing them to generate high-quality samples from a complex distribution. In this survey, we provide an exhaustive overview of existing methods using diffusion models for image editing, covering both theoretical and practical aspects in the field. We delve into a thorough analysis and categorization of these works from multiple perspectives, including learning strategies, user-input conditions, and the array of specific editing tasks that can be accomplished. In addition, we pay special attention to image inpainting and outpainting, and explore both earlier traditional context-driven and current multimodal conditional methods, offering a comprehensive analysis of their methodologies. To further evaluate the performance of text-guided image editing algorithms, we propose a systematic benchmark, EditEval, featuring an innovative metric, LMM Score. Finally, we address current limitations and envision some potential directions for future research.
Yi Huang 0035, Jiancheng Huang, Yifan Liu 0001, Mingfu Yan, Jiaxi Lv, Jianzhuang Liu, Wei Xiong 0008, He Zhang 0004, Liangliang Cao, Shifeng Chen
IEEE Trans. Pattern Anal. Mach. Intell.9
2024 Efficient-3Dim: Learning a Generalizable Single-image Novel-view Synthesizer in One Day
abstract
The task of novel view synthesis aims to generate unseen perspectives of an object or scene from a limited set of input images. Nevertheless, synthesizing novel views from a single image remains a significant challenge. Previous approaches tackle this problem by adopting mesh prediction, multi-plane image construction, or more advanced techniques such as neural radiance fields. Recently, a pre-trained diffusion model that is specifically designed for 2D image synthesis has demonstrated its capability in producing photorealistic novel views, if sufficiently optimized with a 3D finetuning task. Despite greatly improved fidelity and generalizability, training such a powerful diffusion model requires a vast volume of training data and model parameters, resulting in a notoriously long time and high computational costs. To tackle this issue, we propose Efficient-3DiM, a highly efficient yet effective framework to learn a single-image novel-view synthesizer. Motivated by our in-depth analysis of the diffusion model inference process, we propose several pragmatic strategies to reduce training overhead to a manageable scale, including a crafted timestep sampling strategy, a superior 3D feature extractor, and an enhanced training scheme. When combined, our framework can reduce the total training time from 10 days to less than 1 day, significantly accelerating the training process on the same computational platform (an instance with 8 Nvidia A100 GPUs). Comprehensive experiments are conducted to demonstrate the efficiency and generalizability of our proposed method.
Yifan Jiang 0001, Hao Tang 0001, Jen-Hao Rick Chang, Liangchen Song, Zhangyang Wang, Liangliang Cao
ICLR6
2024 Ferret: Refer and Ground Anything Anywhere at Any Granularity
abstract
We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hybrid region representation that integrates discrete coordinates and continuous features jointly to represent a region in the image. To extract the continuous features of versatile regions, we propose a spatial-aware visual sampler, adept at handling varying sparsity across different shapes. Consequently, Ferret can accept diverse region inputs, such as points, bounding boxes, and free-form shapes. To bolster the desired capability of Ferret, we curate GRIT, a comprehensive refer-and-ground instruction tuning dataset including 1.1M samples that contain rich hierarchical spatial knowledge, with an additional 130K hard negative data to promote model robustness. The resulting model not only achieves superior performance in classical referring and grounding tasks, but also greatly outperforms existing MLLMs in region-based and localization-demanded multimodal chatting. Our evaluations also reveal a significantly improved capability of describing image details and a remarkable alleviation in object hallucination.
Haoxuan You, Haotian Zhang 0005, Zhe Gan, Xianzhi Du, Bowen Zhang 0002, Liangliang Cao, Shih-Fu Chang, Yinfei Yang
ICLR7
2023 STAIR: Learning Sparse Text and Image Representation in Grounded Tokens
abstract
Chen Chen, Bowen Zhang, Liangliang Cao, Jiguang Shen, Tom Gunter, Albin Jose, Alexander Toshev, Yantao Zheng, Jonathon Shlens, Ruoming Pang, Yinfei Yang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Chen Chen 0005, Bowen Zhang 0002, Liangliang Cao, Jiguang Shen, Tom Gunter, Albin Madappally Jose, Alexander Toshev, Yantao Zheng, Jonathon Shlens, Ruoming Pang, Yinfei Yang
EMNLP3
2023 RoomDreamer: Text-Driven 3D Indoor Scene Synthesis with Coherent Geometry and Texture
abstract
The techniques for 3D indoor scene capturing are widely used, but the meshes produced leave much to be desired. In this paper, we propose "RoomDreamer", which leverages powerful natural language to synthesize a new room with a different style. Unlike existing image synthesis methods, our work addresses the challenge of synthesizing both geometry and texture aligned to the input scene structure and prompt simultaneously. The key insight is that a scene should be treated as a whole, taking into account both scene texture and geometry. The proposed framework consists of two significant components: Geometry Guided Diffusion and Mesh Optimization. Geometry Guided Diffusion for 3D Scene guarantees the consistency of the scene style by applying the 2D prior to the entire scene simultaneously. Mesh Optimization improves the geometry and texture jointly and eliminates the artifacts in the scanned scene. To validate the proposed method, real indoor scenes scanned with smartphones are used for extensive experiments, through which the effectiveness of our method is demonstrated.
Liangchen Song, Liangliang Cao, Hongyu Xu, Kai Kang 0006, Junsong Yuan 0001
ACM Multimedia2
2022 Improving Confidence Estimation on Out-of-Domain Data for End-to-End Speech Recognition
abstract
As end-to-end automatic speech recognition (ASR) models reach promising performance, various downstream tasks rely on good confidence estimators for these systems. Recent research has shown that model-based confidence estimators have a significant advantage over using the output softmax probabilities. If the input data to the speech recogniser is from mismatched acoustic and linguistic conditions, the ASR performance and the corresponding confidence estimators may exhibit severe degradation. Since confidence models are often trained on the same in-domain data as the ASR, generalising to out-of-domain (OOD) scenarios is challenging. By keeping the ASR model untouched, this paper proposes two approaches to improve the model-based confidence estimators on OOD data: using pseudo transcriptions and an additional OOD language model. With an ASR model trained on LibriSpeech, experiments show that the proposed methods can greatly improve the confidence metrics on TED-LIUM and Switchboard datasets while preserving in-domain performance. Furthermore, the improved confidence estimators are better calibrated on OOD data and can provide a much more reliable criterion for data selection.
Qiujia Li, Yu Zhang 0033, David Qiu, Yanzhang He, Liangliang Cao, Philip C. Woodland
ICASSP5
2022 PriFit: Learning to Fit Primitives Improves Few Shot Point Cloud Segmentation
abstract
Abstract We present PriFit, a semi‐supervised approach for label‐efficient learning of 3D point cloud segmentation networks. PriFit combines geometric primitive fitting with point‐based representation learning. Its key idea is to learn point representations whose clustering reveals shape regions that can be approximated well by basic geometric primitives, such as cuboids and ellipsoids. The learned point representations can then be re‐used in existing network architectures for 3D point cloud segmentation, and improves their performance in the few‐shot setting. According to our experiments on the widely used ShapeNet and PartNet benchmarks, PriFit outperforms several state‐of‐the‐art methods in this setting, suggesting that decomposability into primitives is a useful prior for learning representations predictive of semantic parts. We present a number of ablative experiments varying the choice of geometric primitives and downstream tasks to demonstrate the effectiveness of the method.
Gopal Sharma, Bidya Dash, Aruni Roy Chowdhury, Matheus Gadelha, Marios Loizou, Liangliang Cao, Rui Wang 0003, Erik G. Learned-Miller, Subhransu Maji, Evangelos Kalogerakis
Comput. Graph. Forum6
2021 Improving Streaming Automatic Speech Recognition with Non-Streaming Model Distillation on Unsupervised Data
abstract
Streaming end-to-end automatic speech recognition (ASR) models are widely used on smart speakers and on-device applications. Since these models are expected to transcribe speech with minimal latency, they are constrained to be causal with no future context, compared to their non-streaming counterparts. Consequently, streaming models usually perform worse than non-streaming models. We propose a novel and effective learning method by leveraging a non-streaming ASR model as a teacher to generate transcripts on an arbitrarily large data set, which is then used to distill knowledge into streaming ASR models. This way, we scale the training of streaming models to up to 3 million hours of YouTube audio. Experiments show that our approach can significantly reduce the word error rate (WER) of RNN-T models not only on LibriSpeech but also on YouTube data in four languages. For example, in French, we are able to reduce the WER by 16.4% relatively to a baseline streaming model by leveraging a non-streaming teacher model trained on the same amount of labeled data as the baseline.
Thibault Doutre, Wei Han 0002, Zhiyun Lu, Chung-Cheng Chiu, Ruoming Pang, Arun Narayanan, Ananya Misra, Yu Zhang 0033, Liangliang Cao
ICASSP10
2021 Confidence Estimation for Attention-Based Sequence-to-Sequence Models for Speech Recognition
abstract
For various speech-related tasks, confidence scores from a speech recogniser are a useful measure to assess the quality of transcriptions. In traditional hidden Markov model-based automatic speech recognition (ASR) systems, confidence scores can be reliably obtained from word posteriors in decoding lattices. However, for an ASR system with an auto-regressive decoder, such as an attention-based sequence-to-sequence model, computing word posteriors is difficult. An obvious alternative is to use the decoder softmax probability as the model confidence. In this paper, we first examine how some commonly used regularisation methods influence the softmax-based confidence scores and study the overconfident behaviour of end-to-end models. Then we propose a lightweight and effective approach named confidence estimation module (CEM) on top of an existing end-to-end ASR model. Experiments on LibriSpeech show that CEM can mitigate the overconfidence problem and can produce more reliable confidence scores with and without shallow fusion of a language model. Further analysis shows that CEM generalises well to speech from a moderately mismatched domain and can potentially improve downstream tasks such as semi-supervised learning.
Qiujia Li, David Qiu, Yu Zhang 0033, Bo Li 0028, Yanzhang He, Philip C. Woodland, Liangliang Cao, Trevor Strohman
ICASSP7
2021 Learning Word-Level Confidence for Subword End-To-End ASR
abstract
We study the problem of word-level confidence estimation in subword-based end-to-end (E2E) models for automatic speech recognition (ASR). Although prior works have proposed training auxiliary confidence models for ASR systems, they do not extend naturally to systems that operate on word-pieces (WP) as their vocabulary. In particular, ground truth WP correctness labels are needed for training confidence models, but the non-unique tokenization from word to WP causes inaccurate labels to be generated. This paper proposes and studies two confidence models of increasing complexity to solve this problem. The final model uses self-attention to directly learn word-level confidence without needing subword tokenization, and exploits full context features from multiple hypotheses to improve confidence accuracy. Experiments on Voice Search and long-tail test sets show standard metrics (e.g., NCE, AUC, RMSE) improving substantially. The proposed confidence module also enables a model selection approach to combine an on-device E2E model with a hybrid model on the server to address the rare word recognition problem for the E2E model.
David Qiu, Qiujia Li, Yanzhang He, Yu Zhang 0033, Bo Li 0028, Liangliang Cao, Rohit Prabhavalkar, Deepti Bhatia, Wei Li 0133, Tara N. Sainath, Ian McGraw
ICASSP6
2021 Bridging the Gap Between Streaming and Non-Streaming ASR Systems by Distilling Ensembles of CTC and RNN-T Models
Thibault Doutre, Wei Han 0002, Chung-Cheng Chiu, Ruoming Pang, Olivier Siohan, Liangliang Cao
Interspeech6
2021 Residual Energy-Based Models for End-to-End Speech Recognition
abstract
End-to-end models with auto-regressive decoders have shown impressive results for automatic speech recognition (ASR). These models formulate the sequence-level probability as a product of the conditional probabilities of all individual tokens given their histories. However, the performance of locally normalised models can be sub-optimal because of factors such as exposure bias. Consequently, the model distribution differs from the underlying data distribution. In this paper, the residual energy-based model (R-EBM) is proposed to complement the auto-regressive ASR model to close the gap between the two distributions. Meanwhile, R-EBMs can also be regarded as utterance-level confidence estimators, which may benefit many downstream tasks. Experiments on a 100hr LibriSpeech dataset show that R-EBMs can reduce the word error rates (WERs) by 8.2%/6.7% while improving areas under precision-recall curves of confidence scores by 12.6%/28.4% on test-clean/test-other sets. Furthermore, on a state-of-the-art model using self-supervised learning (wav2vec 2.0), R-EBMs still significantly improves both the WER and confidence estimation performance.
Qiujia Li, Yu Zhang 0033, Bo Li 0028, Liangliang Cao, Philip C. Woodland
Interspeech4
2021 Exploring Targeted Universal Adversarial Perturbations to End-to-End ASR Models
abstract
Although end-to-end automatic speech recognition (e2e ASR) models are widely deployed in many applications, there have been very few studies to understand models' robustness against adversarial perturbations.In this paper, we explore whether a targeted universal perturbation vector exists for e2e ASR models.Our goal is to find perturbations that can mislead the models to predict the given targeted transcript such as "thank you" or empty string on any input utterance.We study two different attacks, namely additive and prepending perturbations, and their performances on the state-of-the-art LAS, CTC and RNN-T models.We find that LAS is the most vulnerable to perturbations among the three models.RNN-T is more robust against additive perturbations, especially on long utterances.And CTC is robust against both additive and prepending perturbations.To attack RNN-T, we find prepending perturbation is more effective than the additive perturbation, and can mislead the models to predict the same short target on utterances of arbitrary length.
Zhiyun Lu, Wei Han 0002, Yu Zhang 0033, Liangliang Cao
Interspeech4
2021 Multi-Task Learning for End-to-End ASR Word and Utterance Confidence with Deletion Prediction
abstract
Confidence scores are very useful for downstream applications of automatic speech recognition (ASR) systems. Recent works have proposed using neural networks to learn word or utterance confidence scores for end-to-end ASR. In those studies, word confidence by itself does not model deletions, and utterance confidence does not take advantage of word-level training signals. This paper proposes to jointly learn word confidence, word deletion, and utterance confidence. Empirical results show that multi-task learning with all three objectives improves confidence metrics (NCE, AUC, RMSE) without the need for increasing the model size of the confidence estimation module. Using the utterance-level confidence for rescoring also decreases the word error rates on Google's Voice Search and Long-tail Maps datasets by 3-5% relative, without needing a dedicated neural rescorer.
David Qiu, Yanzhang He, Qiujia Li, Yu Zhang 0033, Liangliang Cao, Ian McGraw
Interspeech5
2021 RNN-T Models Fail to Generalize to Out-of-Domain Audio: Causes and Solutions
abstract
In recent years, all-neural end-to-end approaches have obtained state-of-the-art results on several challenging automatic speech recognition (ASR) tasks. However, most existing works focus on building ASR models where train and test data are drawn from the same domain. This results in poor generalization characteristics on mismatched-domains: e.g., end-to-end models trained on short segments perform poorly when evaluated on longer utterances. In this work, we analyze the generalization properties of streaming and non-streaming recurrent neural network transducer (RNN-T) based end-to-end models in order to identify model components that negatively affect generalization performance. We propose two solutions: combining multiple regularization techniques during training, and using dynamic overlapping inference. On a long-form YouTube test set, when the non-streaming RNN-T model is trained with shorter segments of data, the proposed combination improves word error rate (WER) from 22.3% to 14.8%; when the streaming RNN-T model trained on short Search queries, the proposed techniques improve WER on the YouTube set from 67.0% to 25.3%. Finally, when trained on Librispeech, we find that dynamic overlapping inference improves WER on YouTube from 99.8% to 33.0%.
Chung-Cheng Chiu, Arun Narayanan, Wei Han 0002, Rohit Prabhavalkar, Yu Zhang 0033, Navdeep Jaitly, Ruoming Pang, Tara N. Sainath, Patrick Nguyen, Liangliang Cao
SLT10
2020 Label-Efficient Learning on Point Clouds Using Approximate Convex Decompositions
Matheus Gadelha, Aruni Roy Chowdhury, Gopal Sharma, Evangelos Kalogerakis, Liangliang Cao, Erik G. Learned-Miller, Rui Wang 0003, Subhransu Maji
ECCV (10)5
2020 Speech Sentiment Analysis via Pre-Trained Features from End-to-End ASR Models
abstract
In this paper, we propose to use pre-trained features from end-to-end ASR models to solve speech sentiment analysis as a down-stream task. We show that end-to-end ASR features, which integrate both acoustic and text information from speech, achieve promising results. We use RNN with self-attention as the sentiment classifier, which also provides an easy visualization through attention weights to help interpret model predictions. We use well benchmarked IEMOCAP dataset and a new large-scale speech sentiment dataset SWBD-sentiment for evaluation. Our approach improves the-state-of-the-art accuracy on IEMOCAP from 66.6% to 71.7%, and achieves an accuracy of 70.10% on SWBD-sentiment with more than 49,500 utterances.
Zhiyun Lu, Liangliang Cao, Yu Zhang 0033, Chung-Cheng Chiu, James Fan
ICASSP2
2020 A Large Scale Speech Sentiment Corpus
abstract
We present a multimodal corpus for sentiment analysis based on the existing Switchboard-1 Telephone Speech Corpus released by the Linguistic Data Consortium. This corpus extends the Switchboard-1 Telephone Speech Corpus by adding sentiment labels from 3 different human annotators for every transcript segment. Each sentiment label can be one of three options: positive, negative, and neutral. Annotators are recruited using Google Cloud’s data labeling service and the labeling task was conducted over the internet. The corpus contains a total of 49500 labeled speech segments covering 140 hours of audio. To the best of our knowledge, this is the largest multimodal Corpus for sentiment analysis that includes both speech and text features.
Zhiyun Lu, Liangliang Cao, Yu Zhang 0033, James Fan
LREC4
2020 Deep Active Learning for Effective Pulmonary Nodule Detection
Jingya Liu, Liangliang Cao, Yingli Tian
MICCAI (6)2
2020 Product image recognition with guidance learning and noisy supervision
Qing Li 0058, Xiaojiang Peng, Liangliang Cao, Wenbin Du, Yu Qiao 0001, Qiang Peng
Comput. Vis. Image Underst.3
2019 Improving Object Detection from Scratch via Gated Feature Reuse
Humphrey Shi, NhatHai Phan, Rogério Feris, Liangliang Cao, Ding Liu 0001, Xinchao Wang, Thomas S. Huang, Marios Savvides
BMVC6
2019 Automatic Adaptation of Object Detectors to New Domains Using Self-Training
abstract
This work addresses the unsupervised adaptation of an existing object detector to a new target domain. We assume that a large number of unlabeled videos from this domain are readily available. We automatically obtain labels on the target data by using high-confidence detections from the existing detector, augmented with hard (misclassified) examples acquired by exploiting temporal cues using a tracker. These automatically-obtained labels are then used for re-training the original model. A modified knowledge distillation loss is proposed, and we investigate several ways of assigning soft-labels to the training examples from the target domain. Our approach is empirically evaluated on challenging face and pedestrian detection tasks: a face detector trained on WIDER-Face, which consists of high-quality images crawled from the web, is adapted to a large-scale surveillance data set; a pedestrian detector trained on clear, daytime images from the BDD-100K driving data set is adapted to all other scenarios such as rainy, foggy, night-time. Our results demonstrate the usefulness of incorporating hard examples obtained from tracking, the advantage of using soft-labels via distillation loss versus hard-labels, and show promising performance as a simple method for unsupervised domain adaptation of object detectors, with minimal dependence on hyper-parameters.
Aruni Roy Chowdhury, Prithvijit Chakrabarty, SouYoung Jin, Huaizu Jiang, Liangliang Cao, Erik G. Learned-Miller
CVPR6
2019 3DFPN-HS ^2 2 : 3D Feature Pyramid Network Based High Sensitivity and Specificity Pulmonary Nodule Detection
Jingya Liu, Liangliang Cao, Oguz Akin, Yingli Tian
MICCAI (6)2
2019 Focal Visual-Text Attention for Memex Question Answering
abstract
Recent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collections such as personal photo albums, we have to look at whole collections with sequences of photos. This paper proposes a new multimodal MemexQA task: given a sequence of photos from a user, the goal is to automatically answer questions that help users recover their memory about an event captured in these photos. In addition to a text answer, a few grounding photos are also given to justify the answer. The grounding photos are necessary as they help users quickly verifying the answer. Towards solving the task, we 1) present the MemexQA dataset, the first publicly available multimodal question answering dataset consisting of real personal photo albums; 2) propose an end-to-end trainable network that makes use of a hierarchical process to dynamically determine what media and what time to focus on in the sequential data to answer the question. Experimental results on the MemexQA dataset demonstrate that our model outperforms strong baselines and yields the most relevant grounding photos on this challenging task.
Junwei Liang 0001, Lu Jiang 0004, Liangliang Cao, Yannis Kalantidis, Li-Jia Li 0001, Alex Hauptmann 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Focal Visual-Text Attention for Visual Question Answering
abstract
Recent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collections such as personal photos, we have to look at whole collections with sequences of photos or videos. When answering questions from a large collection, a natural problem is to identify snippets to support the answer. In this paper, we describe a novel neural network called Focal Visual-Text Attention network (FVTA) for collective reasoning in visual question answering, where both visual and text sequence information such as images and text metadata are presented. FVTA introduces an end-to-end approach that makes use of a hierarchical process to dynamically determine what media and what time to focus on in the sequential data to answer the question. FVTA can not only answer the questions well but also provides the justifications which the system results are based upon to get the answers. FVTA achieves state-of-the-art performance on the MemexQA dataset and competitive results on the MovieQA dataset.
Junwei Liang 0001, Lu Jiang 0004, Liangliang Cao, Li-Jia Li 0001, Alex Hauptmann 0001
CVPR3
2018 Learning Deterministic Policy with Target for Power Control in Wireless Networks
abstract
Inter-Cell Interference Coordination (ICIC) is a promising way to improve energy efficiency in wireless networks, especially where small base stations are densely deployed. However, traditional optimization based ICIC schemes suffer from severe performance degradation with complex interference pattern. To address this issue, we propose a Deep Reinforcement Learning with Deterministic Policy and Target (DRL-DPT) framework for ICIC in wireless networks. DRL- DPT overcomes the main obstacles in applying reinforcement learning and deep learning in wireless networks, i.e. continuous state space, continuous action space and convergence. Firstly, a Deep Neural Network (DNN) is involved as the actor to obtain deterministic power control actions in continuous space. Then, to guarantee the convergence, an online training process is presented, which makes use of a dedicated reward function as the target rule and a policy gradient descent algorithm to adjust DNN weights. Experimental results show that the proposed DRL-DPT framework consistently outperforms existing schemes in terms of energy efficiency and throughput under different wireless interference scenarios. More specifically, it improves up to 15% of energy efficiency with faster convergence rate.
Yujiao Lu, Hancheng Lu, Liangliang Cao, Feng Wu 0001, Daren Zhu
GLOBECOM3
2018 Lip2Audspec: Speech Reconstruction from Silent Lip Movements Video
abstract
In this study, we propose a deep neural network for reconstructing intelligible speech from silent lip movement videos. We use auditory spectrogram as spectral representation of speech and its corresponding sound generation method resulting in a more natural sounding reconstructed speech. Our proposed network consists of an autoencoder to extract bottleneck features from the auditory spectrogram which is then used as target to our main lip reading network comprising of CNN, LSTM and fully connected layers. Our experiments show that the autoencoder is able to reconstruct the original auditory spectrogram with a 98% correlation and also improves the quality of reconstructed speech from the main lip reading network. Our model, trained jointly on different speakers is able to extract individual speaker characteristics and gives promising results of reconstructing intelligible speech with superior word recognition accuracy.
Hassan Akbari, Himani Arora, Liangliang Cao, Nima Mesgarani
ICASSP3
2018 Matrix Factorization on GPUs with Memory Optimization and Approximate Computing
abstract
Matrix factorization (MF) discovers latent features from observations, which has shown great promises in the fields of collaborative filtering, data compression, feature extraction, word embedding, etc. While many problem-specific optimization techniques have been proposed, alternating least square (ALS) remains popular due to its general applicability (e.g. easy to handle positive-unlabeled inputs), fast convergence and parallelization capability. Current MF implementations are either optimized for a single machine or with a need of a large computer cluster but still are insufficent. This is because a single machine provides limited compute power for large-scale data while multiple machines suffer from the network communication bottleneck.
Wei Tan 0001, Shiyu Chang, Liana L. Fong, Cheng Li 0014, Liangliang Cao
ICPP6
2017 Visual Memory QA: Your Personal Photo and Video Search Agent
abstract
The boom of mobile devices and cloud services has led to an explosion of personal photo and video data. However, due to the missing user-generated metadata such as titles or descriptions, it usually takes a user a lot of swipes to find some video on the cell phone. To solve the problem, we present an innovative idea called Visual Memory QA which allow a user not only to search but also to ask questions about her daily life captured in the personal videos. The proposed system automatically analyzes the content of personal videos without user-generated metadata, and offers a conversational interface to accept and answer questions. To the best of our knowledge, it is the first to answer personal questions discovered in personal photos or videos. The example questions are "what was the lat time we went hiking in the forest near San Francisco?"; "did we have pizza last week?"; "with whom did I have dinner in AAAI 2015?".
Lu Jiang 0004, Liangliang Cao, Yannis Kalantidis, Sachin Farfade, Alex Hauptmann 0001
AAAI2
2017 Learning from Noisy Labels with Distillation
abstract
The ability of learning from noisy labels is very useful in many visual recognition tasks, as a vast amount of data with noisy labels are relatively easy to obtain. Traditionally, label noise has been treated as statistical outliers, and techniques such as importance re-weighting and bootstrapping have been proposed to alleviate the problem. According to our observation, the real-world noisy labels exhibit multimode characteristics as the true labels, rather than behaving like independent random outliers. In this work, we propose a unified distillation framework to use “side” information, including a small clean dataset and label relations in knowledge graph, to “hedge the risk” of learning from noisy labels. Unlike the traditional approaches evaluated based on simulated label noises, we propose a suite of new benchmark datasets, in Sports, Species and Artifacts domains, to evaluate the task of learning from noisy labels in the practical setting. The empirical study demonstrates the effectiveness of our proposed method in all the domains.
Yuncheng Li, Jianchao Yang, Yale Song, Liangliang Cao, Jiebo Luo 0001, Li-Jia Li 0001
ICCV4
2017 ACM SIGMM Rising Star Award 2017
abstract
The ACM Special Interest Group on Multimedia (SIGMM) is pleased to present this year's Rising Star Award in multimedia computing, communications and applications to Dr. Liangliang Cao for his significant contributions in large-scale multimedia recognition and social media mining. The ACM SIGMM Rising Star Award recognizes a young researcher who has made outstanding research contributions to the field of multimedia computing, communication and applications during the early part of his or her career. Dr. Cao has published extensively in top multimedia related journals and conferences, including 15 in ACM Multimedia and 3 in ACM ICMR. To date, he has garnered 3400+ citations on over 70 papers and 10 patents. This impressive record in his early stage of career demonstrate the impact of his research and his contributions to our field of Multimedia. In his young research career, Dr. Cao has made unique and significant contributions in industrial settings. Most notably he was the Project key person for ALADDIN, on the IBM-Columbia team, the Lead contributor who made the IBM IMARS system 100 times faster, as well as the Lead contributor of a clothes & fashion search app at Yahoo Taiwan, which has received 80+ local media reports. He is a cofounder of HelloVera.AI where he is working as a CTO and Chief Scientist. ....
Liangliang Cao
ACM Multimedia1
2017 Delving Deep into Personal Photo and Video Search
abstract
The ubiquity of mobile devices and cloud services has led to an unprecedented growth of online personal photo and video collections. Due to the scarcity of personal media search log data, research to date has mainly focused on searching images and videos on the web. However, in order to manage the exploding amount of personal photos and videos, we raise a fundamental question: what are the differences and similarities when users search their own photos versus the photos on the web? To the best of our knowledge, this paper is the first to study personal media search using large-scale real-world search logs. We analyze different types of search sessions mined from Flickr search logs and discover a number of interesting characteristics of personal media search in terms of information needs and click behaviors. The insightful observations will not only be instrumental in guiding future personal media search methods, but also benefit related tasks such as personal photo browsing and recommendation. Our findings suggest there is a significant gap between personal queries and automatically detected concepts, which is responsible for the low accuracy of many personal media search queries. To bridge the gap, we propose the deep query understanding model to learn a mapping from the personal queries to the concepts in the clicked photos. Experimental results verify the efficacy of the proposed method in improving personal media search, where the proposed method consistently outperforms baseline methods.
Lu Jiang 0004, Yannis Kalantidis, Liangliang Cao, Sachin Farfade, Jiliang Tang, Alex Hauptmann 0001
WSDM3
2017 Guest editorial: mobile visual tagging with mobile context
Shuqiang Jiang, Liangliang Cao, Jiebo Luo 0001, Ramesh Jain 0001
Multim. Syst.2
2017 Mining Fashion Outfit Composition Using an End-to-End Deep Learning Approach on Set Data
abstract
Composing fashion outfits involves deep under-standing of fashion standards while incorporating creativity for choosing multiple fashion items (e.g., jewelry, bag, pants, dress). In fashion websites, popular or high-quality fashion outfits are usually designed by fashion experts and followed by large audiences. In this paper, we propose a machine learning system to compose fashion outfits automatically. The core of the proposed automatic composition system is to score fashion outfit candidates based on the appearances and metadata. We propose to leverage outfit popularity on fashion-oriented websites to supervise the scoring component. The scoring component is a multimodal multiinstance deep learning system that evaluates instance aesthetics and set compatibility simultaneously. In order to train and evaluate the proposed composition system, we have collected a large-scale fashion outfit dataset with 195K outfits and 368K fashion items from Polyvore. Although the fashion outfit scoring and composition is rather challenging, we have achieved an AUC of 85% for the scoring component, and an accuracy of 77% for a constrained composition task.
Yuncheng Li, Liangliang Cao, Jiebo Luo 0001
IEEE Trans. Multim.2
2017 Context-Associative Hierarchical Memory Model for Human Activity Recognition and Prediction
abstract
Human activity recognition is a challenging high-level vision task, for which multiple factors, such as subject, object, and their diverse interactions, have to be considered and modeled. Current learning-based methods are limited in the capability to integrate human-level concepts into an easily extensible computational framework. Inspired by the existing human memory model, we present a context-associative approach to recognize activity with human-object interaction. The proposed system can recognize incoming visual content based on the previous experienced activities. The high-level activity is parsed into consecutive subactivities, and we build a context cluster to model the temporal relations. The semantic attributes of the subactivity are organized by a concept hierarchy. Based on the hierarchy, a series of similarity functions are defined to turn the recognition computing into retrievals over the contextual memory, similar to the auto-associative characteristics of human memory. Partially matching in retrieval and stored memory make the activity prediction possible. The dynamical evolution of the brain memory is mimicked to allow decay and reinforcement of the input information, providing a natural way to maintain data and save computational time. We evaluate our approach on three data sets: CAD-120, MHOI, and OPPORTUNITY. The proposed method demonstrates promising results compared with other state-of-the-art techniques.
Lei Wang 0060, Xu Zhao 0001, Yunfei Si, Liangliang Cao, Yuncai Liu
IEEE Trans. Multim.4
2017 Image-Based Appraisal of Real Estate Properties
abstract
Real estate appraisal, which is the process of estimating the price for real estate properties, is crucial for both buyers and sellers as the basis for negotiation and transaction. Traditionally, the repeat sales model has been widely adopted to estimate real estate prices. However, it depends on the design and calculation of a complex economic-related index, which is challenging to estimate accurately. Today, real estate brokers provide easy access to detailed online information on real estate properties to their clients. We are interested in estimating the real estate price from these large amounts of easily accessed data. In particular, we analyze the prediction power of online house pictures, which is one of the key factors for online users to make a potential visiting decision. The development of robust computer vision algorithms makes the analysis of visual content possible. In this paper, we employ a recurrent neural network to predict real estate prices using the state-of-the-art visual features. The experimental results indicate that our model outperforms several other state-of-the-art baseline algorithms in terms of both mean absolute error and mean absolute percentage error.
Quanzeng You, Liangliang Cao, Jiebo Luo 0001
IEEE Trans. Multim.3
2016 Poker-CNN: A Pattern Learning Strategy for Making Draws and Bets in Poker Games Using Convolutional Networks
abstract
Poker is a family of card games that includes many varia- tions. We hypothesize that most poker games can be solved as a pattern matching problem, and propose creating a strong poker playing system based on a unified poker representa- tion. Our poker player learns through iterative self-play, and improves its understanding of the game by training on the results of its previous actions without sophisticated domain knowledge. We evaluate our system on three poker games: single player video poker, two-player Limit Texas Hold’em, and finally two-player 2-7 triple draw poker. We show that our model can quickly learn patterns in these very different poker games while it improves from zero knowledge to a competi- tive player against human experts. The contributions of this paper include: (1) a novel represen- tation for poker games, extendable to different poker vari- ations, (2) a Convolutional Neural Network (CNN) based learning model that can effectively learn the patterns in three different games, and (3) a self-trained system that signif- icantly beats the heuristic-based program on which it is trained, and our system is competitive against human expert players.
Nikolai Yakovenko, Liangliang Cao, Colin Raffel, James Fan
AAAI2
2016 Multi-Scale Fully Convolutional Network for Fast Face Detection
Yancheng Bai, Wenjing Ma, Yucheng Li 0002, Liangliang Cao, Luwei Yang
BMVC4
2016 Video2GIF: Automatic Generation of Animated GIFs from Video
abstract
We introduce the novel problem of automatically generating animated GIFs from video. GIFs are short looping video with no sound, and a perfect combination between image and video that really capture our attention. GIFs tell a story, express emotion, turn events into humorous moments, and are the new wave of photojournalism. We pose the question: Can we automate the entirely manual and elaborate process of GIF creation by leveraging the plethora of user generated GIF content? We propose a Robust Deep RankNet that, given a video, generates a ranked list of its segments according to their suitability as GIF. We train our model to learn what visual content is often selected for GIFs by using over 100K user generated GIFs and their corresponding video sources. We effectively deal with the noisy web data by proposing a novel adaptive Huber loss in the ranking formulation. We show that our approach is robust to outliers and picks up several patterns that are frequently present in popular animated GIFs. On our new large-scale benchmark dataset, we show the advantage of our approach over several state-of-the-art methods.
Michael Gygli, Yale Song, Liangliang Cao
CVPR3
2016 TGIF: A New Dataset and Benchmark on Animated GIF Description
abstract
With the recent popularity of animated GIFs on social media, there is need for ways to index them with rich meta-data. To advance research on animated GIF understanding, we collected a new dataset, Tumblr GIF (TGIF), with 100K animated GIFs from Tumblr and 120K natural language descriptions obtained via crowdsourcing. The motivation for this work is to develop a testbed for image sequence description systems, where the task is to generate natural language descriptions for animated GIFs or video clips. To ensure a high quality dataset, we developed a series of novel quality controls to validate free-form text input from crowd-workers. We show that there is unambiguous association between visual content and natural language descriptions in our dataset, making it an ideal benchmark for the visual content captioning task. We perform extensive statistical analyses to compare our dataset to existing image and video description datasets. Next, we provide baseline results on the animated GIF description task, using three representative techniques: nearest neighbor, statistical machine translation, and recurrent neural networks. Finally, we show that models fine-tuned from our animated GIF description dataset can be helpful for automatic movie description.
Yuncheng Li, Yale Song, Liangliang Cao, Joel R. Tetreault, Larry Goldberg, Alejandro Jaimes, Jiebo Luo 0001
CVPR3
2016 Faster and Cheaper: Parallelizing Large-Scale Matrix Factorization on GPUs
abstract
Matrix factorization (MF) is used by many popular algorithms such as collaborative filtering. GPU with massive cores and high memory bandwidth sheds light on accelerating MF much further when appropriately exploiting its architectural characteristics.
Wei Tan 0001, Liangliang Cao, Liana L. Fong
HPDC2
2016 Building Joint Spaces for Relation Extraction
Chang Wang 0001, Liangliang Cao, James Fan
IJCAI2
2016 Incremental Learning for Fine-Grained Image Recognition
abstract
This paper considers the problem of fine-grained image recognition with a growing vocabulary. Since in many real world applications we often have to add a new object category or visual concept with just a few images to learn from, it is crucial to develop a method that is able to generalize the recognition model from existing classes to new classes. Deep convolutional neural networks are capable of constructing powerful image representations; however, these networks usually rely on a logistic loss function that cannot handle the incremental learning problem. In this paper, we present a new method that can efficiently learn a new class given only a limited number of training examples, which we evaluate on the problems of food and clothing recognition. To illustrate the performance of our proposed method on the task of recognizing different kinds of food, when using only 1.3\% of training examples per category we achieved about 73\% of the performance (as measured by F1-score) compared to when using all available training data.
Liangliang Cao, Jenhao Hsiao, Paloma de Juan, Yuncheng Li, Bart Thomee
ICMR1
2016 GPU-FV: Realtime Fisher Vector and Its Applications in Video Monitoring
abstract
Fisher vector has been widely used in many multimedia retrieval and visual recognition applications with good performance. However, the computation complexity prevents its usage in real-time video monitoring. In this work, we proposed and implemented GPU-FV, a fast Fisher vector extraction method with the help of modern GPUs. The challenge of implementing Fisher vector on GPUs lies in the data dependency in feature extraction and expensive memory access in Fisher vector computing. To handle these challenges, we carefully designed GPU-FV in a way that utilizes the computing power of GPU as much as possible, and applied optimizations such as loop tiling to boost the performance. GPU-FV is about 12 times faster than the CPU version, and 50\% faster than a non-optimized GPU implementation. For standard video input (320*240), GPU-FV can process each frame within 34ms on a model GPU. Our experiments show that GPU-FV obtains a similar recognition accuracy as traditional FV on VOC 2007 and Caltech 256 image sets. We also applied GPU-FV for realtime video monitoring tasks and found that GPU-FV outperforms a number of previous works. Especially, when the number of training examples are small, GPU-FV outperforms the recent popular deep CNN features borrowed from ImageNet.
Wenjing Ma, Liangliang Cao, Lei Yu 0012, Guoping Long, Yucheng Li 0002
ICMR2
2016 Detecting Sarcasm in Multimodal Social Platforms
abstract
Sarcasm is a peculiar form of sentiment expression, where the surface sentiment differs from the implied sentiment. The detection of sarcasm in social media platforms has been applied in the past mainly to textual utterances where lexical indicators (such as interjections and intensifiers), linguistic markers, and contextual information (such as user profiles, or past conversations) were used to detect the sarcastic tone. However, modern social media platforms allow to create multimodal messages where audiovisual content is integrated with the text, making the analysis of a mode in isolation partial. In our work, we first study the relationship between the textual and visual aspects in multimodal posts from three major social media platforms, i.e., Instagram, Tumblr and Twitter, and we run a crowdsourcing task to quantify the extent to which images are perceived as necessary by human annotators. Moreover, we propose two different computational frameworks to detect sarcasm that integrate the textual and visual modalities. The first approach exploits visual semantics trained on an external dataset, and concatenates the semantics features with state-of-the-art textual features. The second method adapts a visual neural network initialized with parameters trained on ImageNet to multimodal sarcastic posts. Results show the positive effect of combining modalities for the detection of sarcasm across platforms and methods.
Rossano Schifanella, Paloma de Juan, Joel R. Tetreault, Liangliang Cao
ACM Multimedia4
2016 Robust Visual-Textual Sentiment Analysis: When Attention meets Tree-structured Recursive Neural Networks
abstract
Sentiment analysis is crucial for extracting social signals from social media content. Due to huge variation in social media, the performance of sentiment classifiers using single modality (visual or textual) still lags behind satisfaction. In this paper, we propose a new framework that integrates textual and visual information for robust sentiment analysis. Different from previous work, we believe visual and textual information should be treated jointly in a structural fashion. Our system first builds a semantic tree structure based on sentence parsing, aimed at aligning textual words and image regions for accurate analysis. Next, our system learns a robust joint visual-textual semantic representation by incorporating 1) an attention mechanism with LSTM (long short term memory) and 2) an auxiliary semantic learning task. Extensive experimental results on several known data sets show that our method outperforms existing the state-of-the-art joint models in sentiment analysis. We also investigate different tree-structured LSTM (T-LSTM) variants and analyze the effect of the attention mechanism in order to provide deeper insight on how the attention mechanism helps the learning of the joint visual-textual sentiment classifier.
Quanzeng You, Liangliang Cao, Hailin Jin, Jiebo Luo 0001
ACM Multimedia2
2016 A hybrid term-term relations analysis approach for topic detection
Chen Zhang 0003, Hao Wang 0005, Liangliang Cao, Wei Wang 0061, Fanjiang Xu
Knowl. Based Syst.3
2015 You are what you tweet...pic! gender prediction based on semantic analysis of social media images
abstract
We propose a method to extract user attributes from the pictures posted in social media feeds, specifically gender information. While traditional approaches rely on text analysis or exploit visual information only from the user profile picture or colors, we propose to look at the distribution of semantics in the pictures coming from the whole feed of a person to estimate gender. In order to compute such semantic distribution, we trained models from existing visual taxonomies to recognize objects, scenes and activities, and applied them to the images in each user's feed. Experiments conducted on a set of ten thousand twitter users and their collection of half a million images revealed that the gender signal can indeed be extracted from the users image feed (75.6% accuracy). Furthermore, the combination of visual cues resulted almost as strong as textual analysis in predicting gender, while providing complementary information that can be employed to further boost gender prediction accuracy to 88% when combined with textual data. As a byproduct of our investigation, we were also able to extrapolate the semantic categories of posted pictures mostly correlated to males and females.
Michele Merler, Liangliang Cao, John R. Smith
ICME2
2015 Medical Synonym Extraction with Concept Space Models
Chang Wang 0001, Liangliang Cao, Bowen Zhou 0006
IJCAI2
2015 Multi-facet Learning using Deep Convolutional Neural Network for Person-Related Categories in Photos
abstract
This paper proposes to leverage multiple facets of person photos to improve the training of deep neural networks. Existing studies usually require a lot of labeled images to train deep convolutional networks. Our study suggests exploring multiple datasets and learning effective representation to learn related visual concepts. The practice of learning from multiple facets implicitly enforces to share features for image recognition. We show deep neural network benefits from the learning of multiple person-related categories in photos. Faceted classification systems learn from multiple resources, and alleviate the overfitting problems in deep learning. Moreover, by exploring multiple taxonomies of an object, it provides a finer annotation for the query images.
Liangliang Cao, Zhicheng Yan 0001, John R. Smith
ICMR1
2015 Automated Axon Segmentation from Highly Noisy Microscopic Videos
abstract
We present a novel method for automated segmentation of axons in extremely noisy videos obtained via two-photon microscopy in awake mice. We formulate segmentation as a pixel-wise classification problem in which a pixel is classified into "axon" or "non-axon" based on its feature vector. In order to deal with high levels of noise, the features of our classifier are derived from spatio-temporal Independent Component Analysis (stICA) which effectively isolates noise from signal components while leveraging temporal coherence from the video. We fit parametric models to represent the distribution of the extracted features and apply a probabilistic classifier over stICA components to determine the label of each pixel. Finally, we show compelling qualitative and quantitative results from very challenging two-photon microscopic, demonstrating the usefulness of our approach. An example time-series of two-photon images with our automated ROI extraction over layed is available with the supplemental materials.
John Bowler, Rogério Feris, Liangliang Cao
WACV3
2015 Max-Confidence Boosting With Uncertainty for Visual Tracking
abstract
The challenges in visual tracking call for a method which can reliably recognize the subject of interests in an environment, where the appearance of both the background and the foreground change with time. Many existing studies model this problem as tracking by classification with online updating of the classification models, however, most of them overlook the ambiguity in visual modeling and do not consider the prior information in the tracking process. In this paper, we present a novel visual tracking method called max-confidence boosting (MCB), which explores a new way of online updating ambiguous visual phenomenon. The MCB framework models uncertainty in prior knowledge utilizing the indeterministic labels, which are used in updating models from previous frames and the new frame. Our proposed MCB tracker allows ambiguity in the tracking process and can effectively alleviate the drift problem. Many experimental results in challenging video sequences verify the success of our method, and our MCB tracker outperforms a number of the state-of-the-art tracking by classification methods.
Wen Guo 0003, Liangliang Cao, Tony X. Han, Shuicheng Yan, Changsheng Xu
IEEE Trans. Image Process.2
2015 A Multifaceted Approach to Social Multimedia-Based Prediction of Elections
abstract
Compared with real-world polling, election prediction based on social media can be far more timely and cost-effective due to the immediate availability of fast evolving Web contents. However, information from social media may suffer from noise and sampling bias that are caused by various factors and thus pose one of biggest challenges in social media-based data analytics. This paper presents a new model, named competitive vector auto regression (CVAR), to build a reliable forecasting system for the US presidential elections and US House race. Our CVAR model is designed to analyze the correlation between image-centric social multimedia and real-world phenomena. By introducing the competition mechanism, CVAR compares the popularity among multiple competing candidates. More importantly , CVAR is able to combine visual information with textual information from rich and multifaceted social multimedia, which helps extract reliable signals and mitigate sampling bias. As a result, our proposed system can 1) accurately predict the election outcome, 2) infer the sentiment of the candidate photos shared in the social media communities, and 3) account for the sentiment of viewer comments towards the candidates on the related images. The experiments on the 2012 US presidential election at both national and state levels, as well as the 2014 US House race, have demonstrated the power and promise of the proposed approach.
Quanzeng You, Liangliang Cao, Yang Cong, Xianchao Zhang 0001, Jiebo Luo 0001
IEEE Trans. Multim.2
2014 Cuteness Recognition and Localization in the Photos of Animals
abstract
Among the flourishing amount of photos in the social media websites, "cute" images of animals are particularly attractive to the Internet users. This paper considers building an automatic model which can distinguish cute images from non-cute ones. To make the recognition results more interpretable, a lot of efforts are made to find which part of the animal appears attractive to the human users. To validate the success of our proposed method, we collect three new datasets of different animals, i.e., cats, dogs, and rabbits with both cute and non-cute images. Our model obtains promising performance in distinguishing cute images from non-cute ones. Moreover, it outperforms the classical models with not only better recognition accuracy, but also more intuitive localization of the cuteness in the images. The contribution of this paper is three-fold: (1) We collect new datasets for cuteness recognition, (2) We extend the powerful Fisher Vector representation to localize cute part in the animal recognition, and (3) Extensive experimental results show that our proposed method can recognize cute animals of cats, dogs, and rabbits.
Liangliang Cao, Jinhui Tang 0001
ACM Multimedia3
2014 GeoMM 2014: the third ACM multimedia workshop ongeotagging and its applications in multimedia
abstract
It is our great pleasure to welcome you to the Third ACM Workshop on Geotagging and Its Applications in Multimedia -- GeoMM'14. This year's event continues the workshops in 2012 and 2013, with the goal of building a forum for the presentation and synthesis of vision and insight from leading experts and practitioners on the developing directions of geotagging research related to multimedia. Following the success in previous years, the GeoMM workshop serves as a venue for the premier research in geotagging and multimedia, and continues to attract submissions from a diverse set of researchers, who address newly arising problems within this emerging field. Five regular papers are presented in this workshop, covering a number of novel applications and new methodologies. An invited paper is also presented to introduce the related MediaEval 2014 Placing task, which consists of 5 million geotagged photos and 25,000 geotagged videos. We believe this workshop will benefit more and more research works in the broad research field.
Liangliang Cao, Gerald Friedland, Lexing Xie
ACM Multimedia1
2014 Modeling Attributes from Category-Attribute Proportions
abstract
Attribute-based representation has been widely used in visual recognition and retrieval due to its interpretability and cross-category generalization properties. However, classic attribute learning requires manually labeling attributes on the images, which is very expensive, and not scalable. In this paper, we propose to model attributes from category-attribute proportions. The proposed framework can model attributes without attribute labels on the images. Specifically, given a multi-class image datasets with N categories, we model an attribute, based on an N-dimensional category-attribute proportion vector, where each element of the vector characterizes the proportion of images in the corresponding category having the attribute. The attribute learning can be formulated as a learning from label proportion (LLP) problem. Our method is based on a newly proposed machine learning algorithm called $\propto$SVM. Finding the category-attribute proportions is much easier than manually labeling images, but it is still not a trivial task. We further propose to estimate the proportions from multiple modalities such as human commonsense knowledge, NLP tools, and other domain knowledge. The value of the proposed approach is demonstrated by various applications including modeling animal attributes, visual sentiment attributes, and scene attributes.
Felix X. Yu, Liangliang Cao, Michele Merler, Noel Codella, Tao Chen 0015, John R. Smith, Shih-Fu Chang
ACM Multimedia2
2014 Learning mid-level features from object hierarchy for image classification
abstract
We propose a new approach for constructing mid-level visual features for image classification. We represent an image using the outputs of a collection of binary classifiers. These binary classifiers are trained to differentiate pairs of object classes in an object hierarchy. Our feature representation implicitly captures the hierarchical structure in object classes. We show that our proposed approach outperforms other baseline methods in image classification.
Somayah Albaradei, Yang Wang 0003, Liangliang Cao, Li-Jia Li 0001
WACV3
2014 A spatial-color layout feature for representing galaxy images
abstract
We propose a spatial-color layout feature specially designed for galaxy images. Inspired by findings on galaxy formation and evolution from Astronomy, the proposed feature captures both global and local morphological information of galaxies. In addition, our feature is scale and rotation invariant. By developing a hashing-based approach with the proposed feature, we implemented an efficient galaxy image retrieval system on a dataset with more than 280 thousand galaxy images from the Sloan Digital Sky Survey project. Given a query image, the proposed system can rank-order all galaxies from the dataset according to relevance in only 35 milliseconds on a single PC. To the best of our knowledge, this is one of the first works on galaxy-specific feature design and large-scale galaxy image retrieval. We evaluated the performance of the proposed feature and the galaxy image retrieval system using web user annotations, showing that the proposed feature outperforms other classic features, including HOG, Gist, LBP, and Color-histograms. The success of our retrieval system demonstrates the advantages of leveraging computer vision techniques in Astronomy problems.
Yin Cui, Yongzhou Xiang, Kun Rong, Rogério Feris, Liangliang Cao
WACV5
2014 Guest Editorial: Special issue on large scale multimedia semantic indexing
Yadong Mu, Yi Yang 0001, Liangliang Cao, Shuicheng Yan, Qi Tian 0001
Comput. Vis. Image Underst.3
2013 Efficient Maximum Appearance Search for Large-Scale Object Detection
abstract
In recent years, efficiency of large-scale object detection has arisen as an important topic due to the exponential growth in the size of benchmark object detection datasets. Most current object detection methods focus on improving accuracy of large-scale object detection with efficiency being an afterthought. In this paper, we present the Efficient Maximum Appearance Search (EMAS) model which is an order of magnitude faster than the existing state-of-the-art large-scale object detection approaches, while maintaining comparable accuracy. Our EMAS model consists of representing an image as an ensemble of densely sampled feature points with the proposed Point wise Fisher Vector encoding method, so that the learnt discriminative scoring function can be applied locally. Consequently, the object detection problem is transformed into searching an image sub-area for maximum local appearance probability, thereby making EMAS an order of magnitude faster than the traditional detection methods. In addition, the proposed model is also suitable for incorporating global context at a negligible extra computational cost. EMAS can also incorporate fusion of multiple features, which greatly improves its performance in detecting multiple object categories. Our experiments show that the proposed algorithm can perform detection of 1000 object classes in less than one minute per image on the Image Net ILSVRC2012 dataset and for 107 object classes in less than 5 seconds per image for the SUN09 dataset using a single CPU.
Qiang Chen 0007, Rogério Feris, Ankur Datta, Liangliang Cao, ZhongYang Huang, Shuicheng Yan
CVPR5
2013 Learning Locally-Adaptive Decision Functions for Person Verification
abstract
This paper considers the person verification problem in modern surveillance and video retrieval systems. The problem is to identify whether a pair of face or human body images is about the same person, even if the person is not seen before. Traditional methods usually look for a distance (or similarity) measure between images (e.g., by metric learning algorithms), and make decisions based on a fixed threshold. We show that this is nevertheless insufficient and sub-optimal for the verification problem. This paper proposes to learn a decision function for verification that can be viewed as a joint model of a distance metric and a locally adaptive thresholding rule. We further formulate the inference on our decision function as a second-order large-margin regularization problem, and provide an efficient algorithm in its dual from. We evaluate our algorithm on both human body verification and face verification problems. Our method outperforms not only the classical metric learning algorithm including LMNN and ITML, but also the state-of-the-art in the computer vision community.
Zhen Li 0028, Shiyu Chang, Feng Liang 0002, Thomas S. Huang, Liangliang Cao, John R. Smith
CVPR5
2013 Designing Category-Level Attributes for Discriminative Visual Recognition
abstract
Attribute-based representation has shown great promises for visual recognition due to its intuitive interpretation and cross-category generalization property. However, human efforts are usually involved in the attribute designing process, making the representation costly to obtain. In this paper, we propose a novel formulation to automatically design discriminative "category-level attributes", which can be efficiently encoded by a compact category-attribute matrix. The formulation allows us to achieve intuitive and critical design criteria (category-separability, learn ability) in a principled way. The designed attributes can be used for tasks of cross-category knowledge transfer, achieving superior performance over well-known attribute dataset Animals with Attributes (AwA) and a large-scale ILSVRC2010 dataset (1.2M images). This approach also leads to state-of-the-art performance on the zero-shot learning task on AwA.
Felix X. Yu, Liangliang Cao, Rogério Feris, John R. Smith, Shih-Fu Chang
CVPR2
2013 Large-scale video event classification using dynamic temporal pyramid matching of visual semantics
abstract
Video event classification and retrieval has recently emerged as a challenging research topic. In addition to the variation in appearance of visual content and the large scale of the collections to be analyzed, this domain presents new and unique challenges in the modeling of the explicit temporal structure and implicit temporal trends of content within the video events. In this study, we present a technique for video event classification that captures temporal information over semantics using a scalable and efficient modeling scheme. An architecture for partitioning videos into a linear temporal pyramid, using segments of equal length and segments determined by the patterns of the underlying data, is applied over a rich underlying semantic description at the frame level using a taxonomy of nearly 1000 concepts containing 500,000 training images. Forward model selection with data bagging is used to prune the space of temporal features and data for efficiency. The system is implemented in the Hadoop Map-Reduce environment for arbitrary scalability. Our method is applied to the TRECVID Multimedia Event Detection 2012 task. Results demonstrate a significant boost in performance of over 50%, in terms of mean average precision, compared to common max or average pooling, and 17.7% compared to more complex pooling strategies that ignore temporal content.
Noel Codella, Gang Hua 0001, Liangliang Cao, Michele Merler, Leiguang Gong, Matthew L. Hill, John R. Smith
ICIP3
2013 Learning by focusing: A new framework for concept recognition and feature selection
abstract
In this paper, we develop a new method for feature selection and category learning. We first introduce two observations from our experiments: (1) It is easier to distinguish two concepts than to learn an isolated concept. (2) To distinguish different concept pairs we can find different selections of optimal features. These two observations may partly explain the success of human vision learning, especially why an infant can simultaneously capture distinguished visual features when learning new concepts. Based on these two observations, we developed a new learning-by-focusing method which first constructs focalized concept discriminators for pairs of concepts, and then builds nonlinear classifiers using the discrimination scores. We build datasets for four concept structure: vehicle, human affliction, sports, and animals, and experiments on all the four datasets verify the success of our new approach.
Liangliang Cao, Leiguang Gong, John R. Kender, Noel Codella, John R. Smith
ICME1
2013 Second ACM multimedia workshop on geotagging and its applications in multimedia (GeoMM 2013)
abstract
The Workshop on Geotagging and Its Applications in Multimedia (GeoMM 2013) focuses on new applications and methods of geotagging and in geo-location support systems. As the location based multimedia becomes more and more popular in the era of Web and mobile applications, the increase in the use of geotagging and improvements in geo-location support systems open up a new dimension for the description, organization and manipulation of multimedia data. This new dimension radically expands the usefulness of multimedia data both for daily users of the Internet and social networking sites as well as for experts in particular application scenarios. The workshop serves as a venue for the premier research in geotagging and multimedia, and continues to attract submissions from a diverse set of researchers, who address newly arising problems within this emerging field.
Liangliang Cao, Gerald Friedland, Pascal Kelm
ACM Multimedia1
2013 Learning latent spatio-temporal compositional model for human action recognition
abstract
Action recognition is an important problem in multimedia understanding. This paper addresses this problem by building an expressive compositional action model. We model one action instance in the video with an ensemble of spatio-temporal compositions: a number of discrete temporal anchor frames, each of which is further decomposed to a layout of deformable parts. In this way, our model can identify a Spatio-Temporal And-Or Graph (STAOG) to represent the latent structure of actions \emph{e.g.} triple jumping, swinging and high jumping. The STAOG model comprises four layers: (i) a batch of leaf-nodes in bottom for detecting various action parts within video patches; (ii) the or-nodes over bottom, i.e. switch variables to activate their children leaf-nodes for structural variability; (iii) the and-nodes within an anchor frame for verifying spatial composition; and (iv) the root-node at top for aggregating scores over temporal anchor frames. Moreover, the contextual interactions are defined between leaf-nodes in both spatial and temporal domains. For model training, we develop a novel weakly supervised learning algorithm which iteratively determines the structural configuration (e.g. the production of leaf-nodes associated with the or-nodes) along with the optimization of multi-layer parameters. By fully exploiting spatio-temporal compositions and interactions, our approach handles well large intra-class action variance (\emph{e.g.} different views, individual appearances, spatio-temporal structures). The experimental results on the challenging databases demonstrate superior performance of our approach over other methods.
Xiaodan Liang, Liangliang Cao
ACM Multimedia3
2013 Massive-scale multimedia semantic modeling
abstract
Visual data is exploding! 500 billion consumer photos are taken each year world-wide, 633 million photos taken per year in NYC alone. 120 new video-hours are uploaded on YouTube per minute. The explosion of digital multimedia data is creating a valuable open source for insights. However, the unconstrained nature of 'image/video in the wild' makes it very challenging for automated computer-based analysis. Furthermore, the most interesting content in the multimedia files is often complex in nature reflecting a diversity of human behaviors, scenes, activities and events. To address these challenges, this tutorial will provide a unified overview of the two emerging techniques: Semantic modeling and Massive scale visual recognition, with a goal of both introducing people from different backgrounds to this exciting field and reviewing state of the art research in the new computational era.
John R. Smith, Liangliang Cao
ACM Multimedia2
2013 Discovering Latent Clusters from Geotagged Beach Images
Yang Wang 0003, Liangliang Cao
MMM (2)2
2013 Introduction to the special section of best papers of ACM multimedia 2012
abstract
No abstract available.
Ioannis Kompatsiaris, Wenjun Zeng 0001, Gang Hua 0001, Liangliang Cao
ACM Trans. Multim. Comput. Commun. Appl.4
2012 Scene Aligned Pooling for Complex Video Recognition
Liangliang Cao, Yadong Mu, Apostol Natsev, Shih-Fu Chang, Gang Hua 0001, John R. Smith
ECCV (2)1
2012 Video Event Detection Using Temporal Pyramids of Visual Semantics with Kernel Optimization and Model Subspace Boosting
abstract
In this study, we present a system for video event classification that generates a temporal pyramid of static visual semantics using minimum-value, maximum-value, and average-value aggregation techniques. Kernel optimization and model subspace boosting are then applied to customize the pyramid for each event. SVM models are independently trained for each level in the pyramid using kernel selection according to 3-fold cross-validation. Kernels that both enforce static temporal order and permit temporal alignment are evaluated. Model subspace boosting is used to select the best combination of pyramid levels and aggregation techniques for each event. The NIST TRECVID Multimedia Event Detection (MED) 2011 dataset was used for evaluation. Results demonstrate that kernel optimizations using both temporally static and dynamic kernels together achieves better performance than any one particular method alone. In addition, model sub-space boosting reduces the size of the model by 80%, while maintaining 96% of the performance gain.
Noel Codella, Apostol Natsev, Gang Hua 0001, Matthew L. Hill, Liangliang Cao, Leiguang Gong, John R. Smith
ICME5
2012 GeoMM'12: ACM international workshop on geotagging and its applications in multimedia
abstract
Geotagging is the process of adding geographical identification metadata to various media files such as photos, videos, websites, messages, and tweets. It is not limited to GPS sensor data but an extension of current multimedia files with a wide variety of location-specific information. The GeoMM'12 workshop presents research on recent research on geotagging within the context of multimedia analysis. This workshop aims to not only provide more cutting edge algorithms, but also motivate novel applications in this promising field.
Liangliang Cao, Gerald Friedland, Martha A. Larson
ACM Multimedia1
2012 Submodular video hashing: a unified framework towards video pooling and indexing
abstract
This paper develops a novel framework for efficient large-scale video retrieval. We aim to find video according to higher level similarities, which is beyond the scope of traditional near duplicate search. Following the popular hashing technique we employ compact binary codes to facilitate nearest neighbor search. Unlike the previous methods which capitalize on only one type of hash code for retrieval, this paper combines heterogeneous hash codes to effectively describe the diverse and multi-scale visual contents in videos. Our method integrates feature pooling and hashing in a single framework. In the pooling stage, we cast video frames into a set of pre-specified components, which capture a variety of semantics of video contents. In the hashing stage, we represent each video component as a compact hash code, and combine multiple hash codes into hash tables for effective search. To speed up the retrieval while retaining most informative codes, we propose a graph-based influence maximization method to bridge the pooling and hashing stages. We show that the influence maximization problem is submodular, which allows a greedy optimization method to achieve a nearly optimal solution. Our method works very efficiently, retrieving thousands of video clips from TRECVID dataset in about 0.001 second. For a larger scale synthetic dataset with 1M samples, it uses less than 1 second in response to 100 queries. Our method is extensively evaluated in both unsupervised and supervised scenarios, and the results on TRECVID Multimedia Event Detection and Columbia Consumer Video datasets demonstrate the success of our proposed technique.
Liangliang Cao, Zhenguo Li, Yadong Mu, Shih-Fu Chang
ACM Multimedia1
2012 RankCompete: Simultaneous ranking and clustering of information networks
Liangliang Cao, Xin Jin 0001, Zhijun Yin, Andrey Del Pozo, Jiebo Luo 0001, Jiawei Han 0001, Thomas S. Huang
Neurocomputing1
2012 Web-Scale Multimedia Information Networks
abstract
The abundance of multimedia data on the Web presents both challenges (how to annotate, search, and mine) and opportunities (crawling the Web to create large structured multimedia data bases which can be used to do inference effectively). Because of the huge data volume, considering all semantic concepts as on the same (flat) level is not viable. In this paper, we introduce a unified STRUCTURED representation called multimedia information networks (MINets), which incorporates ontology and cross-media links, covering both content and context knowledge. Ontology and cross-media structures are constructed and expanded by automatically constructing MINets from web-scale data by state-of-the-art information extraction and knowledge-based population techniques. The resultant MINet will contain a wide range of linkages, including logical, statistical, and semantic relations among informative concept nodes, which connects proliferative ontology as well as cross-media web-scale resources together. The raw data collected in construction phase often contain much noisy, incomplete, or even conflicting information which could be detrimental to information extraction and utilization. Then, the redundant link structure can be utilized to distill MINets and improve quality of information (QoI). Moreover, advanced inference theory and system can be built upon the linked MINets, and then high-level ontological knowledge can be inferred and integrated in a logically harmonious network structure in MINets which is consistent with human cognition. Even more, as information channels, the ontology and cross-media links in MINets connect informative knowledge resources together, which makes it possible to increase the portability of information between different resources to increase information utilization levels.
Guo-Jun Qi, Min-Hsuan Tsai, Shen-Fu Tsai, Liangliang Cao, Thomas S. Huang
Proc. IEEE4
2012 Latent Community Topic Analysis: Integration of Community Discovery with Topic Modeling
abstract
This article studies the problem of latent community topic analysis in text-associated graphs. With the development of social media, a lot of user-generated content is available with user networks. Along with rich information in networks, user graphs can be extended with text information associated with nodes. Topic modeling is a classic problem in text mining and it is interesting to discover the latent topics in text-associated graphs. Different from traditional topic modeling methods considering links, we incorporate community discovery into topic analysis in text-associated graphs to guarantee the topical coherence in the communities so that users in the same community are closely linked to each other and share common latent topics. We handle topic modeling and community discovery in the same framework. In our model we separate the concepts of community and topic, so one community can correspond to multiple topics and multiple communities can share the same topic. We compare different methods and perform extensive experiments on two real datasets. The results confirm our hypothesis that topics could help understand community structure, while community structure could help model topics.
Zhijun Yin, Liangliang Cao, Quanquan Gu, Jiawei Han 0001
ACM Trans. Intell. Syst. Technol.2
2012 Hierarchical Filtered Motion for Action Recognition in Crowded Videos
abstract
Action recognition with cluttered and moving background is a challenging problem. One main difficulty lies in the fact that the motion field in an action region is contaminated by the background motions. We propose a hierarchical filtered motion (HFM) method to recognize actions in crowded videos by the use of motion history image (MHI) as basic representations of motion because of its robustness and efficiency. First, we detect interest points as the two-dimensional Harris corners with recent motion, e.g., locations with high intensities in the MHI. Then, a global spatial motion smoothing filter is applied to the gradients of the MHI to eliminate isolated unreliable or noisy motions. At each interest point, a local motion field filter is applied to the smoothed gradients of the MHI by computing structure proximity between any pixel in the local region and the interest point. Thus, the motion at a pixel is enhanced or weakened based on its structure proximity with the interest point. To validate its effectiveness, we characterize the spatial and temporal features by histograms of oriented gradient in the intensity image and the MHI, respectively, and use a Gaussian-mixture-model-based classifier for action recognition. The performance of the proposed approach achieves the state-of-the-art results on the KTH dataset that has clean background. More importantly, we perform cross-dataset action classification and detection experiments, where the KTH dataset is used for training, while the microsoft research (MSR) action dataset II that consists of crowded videos with people moving in the background is used for testing. Our experiments show that the proposed HFM method significantly outperforms existing techniques.
Yingli Tian, Liangliang Cao, Zicheng Liu 0001, Zhengyou Zhang
IEEE Trans. Syst. Man Cybern. Part C2
2011 Large-scale image classification: Fast feature extraction and SVM training
abstract
Most research efforts on image classification so far have been focused on medium-scale datasets, which are often defined as datasets that can fit into the memory of a desktop (typically 4G~48G). There are two main reasons for the limited effort on large-scale image classification. First, until the emergence of ImageNet dataset, there was almost no publicly available large-scale benchmark data for image classification. This is mostly because class labels are expensive to obtain. Second, large-scale classification is hard because it poses more challenges than its medium-scale counterparts. A key challenge is how to achieve efficiency in both feature extraction and classifier training without compromising performance. This paper is to show how we address this challenge using ImageNet dataset as an example. For feature extraction, we develop a Hadoop scheme that performs feature extraction in parallel using hundreds of mappers. This allows us to extract fairly sophisticated features (with dimensions being hundreds of thousands) on 1.2 million images within one day. For SVM training, we develop a parallel averaging stochastic gradient descent (ASGD) algorithm for training one-against-all 1000-class SVM classifiers. The ASGD algorithm is capable of dealing with terabytes of training data and converges very fast-typically 5 epochs are sufficient. As a result, we achieve state-of-the-art performance on the ImageNet 1000-class classification, i.e., 52.9% in classification accuracy and 71.8% in top 5 hit rate.
Yuanqing Lin, Fengjun Lv, Shenghuo Zhu, Ming Yang 0007, Timothée Cour, Kai Yu 0001, Liangliang Cao, Thomas S. Huang
CVPR7
2011 LPTA: A Probabilistic Model for Latent Periodic Topic Analysis
abstract
This paper studies the problem of latent periodic topic analysis from time stamped documents. The examples of time stamped documents include news articles, sales records, financial reports, TV programs, and more recently, posts from social media websites such as Flickr, Twitter, and Face book. Different from detecting periodic patterns in traditional time series database, we discover the topics of coherent semantics and periodic characteristics where a topic is represented by a distribution of words. We propose a model called LPTA (Latent Periodic Topic Analysis) that exploits the periodicity of the terms as well as term co-occurrences. To show the effectiveness of our model, we collect several representative datasets including Seminar, DBLP and Flickr. The results show that our model can discover the latent periodic topics effectively and leverage the information from both text and time well.
Zhijun Yin, Liangliang Cao, Jiawei Han 0001, ChengXiang Zhai, Thomas S. Huang
ICDM2
2011 Compositional object pattern: a new model for album event recognition
abstract
In this paper, we study the problem of recognizing events in personal photo albums. In consumer photo collections or online photo communities, photos are usually organized in albums according to their events. However, interpreting photo albums is more complicated than the traditional problem of understanding single photos, because albums generally exhibit much more varieties than single image. To solve this challenge, we propose a novel representation, called Compositional Object Pattern, which characterizes object level pattern conveying much richer semantic than low level visual feature. To interpret the rich semantics in albums, we mine frequent object patterns in the training set, and then rank them by their discriminating power. The album feature is then set as the frequencies of these frequent and discriminative patterns, called Compositional Object Pattern Frequency(COPF). We show with experimental result that our algorithm is capable of recognizing holidays with accuracy higher than the baseline method.
Shen-Fu Tsai, Liangliang Cao, Thomas S. Huang
ACM Multimedia2
2011 Learning to Search Efficiently in High Dimensions
abstract
High dimensional similarity search in large scale databases becomes an important challenge due to the advent of Internet. For such applications, specialized data structures are required to achieve computational efficiency. Traditional approaches relied on algorithmic constructions that are often data independent (such as Locality Sensitive Hashing) or weakly dependent (such as kd-trees, k-means trees). While supervised learning algorithms have been applied to related problems, those proposed in the literature mainly focused on learning hash codes optimized for compact embedding of the data rather than search efficiency. Consequently such an embedding has to be used with linear scan or another search algorithm. Hence learning to hash does not directly address the search efficiency issue. This paper considers a new framework that applies supervised learning to directly optimize a data structure that supports efficient large scale search. Our approach takes both search quality and computational cost into consideration. Specifically, we learn a boosted search forest that is optimized using pair-wise similarity labeled examples. The output of this search forest can be efficiently converted into an inverted indexing data structure, which can leverage modern text search infrastructure to achieve both scalability and efficiency. Experimental results show that our approach significantly outperforms the start-of-the-art learning to hash methods (such as spectral hashing), as well as state-of-the-art high dimensional search algorithms (such as LSH and k-means trees).
Zhen Li 0028, Huazhong Ning, Liangliang Cao, Tong Zhang 0001, Yihong Gong, Thomas S. Huang
NIPS3
2011 Diversified Trajectory Pattern Ranking in Geo-tagged Social Media
abstract
Social media such as those residing in the popular photo sharing websites is attracting increasing attention in recent years. As a type of user-generated data, wisdom of the crowd is embedded inside such social media. In particular, millions of users upload to Flickr their photos, many associated with temporal and geographical information. In this paper, we investigate how to rank the trajectory patterns mined from the uploaded photos with geotags and timestamps. The main objective is to reveal the collective wisdom recorded in the seemingly isolated photos and the individual travel sequences reflected by the geo-tagged photos. Instead of focusing on mining frequent trajectory patterns from geo-tagged social media, we put more effort into ranking the mined trajectory patterns and diversifying the ranking results. Through leveraging the relationships among users, locations and trajectories, we rank the trajectory patterns. We then use an exemplar-based algorithm to diversify the results in order to discover the representative trajectory patterns. We have evaluated the proposed framework on 12 different cities using a Flickr dataset and demonstrated its effectiveness.
Zhijun Yin, Liangliang Cao, Jiawei Han 0001, Jiebo Luo 0001, Thomas S. Huang
SDM2
2011 Geographical topic discovery and comparison
abstract
This paper studies the problem of discovering and comparing geographical topics from GPS-associated documents. GPS-associated documents become popular with the pervasiveness of location-acquisition technologies. For example, in Flickr, the geo-tagged photos are associated with tags and GPS locations. In Twitter, the locations of the tweets can be identified by the GPS locations from smart phones. Many interesting concepts, including cultures, scenes, and product sales, correspond to specialized geographical distributions. In this paper, we are interested in two questions: (1) how to discover different topics of interests that are coherent in geographical regions? (2) how to compare several topics across different geographical locations? To answer these questions, this paper proposes and compares three ways of modeling geographical topics: location-driven model, text-driven model, and a novel joint model called LGTA (Latent Geographical Topic Analysis) that combines location and text. To make a fair comparison, we collect several representative datasets from Flickr website including Landscape, Activity, Manhattan, National park, Festival, Car, and Food. The results show that the first two methods work in some datasets but fail in others. LGTA works well in all these datasets at not only finding regions of interests but also providing effective comparisons of the topics across different locations. The results confirm our hypothesis that the geographical distributions can help modeling topics, while topics provide important cues to group different geographical regions.
Zhijun Yin, Liangliang Cao, Jiawei Han 0001, ChengXiang Zhai, Thomas S. Huang
WWW2
2010 Visual cube and on-line analytical processing of images
abstract
On-Line Analytical Processing (OLAP) has shown great success in many industry applications, including sales, marketing, management, financial data analysis, etc. In this paper, we propose Visual Cube and multi-dimensional OLAP of image collections, such as web images indexed in search engines (e.g., Google and Bing), product images (e.g. Amazon) and photos shared on social networks (e.g., Facebook and Flickr). It provides online responses to user requests with summarized statistics of image information and handles rich semantics related to image visual features. A clustering structure measure is proposed to help users freely navigate and explore images. Efficient algorithms are developed to construct Visual Cube. In addition, we introduce the new issue of Cell Overlapping in data cube and present efficient solutions for Visual Cube computation and OLAP operations. Extensive experiments are conducted and the results show good performance of our algorithms.
Xin Jin 0001, Jiawei Han 0001, Liangliang Cao, Jiebo Luo 0001, Bolin Ding, Cindy Xide Lin
CIKM3
2010 Cross-dataset action detection
abstract
In recent years, many research works have been carried out to recognize human actions from video clips. To learn an effective action classifier, most of the previous approaches rely on enough training labels. When being required to recognize the action in a different dataset, these approaches have to re-train the model using new labels. However, labeling video sequences is a very tedious and time-consuming task, especially when detailed spatial locations and time durations are required. In this paper, we propose an adaptive action detection approach which reduces the requirement of training labels and is able to handle the task of cross-dataset action detection with few or no extra training labels. Our approach combines model adaptation and action detection into a Maximum a Posterior (MAP) estimation framework, which explores the spatial-temporal coherence of actions and makes good use of the prior information which can be obtained without supervision. Our approach obtains state-of-the-art results on KTH action dataset using only 50% of the training labels in tradition approaches. Furthermore, we show that our approach is effective for the cross-dataset detection which adapts the model trained on KTH to two other challenging datasets.
Liangliang Cao, Zicheng Liu 0001, Thomas S. Huang
CVPR1
2010 A worldwide tourism recommendation system based on geotaggedweb photos
abstract
This work aims to build a system to suggest tourist destinations based on visual matching and minimal user input. A user can provide either a photo of the desired scenary or a keyword describing the place of interest, and the system will look into its database for places that share the visual characteristics. To that end, we first cluster a large-scale geotagged web photo collection into groups by location and then find the representative images for each group. Tourist destination recommendations are produced by comparing the query against the representative tags or representative images under the premise of "if you like that place, you may also like these places".
Liangliang Cao, Jiebo Luo 0001, Andrew C. Gallagher, Xin Jin 0001, Jiawei Han 0001, Thomas S. Huang
ICASSP1
2010 Accurate and efficient reconstruction of 3D faces from stereo images
abstract
In this paper, we propose a novel algorithm for reconstructing the 3D shape and texture of human faces from two stereo images, which are captured from calibrated cameras. Our approach works in a sparse to dense manner: we first build a coarse shape estimation based on 3D keypoints, and then use a linear morphable model to efficiently match the detail shape and texture. Compared with the previous works, our algorithm can reconstruct the 3D face shape in a speed comparable with that of the fastest algorithm available, but gives a higher accuracy. It can also recover the texture with more complete, realistic looking. Our results show that the new algorithm possesses significant characteristics of a 3D face model reconstruction system, and is especially useful for face recognition and animation applications in practice.
Vuong Le, Hao Tang 0001, Liangliang Cao, Thomas S. Huang
ICIP3
2010 Action detection using multiple spatial-temporal interest point features
abstract
This paper considers the problem of detecting actions from cluttered videos. Compared with the classical action recognition problem, this paper aims to estimate not only the scene category of a given video sequence, but also the spatial-temporal locations of the action instances. In recent years, many feature extraction schemes have been designed to describe various aspects of actions. However, due to the difficulty of action detection, e.g., the cluttered background and potential occlusions, a single type of features cannot solve the action detection problems perfectly in cluttered videos. In this paper, we attack the detection problem by combining multiple Spatial-Temporal Interest Point (STIP) features, which detect salient patches in the video domain, and describe these patches by feature of local regions. The difficulty of combining multiple STIP features for action detection is two folds: First, the number of salient patches detected by different STIP methods varies across different salient patches. How to combine such features is not considered by existing fusion methods. Second, the detection in the videos should be efficient, which excludes many slow machine learning algorithms. To handle these two difficulties, we propose a new approach which combines Gaussian Mixture Model with Branch-and-Bound search to efficiently locate the action of interest. We build a new challenging dataset for our action detection task, and our algorithm obtains impressive results. On classical KTH dataset, our method outperforms the state-of-the-art methods.
Liangliang Cao, Yingli Tian, Zicheng Liu 0001, Benjamin Z. Yao, Zhengyou Zhang, Thomas S. Huang
ICME1
2010 The wisdom of social multimedia: using flickr for prediction and forecast
abstract
Social multimedia hosting and sharing websites, such as Flickr, Facebook, Youtube, Picasa, ImageShack and Photobucket, are increasingly popular around the globe. A major trend in the current studies on social multimedia is using the social media sites as a source of huge amount of labeled data for solving large scale computer science problems in computer vision, data mining and multimedia. In this paper, we take a new path to explore the global trends and sentiments that can be drawn by analyzing the sharing patterns of uploaded and downloaded social multimedia. In a sense, each time an image or video is uploaded or viewed, it constitutes an implicit vote for (or against) the subject of the image. This vote carries along with it a rich set of associated data including time and (often) location information. By aggregating such votes across millions of Internet users, we reveal the wisdom that is embedded in social multimedia sites for social science applications such as politics, economics, and marketing. We believe that our work opens a brand new arena for the multimedia research community with a potentially big impact on society and social sciences.
Xin Jin 0001, Andrew C. Gallagher, Liangliang Cao, Jiebo Luo 0001, Jiawei Han 0001
ACM Multimedia3
2010 A Study on Sampling Strategies in Space-Time Domain for Recognition Applications
Mert Dikmen, Dennis J. Lin, Andrey Del Pozo, Liangliang Cao, Yun Fu 0001, Thomas S. Huang
MMM4
2010 RankCompete: simultaneous ranking and clustering of web photos
abstract
With the explosive growth of digital cameras and online media, it has become crucial to design efficient methods that help users browse and search large image collections. The recent VisualRank algorithm [4] employs visual similarity to represent the link structure in a graph so that the classic PageRank algorithm can be applied to select the most relevant images. However, measuring visual similarity is difficult when there exist diversified semantics in the image collection, and the results from VisualRank cannot supply good visual summarization with diversity. This paper proposes to rank the images in a structural fashion, which aims to discover the diverse structure embedded in photo collections, and rank the images according to their similarity among local neighborhoods instead of across the entire photo collection. We design a novel algorithm named RankCompete, which generalizes the PageRank algorithm for the task of simultaneous ranking and clustering. The experimental results show that RankCompete outperforms VisualRank and provides an efficient but effective tool for organizing web photos.
Liangliang Cao, Andrey Del Pozo, Xin Jin 0001, Jiebo Luo 0001, Jiawei Han 0001, Thomas S. Huang
WWW1
2010 Image Segmentation by MAP-ML Estimations
abstract
Image segmentation plays an important role in computer vision and image analysis. In this paper, image segmentation is formulated as a labeling problem under a probability maximization framework. To estimate the label configuration, an iterative optimization scheme is proposed to alternately carry out the maximum a posteriori (MAP) estimation and the maximum likelihood (ML) estimation. The MAP estimation problem is modeled with Markov random fields (MRFs) and a graph cut algorithm is used to find the solution to the MAP estimation. The ML estimation is achieved by computing the means of region features in a Gaussian model. Our algorithm can automatically segment an image into regions with relevant textures or colors without the need to know the number of regions in advance. Its results match image edges very well and are consistent with human perception. Comparing to six state-of-the-art algorithms, extensive experiments have shown that our algorithm performs the best.
Shifeng Chen, Liangliang Cao, Yueming Wang 0001, Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Image Process.2
2009 Heterogeneous feature machines for visual recognition
abstract
With the recent efforts made by computer vision researchers, more and more types of features have been designed to describe various aspects of visual characteristics. Modeling such heterogeneous features has become an increasingly critical issue. In this paper, we propose a machinery called the Heterogeneous Feature Machine (HFM) to effectively solve visual recognition tasks in need of multiple types of features. Our HFM builds a kernel logistic regression model based on similarities that combine different features and distance metrics. Different from existing approaches that use a linear weighting scheme to combine different features, HFM does not require the weights to remain the same across different samples, and therefore can effectively handle features of different types with different metrics. To prevent the model from overfitting, we employ the so-called group LASSO constraints to reduce model complexity. In addition, we propose a fast algorithm based on co-ordinate gradient descent to efficiently train a HFM. The power of the proposed scheme is demonstrated across a wide variety of visual recognition tasks including scene, event and action recognition.
Liangliang Cao, Jiebo Luo 0001, Feng Liang 0002, Thomas S. Huang
ICCV1
2009 Action detection in complex scenes with spatial and temporal ambiguities
abstract
In this paper, we investigate the detection of semantic human actions in complex scenes. Unlike conventional action recognition in well-controlled environments, action detection in complex scenes suffers from cluttered backgrounds, heavy crowds, occluded bodies, and spatial-temporal boundary ambiguities caused by imperfect human detection and tracking. Conventional algorithms are likely to fail with such spatial-temporal ambiguities. In this work, the candidate regions of an action are treated as a bag of instances. Then a novel multiple-instance learning framework, named SMILE-SVM (Simulated annealing Multiple Instance LEarning Support Vector Machines), is presented for learning human action detector based on imprecise action locations. SMILE-SVM is extensively evaluated with satisfactory performances on two tasks: 1) human action detection on a public video action database with cluttered backgrounds, and 2) a real world problem of detecting whether the customers in a shopping mall show an intention to purchase the merchandise on shelf (even if they didn't buy it eventually). In addition, the complementary nature of motion and appearance features in action detection are also validated, demonstrating a boosted performance in our experiments.
Yuxiao Hu 0001, Liangliang Cao, Fengjun Lv, Shuicheng Yan, Yihong Gong, Thomas S. Huang
ICCV2
2009 Enhancing semantic and geographic annotation of web images via logistic canonical correlation regression
abstract
Photo community sites such as Flickr and Picasa Web Album host a massive amount of personal photos with millions of new photos uploaded every month. These photos constitute an overwhelming source of images that require effective management. There is an increasingly imperative need for semantic annotation of these web images. This paper addresses the problem by considering two kinds of annotation: semantic annotation and geographic annotation. Both are useful for image search and retrieval and for facilitating communities and social networks. This paper proposes a novel method of Logistic Canonical Correlation Regression (LCCR) for the annotation task. This model exploits the canonical correlation between heterogeneous features and an annotation lexicon of interest, and builds a generalized annotation engine based on canonical correlations in order to produce enhanced annotation for web images. We validate the effectiveness of our algorithm using a dataset of over 380,000 images tagged with GPS coordinates.
Liangliang Cao, Jie Yu 0001, Jiebo Luo 0001, Thomas S. Huang
ACM Multimedia1
2009 GAD: General Activity Detection for Fast Clustering on Large Data
abstract
In this paper, we propose GAD (General Activity Detection) for fast clustering on large scale data. Within this framework we design a set of algorithms for different scenarios: (1) Exact GAD algorithm E-GAD, which is much faster than K-Means and gets the same clustering result. (2) Approximate GAD algorithms with different assumptions, which are faster than E-GAD while achieving different degrees of approximation. (3) GAD based algorithms to handle the “large clusters” problem which appears in many large scale clustering applications. Two existing activity detection algorithms GT and CGAUTC are special cases under the framework. The most important contribution of our work is that the framework is the general solution to exploit activity detection for fast clustering in both exact and approximate senarios, and our proposed algorithms within the framework can achieve very high speed. Extensive experiments have been conducted on several large datasets from various real world applications; the results show that our proposed algorithms are effective and efficient.
Xin Jin 0001, Sangkyum Kim, Jiawei Han 0001, Liangliang Cao, Zhijun Yin
SDM4
2009 Responses to the Comments on "What the Back of the Object Looks Like: 3D Reconstruction from Line Drawings without Hidden Lines"
abstract
Varley (2009) made comments on our paper in (L. Cao et al., 2008) section by section. We answer them in this response paper.
Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 Responses to the Comments on "Plane-Based Optimization for 3D Object Reconstruction from Single Line Drawings"
abstract
We disagree with the comments made by Varley [1] on our previous paper [2]. In this paper, we respond to his comments and show that they are not correct.
Jianzhuang Liu, Liangliang Cao, Zhenguo Li, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Image Annotation Within the Context of Personal Photo Collections Using Hierarchical Event and Scene Models
abstract
Most image annotation systems consider a single photo at a time and label photos individually. In this work, we focus on collections of personal photos and exploit the contextual information naturally implied by the associated GPS and time metadata. First, we employ a constrained clustering method to partition a photo collection into event-based subcollections, considering that the GPS records may be partly missing (a practical issue). We then use conditional random field (CRF) models to exploit the correlation between photos based on 1) time-location constraints and 2) the relationship between collection-level annotation (i.e., events) and image-level annotation (i.e., scenes). With the introduction of such a multilevel annotation hierarchy, our system addresses the problem of annotating consumer photo collections that requires a more hierarchical description of the customers' activities than do the simpler image annotation tasks. The efficacy of the proposed system is validated by extensive evaluation using a sizable geotagged personal photo collection database, which consists of over 100 photo collections and is manually labeled for 12 events and 12 scenes to create ground truth.
Liangliang Cao, Jiebo Luo 0001, Henry A. Kautz, Thomas S. Huang
IEEE Trans. Multim.1
2008 Annotating collections of photos using hierarchical event and scene models
abstract
Most image annotation systems consider a single photo at a time and label photos individually. In this work, we focus on collections of personal photos and explore the associated GPS and time information for semantic annotation. First, we employ a constrained clustering method to partition a photo collection into event-based sub-collections, considering that the GPS records may be partly missing (a practical issue). We then use conditional random field (CRF) models to exploit the correlation between photos based on (1) time-location constraints and (2) the relationship between collection-level annotation (i.e., events) and image-level annotation (i.e., scenes). With the introduction of such a multi-level annotation hierarchy, our system addresses the problem of annotating consumer photo collections that requires a more hierarchical description of the customers’ activities than do the simpler image annotation tasks. The efficacy of the proposed system is validated using a geotagged customer photo collection database, which consists of over 100 folders and is labeled for 12 events and 12 scenes.
Liangliang Cao, Jiebo Luo 0001, Henry A. Kautz, Thomas S. Huang
CVPR1
2008 Gender recognition from body
abstract
This paper studies the problem of recognizing gender from full body images. This problem has not been addressed before, partly because of the variant nature of human bodies and clothing that can bring tough difficulties. However, gender recognition has high application potentials, e.g. security surveillance and customer statistics collection in restaurants, supermarkets, and even building entrances. In this paper, we build a system of recognizing gender from full body images, taken from frontal or back views. Our contributions are three-fold. First, to handle the variety of human body characteristics, we represent each image by a collection of patch features, which model different body parts and provide a set of clues for gender recognition. To combine the clues, we build an ensemble learning algorithm from those body parts to recognize gender from fixed view body images (frontal or back). Second, we relax the fixed view constraint and show the possibility to train a flexible classifier for mixed view images with the almost same accuracy as the fixed view case. At last, our approach is shown to be robust to small alignment errors, which is preferred in many applications.
Liangliang Cao, Mert Dikmen, Yun Fu 0001, Thomas S. Huang
ACM Multimedia1
2008 Annotating photo collections by label propagation according to multiple similarity cues
abstract
This paper considers the emerging problem of annotating personal photo collections that are taken by digital cameras and may have been subsequently organized by customers. Unlike the images from the web searching engine or commercial image banks (e.g. the Corel database), the photos in the same personal collection are related to each other in time, location, and content. Advanced technologies can record the GPS coordinates for each photo, and thus provide a richer source of context to model and enforce the correlation between the photos in the same collection. Recognizing the well-known limitations ("semantic gap") of visual recognition algorithms, we exploit the correlation between the photos to enhance the annotation performance. In our approach, high-confidence annotation labels are first obtained for certain photos and then propagated to the remaining photos in the same collection, according to time, location, and visual proximity (or similarity). A novel generative probabilistic model is employed, which outperforms the pervious linear propagation scheme. Experimental results have shown the advantages of the proposed annotation scheme.
Liangliang Cao, Jiebo Luo 0001, Thomas S. Huang
ACM Multimedia1
2008 Image annotation using personal calendars as context
abstract
10.1145/1459359.1459458
Andrew C. Gallagher, Carman Neustaedter, Liangliang Cao, Jiebo Luo 0001, Tsuhan Chen
ACM Multimedia3
2008 What the Back of the Object Looks Like: 3D Reconstruction from Line Drawings without Hidden Lines
abstract
The human vision system can interpret a single 2D line drawing as a 3D object without much difficulty even if the hidden lines of the object are invisible. Many reconstruction methods have been proposed to emulate this ability, but they cannot recover the complete object if the hidden lines of the object are not shown. This paper proposes a novel approach to reconstructing a complete 3D object, including the shape of the back of the object, from a line drawing without hidden lines. First, we develop theoretical constraints and an algorithm for the inference of the topology of the invisible edges and vertices of an object. Then we present a reconstruction method based on perceptual symmetry and planarity of the object. We show a number of examples to demonstrate the success of our approach.
Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Plane-Based Optimization for 3D Object Reconstruction from Single Line Drawings
abstract
In previous optimization-based methods of 3D planar-faced object reconstruction from single 2D line drawings, the missing depths of the vertices of a line drawing (and other parameters in some methods) are used as the variables of the objective functions. A 3D object with planar faces is derived by finding values for these variables that minimize the objective functions. These methods work well for simple objects with a small number N of variables. As N grows, however, it is very difficult for them to find expected objects. This is because with the nonlinear objective functions in a space of large dimension N, the search for optimal solutions can easily get trapped into local minima. In this paper, we use the parameters of the planes that pass through the planar faces of an object as the variables of the objective function. This leads to a set of linear constraints on the planes of the object, resulting in a much lower dimensional nullspace where optimization is easier to achieve. We prove that the dimension of this nullspace is exactly equal to the minimum number of vertex depths which define the 3D object. Since a practical line drawing is usually not an exact projection of a 3D object, we expand the nullspace to a larger space based on the singular value decomposition of the projection matrix of the line drawing. In this space, robust 3D reconstruction can be achieved. Compared with two most related methods, our method not only can reconstruct more complex 3D objects from 2D line drawings, but also is computationally more efficient.
Jianzhuang Liu, Liangliang Cao, Zhenguo Li, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Iterative MAP and ML Estimations for Image Segmentation
abstract
Image segmentation plays an important role in computer vision and image analysis. In this paper, the segmentation problem is formulated as a labeling problem under a probability maximization framework. To estimate the label configuration, an iterative optimization scheme is proposed to alternately carry out the maximum a posteriori (MAP) estimation and the maximum-likelihood (ML) estimation. The MAP estimation problem is modeled with Markov random fields (MRFs). A graph-cut algorithm is used to find the solution to the MAP-MRF estimation. The ML estimation is achieved by finding the means of region features. Our algorithm can automatically segment an image into regions with relevant textures or colors without the need to know the number of regions in advance. In addition, under the same framework, it can be extended to another algorithm that extracts objects of a particular class from a group of images. Extensive experiments have shown the effectiveness of our approach.
Shifeng Chen, Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
CVPR2
2007 Spatially Coherent Latent Topic Model for Concurrent Segmentation and Classification of Objects and Scenes
abstract
We present a novel generative model for simultaneously recognizing and segmenting object and scene classes. Our model is inspired by the traditional bag of words representation of texts and images as well as a number of related generative models, including probabilistic Latent Semantic Analysis (pLSA) and Latent Dirichlet Allocation (LDA). A major drawback of the pLSA and LDA models is the assumption that each patch in the image is independently generated given its corresponding latent topic. While such representation provides an efficient computational method, it lacks the power to describe the visually coherent images and scenes. Instead, we propose a spatially coherent latent topic model (Spatial-LTM). Spatial-LTM represents an image containing objects in a hierarchical way by over-segmented image regions of homogeneous appearances and the salient image patches within the regions. Only one single latent topic is assigned to the image patches within each region, enforcing the spatial coherency of the model. This idea gives rise to the following merits of Spatial-LTM: (1) Spatial-LTM provides a unified representation for spatially coherent bag of words topic models; (2) Spatial-LTM can simultaneously segment and classify objects, even in the case of occlusion and multiple instances; and (3) Spatial-LTM can be trained either unsupervised or supervised, as well as when partial object labels are provided. We verify the success of our model in a number of segmentation and classification experiments.
Liangliang Cao, Li Fei-Fei 0001
ICCV1
2006 Degen Generalized Cylinders and Their Properties
Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
ECCV (1)1
2006 3D object retrieval using 2D line drawing and graph based relevance reedback
abstract
This paper aims to provide a user-friendly interface for 3D object retrieval. In previous 3D retrieval systems, the user mainly uses two methods to input a query: providing an existing 3D objects, or providing partial shape information of desired objects such as text and 2D shapes. The first method fails when the user does not have a similar 3D object in hand, and the second method cannot sufficiently describe 3D shapes of objects. We believe that the best way is to have a good interface that can convert a 2D sketch drawn by the user into a 3D object as the query. A 2D line drawing is easy to be drawn and is the simplest and most direct way of illustrating a 3D object. In this paper, we develop an interface of 3D object reconstruction from line drawings, which allows the user to draw line drawings of objects with both planar and curved surfaces. In addition, in order to refine the retrieved results, we develop a relevance feedback algorithm based on a novel graph discriminant analysis. Compared with recently published relevance feedback algorithms, our algorithm achieves better retrieval performance.
Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia1
2005 3D Object Reconstruction from a Single 2D Line Drawing without Hidden Lines
abstract
The human vision system can interpret a single 2D line drawing as a 3D object without much difficulty even if the hidden lines of the object are invisible. Several reconstruction approaches have tried to emulate this ability, but they cannot recover the complete object if the hidden lines of the object are not shown. This paper proposes a novel approach for reconstructing complete 3D objects from line drawings without hidden lines. First, we develop some constraints and properties for the inference of the topology of the invisible edges and vertices of an object. Then we present a reconstruction method based on perceptual symmetry and planarity of the object. We give a number of examples to demonstrate the ability of our approach.
Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
ICCV1