EDBT 2026 Demo / reviewers in the wild / expert
Hanqing Lu
dblp:39/6752 · also Han-Qing Lu
· DBLP profile ↗
442ranked-venue papers
3as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 313 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 188 · 2 first-author · 25 since 2021Databases, data management, data science and information retrieval · 27 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-authorComputer networks · 5Systems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User IntentsabstractZiyi Wang, Yuxuan Lu, Yimeng Zhang, Pei Chen, Ziwei Dong, Jing Huang, Jiri Gesi, Xianfeng Tang, Chen Luo, Qun Liu, Yisi Sang, Hanqing Lu, Manling Li, Jin Lai, Dakuo Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuxuan Lu 0003, Ziwei Dong, Jiri Gesi, Xianfeng Tang, Chen Luo 0003, Yisi Sang, Hanqing Lu, Manling Li, Jin Lai, Dakuo Wang |
ACL (1) | 12 |
| 2025 | Towards Context-Robust LLMs: A Gated Representation Fine-tuning ApproachabstractLarge Language Models (LLMs) enhanced with external contexts, such as through retrieval-augmented generation (RAG), often face challenges in handling imperfect evidence.They tend to over-rely on external knowledge, making them vulnerable to misleading and unhelpful contexts.To address this, we propose the concept of context-robust LLMs, which can effectively balance internal knowledge with external context, similar to human cognitive processes.Specifically, context-robust LLMs should rely on external context only when lacking internal knowledge, identify contradictions between internal and external knowledge, and disregard unhelpful contexts.To achieve this goal, we introduce Grft, a lightweight and plug-and-play gated representation fine-tuning approach.Grft consists of two key components: a gating mechanism to detect and filter problematic inputs, and low-rank representation adapters to adjust hidden representations.By training a lightweight intervention function with only 0.0004% of model size on fewer than 200 examples, Grft can effectively adapts LLMs towards context-robust behaviors. Shenglai Zeng, Kai Guo 0003, Hanqing Lu, Yue Xing 0002, Hui Liu 0031 |
ACL (1) | 5 |
| 2025 | Mitigating the Privacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic DataabstractShenglai Zeng, Jiankun Zhang, Pengfei He, Jie Ren, Tianqi Zheng, Hanqing Lu, Han Xu, Hui Liu, Yue Xing, Jiliang Tang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Shenglai Zeng, Jiankun Zhang 0001, Jie Ren 0019, Hanqing Lu, Han Xu 0002, Hui Liu 0003, Yue Xing 0002, Jiliang Tang |
EMNLP | 6 |
| 2025 | Towards Knowledge Checking in Retrieval-augmented Generation: A Representation PerspectiveabstractShenglai Zeng, Jiankun Zhang, Bingheng Li, Yuping Lin, Tianqi Zheng, Dante Everaert, Hanqing Lu, Hui Liu, Hui Liu, Yue Xing, Monica Xiao Cheng, Jiliang Tang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shenglai Zeng, Jiankun Zhang 0001, Bingheng Li, Yuping Lin, Dante Everaert, Hanqing Lu, Hui Liu 0033, Hui Liu 0031, Yue Xing 0002, Monica Xiao Cheng, Jiliang Tang |
NAACL (Long Papers) | 7 |
| 2025 | Hierarchical Contrastive Learning for Semantic SegmentationabstractRecently, pixel-to-pixel contrastive learning in single-scale feature space has been widely studied in semantic segmentation to learn a unified feature expression for pixels of the same category. However, the unified representation is too extreme, and the receptive field of each single-scale pixel is limited, which is insufficient to reflect the representative features of the category. To address these problems, this article extends the single-scale feature space to that of multiscale and proposes a hierarchical contrastive learning (Hi-CL) method to explore pixel-to-component semantic relationships. First, we generate multiscale candidate samples by applying several pooling windows with different sizes on a feature map, where different windows may represent different parts of the objects in the image. Then, we prune the sample set through threshold-based criteria to select appropriate samples for feature representation learning. Finally, Hi-CL is performed to learn the pixel-to-component consistency with the pruned samples. Our method is easy to be applied on existing semantic segmentation models and obtains consistent improvement. Furthermore, we achieve state-of-the-art results on three popular benchmarks, including Cityscapes, ADE20K, and COCO Stuff datasets. Jie Jiang 0016, Xingjian He, Weining Wang 0001, Hanqing Lu, Jing Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Exploring Query Understanding for Amazon Product SearchabstractOnline shopping platforms, such as Amazon, offer services to billions of people worldwide. Unlike web search or other search engines, product search engines have their unique characteristics, primarily featuring short queries which are mostly a combination of product attributes and structured product search space. The uniqueness of product search underscores the crucial importance of the query understanding component. However, there are limited studies focusing on exploring this impact within real-world product search engines. In this work, we aim to bridge this gap by conducting a comprehensive study and sharing our year-long journey investigating how the query understanding service impacts Amazon Product Search. Firstly, we explore how query understanding-based ranking features influence the ranking process. Next, we delve into how the query understanding system contributes to understanding the performance of a ranking model. Building on the insights gained from our study on the evaluation of the query understanding-based ranking model, we propose a query understanding-based multi-task learning framework for ranking. We present our studies and investigations on Amazon Search. Chen Luo 0003, Xianfeng Tang, Hanqing Lu, Yaochen Xie, Hui Liu 0003, Zhenwei Dai, Limeng Cui, Ashutosh Joshi, Sreyashi Nag, Yang Li 0055, Rahul Goutam, Jiliang Tang, Qi He 0002 |
IEEE Big Data | 3 |
| 2024 | Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and BeyondabstractDeveloping a universal model that can effectively harness heterogeneous resources and respond to a wide range of personalized needs has been a longstanding community aspiration. Our daily choices, especially in domains like fashion and retail, are substantially shaped by multi-modal data, such as pictures and textual descriptions. These modalities not only offer intuitive guidance but also cater to personalized user preferences. However, the predominant personalization approaches mainly focus on ID or text-based recommendation problems, failing to comprehend the information spanning various tasks or modalities. In this paper, our goal is to establish a Unified paradigm for Multi-modal Personalization systems (UniMP), which effectively leverages multi-modal data while eliminating the complexities associated with task- and modality-specific customization. We argue that the advancements in foundational generative modeling have provided the flexibility and effectiveness necessary to achieve the objective. In light of this, we develop a generic and extensible personalization generative framework, that can handle a wide range of personalized needs including item recommendation, product search, preference prediction, explanation generation, and further user-guided image generation. Our methodology enhances the capabilities of foundational language models for personalized tasks by seamlessly ingesting interleaved cross-modal user history information, ensuring a more precise and customized experience for users. To train and evaluate the proposed multi-modal personalized tasks, we also introduce a novel and comprehensive benchmark covering a variety of user requirements. Our experiments on the real-world benchmark showcase the model's potential, outperforming competitive methods specialized for each task. Tianxin Wei, Bowen Jin, Ruirui Li 0002, Hansi Zeng, Jianhui Sun, Qingyu Yin, Hanqing Lu, Suhang Wang, Jingrui He, Xianfeng Tang |
ICLR | 8 |
| 2024 | Language Models as Semantic IndexersabstractSemantic identifier (ID) is an important concept in information retrieval that aims to preserve the semantics of objects such as documents and items inside their IDs. Previous studies typically adopt a two-stage pipeline to learn semantic IDs by first procuring embeddings using off-the-shelf text encoders and then deriving IDs based on the embeddings. However, each step introduces potential information loss, and there is usually an inherent mismatch between the distribution of embeddings within the latent space produced by text encoders and the anticipated distribution required for semantic indexing. It is non-trivial to design a method that can learn the document’s semantic representations and its hierarchical structure simultaneously, given that semantic IDs are discrete and sequentially structured, and the semantic supervision is deficient. In this paper, we introduce LMIndexer, a self-supervised framework to learn semantic IDs with a generative language model. We tackle the challenge of sequential discrete ID by introducing a semantic indexer capable of generating neural sequential discrete representations with progressive training and contrastive learning. In response to the semantic supervision deficiency, we propose to train the model with a self-supervised document reconstruction objective. We show the high quality of the learned IDs and demonstrate their effectiveness on three tasks including recommendation, product search, and document retrieval on five datasets from various domains. Code is available at https://github.com/PeterGriffinJin/LMIndexer. Bowen Jin, Hansi Zeng, Guoyin Wang 0001, Xiusi Chen, Tianxin Wei, Ruirui Li 0002, Zheng Li 0018, Hanqing Lu, Suhang Wang, Jiawei Han 0001, Xianfeng Tang |
ICML | 10 |
| 2023 | Asynchronous Event Processing with Local-Shift Graph Convolutional NetworkabstractEvent cameras are bio-inspired sensors that produce sparse and asynchronous event streams instead of frame-based images at a high-rate. Recent works utilizing graph convolutional networks (GCNs) have achieved remarkable performance in recognition tasks, which model event stream as spatio-temporal graph. However, the computational mechanism of graph convolution introduces redundant computation when aggregating neighbor features, which limits the low-latency nature of the events. And they perform a synchronous inference process, which can not achieve a fast response to the asynchronous event signals. This paper proposes a local-shift graph convolutional network (LSNet), which utilizes a novel local-shift operation equipped with a local spatio-temporal attention component to achieve efficient and adaptive aggregation of neighbor features. To improve the efficiency of pooling operation in feature extraction, we design a node-importance based parallel pooling method (NIPooling) for sparse and low-latency event data. Based on the calculated importance of each node, NIPooling can efficiently obtain uniform sampling results in parallel, which retains the diversity of event streams. Furthermore, for achieving a fast response to asynchronous event signals, an asynchronous event processing procedure is proposed to restrict the network nodes which need to recompute activations only to those affected by the new arrival event. Experimental results show that the computational cost can be reduced by nearly 9 times through using local-shift operation and the proposed asynchronous procedure can further improve the inference efficiency, while achieving state-of-the-art performance on gesture recognition and object recognition. Linhui Sun, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
AAAI | 4 |
| 2023 | Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text GenerationabstractModeling customer shopping intentions is a crucial task for e-commerce, as it directly impacts user experience and engagement. Thus, accurately understanding customer preferences is essential for providing personalized recommendations. Session-based recommendation, which utilizes customer session data to predict their next interaction, has become increasingly popular. However, existing session datasets have limitations in terms of item attributes, user diversity, and dataset scale. As a result, they cannot comprehensively capture the spectrum of user behaviors and preferences.To bridge this gap, we present the Amazon Multilingual Multi-locale Shopping Session Dataset, namely Amazon-M2. It is the first multilingual dataset consisting of millions of user sessions from six different locales, where the major languages of products are English, German, Japanese, French, Italian, and Spanish.Remarkably, the dataset can help us enhance personalization and understanding of user preferences, which can benefit various existing tasks as well as enable new tasks. To test the potential of the dataset, we introduce three tasks in this work:(1) next-product recommendation, (2) next-product recommendation with domain shifts, and (3) next-product title generation.With the above tasks, we benchmark a range of algorithms on our proposed dataset, drawing new insights for further research and practice. In addition, based on the proposed dataset and tasks, we hosted a competition in the KDD CUP 2023 https://www.aicrowd.com/challenges/amazon-kdd-cup-23-multilingual-recommendation-challenge and have attracted thousands of users and submissions. The winning solutions and the associated workshop can be accessed at our website~https://kddcup23.github.io/. Wei Jin 0009, Haitao Mao, Zheng Li 0018, Haoming Jiang, Chen Luo 0003, Hongzhi Wen, Haoyu Han 0001, Hanqing Lu, Ruirui Li 0002, Monica Xiao Cheng, Rahul Goutam, Karthik Subbian, Suhang Wang, Yizhou Sun, Jiliang Tang, Xianfeng Tang |
NeurIPS | 8 |
| 2023 | Question-Guided Erasing-Based Spatiotemporal Attention Learning for Video Question AnsweringabstractSpatiotemporal attention learning for video question answering (VideoQA) has always been a challenging task, where existing approaches treat the attention parts and the nonattention parts in isolation. In this work, we propose to enforce the correlation between the attention parts and the nonattention parts as a distance constraint for discriminative spatiotemporal attention learning. Specifically, we first introduce a novel attention-guided erasing mechanism in the traditional spatiotemporal attention to obtain multiple aggregated attention features and nonattention features and then learn to separate the attention and the nonattention features with an appropriate distance. The distance constraint is enforced by a metric learning loss, without increasing the inference complexity. In this way, the model can learn to produce more discriminative spatiotemporal attention distribution on videos, thus enabling more accurate question answering. In order to incorporate the multiscale spatiotemporal information that is beneficial for video understanding, we additionally develop a pyramid variant on basis of the proposed approach. Comprehensive ablation experiments are conducted to validate the effectiveness of our approach, and state-of-the-art performance is achieved on several widely used datasets for VideoQA. Fei Liu 0047, Jing Liu 0001, Richang Hong, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph AlignmentabstractZijie Huang, Zheng Li, Haoming Jiang, Tianyu Cao, Hanqing Lu, Bing Yin, Karthik Subbian, Yizhou Sun, Wei Wang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Zijie Huang 0002, Zheng Li 0018, Haoming Jiang, Tianyu Cao 0001, Hanqing Lu, Karthik Subbian, Yizhou Sun, Wei Wang 0010 |
ACL (1) | 5 |
| 2022 | MENet: A Memory-Based Network with Dual-Branch for Efficient Event Stream Processing
Linhui Sun, Yifan Zhang 0001, Ke Cheng 0002, Jian Cheng 0001, Hanqing Lu |
ECCV (24) | 5 |
| 2022 | Can Clicks Be Both Labels and Features?: Unbiased Behavior Feature Collection and Uncertainty-aware Learning to RankabstractUsing implicit feedback collected from user clicks as training labels for learning-to-rank algorithms is a well-developed paradigm that has been extensively studied and used in modern IR systems. Using user clicks as ranking features, on the other hand, has not been fully explored in existing literature. Despite its potential in improving short-term system performance, whether the incorporation of user clicks as ranking features is beneficial for learning-to-rank systems in the long term is still questionable. Two of the most important problems are (1) the explicit bias introduced by noisy user behavior, and (2) the implicit bias, which we refer to as the exploitation bias, introduced by the dynamic training and serving of learning-to-rank systems with behavior features. In this paper, we explore the possibility of incorporating user clicks as both training labels and ranking features for learning to rank. We formally investigate the problems in feature collection and model training, and propose a counterfactual feature projection function and a novel uncertainty-aware learning to rank framework. Experiments on public datasets show that ranking models learned with the proposed framework can significantly outperform models built with raw click features and algorithms that rank items without considering model uncertainty. Tao Yang 0030, Chen Luo 0003, Hanqing Lu, Parth Gupta, Qingyao Ai |
SIGIR | 3 |
| 2022 | Super-resolution semantic segmentation with relation calibrating network
Jie Jiang 0016, Jing Liu 0001, Jun Fu 0005, Weining Wang 0001, Hanqing Lu |
Pattern Recognit. | 5 |
| 2022 | Action recognition via pose-based graph convolutional networks with intermediate dense supervision
Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
Pattern Recognit. | 4 |
| 2022 | Dynamic Orthogonal Projection Constrained Discriminative TrackingabstractDue to the end-to-end feature learning with convolutional neural networks (CNNs), modern discriminative trackers improve the state of the art significantly. To achieve a strong discrimination, the learned features are usually high-dimensional, resulting in a massive number of parameters contained in the discriminative model and the increase of risk of over-fitting in the online tracking. In this letter, we try to alleviate the risk of over-fitting by means of the adaptive dimensionality reduction (DR) through CNNs. Specifically, an orthogonality constrained ridge regression model is proposed to reduce the dimensionality of features, and a dynamic sub-network (DOPNet) is designed to learn to perform DR. After trained with an orthogonality loss and a regression one, DOPNet generates a set of orthogonal bases (i. e., weights in FC layers) dynamically to reduce the feature dimensionality for a discriminative model in the online tracking. Based on the novel discriminative model and DOPNet, an effective and efficient tracker, DOPTracker, is developed. DOPTracker achieves the state-of-the-art results on four benchmarks, OTB-2015, VOT-2018, NfS, and GOT-10 k while running at 30 FPS. Ming Tang 0001, Guibo Zhu, Jinqiao Wang, Hanqing Lu |
IEEE Signal Process. Lett. | 5 |
| 2022 | An Efficient Sampling-Based Attention Network for Semantic SegmentationabstractSelf-attention is widely explored to model long-range dependencies in semantic segmentation. However, this operation computes pair-wise relationships between the query point and all other points, leading to prohibitive complexity. In this paper, we propose an efficient Sampling-based Attention Network which combines a novel sample method with an attention mechanism for semantic segmentation. Specifically, we design a Stochastic Sampling-based Attention Module (SSAM) to capture the relationships between the query point and a stochastic sampled representative subset from a global perspective, where the sampled subset is selected by a Stochastic Sampling Module. Compared to self-attention, our SSAM achieves comparable segmentation performance while significantly reducing computational redundancy. In addition, with the observation that not all pixels are interested in the contextual information, we design a Deterministic Sampling-based Attention Module (DSAM) to sample features from a local region for obtaining the detailed information. Extensive experiments demonstrate that our proposed method can compete or perform favorably against the state-of-the-art methods on the Cityscapes, ADE20K, COCO Stuff, and PASCAL Context datasets. Xingjian He, Jing Liu 0001, Weining Wang 0001, Hanqing Lu |
IEEE Trans. Image Process. | 4 |
| 2022 | Global-Guided Selective Context Network for Scene ParsingabstractRecent studies on semantic segmentation are exploiting contextual information to address the problem of inconsistent parsing prediction in big objects and ignorance in small objects. However, they utilize multilevel contextual information equally across pixels, overlooking those different pixels may demand different levels of context. Motivated by the above-mentioned intuition, we propose a novel global-guided selective context network (GSCNet) to adaptively select contextual information for improving scene parsing. Specifically, we introduce two global-guided modules, called global-guided global module (GGM) and global-guided local module (GLM), to, respectively, select global context (GC) and local context (LC) for pixels. When given an input feature map, GGM jointly employs the input feature map and its globally pooled feature to learn its global contextual demand based on which per-pixel GC is selected. While GLM adopts low-level feature from the adjacent stage as LC and synthetically models the input feature map, its globally pooled feature and LC to generate local contextual demand, based on which per-pixel LC is selected. Furthermore, we combine these two modules as a selective context block and import such SCBs in different levels of the network to propagate contextual information in a coarse-to-fine manner. Finally, we conduct extensive experiments to verify the effectiveness of our proposed model and achieve state-of-the-art performance on four challenging scene parsing data sets, i.e., Cityscapes, ADE20K, PASCAL Context, and COCO Stuff. Especially, GSCNet-101 obtains 82.6% on Cityscapes test set without using coarse data and 56.22% on ADE20K test set. Jie Jiang 0016, Jing Liu 0001, Jun Fu 0005, Zechao Li, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | Consistent-Separable Feature Representation for Semantic SegmentationabstractCross-entropy loss combined with softmax is one of the most commonly used supervision components in most existing segmentation methods. The softmax loss is typically good at optimizing the inter-class difference, but not good at reducing the intra-class variation, which can be suboptimal for semantic segmentation task. In this paper, we propose a Consistent-Separable Feature Representation Network to model the Consistent-Separable (C-S) features, which are intra-class consistent and inter-class separable, improving the discriminative power of the deep features. Specifically, we develop a Consistent-Separable Feature Learning Module to obtain C-S features through a new loss, called Class-Aware Consistency loss. This loss function is proposed to force the deep features to be consistent among the same class and apart between different classes. Moreover, we design an Adaptive feature Aggregation Module to fuse the C-S features and original features from backbone for the better semantic prediction. We show that compared with various baselines, the proposed method brings consistent performance improvement. Our proposed approach achieves state-of-the-art performance on Cityscapes (82.6% mIoU in test set), ADE20K (46.65% mIoU in validation set), COCO Stuff (41.3% mIoU in validation set) and PASCAL Context (55.9% mIoU in test set). Xingjian He, Jing Liu 0001, Jun Fu 0005, Jinqiao Wang, Hanqing Lu |
AAAI | 6 |
| 2021 | QUEACO: Borrowing Treasures from Weakly-labeled Behavior Data for Query Attribute Value ExtractionabstractWe study the problem of query attribute value extraction, which aims to identify named entities from user queries as diverse surface form attribute values and afterward transform them into formally canonical forms. Such a problem consists of two phases: named entity recognition (NER) and attribute value normalization (AVN). However, existing works only focus on the NER phase but neglect equally important AVN. To bridge this gap, this paper proposes a unified query attribute value extraction system in e-commerce search named QUEACO, which involves both two phases. Moreover, by leveraging large-scale weakly-labeled behavior data, we further improve the extraction performance with less supervision cost. Specifically, for the NER phase, QUEACO adopts a novel teacher-student network, where a teacher network that is trained on the strongly-labeled data generates pseudo-labels to refine the weakly-labeled data for training a student network. Meanwhile, the teacher network can be dynamically adapted by the feedback of the student's performance on strongly-labeled data to maximally denoise the noisy supervisions from the weak labels. For the AVN phase, we also leverage the weakly-labeled query-to-attribute behavior data to normalize surface form attribute values from queries into canonical forms from products. Extensive experiments on a real-world large-scale E-commerce dataset demonstrate the effectiveness of QUEACO. Danqing Zhang, Zheng Li 0018, Tianyu Cao 0001, Chen Luo 0003, Hanqing Lu, Yiwei Song, Tuo Zhao, Qiang Yang 0001 |
CIKM | 6 |
| 2021 | Improving Multiple Object Tracking With Single Object TrackingabstractDespite considerable similarities between multiple object tracking (MOT) and single object tracking (SOT) tasks, modern MOT methods have not benefited from the development of SOT ones to achieve satisfactory performance. The major reason for this situation is that it is inappropriate and inefficient to apply multiple SOT models directly to the MOT task, although advanced SOT methods are of the strong discriminative power and can run at fast speeds.In this paper, we propose a novel and end-to-end trainable MOT architecture that extends CenterNet by adding an SOT branch for tracking objects in parallel with the existing branch for object detection, allowing the MOT task to benefit from the strong discriminative power of SOT methods in an effective and efficient way. Unlike most existing SOT methods which learn to distinguish the target object from its local backgrounds, the added SOT branch trains a separate SOT model per target online to distinguish the target from its surrounding targets, assigning SOT models the novel discrimination. Moreover, similar to the detection branch, the SOT branch treats objects as points, making its online learning efficient even if multiple targets are processed simultaneously. Without tricks, the proposed tracker achieves MOTAs of 0.710 and 0.686, IDF1s of 0.719 and 0.714, on MOT17 and MOT20 benchmarks, respectively, while running at 16 FPS on MOT17. Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Guibo Zhu, Jinqiao Wang, Hanqing Lu |
CVPR | 6 |
| 2021 | AdaSGN: Adapting Joint Number and Model Size for Efficient Skeleton-Based Action RecognitionabstractExisting methods for skeleton-based action recognition mainly focus on improving the recognition accuracy, whereas the efficiency of the model is rarely considered. Recently, there are some works trying to speed up the skeleton modeling by designing light-weight modules. However, in addition to the model size, the amount of the data involved in the calculation is also an important factor for the running speed, especially for the skeleton data where most of the joints are redundant or non-informative to identify a specific skeleton. Besides, previous works usually employ one fix-sized model for all the samples regardless of the difficulty of recognition, which wastes computations for easy samples. To address these limitations, a novel approach, called AdaSGN, is proposed in this paper, which can reduce the computational cost of the inference process by adaptively controlling the input number of the joints of the skeleton on-the-fly. Moreover, it can also adaptively select the optimal model size for each sample to achieve a better trade-off between the accuracy and the efficiency. We conduct extensive experiments on three challenging datasets, namely, NTU-60, NTU-120 and SHREC, to verify the superiority of the proposed approach, where AdaSGN achieves comparable or even higher performance with much lower GFLOPs compared with the baseline method. Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
ICCV | 4 |
| 2021 | HAIR: Hierarchical Visual-Semantic Relational Reasoning for Video Question AnsweringabstractRelational reasoning is at the heart of video question answering. However, existing approaches suffer from several common limitations: (1) they only focus on either object-level or frame-level relational reasoning, and fail to integrate the both; and (2) they neglect to leverage semantic knowledge for relational reasoning. In this work, we propose a Hierarchical VisuAl-Semantic RelatIonal Reasoning (HAIR) framework to address these limitations. Specifically, we present a novel graph memory mechanism to perform relational reasoning, and further develop two types of graph memory: a) visual graph memory that leverages visual information of video for relational reasoning; b) semantic graph memory that is specifically designed to explicitly leverage semantic knowledge contained in the classes and attributes of video objects, and perform relational reasoning in the semantic space. Taking advantage of both graph memory mechanisms, we build a hierarchical framework to enable visual-semantic relational reasoning from object level to frame level. Experiments on four challenging benchmark datasets show that the proposed framework leads to state-of-the-art performance, with fewer parameters and faster inference speed. Besides, our approach also shows superior performance on other video+language task. Fei Liu 0047, Jing Liu 0001, Weining Wang 0001, Hanqing Lu |
ICCV | 4 |
| 2021 | High-Performance Discriminative Tracking with TransformersabstractEnd-to-end discriminative trackers improve the state of the art significantly, yet the improvement in robustness and efficiency is restricted by the conventional discriminative model, i.e., least-squares based regression. In this paper, we present DTT, a novel single-object discriminative tracker, based on an encoder-decoder Transformer architecture. By self- and encoder-decoder attention mechanisms, our approach is able to exploit the rich scene information in an end-to-end manner, effectively removing the need for hand-designed discriminative models. In online tracking, given a new test frame, dense prediction is performed at all spatial positions. Not only location, but also bounding box of the target object is obtained in a robust fashion, streamlining the discriminative tracking pipeline. DTT is conceptually simple and easy to implement. It yields state-of-the-art performance on four popular benchmarks including GOT-10k, LaSOT, NfS, and TrackingNet while running at over 50 FPS, confirming its effectiveness and efficiency. We hope DTT may provide a new perspective for single-object visual tracking. Ming Tang 0001, Linyu Zheng, Guibo Zhu, Jinqiao Wang, Xuetao Feng, Hanqing Lu |
ICCV | 8 |
| 2021 | High-Performance Discriminative Tracking with Target-Aware Feature Embeddings
Ming Tang 0001, Linyu Zheng, Guibo Zhu, Jinqiao Wang, Hanqing Lu |
PRCV (1) | 6 |
| 2021 | Macro-micro mutual learning inside compositional model for human pose estimation
Yingying Chen 0003, Congqi Cao, Yakui Chu, Jinqiao Wang, Hanqing Lu |
Neurocomputing | 6 |
| 2021 | STN-enhanced message passing guided by adversarial learning for human pose estimation
Yingying Chen 0003, Congqi Cao, Jinqiao Wang, Hanqing Lu |
Neurocomputing | 5 |
| 2021 | Extremely Lightweight Skeleton-Based Action Recognition With ShiftGCN++abstractIn skeleton-based action recognition, graph convolutional networks (GCNs) have achieved remarkable success. However, there are two shortcomings of current GCN-based methods. Firstly, the computation cost is pretty heavy, typically over 15 GFLOPs for one action sample. Some recent works even reach ~100 GFLOPs. Secondly, the receptive fields of both spatial graph and temporal graph are inflexible. Although recent works introduce incremental adaptive modules to enhance the expressiveness of spatial graph, their efficiency is still limited by regular GCN structures. In this paper, we propose a shift graph convolutional network (ShiftGCN) to overcome both shortcomings. ShiftGCN is composed of novel shift graph operations and lightweight point-wise convolutions, where the shift graph operations provide flexible receptive fields for both spatial graph and temporal graph. To further boost the efficiency, we introduce four techniques and build a more lightweight skeleton-based action recognition model named ShiftGCN++. ShiftGCN++ is an extremely computation-efficient model, which is designed for low-power and low-cost devices with very limited computing power. On three datasets for skeleton-based action recognition, ShiftGCN notably exceeds the state-of-the-art methods with over 10× less FLOPs and 4× practical speedup. ShiftGCN++ further boosts the efficiency of ShiftGCN, which achieves comparable performance with 6× less FLOPs and 2× practical speedup. Ke Cheng 0002, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
IEEE Trans. Image Process. | 5 |
| 2021 | Semi-Supervised Scene Text RecognitionabstractScene text recognition has been widely researched with supervised approaches. Most existing algorithms require a large amount of labeled data and some methods even require character-level or pixel-wise supervision information. However, labeled data is expensive, unlabeled data is relatively easy to collect, especially for many languages with fewer resources. In this paper, we propose a novel semi-supervised method for scene text recognition. Specifically, we design two global metrics, i.e., edit reward and embedding reward, to evaluate the quality of generated string and adopt reinforcement learning techniques to directly optimize these rewards. The edit reward measures the distance between the ground truth label and the generated string. Besides, the image feature and string feature are embedded into a common space and the embedding reward is defined by the similarity between the input image and generated string. It is natural that the generated string should be the nearest with the image it is generated from. Therefore, the embedding reward can be obtained without any ground truth information. In this way, we can effectively exploit a large number of unlabeled images to improve the recognition performance without any additional laborious annotations. Extensive experimental evaluations on the five challenging benchmarks, the Street View Text, IIIT5K, and ICDAR datasets demonstrate the effectiveness of the proposed approach, and our method significantly reduces annotation effort while maintaining competitive recognition performance. Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu |
IEEE Trans. Image Process. | 4 |
| 2021 | Visual Question Answering With Dense Inter- and Intra-Modality InteractionsabstractLearning effective interactions between multi-modal features is at the heart of visual question answering (VQA). A common defect of the existing VQA approaches is that they only consider a very limited amount of inter-modality interactions, which may be not enough to model latent complex image-question relations that are necessary for accurately answering questions. Besides, most methods neglect the modeling of the intra-modality interactions that is also important to VQA. In this work, we propose a novel DenIII framework for modeling dense inter- and intra-modality interactions. It densely connects all pairwise layers of the network via the proposed Inter- and Intra-modality Attention Connectors, capturing fine-grained interplay across all hierarchical levels. The Inter-modality Attention Connector efficiently connects the multi-modality features at any two layers with bidirectional attention, capturing the inter-modality interactions. While the Intra-modality Attention Connector connects the features of the same modality with unidirectional attention, and models the intra-modality interactions. Extensive ablation studies and visualizations validate the effectiveness of our method, and DenIII achieves state-of-the-art or competitive performance on three publicly available datasets. Fei Liu 0047, Jing Liu 0001, Zhiwei Fang, Richang Hong, Hanqing Lu |
IEEE Trans. Multim. | 5 |
| 2021 | Scene Segmentation With Dual Relation-Aware Attention NetworkabstractIn this article, we propose a Dual Relation-aware Attention Network (DRANet) to handle the task of scene segmentation. How to efficiently exploit context is essential for pixel-level recognition. To address the issue, we adaptively capture contextual information based on the relation-aware attention mechanism. Especially, we append two types of attention modules on the top of the dilated fully convolutional network (FCN), which model the contextual dependencies in spatial and channel dimensions, respectively. In the attention modules, we adopt a self-attention mechanism to model semantic associations between any two pixels or channels. Each pixel or channel can adaptively aggregate context from all pixels or channels according to their correlations. To reduce the high cost of computation and memory caused by the abovementioned pairwise association computation, we further design two types of compact attention modules. In the compact attention modules, each pixel or channel is built into association only with a few numbers of gathering centers and obtains corresponding context aggregation over these gathering centers. Meanwhile, we add a cross-level gating decoder to selectively enhance spatial details that boost the performance of the network. We conduct extensive experiments to validate the effectiveness of our network and achieve new state-of-the-art segmentation performance on four challenging scene segmentation data sets, i.e., Cityscapes, ADE20K, PASCAL Context, and COCO Stuff data sets. In particular, a Mean IoU score of 82.9% on the Cityscapes test set is achieved without using extra coarse annotated data. Jun Fu 0005, Jing Liu 0001, Jie Jiang 0016, Yong Li 0034, Yongjun Bao, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | Progressive Bi-C3D Pose Grammar for Human Pose EstimationabstractIn this paper, we propose a progressive pose grammar network learned with Bi-C3D (Bidirectional Convolutional 3D) for human pose estimation. Exploiting the dependencies among the human body parts proves effective in solving the problems such as complex articulation, occlusion and so on. Therefore, we propose two articulated grammars learned with Bi-C3D to build the relationships of the human joints and exploit the contextual information of human body structure. Firstly, a local multi-scale Bi-C3D kinematics grammar is proposed to promote the message passing process among the locally related joints. The multi-scale kinematics grammar excavates different levels human context learned by the network. Moreover, a global sequential grammar is put forward to capture the long-range dependencies among the human body joints. The whole procedure can be regarded as a local-global progressive refinement process. Without bells and whistles, our method achieves competitive performance on both MPII and LSP benchmarks compared with previous methods, which confirms the feasibility and effectiveness of C3D in information interactions. Yingying Chen 0003, Jinqiao Wang, Hanqing Lu |
AAAI | 4 |
| 2020 | Decoupled Spatial-Temporal Attention Network for Skeleton-Based Action-Gesture Recognition
Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
ACCV (5) | 4 |
| 2020 | Skeleton-Based Action Recognition With Shift Graph Convolutional NetworkabstractAction recognition with skeleton data is attracting more attention in computer vision. Recently, graph convolutional networks (GCNs), which model the human body skeletons as spatiotemporal graphs, have obtained remarkable performance. However, the computational complexity of GCN-based methods are pretty heavy, typically over 15 GFLOPs for one action sample. Recent works even reach about 100 GFLOPs. Another shortcoming is that the receptive fields of both spatial graph and temporal graph are inflexible. Although some works enhance the expressiveness of spatial graph by introducing incremental adaptive modules, their performance is still limited by regular GCN structures. In this paper, we propose a novel shift graph convolutional network (Shift-GCN) to overcome both shortcomings. Instead of using heavy regular graph convolutions, our Shift-GCN is composed of novel shift graph operations and lightweight point-wise convolutions, where the shift graph operations provide flexible receptive fields for both spatial graph and temporal graph. On three datasets for skeleton-based action recognition, the proposed Shift-GCN notably exceeds the state-of-the-art methods with more than 10 times less computational complexity. Ke Cheng 0002, Yifan Zhang 0001, Weihan Chen, Jian Cheng 0001, Hanqing Lu |
CVPR | 6 |
| 2020 | Normalized and Geometry-Aware Self-Attention Network for Image CaptioningabstractSelf-attention (SA) network has shown profound value in image captioning. In this paper, we improve SA from two aspects to promote the performance of image captioning. First, we propose Normalized Self-Attention (NSA), a reparameterization of SA that brings the benefits of normalization inside SA. While normalization is previously only applied outside SA, we introduce a novel normalization method and demonstrate that it is both possible and beneficial to perform it on the hidden activations inside SA. Second, to compensate for the major limit of Transformer that it fails to model the geometry structure of the input objects, we propose a class of Geometry-aware Self-Attention (GSA) that extends SA to explicitly and efficiently consider the relative geometry relations between the objects in the image. To construct our image captioning model, we combine the two modules and apply it to the vanilla self-attention network. We extensively evaluate our proposals on MS-COCO image captioning dataset and superior results are achieved when comparing to state-of-the-art approaches. Further experiments on three challenging tasks, i.e. video captioning, machine translation, and visual question answering, show the generality of our methods. Longteng Guo, Jing Liu 0001, Shichen Lu, Hanqing Lu |
CVPR | 6 |
| 2020 | Decoupling GCN with DropGraph Module for Skeleton-Based Action Recognition
Ke Cheng 0002, Yifan Zhang 0001, Congqi Cao, Lei Shi 0018, Jian Cheng 0001, Hanqing Lu |
ECCV (24) | 6 |
| 2020 | Learning Feature Embeddings for Discriminant Model Based Tracking
Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu |
ECCV (15) | 5 |
| 2020 | Occlusion-Aware Siamese Network for Human Pose Estimation
Yingying Chen 0003, Yunze Gao, Jinqiao Wang, Hanqing Lu |
ECCV (20) | 5 |
| 2020 | Point Set Attention Network For Semantic SegmentationabstractSelf-attention mechanism which aggregates features by capturing long-range relation between each pixel and its context, has been widely used in semantic segmentation task. However, only considering the similarity between pixels is easily disturbed by noisy pixels. In this paper, we propose a Point Set Attention Network (PSANet) for improving self-attention mechanism by correcting the noisy pixels. Specifically, we emphasize to contribute mutual improvement between pixels of the same class. For the pixel at a certain position, we consider current pixel and pixels in its neighborhood as a point set and generate a context-aware mask for selecting pixels belong to the same class as current pixel. Then the selected pixels are aggregated for a center pixel, with which current pixel is updated. In this way, pixels belong to the same class may have similar feature representation, thus promoting intra-class mutual improvement. Next, we compute the relation between each updated pixel and its context features, which often represent all pixels in the feature map. Finally, we aggregate context features on each pixel according to their relations with the pixel. We conduct extensive experiments to validate the effectiveness of our network and achieve outstanding performance on Cityscapes and PASCAL Context datasets. Jie Jiang 0016, Jing Liu 0001, Jun Fu 0005, Hanqing Lu |
ICIP | 5 |
| 2020 | Rethinking The Pid Optimizer For Stochastic Optimization Of Deep NetworksabstractStochastic gradient descent with momentum (SGD-Momentum) always causes the overshoot problem due to the integral action of the momentum term. Recently, an ID optimizer is proposed to solve the overshoot problem with the help of derivative information. However, the derivative term suffers from the interference of the high-frequency noise, especially for the stochastic gradient descent method that uses minibatch data in each update step. In this work, we propose a complete PID optimizer, which weakens the effect of the D term and adds a P term to more stably alleviate the overshoot problem. To further reduce the interference of the high-frequency noise, two effective and efficient methods are proposed to stabilize the training process. Extensive experiments on three widely used benchmark datasets with different scales, i.e., MNIST, Cifar10 and TinyImageNet, demonstrate the superiority of our proposed PID optimizer on various popular deep neural networks. Lei Shi 0018, Yifan Zhang 0001, Wanguo Wang, Jian Cheng 0001, Hanqing Lu |
ICME | 5 |
| 2020 | High-Speed And Accurate Scale Estimation For Visual Tracking With Gaussian Process RegressionabstractRecent years have seen remarkable progress in the visual tracking domain. However, it remains a challenging task to estimate the scale of target efficiently and accurately. In this paper, we present a novel and high-performance scale estimation approach for tracking-by-detection framework. The proposed approach, named GPAS, formulates the scale estimation as a Gaussian process regression problem based on scale pyramid representation. In general, it enjoys the following there advantages. (i) Efficient. It only takes 2ms to estimate the scale of a target on a single CPU. (ii) Accurate. Without bells and whistles, its accuracy surpasses all previous hand-crafted features based scale estimation methods by large margins. (iii) Generic. It can be incorporated into any tracking-by-detection framework based trackers easily. Experiment results show that compared to the latest and classical scale estimation method, fDSST, our GPAS significantly improves the performance by 6.2% in mean distance precision, 8.9% in mean overlap precision, and 5.5% in mean AUC on 28 sequences of OTB2013 with significant scale variations. Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu |
ICME | 5 |
| 2020 | Motion Complementary Network for Efficient Action RecognitionabstractBoth two-stream ConvNet and 3D ConvNet are widely used in action recognition. However, both methods are not efficient for deployment: calculating optical flow is very slow, while 3D convolution is computationally expensive. Our key insight is that the motion information from optical flow maps is complementary to the motion information from 3D ConvNet. Instead of simply combining these two methods, we propose two novel techniques to enhance the performance with less computational cost: fixed-motion-accumulation and balanced-motion-policy. With these two techniques, we propose a novel framework called Efficient Motion Complementary Network(EMC-Net) that enjoys both high efficiency and high performance. We conduct extensive experiments on Kinetics, UCF101, and Jester datasets. We achieve notably higher performance while consuming 4.7× less computation than I3D, 11.6× less computation than ECO, 17.8× less computation than R(2+1)D. On Kinetics dataset, we achieve 2.6% better performance than the recent proposed TSM with 1.4× fewer FLOPs and 10ms faster on K80 GPU. Ke Cheng 0002, Yifan Zhang 0001, Chenghua Li, Jian Cheng 0001, Hanqing Lu |
ICPR | 5 |
| 2020 | PEAN: 3D Hand Pose Estimation Adversarial NetworkabstractDespite recent emerging research attention, 3D hand pose estimation still suffers from the problems of predicting inaccurate or invalid poses which conflict with physical and kinematic constraints. To address these problems, we propose a novel 3D hand pose estimation adversarial network (PEAN) which can implicitly utilize such constraints to regularize the prediction in an adversarial learning framework. PEAN contains two parts: a 3D hierarchical estimation network (3DHNet) to predict hand pose, which decouples the task into multiple subtasks with a hierarchical structure; a pose discrimination network (PDNet) to judge the reasonableness of the estimated 3D hand pose, which back-propagates the constraints to the estimation network. During the adversarial learning process, PDNet is expected to distinguish the estimated 3D hand pose and the ground truth, while 3DHNet is expected to estimate more valid pose to confuse PDNet. In this way, 3DHNet is capable of generating 3D poses with accurate positions and adaptively adjusting the invalid poses without additional prior knowledge. Experiments show that the proposed 3DHNet does a good job in predicting hand poses, and introducing PDNet to 3DHNet does further improve the accuracy and reasonableness of the predicted results. As a result, the proposed PEAN achieves the state-of-the-art performance on three public hand pose estimation datasets. Linhui Sun, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
ICPR | 4 |
| 2020 | Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent LearningabstractMost image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently, non-autoregressive decoding has been proposed in machine translation to speed up the inference time by generating all words in parallel. Typically, these models use the word-level cross-entropy loss to optimize each word independently. However, such a learning process fails to consider the sentence-level consistency, thus resulting in inferior generation quality of these non-autoregressive models. In this paper, we propose a Non-Autoregressive Image Captioning (NAIC) model with a novel training paradigm: Counterfactuals-critical Multi-Agent Learning (CMAL). CMAL formulates NAIC as a multi-agent reinforcement learning system where positions in the target sequence are viewed as agents that learn to cooperatively maximize a sentence-level reward. Besides, we propose to utilize massive unlabeled images to boost captioning performance. Extensive experiments on MSCOCO image captioning benchmark show that our NAIC model achieves a performance comparable to state-of-the-art autoregressive models, while brings 13.9x decoding speedup. Longteng Guo, Jing Liu 0001, Xingjian He, Jie Jiang 0016, Hanqing Lu |
IJCAI | 6 |
| 2020 | Dual Hierarchical Temporal Convolutional Network with QA-Aware Dynamic Normalization for Video Story Question AnsweringabstractVideo story question answering (video story QA) is a challenging problem, as it requires a joint understanding of diverse data sources (i.e., video, subtitle, question, and answer choices). Existing approaches for video story QA have several common defects: (1) single temporal scale; (2) static and rough multimodal interaction; and (3) insufficient (or shallow) exploitation of both question and answer choices. In this paper, we propose a novel framework named Dual Hierarchical Temporal Convolutional Network (DHTCN) to address the aforementioned defects together. The proposed DHTCN explores multiple temporal scales by building hierarchical temporal convolutional network. In each temporal convolutional layer, two key components, namely AttLSTM and QA-Aware Dynamic Normalization, are introduced to capture the temporal dependency and the multimodal interaction in a dynamic and fine-grained manner. To enable sufficient exploitation of both question and answer choices, we increase the depth of QA pairs with a stack of non-linear layers, and exploit QA pairs in each layer of the network. Extensive experiments are conducted on two widely used datasets: TVQA and MovieQA, demonstrating the effectiveness of DHTCN. Our model obtains state-of-the-art results on the both datasets. Fei Liu 0047, Jing Liu 0001, Richang Hong, Hanqing Lu |
ACM Multimedia | 5 |
| 2020 | Progressive rectification network for irregular text recognition
Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu |
Sci. China Inf. Sci. | 4 |
| 2020 | Siamese Deformable Cross-Correlation Network for Real-Time Visual Tracking
Linyu Zheng, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang, Hanqing Lu |
Neurocomputing | 5 |
| 2020 | Food det: Detecting foods in refrigerator with supervised transformer network
Yousong Zhu, Xu Zhao 0003, Chaoyang Zhao, Jinqiao Wang, Hanqing Lu |
Neurocomputing | 5 |
| 2020 | Contextual deconvolution network for semantic segmentation
Jun Fu 0005, Jing Liu 0001, Yong Li 0034, Yongjun Bao, Weipeng Yan, Zhiwei Fang, Hanqing Lu |
Pattern Recognit. | 7 |
| 2020 | Gesture recognition based on deep deformable 3D convolutional neural networks
Yifan Zhang 0001, Lei Shi 0018, Yi Wu 0001, Ke Cheng 0002, Jian Cheng 0001, Hanqing Lu |
Pattern Recognit. | 6 |
| 2020 | Skeleton-Based Action Recognition With Multi-Stream Adaptive Graph Convolutional NetworksabstractGraph convolutional networks (GCNs), which generalize CNNs to more generic non-Euclidean structures, have achieved remarkable performance for skeleton-based action recognition. However, there still exist several issues in the previous GCN-based models. First, the topology of the graph is set heuristically and fixed over all the model layers and input data. This may not be suitable for the hierarchy of the GCN model and the diversity of the data in action recognition tasks. Second, the second-order information of the skeleton data, i.e., the length and orientation of the bones, is rarely investigated, which is naturally more informative and discriminative for the human action recognition. In this work, we propose a novel multi-stream attention-enhanced adaptive graph convolutional neural network (MS-AAGCN) for skeleton-based action recognition. The graph topology in our model can be either uniformly or individually learned based on the input data in an end-to-end manner. This data-driven approach increases the flexibility of the model for graph construction and brings more generality to adapt to various data samples. Besides, the proposed adaptive graph convolutional layer is further enhanced by a spatial-temporal-channel attention module, which helps the model pay more attention to important joints, frames and features. Moreover, the information of both the joints and bones, together with their motion information, are simultaneously modeled in a multi-stream framework, which shows notable improvement for the recognition accuracy. Extensive experiments on the two large-scale datasets, NTU-RGBD and Kinetics-Skeleton, demonstrate that the performance of our model exceeds the state-of-the-art with a significant margin. Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
IEEE Trans. Image Process. | 4 |
| 2020 | Show, Tell, and Polish: Ruminant Decoding for Image CaptioningabstractThe encoder-decoder framework has been the base of popular image captioning models, which typically predicts the target sentence based on the encoded source image one word at a time in sequence. However, such a single-pass decoding framework encounters two problems. First, mistakes in the predicted words cannot be corrected and may propagate to the entire sentence. Second, because the single-pass decoder cannot access the following un-generated words, it can only perform local planning to choose every single word according to the preceding words, while lacks the global planning ability as for maintaining the semantic consistency and fluency of the whole sentence. In order to address the above two problems, in this work, we design a ruminant captioning framework which contains an image encoder, a base decoder, and a ruminant decoder. Specifically, the outputs of the former/base decoder are utilized as the global information to guide the words prediction of the latter/ruminant decoder, in an attempt to mimic human polishing process. We enable jointly training of the whole framework and overcome the non-differential problem of discrete words by designing a novel reinforcement learning based optimization algorithm. Experiments on two datasets (MS COCO and Flickr30 k) demonstrate that our ruminant decoding method can bring significant improvements over traditional single-pass decoding based models and achieves state-of-the-art performance. Longteng Guo, Jing Liu 0001, Shichen Lu, Hanqing Lu |
IEEE Trans. Multim. | 4 |
| 2019 | Gate-based Bidirectional Interactive Decoding Network for Scene Text RecognitionabstractScene text recognition has attracted rapidly increasing attention from the research community. Recent dominant approaches typically follow an attention-based encoder-decoder framework that uses a unidirectional decoder to perform decoding in a left-to-right manner, but ignoring equally important right-to-left grammar information. In this paper, we propose a novel Gate-based Bidirectional Interactive Decoding Network (GBIDN) for scene text recognition. Firstly, the backward decoder performs decoding from right to left and generates the reverse language context. After that, the forward decoder simultaneously utilizes the visual context from image encoder and the reverse language context from backward decoder through two attention modules. In this way, the bidirectional decoders perform effective interaction to fully fuse the bidirectional grammar information and further improve the decoding quality. Besides, in order to relieve the adverse effect of noises, we devise a gated context mechanism to adaptively make use of the visual context and reverse language context. Extensive experiments on various challenging benchmarks demonstrate the effectiveness of our method. Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu |
CIKM | 4 |
| 2019 | Dual Attention Network for Scene SegmentationabstractIn this paper, we address the scene segmentation task by capturing rich contextual dependencies based on the self-attention mechanism. Unlike previous works that capture contexts by multi-scale features fusion, we propose a Dual Attention Networks (DANet) to adaptively integrate local features with their global dependencies. Specifically, we append two types of attention modules on top of traditional dilated FCN, which model the semantic interdependencies in spatial and channel dimensions respectively. The position attention module selectively aggregates the features at each position by a weighted sum of the features at all positions. Similar features would be related to each other regardless of their distances. Meanwhile, the channel attention module selectively emphasizes interdependent channel maps by integrating associated features among all channel maps. We sum the outputs of the two attention modules to further improve feature representation which contributes to more precise segmentation results. We achieve new state-of-the-art segmentation performance on three challenging scene segmentation datasets, i.e., Cityscapes, PASCAL Context and COCO Stuff dataset. In particular, a Mean IoU score of 81.5% on Cityscapes test set is achieved without using coarse data. Jun Fu 0005, Jing Liu 0001, Haijie Tian, Yong Li 0034, Yongjun Bao, Zhiwei Fang, Hanqing Lu |
CVPR | 7 |
| 2019 | MSCap: Multi-Style Image Captioning With Unpaired Stylized TextabstractIn this paper, we propose an adversarial learning network for the task of multi-style image captioning (MSCap) with a standard factual image caption dataset and a multi-stylized language corpus without paired images. How to learn a single model for multi-stylized image captioning with unpaired data is a challenging and necessary task, whereas rarely studied in previous works. The proposed framework mainly includes four contributive modules following a typical image encoder. First, a style dependent caption generator to output a sentence conditioned on an encoded image and a specified style. Second, a caption discriminator is presented to distinguish the input sentence to be real or not. The discriminator and the generator are trained in an adversarial manner to enable more natural and human-like captions. Third, a style classifier is employed to discriminate the specific style of the input sentence. Besides, a back-translation module is designed to enforce the generated stylized captions are visually grounded, with the intuition of the cycle consistency for factual caption and stylized caption. We enable an end-to-end optimization of the whole model with differentiable softmax approximation. At last, we conduct comprehensive experiments using a combined dataset containing four caption styles to demonstrate the outstanding performance of our proposed method. Longteng Guo, Jing Liu 0001, Jiangwei Li, Hanqing Lu |
CVPR | 5 |
| 2019 | Skeleton-Based Action Recognition With Directed Graph Neural NetworksabstractThe skeleton data have been widely used for the action recognition tasks since they can robustly accommodate dynamic circumstances and complex backgrounds. In existing methods, both the joint and bone information in skeleton data have been proved to be of great help for action recognition tasks. However, how to incorporate these two types of data to best take advantage of the relationship between joints and bones remains a problem to be solved. In this work, we represent the skeleton data as a directed acyclic graph based on the kinematic dependency between the joints and bones in the natural human body. A novel directed graph neural network is designed specially to extract the information of joints, bones and their relations and make prediction based on the extracted features. In addition, to better fit the action recognition task, the topological structure of the graph is made adaptive based on the training process, which brings notable improvement. Moreover, the motion information of the skeleton sequence is exploited and combined with the spatial information to further enhance the performance in a two-stream framework. Our final model is tested on two large-scale datasets, NTU-RGBD and Skeleton-Kinetics, and exceeds state-of-the-art performance on both of them. Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
CVPR | 4 |
| 2019 | Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action RecognitionabstractIn skeleton-based action recognition, graph convolutional networks (GCNs), which model the human body skeletons as spatiotemporal graphs, have achieved remarkable performance. However, in existing GCN-based methods, the topology of the graph is set manually, and it is fixed over all layers and input samples. This may not be optimal for the hierarchical GCN and diverse samples in action recognition tasks. In addition, the second-order information (the lengths and directions of bones) of the skeleton data, which is naturally more informative and discriminative for action recognition, is rarely investigated in existing methods. In this work, we propose a novel two-stream adaptive graph convolutional network (2s-AGCN) for skeleton-based action recognition. The topology of the graph in our model can be either uniformly or individually learned by the BP algorithm in an end-to-end manner. This data-driven method increases the flexibility of the model for graph construction and brings more generality to adapt to various data samples. Moreover, a two-stream framework is proposed to model both the first-order and the second-order information simultaneously, which shows notable improvement for the recognition accuracy. Extensive experiments on the two large-scale datasets, NTU-RGBD and Kinetics-Skeleton, demonstrate that the performance of our model exceeds the state-of-the-art with a significant margin. Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
CVPR | 4 |
| 2019 | Adaptive Context Network for Scene ParsingabstractRecent works attempt to improve scene parsing performance by exploring different levels of contexts, and typically train a well-designed convolutional network to exploit useful contexts across all pixels equally. However, in this paper, we find that the context demands are varying from different pixels or regions in each image. Based on this observation, we propose an Adaptive Context Network (ACNet) to capture the pixel-aware contexts by a competitive fusion of global context and local context according to different per-pixel demands. Specifically, when given a pixel, the global context demand is measured by the similarity between the global feature and its local feature, whose reverse value can also be used to measure the local context demand. We model the two demanding measurements by the proposed global context module and local context module, respectively, to generate their adaptive contextual features. Furthermore, we import multiple such modules to build several adaptive context blocks in different levels of network to obtain a coarse-to-fine result. Finally, comprehensive experimental evaluations demonstrate the effectiveness of the proposed ACNet, and new state-of-the-arts performances are achieved on all four public datasets, i.e. Cityscapes, ADE20K, PASCAL Context, and COCO Stuff. Jun Fu 0005, Jing Liu 0001, Yong Li 0034, Yongjun Bao, Jinhui Tang 0001, Hanqing Lu |
ICCV | 7 |
| 2019 | Fast-deepKCF Without Boundary EffectabstractIn recent years, correlation filter based trackers (CF trackers) have received much attention because of their top performance. Most CF trackers, however, suffer from low frame-per-second (fps) in pursuit of higher localization accuracy by relaxing the boundary effect or exploiting the high-dimensional deep features. In order to achieve real-time tracking speed while maintaining high localization accuracy, in this paper, we propose a novel CF tracker, fdKCF*, which casts aside the popular acceleration tool, i.e., fast Fourier transform, employed by all existing CF trackers, and exploits the inherent high-overlap among real (i.e., noncyclic) and dense samples to efficiently construct the kernel matrix. Our fdKCF* enjoys the following three advantages. (i) It is efficiently trained in kernel space and spatial domain without the boundary effect. (ii) Its fps is almost independent of the number of feature channels. Therefore, it is almost real-time, i.e., 24 fps on OTB-2015, even though the high-dimensional deep features are employed. (iii) Its localization accuracy is state-of-the-art. Extensive experiments on four public benchmarks, OTB-2013, OTB-2015, VOT2016, and VOT2017, show that the proposed fdKCF* achieves the state-of-the-art localization performance with remarkably faster speed than C-COT and ECO. Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu |
ICCV | 5 |
| 2019 | Cascade Attention Network for Person Re-IdentificationabstractPerson re-identification is a challenging task due to the viewpoint, illumination and pose variations. Recent works focus on extracting part-level features to offer beneficial fine-grained information. However, the part misalignment as well as the multi-stage training process limits their performance. Inspired by the human visual attention mechanism, this paper builds a cascade attention network(CAN) to learn the discriminative person features in a coarse-to-fine manner. Firstly, we employ the human semantic parsing module to generate coarse-grained part-level attention, which corresponds to the division of human body parts and can effectively filter the background noise. Then, to extract the local detailed features within each part, we introduce spatial-channel attention module to generate fine-grained pixel-level attention, which can further highlight the distinctive characteristics and repress the irrelevant ones. Finally, we can obtain an efficient person feature descriptor by combining both the global and local features. The whole learning process is conducted end-to-end. Experimental results show that the proposed method not only considerably outperforms its counter part but also achieves competitive performance on Market-1501 and DukeMTMC. Haiyun Guo, Huiyao Wu, Chaoyang Zhao, Huichen Zhang, Jinqiao Wang, Hanqing Lu |
ICIP | 6 |
| 2019 | Language and Visual Relations Encoding for Visual Question AnsweringabstractVisual Question Answering (VQA) involves complex relations of two modalities, including the relations between words and between image regions. Thus, encoding these relations is important to accurate VQA. In this paper, we propose two modules to encode the two types of relations respectively. The language relation encoding module is proposed to encode multi-scale relations between words via a novel masked self-attention. The visual relation encoding module is proposed to encode the relations between image regions. It computes the response at a position as a weighted sum of the features at other positions in the feature maps. Extensive experiments demonstrate the effectiveness of each modules. Our model achieves state-of-the-art performance on the VQA 1.0 dataset. Fei Liu 0047, Jing Liu 0001, Zhiwei Fang, Hanqing Lu |
ICIP | 4 |
| 2019 | Gesture Recognition Using Spatiotemporal Deformable Convolutional RepresentationabstractDynamic gesture recognition, which plays an essential role in human-computer interaction, has been widely investigated but not yet addressed. The interference of the varied and complex background makes the classifier easily be misguided due to the relatively smaller size of the hands and arms compared with the full scenes. In this paper, we address the problem by proposing a novel spatiotemporal deformable convolutional neural network for end-to-end learning. To eliminate the background interference, a light-weight spatiotemporal deformable convolution module is specially designed to augment the spatiotemporal sampling locations of 3D convolution by learning additional offsets according to the preceding feature map. The proposed method is evaluated on two challenging datasets, EgoGesture and Jester, and achieves the state-of-the-art performance on both of the two datasets. The code and trained models will be released for better communication and future work. Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
ICIP | 5 |
| 2019 | Bi-Directional Message Passing Based Scanet for Human Pose EstimationabstractArticulated human pose estimation is one of the fundamental computer vision problems. In this paper, a Bi-directional Message Passing(BDMP) module is proposed to fuse convolutional features of different scales in the up-sampling process of the hourglass model for human pose estimation. Moreover, a novel module which integrates Spatial and Channelwise Attention Network(SCANet) is proposed to refine the features obtained from the message passing stage. We design a Semantics-aware Channel-wise Attention(SACWA) module to reduce the feature redundancy and enrich the semantic information simultaneously. A Sharper Spatial Attention(SSA) module based on the Gumbel-Softmax sampling is proposed to exclude the interference from cluttered background and overcomes the gradient degradation induced by the softmax normalization. The proposed framework achieves leading position on MPII benchmark against the state-of-the-arts methods with much less parameters. Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001, Hanqing Lu |
ICME | 5 |
| 2019 | Mask Guided Knowledge Distillation for Single Shot DetectorabstractIn this paper, we explore the idea of distilling small networks for object detection task. More specifically, we propose a two-stage approach to learn more compact and efficient detectors under the single-shot object detection framework by leveraging knowledge distillation. During the 1st stage, we learn the feature maps of the student model for each of the prediction head from the teacher model. Instead of fitting the whole feature map directly, here we propose the mask guided structure including not only the entire feature map (i.e. global features) but also region features covered by the object (i.e. local features), which can significantly improve the performance of the student network. For the 2nd stage, the ground-truth is used to further refine the performance. Experimental results on PASCAL VOC and KITTI dataset demonstrate the effectiveness of our proposed approach. We achieve 56.88% mAP on VOC2007 at 143 FPS with the backbone of 1/8 VGG16. Yousong Zhu, Chaoyang Zhao, Chenxia Han, Jinqiao Wang, Hanqing Lu |
ICME | 5 |
| 2019 | Reading selectively via Binary Input Gated Recurrent UnitabstractRecurrent Neural Networks (RNNs) have shown great promise in sequence modeling tasks. Gated Recurrent Unit (GRU) is one of the most used recurrent structures, which makes a good trade-off between performance and time spent. However, its practical implementation based on soft gates only partially achieves the goal to control information flow. We can hardly explain what the network has learnt internally. Inspired by human reading, we introduce binary input gated recurrent unit (BIGRU), a GRU based model using a binary input gate instead of the reset gate in GRU. By doing so, our model can read selectively during interference. In our experiments, we show that BIGRU mainly ignores the conjunctions, adverbs and articles that do not make a big difference to the document understanding, which is meaningful for us to further understand how the network works. In addition, due to reduced interference from redundant information, our model achieves better performances than baseline GRU in all the testing tasks. Peisong Wang 0001, Hanqing Lu, Jian Cheng 0001 |
IJCAI | 3 |
| 2019 | Densely Connected Attention Flow for Visual Question AnsweringabstractLearning effective interactions between multi-modal features is at the heart of visual question answering (VQA). A common defect of the existing VQA approaches is that they only consider a very limited amount of interactions, which may be not enough to model latent complex image-question relations that are necessary for accurately answering questions. Therefore, in this paper, we propose a novel DCAF (Densely Connected Attention Flow) framework for modeling dense interactions. It densely connects all pairwise layers of the network via Attention Connectors, capturing fine-grained interplay between image and question across all hierarchical levels. The proposed Attention Connector efficiently connects the multi-modal features at any two layers with symmetric co-attention, and produces interaction-aware attention features. Experimental results on three publicly available datasets show that the proposed method achieves state-of-the-art performance. Fei Liu 0047, Jing Liu 0001, Zhiwei Fang, Richang Hong, Hanqing Lu |
IJCAI | 5 |
| 2019 | Aligning Linguistic Words and Visual Semantic Units for Image CaptioningabstractImage captioning attempts to generate a sentence composed of several linguistic words, which are used to describe objects, attributes, and interactions in an image, denoted as visual semantic units in this paper. Based on this view, we propose to explicitly model the object interactions in semantics and geometry based on Graph Convolutional Networks (GCNs), and fully exploit the alignment between linguistic words and visual semantic units for image captioning. Particularly, we construct a semantic graph and a geometry graph, where each node corresponds to a visual semantic unit, i.e., an object, an attribute, or a semantic (geometrical) interaction between two objects. Accordingly, the semantic (geometrical) context-aware embeddings for each unit are obtained through the corresponding GCN learning processers. At each time step, a context gated attention module takes as inputs the embeddings of the visual semantic units and hierarchically align the current word with these units by first deciding which type of visual semantic unit (object, attribute, or interaction) the current word is about, and then finding the most correlated visual semantic units under this type. Extensive experiments are conducted on the challenging MS-COCO image captioning dataset, and superior results are reported when comparing to state-of-the-art approaches. Longteng Guo, Jing Liu 0001, Jinhui Tang 0001, Jiangwei Li, Hanqing Lu |
ACM Multimedia | 6 |
| 2019 | Erasing-based Attention Learning for Visual Question AnsweringabstractAttention learning for visual question answering remains a challenging task, where most existing methods treat the attention and the non-attention parts in isolation. In this paper, we propose to enforce the correlation between the attention and the non-attention parts as a constraint for attention learning. We first adopt an attention-guided erasing scheme to obtain the attention and the non-attention parts respectively, and then learn to separate the attention and the non-attention parts by an appropriate distance margin in a feature embedding space. Furthermore, we associate a typical classification loss with the above distance constraint to learn a more discriminative attention map for answer prediction. The proposed approach does not introduce extra model parameters or inference complexity, and can be combined with any attention-based models. Extensive ablation experiments validate the effectiveness of our method, and new state-of-the-art or competitive results on four publicly available datasets are achieved. Fei Liu 0047, Jing Liu 0001, Richang Hong, Hanqing Lu |
ACM Multimedia | 4 |
| 2019 | Reading scene text with fully convolutional sequence modeling
Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001, Hanqing Lu |
Neurocomputing | 5 |
| 2019 | Improving visual question answering using dropout and enhanced question encoder
Zhiwei Fang, Jing Liu 0001, Yong Li 0034, Yanyuan Qiao, Hanqing Lu |
Pattern Recognit. | 5 |
| 2019 | Skeleton-Based Action Recognition With Gated Convolutional Neural NetworksabstractFor skeleton-based action recognition, most of the existing works used recurrent neural networks. Using convolutional neural networks (CNNs) is another attractive solution considering their advantages in parallelization, effectiveness in feature learning, and model base sufficiency. Besides these, skeleton data are low-dimensional features. It is natural to arrange a sequence of skeleton features chronologically into an image, which retains the original information. Therefore, we solve the sequence learning problem as an image classification task using CNNs. For better learning ability, we build a classification network with stacked residual blocks and having a special design called linear skip gated connection which can benefit information propagation across multiple residual blocks. When arranging the coordinates of body joints in one frame into a skeleton feature, we systematically investigate the performance of part-based, chain-based, and traversal-based orders. Furthermore, a fully convolutional permutation network is designed to learn an optimized order for data rearrangement. Without any bells and whistles, our proposed model achieves state-of-the-art performance on two challenging benchmark datasets, outperforming existing methods significantly. Congqi Cao, Cuiling Lan, Yifan Zhang 0001, Wenjun Zeng 0001, Hanqing Lu, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Pixelwise Deep Sequence Learning for Moving Object DetectionabstractMoving object detection is an essential, well-studied but still open problem in computer vision and plays a fundamental role in many applications. Traditional approaches usually reconstruct background images with hand-crafted visual features, such as color, texture, and edge. Due to lack of prior knowledge or semantic information, it is difficult to deal with complicated and rapid changing scenes. To exploit the temporal structure of the pixel-level semantic information, in this paper, we propose an end-to-end deep sequence learning architecture for moving object detection. First, the video sequences are input into a deep convolutional encoder-decoder network for extracting pixel-wise semantic features. Then, to exploit the temporal context, we propose a novel attention long short-term memory (Attention ConvLSTM) to model pixelwise changes over time. A spatial transformer network and a conditional random field layer are finally appended to reduce the sensitivity to camera motion and smooth the foreground boundaries. A multi-task loss is proposed to jointly optimization for frame-based classification and temporal prediction in an end-to-end network. Experimental results on CDnet 2014 and LASIESTA show 12.15% and 16.71% improvement to the state of the art, respectively. Yingying Chen 0003, Jinqiao Wang, Bingke Zhu, Ming Tang 0001, Hanqing Lu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Feature Distilled TrackingabstractFeature extraction and representation is one of the most important components for fast, accurate, and robust visual tracking. Very deep convolutional neural networks (CNNs) provide effective tools for feature extraction with good generalization ability. However, extracting features using very deep CNN models needs high performance hardware due to its large computation complexity, which prohibits its extensions in real-time applications. To alleviate this problem, we aim at obtaining small and fast-to-execute shallow models based on model compression for visual tracking. Specifically, we propose a small feature distilled network (FDN) for tracking by imitating the intermediate representations of a much deeper network. The FDN extracts rich visual features with higher speed than the original deeper network. To further speed-up, we introduce a shift-and-stitch method to reduce the arithmetic operations, while preserving the spatial resolution of the distilled feature maps unchanged. Finally, a scale adaptive discriminative correlation filter is learned on the distilled feature for visual tracking to handle scale variation of the target. Comprehensive experimental results on object tracking benchmark datasets show that the proposed approach achieves 5× speed-up with competitive performance to the state-of-the-art deep trackers. Guibo Zhu, Jinqiao Wang, Peisong Wang 0001, Yi Wu 0001, Hanqing Lu |
IEEE Trans. Cybern. | 5 |
| 2019 | Attention CoupleNet: Fully Convolutional Attention Coupling Network for Object DetectionabstractThe field of object detection has made great progress in recent years. Most of these improvements are derived from using a more sophisticated convolutional neural network. However, in the case of humans, the attention mechanism, global structure information, and local details of objects all play an important role for detecting an object. In this paper, we propose a novel fully convolutional network, named as Attention CoupleNet, to incorporate the attention-related information and global and local information of objects to improve the detection performance. Specifically, we first design a cascade attention structure to perceive the global scene of the image and generate class-agnostic attention maps. Then the attention maps are encoded into the network to acquire object-aware features. Next, we propose a unique fully convolutional coupling structure to couple global structure and local parts of the object to further formulate a discriminative feature representation. To fully explore the global and local properties, we also design different coupling strategies and normalization ways to make full use of the complementary advantages between the global and local information. Extensive experiments demonstrate the effectiveness of our approach. We achieve state-of-the-art results on all three challenging data sets, i.e., a mAP of 85.7% on VOC07, 84.3% on VOC12, and 35.4% on COCO. Codes are publicly available at https://github.com/tshizys/CoupleNet. Yousong Zhu, Chaoyang Zhao, Haiyun Guo, Jinqiao Wang, Xu Zhao 0003, Hanqing Lu |
IEEE Trans. Image Process. | 6 |
| 2019 | Dynamic Collaborative TrackingabstractCorrelation filter has been demonstrated remarkable success for visual tracking recently. However, most existing methods often face model drift caused by several factors, such as unlimited boundary effect, heavy occlusion, fast motion, and distracter perturbation. To address the issue, this paper proposes a unified dynamic collaborative tracking framework that can perform more flexible and robust position prediction. Specifically, the framework learns the object appearance model by jointly training the objective function with three components: target regression submodule, distracter suppression submodule, and maximum margin relation submodule. The first submodule mainly takes advantage of the circulant structure of training samples to obtain the distinguishing ability between the target and its surrounding background. The second submodule optimizes the label response of the possible distracting region close to zero for reducing the peak value of the confidence map in the distracting region. Inspired by the structure output support vector machines, the third submodule is introduced to utilize the differences between target appearance representation and distracter appearance representation in the discriminative mapping space for alleviating the disturbance of the most possible hard negative samples. In addition, a CUR filter as an assistant detector is embedded to provide effective object candidates for alleviating the model drift problem. Comprehensive experimental results show that the proposed approach achieves the state-of-the-art performance in several public benchmark data sets. Guibo Zhu, Zhaoxiang Zhang 0001, Jinqiao Wang, Yi Wu 0001, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | BTDP: Toward Sparse Fusion with Block Term Decomposition Pooling for Visual Question AnsweringabstractBilinear models are very powerful in multimodal fusion tasks like Visual Question Answering. The predominant bilinear methods can all be seen as a kind of tensor-based decomposition operation that contains a key kernel called “core tensor.” Current approaches usually focus on reducing the computation complexity by applying low-rank constraint on the core tensor. In this article, we propose a novel bilinear architecture called Block Term Decomposition Pooling (BTDP), which not only maintains the advantages of previous bilinear methods but also conducts sparse bilinear interactions between modalities. Our method is based on Block Term Decompositions theory of tensor, which will result in a sparse and learnable block-diagonal core tensor for multimodal fusion. We prove that using such a block-diagonal core tensor is equivalent to conducting many “tiny” bilinear operations in different feature spaces. Thus, introducing sparsity into the bilinear operation can significantly increase the performance of feature fusion and improve VQA models. What is more, our BTDP is very flexible in design. We develop several variants of BTDP and discuss the effects of the diagonal blocks of core tensor. Extensive experiments on two challenging VQA-v1 and VQA-v2 datasets show that our BTDP method outperforms current bilinear models, achieving state-of-the-art performance. Zhiwei Fang, Jing Liu 0001, Xueliang Liu, Qu Tang, Yong Li 0034, Hanqing Lu |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2018 | Learning Coarse-to-Fine Structured Feature Embedding for Vehicle Re-IdentificationabstractVehicle re-identification (re-ID) is to identify the same vehicle across different cameras. It’s a significant but challenging topic, which has received little attention due to the complex intra-class and inter-class variation of vehicle images and the lack of large-scale vehicle re-ID dataset. Previous methods focus on pulling images from different vehicles apart but neglect the discrimination between vehicles from different vehicle models, which is actually quite important to obtain a correct ranking order for vehicle re-ID. In this paper, we learn a structured feature embedding for vehicle re-ID with a novel coarse-to-fine ranking loss to pull images of the same vehicle as close as possible and achieve discrimination between images from different vehicles as well as vehicles from different vehicle models. In the learnt feature space, both intra-class compactness and inter-class distinction are well guaranteed and the Euclidean distance between features directly reflects the semantic similarity of vehicle images. Furthermore, we build so far the largest vehicle re-ID dataset "Vehicle-1M," which involves nearly 1 million images captured in various surveillance scenarios. Experimental results on "Vehicle-1M" and "VehicleID" demonstrate the superiority of our proposed approach. Haiyun Guo, Chaoyang Zhao, Zhiwei Liu 0004, Jinqiao Wang, Hanqing Lu |
AAAI | 5 |
| 2018 | Answer Distillation for Visual Question Answering
Zhiwei Fang, Jing Liu 0001, Qu Tang, Yong Li 0034, Hanqing Lu |
ACCV (1) | 5 |
| 2018 | Dense Chained Attention Network for Scene Text RecognitionabstractReading text in the wild is a challenging task in computer vision. Scene text suffers from various background noise, including shadow, irrelevant symbols and background texture. In order to reduce the disturbance of background noise, we propose a dense chained attention network with stacked attention modules for scene text recognition. Each attention module learns the attention map that is adapted to corresponding features to enhance the foreground text and suppress the background noise. Besides, the attention branch is designed with the convolution-deconvolution structure which rapidly captures global information to guide the discriminative feature selection. We stack multiple attention modules to gradually refine the attention maps and capture both the low-level appearance feature and the high-level semantic information. Extensive experiments on the standard benchmarks, the Street View Text, IIIT5K, and ICDAR datasets validate the superiority of the proposed method. The dense chained attention network achieves state-of-the-art or highly competitive recognition performance. Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001, Hanqing Lu |
ICIP | 5 |
| 2018 | Tree Hierarchical CNNs for Object ParsingabstractObject parsing is a challenging topic in computer vision, which is to distinguish all parts of visual objects. Although lots of works have been proposed, it is difficult to segment complicated objects from complex scenes. Therefore, in this paper we propose a tree hierarchical CNNs for object parsing. Rather than segment all parts of objects at once, we segment object parts step by step in a tree hierarchy and then merge the results together with a full convolutional network. In the tree hierarchy, the segmentation errors of the previous layers of the network outputs could be passed down to following layers and result in accumulated errors. In order to reduce the accumulated errors, we adopt a new part-aware fusion strategy, which fuses global-level feature maps from fully convolutional networks as well as the part-level object feature maps from the output of previous layer. It also contributes to improve the integrity and robustness of object parsing. Finally, the experiments on published datasets show the superiority of the proposed approach, especially for neighboring objects in complex scene. Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001, Hanqing Lu |
ICIP | 6 |
| 2018 | Enhancing Visual Question Answering Using DropoutabstractUsing dropout in Visual Question Answering (VQA) is a common practice to prevent overfitting. However, in multi-path networks, the current way to use dropout may cause two problems: the co-adaptations of neurons and the explosion of output variance. In this paper, we propose the coherent dropout and the siamese dropouy to solve the two problems, respectively. Specifically, in coherent dropout, all relevant dropout layers in multiple paths are forced to work coherently to maximize the ability of preventing neuron co-adaptations. We show that the coherent dropout is simple in implementation but very effective to overcome overfitting. As for the explosion of output variance, we develop a siamese dropout mechanism to explicitly minimize the difference between the two output vectors produced from the same input data during training phase. Such mechanism can reduce the gap between training and inference phases and make the VQA model more robust. Extensive experiments are conducted to verify the effectiveness of coherent dropout and siamese dropout. And the results also show that our methods can bring additional improvements on the state-of-the-art VQA models. Zhiwei Fang, Jing Liu 0001, Yanyuan Qiao, Qu Tang, Yong Li 0034, Hanqing Lu |
ACM Multimedia | 6 |
| 2018 | Recent advances in efficient computation of deep convolutional neural networksabstractDeep neural networks have evolved remarkably over the past few years and they are currently the fundamental tools of many intelligent systems. At the same time, the computational complexity and resource consumption of these networks continue to increase. This poses a significant challenge to the deployment of such networks, especially in real-time applications or on resource-limited devices. Thus, network acceleration has become a hot topic within the deep learning community. As for hardware implementation of deep neural networks, a batch of accelerators based on a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC) have been proposed in recent years. In this paper, we provide a comprehensive survey of recent advances in network acceleration, compression, and accelerator design from both algorithm and hardware points of view. Specifically, we provide a thorough analysis of each of the following topics: network pruning, low-rank approximation, network quantization, teacher–student networks, compact network design, and hardware accelerators. Finally, we introduce and discuss a few possible future directions. Jian Cheng 0001, Peisong Wang 0001, Gang Li 0015, Qinghao Hu 0001, Hanqing Lu |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2018 | Body Joint Guided 3-D Deep Convolutional Descriptors for Action Recognitionabstract3-D convolutional neural networks (3-D CNNs) have been established as a powerful tool to simultaneously learn features from both spatial and temporal dimensions, which is suitable to be applied to video-based action recognition. In this paper, we propose not to directly use the activations of fully connected layers of a 3-D CNN as the video feature, but to use selective convolutional layer activations to form a discriminative descriptor for video. It pools the feature on the convolutional layers under the guidance of body joint positions. Two schemes of mapping body joints into convolutional feature maps for pooling are discussed. The body joint positions can be obtained from any off-the-shelf skeleton estimation algorithm. The helpfulness of the body joint guided feature pooling with inaccurate skeleton estimation is systematically evaluated. To make it end-to-end and do not rely on any sophisticated body joint detection algorithm, we further propose a two-stream bilinear model which can learn the guidance from the body joints and capture the spatio-temporal features simultaneously. In this model, the body joint guided feature pooling is conveniently formulated as a bilinear product operation. Experimental results on three real-world datasets demonstrate the effectiveness of body joint guided pooling which achieves promising performance. Congqi Cao, Yifan Zhang 0001, Chunjie Zhang 0001, Hanqing Lu |
IEEE Trans. Cybern. | 4 |
| 2018 | EgoGesture: A New Dataset and Benchmark for Egocentric Hand Gesture RecognitionabstractGesture is a natural interface in human-computer interaction, especially interacting with wearable devices, such as VR/AR helmet and glasses. However, in the gesture recognition community, it lacks of suitable datasets for developing egocentric (first-person view) gesture recognition methods, in particular in the deep learning era. In this paper, we introduce a new benchmark dataset named EgoGesture with sufficient size, variation, and reality to be able to train deep neural networks. This dataset contains more than 24 000 gesture samples and 3 000 000 frames for both color and depth modalities from 50 distinct subjects. We design 83 different static and dynamic gestures focused on interaction with wearable devices and collect them from six diverse indoor and outdoor scenes, respectively, with variation in background and illumination. We also consider the scenario when people perform gestures while they are walking. The performances of several representative approaches are systematically evaluated on two tasks: gesture classification in segmented data and gesture spotting and recognition in continuous data. Our empirical study also provides an in-depth analysis on input modality selection and domain adaptation between different scenes. Yifan Zhang 0001, Congqi Cao, Jian Cheng 0001, Hanqing Lu |
IEEE Trans. Multim. | 4 |
| 2018 | Social-Aware Movie Recommendation via Multimodal Network LearningabstractWith the rapid development of Internet movie industry social-aware movie recommendation systems (SMRs) have become a popular online web service that provide relevant movie recommendations to users. In this effort many existing movie recommendation approaches learn a user ranking model from user feedback with respect to the movie's content. Unfortunately this approach suffers from the sparsity problem inherent in SMR data. In the present work we address the sparsity problem by learning a multimodal network representation for ranking movie recommendations. We develop a heterogeneous SMR network for movie recommendation that exploits the textual description and movie-poster image of each movie as well as user ratings and social relationships. With this multimodal data we then present a heterogeneous information network learning framework called SMR-multimodal network representation learning (MNRL) for movie recommendation. To learn a ranking metric from the heterogeneous information network we also developed a multimodal neural network model. We evaluated this model on a large-scale dataset from a real world SMR Web site and we find that SMR-MNRL achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Hanqing Lu, Tim Weninger, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IEEE Trans. Multim. | 3 |
| 2018 | Collaborative Deconvolutional Neural Networks for Joint Depth Estimation and Semantic SegmentationabstractSemantic segmentation and single-view depth estimation are two fundamental problems in computer vision. They exploit the semantic and geometric properties of images, respectively, and are thus complementary in scene understanding. In this paper, we propose a collaborative deconvolutional neural network (C-DCNN) to jointly model these two problems for mutual promotion. The C-DCNN consists of two DCNNs, of which each is for one task. The DCNNs provide a finer resolution reconstruction method and are pretrained with hierarchical supervision. The feature maps from these two DCNNs are integrated via a pointwise bilinear layer, which fuses the semantic and depth information and produces higher order features. Then, the integrated features are fed into two sibling classification layers to simultaneously learn for semantic segmentation and depth estimation. In this way, we combine the semantic and depth features in a unified deep network and jointly train them to benefit each other. Specifically, during network training, we process depth estimation as a classification problem where a soft mapping strategy is proposed to map the continuous depth values into discrete probability distributions and the cross entropy loss is used. Besides, a fully connected conditional random field is also used as postprocessing to further improve the performance of semantic segmentation, where the proximity relations of pixels on position, intensity, and depth are jointly considered. We evaluate our approach on two challenging benchmarks: NYU Depth V2 and SUN RGB-D. It is demonstrated that our approach effectively utilizes these two kinds of information and achieves state-of-the-art results on both the semantic segmentation and depth estimation tasks. Jing Liu 0001, Yong Li 0034, Jun Fu 0005, Jiangyun Li, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2017 | Community-Based Question Answering via Contextual Ranking Metric Network LearningabstractThe exponential growth of information on Community-based Question Answering (CQA) sites has raised the challenges for the accurate matching of high-quality answers to the given questions. Many existing approaches learn the matching model mainly based on the semantic similarity between questions and answers, which can not effectively handle the ambiguity problem of questions and the sparsity problem of CQA data. In this paper, we propose to solve these two problems by exploiting users' social contexts. Specifically, we propose a novel framework for CQA task by exploiting both the question-answer content in CQA site and users' social contexts. The experiment on real-world dataset shows the effectiveness of our method. Hanqing Lu, Ming Kong 0001 |
AAAI | 1 |
| 2017 | Community-Based Question Answering via Asymmetric Multi-Faceted Ranking Network LearningabstractNowadays the community-based question answering (CQA) sites become the popular Internet-based web service, which have accumulated millions of questions and their posted answers over time. Thus, question answering becomes an essential problem in CQA sites, which ranks the high-quality answers to the given question. Currently, most of the existing works study the problem of question answering based on the deep semantic matching model to rank the answers based on their semantic relevance, while ignoring the authority of answerers to the given question. In this paper, we consider the problem of community-based question answering from the viewpoint of asymmetric multi-faceted ranking network embedding. We propose a novel asymmetric multi-faceted ranking network learning framework for community-based question answering by jointly exploiting the deep semantic relevance between question-answer pairs and the answerers' authority to the given question. We then develop an asymmetric ranking network learning method with deep recurrent neural networks by integrating both answers' relative quality rank to the given question and the answerers' following relations in CQA sites. The extensive experiments on a large-scale dataset from a real world CQA site show that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Hanqing Lu, Vincent Wenchen Zheng, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
AAAI | 2 |
| 2017 | Egocentric Gesture Recognition Using Recurrent 3D Convolutional Neural Networks with Spatiotemporal Transformer ModulesabstractGesture is a natural interface in interacting with wearable devices such as VR/AR helmet and glasses. The main challenge of gesture recognition in egocentric vision arises from the global camera motion caused by the spontaneous head movement of the device wearer. In this paper, we address the problem by a novel recurrent 3D convolutional neural network for end-to-end learning. We specially design a spatiotemporal transformer module with recurrent connections between neighboring time slices which can actively transform a 3D feature map into a canonical view in both spatial and temporal dimensions. To validate our method, we introduce a new dataset with sufficient size, variation and reality, which contains 83 gestures designed for interaction with wearable devices, and more than 24,000 RGB-D gesture samples from 50 subjects captured in 6 scenes. On this dataset, we show that the proposed network outperforms competing state-of-the-art algorithms. Moreover, our method can achieve state-of-the-art performance on the challenging GTEA egocentric action dataset. Congqi Cao, Yifan Zhang 0001, Yi Wu 0001, Hanqing Lu, Jian Cheng 0001 |
ICCV | 4 |
| 2017 | CoupleNet: Coupling Global Structure with Local Parts for Object Detection
Yousong Zhu, Chaoyang Zhao, Jinqiao Wang, Xu Zhao 0003, Yi Wu 0001, Hanqing Lu |
ICCV | 6 |
| 2017 | Densely connected deconvolutional network for semantic segmentationabstractRecent progress in semantic segmentation has been driven by improving the spatial resolution under Fully Convolutional Networks(FCNs). To address this problem, we propose a Densely Connected Deconvolutional Network (DCDN) for semantic segmentation. In DCDN, multiple shallow deconvolutional networks, which are called as DCDN units, are stacked one by one to make the structure deeper and guarantee the fine recovery of localization information, meanwhile, the inter-unit and intra-unit dense connections are designed to make the network easy to train since the connections improve the flow of information and gradients throughout the network. Besides, the intermediate supervisions are applied to each DCDN unit to ensure the fast convergence. Extensive experiments on two urban scene datasets, i.e., CamVid and GATECH, demonstrate that the proposed model achieves better performance than some state-of-the-art methods without using any post-processing, pretrained model, nor temporal information, whilst requiring less parameters. Jun Fu 0005, Jing Liu 0001, Hanqing Lu |
ICIP | 4 |
| 2017 | Sparse Overlap Cross-Platform Recommendation Via Adaptive Similarity Structure Regularization
Hanqing Lu, Chaochao Chen 0001, Qinyue Jiang |
ICWSM | 1 |
| 2017 | Microblog Sentiment Classification via Recurrent Random Walk Network LearningabstractMicroblog Sentiment Classification (MSC) is a challenging task in microblog mining, arising in many applications such as stock price prediction and crisis management. Currently, most of the existing approaches learn the user sentiment model from their posted tweets in microblogs, which suffer from the insufficiency of discriminative tweet representation. In this paper, we consider the problem of microblog sentiment classification from the viewpoint of heterogeneous MSC network embedding. We propose a novel recurrent random walk network learning framework for the problem by exploiting both users’ posted tweets and their social relations in microblogs. We then introduce the deep recurrent neural networks with random-walk layer for heterogeneous MSC network embedding, which can be trained end-to-end from the scratch. Weemploytheback-propagationmethodfortraining the proposed recurrent random walk network model. The extensive experiments on the large-scale public datasets from Twitter show that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Hanqing Lu, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 2 |
| 2017 | Sketch-based Image Retrieval using Generative Adversarial NetworksabstractFor sketch-based image retrieval (SBIR), we propose a generative adversarial network trained on a large number of sketches and their corresponding real images. To imitate human search process, we attempt to match candidate images with theimaginary image in user single s mind instead of the sketch query, i.e., not only the shape information of sketches but their possible content information are considered in SBIR. Specifically, a conditional generative adversarial network (cGAN) is employed to enrich the content information of sketches and recover the imaginary images, and two VGG-based encoders, which work on real and imaginary images respectively, are used to constrain their perceptual consistency from the view of feature representations. During SBIR, we first generate an imaginary image from a given sketch via cGAN, and then take the output of the learned encoder for imaginary images as the feature of the query sketch. Finally, we build an interactive SBIR system that shows encouraging performance. Longteng Guo, Jing Liu 0001, Zhonghua Luo, Hanqing Lu |
ACM Multimedia | 6 |
| 2017 | Pseudo Label based Unsupervised Deep Discriminative Hashing for Image RetrievalabstractHashing methods play an important role in large scale image retrieval. Traditional hashing methods use hand-crafted features to learn hash functions, which can not capture the high level semantic information. Deep hashing algorithms use deep neural networks to learn feature representation and hash functions simultaneously. Most of these algorithms exploit supervised information to train the deep network. However, supervised information is expensive to obtain. In this paper, we propose a pseudo label based unsupervised deep discriminative hashing algorithm. First, we cluster images via K-means and the cluster labels are treated as pseudo labels. Then we train a deep hashing network with pseudo labels by minimizing the classification loss and quantization loss. Experiments on two datasets demonstrate that our unsupervised deep discriminative hashing method outperforms the state-of-art unsupervised hashing methods. Qinghao Hu 0001, Jiaxiang Wu 0001, Jian Cheng 0001, Lifang Wu, Hanqing Lu |
ACM Multimedia | 5 |
| 2017 | RSVP: A Real-Time Surveillance Video Parsing System with Single Frame SupervisionabstractIn this demo, we present a real-time surveillance video parsing (RSVP) system to parse surveillance videos. Surveillance video parsing, which aims to segment the video frames into several labels, e.g., face, pants, left-legs, has wide applications, especially in security filed. However, it is very tedious and time-consuming to annotate all the frames in a video. We design a RSVP system to parse the surveillance videos in real-time. The RSVP system requires only one labeled frame in training stage. The RSVP system jointly considers the segmentation of preceding frames when parsing one particular frame within the video. The RSVP system is proved to be effective and efficient in real applications. Guanghui Ren, Ruihe Qian, Yao Sun 0004, Changhu Wang, Hanqing Lu, Si Liu 0001 |
ACM Multimedia | 6 |
| 2017 | Decoding with Value Networks for Neural Machine TranslationabstractNeural Machine Translation (NMT) has become a popular technology in recent years, and beam search is its de facto decoding method due to the shrunk search space and reduced computational complexity. However, since it only searches for local optima at each time step through one-step forward looking, it usually cannot output the best target sentence. Inspired by the success and methodology of AlphaGo, in this paper we propose using a prediction network to improve beam search, which takes the source sentence $x$, the currently available decoding output $y_1,\cdots, y_{t-1}$ and a candidate word $w$ at step $t$ as inputs and predicts the long-term value (e.g., BLEU score) of the partial target sentence if it is completed by the NMT model. Following the practice in reinforcement learning, we call this prediction network \emph{value network}. Specifically, we propose a recurrent structure for the value network, and train its parameters from bilingual data. During the test time, when choosing a word $w$ for decoding, we consider both its conditional probability given by the NMT model and its long-term value predicted by the value network. Experiments show that such an approach can significantly improve the translation accuracy on several translation tasks. Di He 0001, Hanqing Lu, Yingce Xia, Tao Qin 0001, Liwei Wang 0001, Tie-Yan Liu |
NIPS | 2 |
| 2017 | Learning Max-Margin GeoSocial Multimedia Network Representations for Point-of-Interest SuggestionabstractWith the rapid development of mobile devices, point-of-interest (POI) suggestion has become a popular online web service, which provides attractive and interesting locations to users. In order to provide interesting POIs, many existing POI recommendation works learn the latent representations of users and POIs from users' past visiting POIs, which suffers from the sparsity problem of POI data. In this paper, we consider the problem of POI suggestion from the viewpoint of learning geosocial multimedia network representations. We propose a novel max-margin metric geosocial multimedia network representation learning framework by exploiting users' check-in behavior and their social relations. We then develop a random-walk based learning method with max-margin metric network embedding. We evaluate the performance of our method on a large-scale geosocial multimedia network dataset and show that our method achieves the best performance than other state-of-the-art solutions. Zhou Zhao 0001, Hanqing Lu, Min Yang 0007, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang |
SIGIR | 3 |
| 2017 | Image enhancement for outdoor long-range surveillance using IQ-learning multiscale RetinexabstractThe visible light camera‐based long‐range surveillance always suffers from the complex atmosphere. When applying some traditional image enhancement methods, the computational effects behave limited because of their poor environment adaptability. To conquer that problem, a blind image quality (IQ) learning‐based multiscale Retinex, i.e. the IQ‐learning multiscale Retinex, is proposed. First, a series of typical degenerated images are collected. Second, several blind IQ evaluation metrics are computed for the dataset above. They are the image brightness degree, the image region contrast degree, the image edge blur degree, the image colour quality degree, and the image noise degree. Third, a wavelet transform multi‐scale Retinex (WT_MSR) is used to carry out the basic image enhancement. A kind of optimal enhancement is implemented by the subjective evaluation and tuning of multiple optimal control parameters (MOCPs) of WT_MSR for these degenerated dataset. Fourth, the back propagation neural network (BPNN) is used to build a connection between the IQ metrics and the MOCPs. Finally, when a new image is captured, this system will compute its IQ metrics and estimate the MOCPs for the WT_MSR by BPNN; then a kind of optimal enhancement can be realised. Many outdoor applications have shown the effectiveness of proposed method. Haoting Liu, Hanqing Lu |
IET Image Process. | 2 |
| 2017 | Cross-media analysis and reasoning: advances and directionsabstractCross-media analysis and reasoning is an active research area in computer science, and a promising direction for artificial intelligence. However, to the best of our knowledge, no existing work has summarized the state-of-the-art methods for cross-media analysis and reasoning or presented advances, challenges, and future directions for the field. To address these issues, we provide an overview as follows: (1) theory and model for cross-media uniform representation; (2) cross-media correlation understanding and deep mining; (3) cross-media knowledge graph construction and learning methodologies; (4) cross-media knowledge evolution and reasoning; (5) cross-media description and generation; (6) cross-media intelligent engines; and (7) cross-media intelligent applications. By presenting approaches, advances, and future directions in cross-media analysis and reasoning, our goal is not only to draw more attention to the state-of-the-art advances in the field, but also to provide technical insights by discussing the challenges and research directions in these areas. Yuxin Peng 0001, Wenwu Zhu 0001, Yao Zhao 0001, Changsheng Xu, Qingming Huang, Hanqing Lu, Tiejun Huang 0001, Wen Gao 0001 |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2017 | Automatic group activity annotation for mobile videos
Chaoyang Zhao, Jinqiao Wang, Jianqiang Li 0002, Hanqing Lu |
Multim. Syst. | 4 |
| 2017 | Learning discriminative context models for concurrent collective activity recognition
Chaoyang Zhao, Jinqiao Wang, Hanqing Lu |
Multim. Tools Appl. | 3 |
| 2017 | Hierarchically Supervised Deconvolutional Network for Semantic Video Segmentation
Jing Liu 0001, Yong Li 0034, Jun Fu 0005, Min Xu 0001, Hanqing Lu |
Pattern Recognit. | 6 |
| 2016 | MC-HOG Correlation Tracking with Saliency ProposalabstractDesigning effective feature and handling the model drift problem are two important aspects for online visual tracking. For feature representation, gradient and color features are most widely used, but how to effectively combine them for visual tracking is still an open problem. In this paper, we propose a rich feature descriptor, MC-HOG, by leveraging rich gradient information across multiple color channels or spaces. Then MC-HOG features are embedded into the correlation tracking framework to estimate the state of the target. For handling the model drift problem caused by occlusion or distracter, we propose saliency proposals as prior information to provide candidates and reduce background interference. In addition to saliency proposals, a ranking strategy is proposed to determine the importance of these proposals by exploiting the learnt appearance filter, historical preserved object samples and the distracting proposals. In this way, the proposed approach could effectively explore the color-gradient characteristics and alleviate the model drift problem. Extensive evaluations performed on the benchmark dataset show the superiority of the proposed method. Guibo Zhu, Jinqiao Wang, Yi Wu 0001, Xiaoyu Zhang 0002, Hanqing Lu |
AAAI | 5 |
| 2016 | Scale-Adaptive Deconvolutional Regression Network for Pedestrian Detection
Yousong Zhu, Jinqiao Wang, Chaoyang Zhao, Haiyun Guo, Hanqing Lu |
ACCV (2) | 5 |
| 2016 | Domain-sensitive Recommendation with user-item subgroup analysisabstractIn this paper, we propose a Domain-sensitive Recommendation (DsRec) algorithm, to make the rating prediction by exploring the user-item subgroup analysis simultaneously, in which a user-item subgroup is deemed as a domain consisting of a subset of items with similar attributes and a subset of users who have interests in these items. The proposed framework of DsRec includes three components: a matrix factorization model for the observed rating reconstruction, a bi-clustering model for the user-item subgroup analysis, and two regularization terms to connect the above two components into a unified formulation. Extensive experiments on three real-world datasets show that our method achieves the better performance over some state-of-the-art methods. Jing Liu 0001, Zechao Li, Xi Zhang 0018, Hanqing Lu |
ICDE | 5 |
| 2016 | Person re-identification via rich color-gradient featureabstractPerson re-identification refers to match the same pedestrian across disjoint views in non-overlapping camera networks. Lots of local and global features in the literature are put forward to solve the matching problem, where color feature is robust to viewpoint variance and gradient feature provides a rich representation robust to illumination change. However, how to effectively combine the color and gradient features is an open problem. In this paper, to effectively leverage the color-gradient property in multiple color spaces, we propose a novel Second Order Histogram feature (SOH) for person reidentification in large surveillance dataset. Firstly, we utilize discrete encoding to transform commonly used color space into Encoding Color Space (ECS), and calculate the statistical gradient features on each color channel. Then, a second order statistical distribution is calculated on each cell map with a spatial partition. In this way, the proposed SOH feature effectively leverages the statistical property of gradient and color as well as reduces the redundant information. Finally, a metric learned by KISSME [1] with Mahalanobis distance is used for person matching. Experimental results on three public datasets, VIPeR, CAVIAR and CUHK01, show the promise of the proposed approach. Lingxiang Wu, Jinqiao Wang, Guibo Zhu, Min Xu 0001, Hanqing Lu |
ICME | 5 |
| 2016 | Deep learning driven hypergraph representation for image-based emotion recognitionabstractIn this paper, we proposed a bi-stage framework for image-based emotion recognition by combining the advantages of deep convolutional neural networks (D-CNN) and hypergraphs. To exploit the representational power of D-CNN, we remodeled its last hidden feature layer as the `attribute' layer in which each hidden unit produces probabilities on a specific semantic attribute. To describe the high-order relationship among facial images, each face was assigned to various hyperedges according to the computed probabilities on different D-CNN attributes. In this way, we tackled the emotion prediction problem by a transductive learning approach, which tends to assign the same label to faces that share many incidental hyperedges (attributes), with the constraints that predicted labels of training samples should be similar to their ground truth labels. We compared the proposed approach to state-of-the-art methods and its effectiveness was demonstrated by extensive experimentation. Yuchi Huang, Hanqing Lu |
ICMI | 2 |
| 2016 | Hybrid hypergraph construction for facial expression recognitionabstractIn this paper, we proposed a novel framework for facial expression recognition, in which face images were taken as vertices in a hypergraph and the task of expression recognition was formulated as the problem of hypergraph based inference. A hybrid strategy was developed to construct hyperedges: we generated probabilities of facial action units by deep convolutional networks and took each action unit as an ‘attribute’ to represent a hyperedge; we also formed hyperedges by using embedded network features before the last full connected layer to perform local clustering. In this way, each face image was assigned to various hyperedges by exploiting the representational power of deep convolutional networks. Our facial expression recognition system generates expression labels by a hypergraph based transductive inference approach, which tends to assign the same label to vertices that share many incidental hyperedges, with the constraints that predicted labels of training images should be similar to their ground truth labels. We compared the proposed approach to state-of-the-art methods and its effectiveness was demonstrated by extensive experimentation. Yuchi Huang, Hanqing Lu |
ICPR | 2 |
| 2016 | Action Recognition with Joints-Pooled 3D Deep Convolutional Descriptors
Congqi Cao, Yifan Zhang 0001, Chunjie Zhang 0001, Hanqing Lu |
IJCAI | 4 |
| 2016 | Object-aware Deep Network for Commodity Image RetrievalabstractRecent years, with the development of e-commerce and population of mobile phones, image-based commodity retrieval has attracted much attention. This paper proposed a deep framework for commodity image retrieval(CMIR) from the view that they are same designed commodities. Our framework can catch as many design details as possible by exploring object detection and ranking sensitive feature learning, while the former is performed based on Faster R-CNN, and the later is learned with a multi-task Siamese Network. Besides, we refine the processing speed of the framework to make it a live system. Our framework is implemented on an android application based on Client/Server structure model whose server response time is about 150 ms per query. Zhiwei Fang, Jing Liu 0001, Yong Li 0034, Jinhui Tang 0001, Hanqing Lu |
ICMR | 7 |
| 2016 | Objectness-aware Semantic SegmentationabstractRecent advances in semantic segmentation are driven by the success of fully convolutional neural network (FCN). However, the coarse label map from the network and the object discrimination ability for semantic segmentation weaken the performance of those FCN-based models. To address these issues, we propose an objectness-aware semantic segmentation framework (OA-Seg) by jointly learning an object proposal network (OPN) and a lightweight deconvolutional neural network (Light-DCNN). First, OPN is learned based on a fully convolutional architecture to simultaneously predict object bounding boxes and their objectness scores. Second, we design a Light-DCNN to provide a finer upsampling way than FCN. The Light-DCNN is constructed with convolutional layers in VGG-net and their mirrored deconvolutional structure, where all fully-connected layers are removed. And hierarchical classification layers are added to multi-scale deconvolutional features to introduce more contextual information for pixel-wise label prediction. Compared with previous works, our approach performs an obvious decrease on model size and convergence time. Thorough evaluations are performed on the PASCAL VOC 2012 benchmark, and our model yields impressive results on its validation data (70.3% mean IoU) and test data (74.1% mean IoU). Jing Liu 0001, Yong Li 0034, Hanqing Lu |
ACM Multimedia | 5 |
| 2016 | Partial Multi-Modal Sparse Coding via Adaptive Similarity Structure RegularizationabstractMulti-modal sparse coding has played an important role in many multimedia applications, where data are usually with multiple modalities. Recently, various multi-modal sparse coding approaches have been proposed to learn sparse codes of multi-modal data, which assume that data appear in all modalities, or at least there is one modality containing all data. However, in real applications, it is often the case that some modalities of the data may suffer from missing information and thus result in partial multi-modality data. In this paper, we propose to solve the partial multi-modal sparse coding problem via multi-modal similarity structure regularization. Specifically, we propose a partial multi-modal sparse coding framework termed Adaptive Partial Multi-Modal Similarity Structure Regularization for Sparse Coding (AdaPM2SC), which preserves the similarity structure within the same modality and between different modalities. Experimental results conducted on two real-world datasets demonstrate that AdaPM2SC significantly outperforms the state-of-the-art methods under partial multi-modality scenario. Zhou Zhao 0001, Hanqing Lu, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
ACM Multimedia | 2 |
| 2016 | Image Classification Using Spatial Difference Descriptor Under Spatial Pyramid Matching Framework
Jiucheng Xu, Yifan Zhang 0001, Chunjie Zhang 0001, Hongsheng Yin 0001, Hanqing Lu |
MMM (1) | 6 |
| 2016 | Learning weighted part models for object tracking
Chaoyang Zhao, Jinqiao Wang, Guibo Zhu, Yi Wu 0001, Hanqing Lu |
Comput. Vis. Image Underst. | 5 |
| 2016 | Clustering based ensemble correlation tracking
Guibo Zhu, Jinqiao Wang, Hanqing Lu |
Comput. Vis. Image Underst. | 3 |
| 2016 | Multiple deep features learning for object retrieval in surveillance videosabstractEfficient indexing and retrieving objects of interest from large‐scale surveillance videos are a significant and challenging topic. In this study, the authors present an effective multiple deep features learning approach for object retrieval in surveillance videos. Based on the discriminative convolutional neural network (CNN), they can learn multiple deep features to comprehensively describe the visual object. To be specific, they utilise the CNN model pre‐trained on ImageNet ILSVRC12 and fine‐tuned on our dataset to abstract structure information. In addition, they train another CNN model supervised by 11 colour names to deliver the colour information. To improve the retrieval performance, the deep features are encoded into short binary codes by locality‐sensitive hash and fused to fast retrieve the object of interest. Retrieval experiments are performed on a dataset of 100k objects extracted from multi‐camera surveillance videos. Comparison results with other common visual features show the effectiveness of the proposed approach. Haiyun Guo, Jinqiao Wang, Hanqing Lu |
IET Comput. Vis. | 3 |
| 2016 | Object co-segmentation via salient and common regions discovery
Yong Li 0034, Jing Liu 0001, Zechao Li, Hanqing Lu, Songde Ma |
Neurocomputing | 4 |
| 2016 | Social recommendation via multi-view user preference learning
Hanqing Lu, Chaochao Chen 0001, Ming Kong 0001, Hanyi Zhang, Zhou Zhao 0001 |
Neurocomputing | 1 |
| 2016 | ActiveAd: A novel framework of linking ad videos to online products
Jinqiao Wang, Min Xu 0001, Hanqing Lu, Ian S. Burnett |
Neurocomputing | 3 |
| 2016 | Chat with illustration
Jing Liu 0001, Hanqing Lu |
Multim. Syst. | 3 |
| 2016 | Enriching one-class collaborative filtering with content information from social media
Jian Cheng 0001, Xi Zhang 0018, Qinshan Liu, Hanqing Lu |
Multim. Syst. | 5 |
| 2016 | A unified model sharing framework for moving object detection
Yingying Chen 0003, Jinqiao Wang, Min Xu 0001, Xiangjian He, Hanqing Lu |
Signal Process. | 5 |
| 2016 | Real-time people counting for indoor scenes
Jinqiao Wang, Huazhong Xu, Hanqing Lu |
Signal Process. | 4 |
| 2016 | Adaptive Content Condensation Based on Grid Optimization for Thumbnail Image GenerationabstractAn ideal thumbnail generator should effectively condense unimportant regions and keep the important content undeformed, completed, and at a proper scale, i.e., accuracy, completeness, and sufficiency. Each retargeting method has its own advantage for resizing arbitrary images. However, they often ignore the completeness and sufficiency for information presentation in thumbnails. In this paper, we formulate thumbnail generation as an image content condensation problem and propose a unified grid optimization framework to fuse multiple operators. From the view of accuracy, completeness, and sufficiency for information presentation, we exploit complementary relationships among three condensation operators and fuse them into a unified grid-based convex programming problem, which could be solved simultaneously and efficiently through numerical optimization. Besides warping energy to preserve the geometric structure of important objects, we put forward two grid-based energy terms to keep the completeness of important objects and retain them at a proper size. Finally, an adaptive procedure is proposed to dynamically adjust the contribution of loss functions for achieving optimal content condensation. Both qualitative and quantitative comparison results demonstrate that the proposed method achieves an excellent tradeoff among accuracy, completeness, and sufficiency of information preservation. The experimental results show that our approach is obviously superior to the state-of-the-art techniques. Jinqiao Wang, Yingying Chen 0003, Tao Mei 0001, Min Xu 0001, La Zhang, Hanqing Lu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2016 | A Coupled Hidden Conditional Random Field Model for Simultaneous Face Clustering and Naming in VideosabstractFor face naming in TV series or movies, a typical way is using subtitles/script alignment to get the time stamps of the names, and tagging them to the faces. We study the problem of face naming in videos when subtitles are not available. To this end, we divide the problem into two tasks: face clustering which groups the faces depicting a certain person into a cluster, and name assignment which associates a name to each face. Each task is formulated as a structured prediction problem and modeled by a hidden conditional random field (HCRF) model. We argue that the two tasks are correlated problems whose outputs can provide prior knowledge of the target prediction for each other. The two HCRFs are coupled in a unified graphical model called coupled HCRF where the joint dependence of the cluster labels and face name association is naturally embedded in the correlation between the two HCRFs. We provide an effective algorithm to optimize the two HCRFs iteratively and the performance of the two tasks on real-world data set can be both improved. Yifan Zhang 0001, Baoyuan Wu, Hanqing Lu |
IEEE Trans. Image Process. | 5 |
| 2016 | Multi-View 3D Object Retrieval With Deep Embedding NetworkabstractIn multi-view 3D object retrieval, each object is characterized by a group of 2D images captured from different views. Rather than using hand-crafted features, in this paper, we take advantage of the strong discriminative power of convolutional neural network to learn an effective 3D object representation tailored for this retrieval task. Specifically, we propose a deep embedding network jointly supervised by classification loss and triplet loss to map the high-dimensional image space into a low-dimensional feature space, where the Euclidean distance of features directly corresponds to the semantic similarity of images. By effectively reducing the intra-class variations while increasing the inter-class ones of the input images, the network guarantees that similar images are closer than dissimilar ones in the learned feature space. Besides, we investigate the effectiveness of deep features extracted from different layers of the embedding network extensively and find that an efficient 3D object representation should be a tradeoff between global semantic information and discriminative local characteristics. Then, with the set of deep features extracted from different views, we can generate a comprehensive description for each 3D object and formulate the multi-view 3D object retrieval as a set-to-set matching problem. Extensive experiments on SHREC'15 data set demonstrate the superiority of our proposed method over the previous state-of-the-art approaches with over 12% performance improvement. Haiyun Guo, Jinqiao Wang, Yue Gao 0002, Jianqiang Li 0002, Hanqing Lu |
IEEE Trans. Image Process. | 5 |
| 2016 | Multimedia News Summarization in SearchabstractIt is a necessary but challenging task to relieve users from the proliferative news information and allow them to quickly and comprehensively master the information of the whats and hows that are happening in the world every day. In this article, we develop a novel approach of multimedia news summarization for searching results on the Internet, which uncovers the underlying topics among query-related news information and threads the news events within each topic to generate a query-related brief overview. First, the hierarchical latent Dirichlet allocation (hLDA) model is introduced to discover the hierarchical topic structure from query-related news documents, and a new approach based on the weighted aggregation and max pooling is proposed to identify one representative news article for each topic. One representative image is also selected to visualize each topic as a complement to the text information. Given the representative documents selected for each topic, a time-bias maximum spanning tree (MST) algorithm is proposed to thread them into a coherent and compact summary of their parent topic. Finally, we design a friendly interface to present users with the hierarchical summarization of their required news information. Extensive experiments conducted on a large-scale news dataset collected from multiple news Web sites demonstrate the encouraging performance of the proposed solution for news summarization in news retrieval. Zechao Li, Jinhui Tang 0001, Xueming Wang, Jing Liu 0001, Hanqing Lu |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2016 | Domain-Sensitive Recommendation with User-Item Subgroup AnalysisabstractCollaborative Filtering (CF) is one of the most successful recommendation approaches to cope with information overload in the real world. However, typical CF methods equally treat every user and item, and cannot distinguish the variation of user's interests across different domains. This violates the reality that user's interests always center on some specific domains, and the users having similar tastes on one domain may have totally different tastes on another domain. Motivated by the observation, in this paper, we propose a novel Domain-sensitive Recommendation (DsRec) algorithm, to make the rating prediction by exploring the user-item subgroup analysis simultaneously, in which a user-item subgroup is deemed as a domain consisting of a subset of items with similar attributes and a subset of users who have interests in these items. The proposed framework of DsRec includes three components: a matrix factorization model for the observed rating reconstruction, a bi-clustering model for the user-item subgroup analysis, and two regularization terms to connect the above two components into a unified formulation. Extensive experiments on Movielens-100K and two real-world product review datasets show that our method achieves the better performance in terms of prediction accuracy criterion over the state-of-the-art methods. Jing Liu 0001, Zechao Li, Xi Zhang 0018, Hanqing Lu |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2016 | User Preference Learning for Online Social RecommendationabstractA social recommendation system has attracted a lot of attention recently in the research communities of information retrieval, machine learning, and data mining. Traditional social recommendation algorithms are often based on batch machine learning methods which suffer from several critical limitations, e.g., extremely expensive model retraining cost whenever new user ratings arrive, unable to capture the change of user preferences over time. Therefore, it is important to make social recommendation system suitable for real-world online applications where data often arrives sequentially and user preferences may change dynamically and rapidly. In this paper, we present a new framework of online social recommendation from the viewpoint of online graph regularized user preference learning (OGRPL), which incorporates both collaborative user-item relationship as well as item content features into an unified preference learning process. We further develop an efficient iterative procedure, OGRPL-FW which utilizes the Frank-Wolfe algorithm, to solve the proposed online optimization problem. We conduct extensive experiments on several large-scale datasets, in which the encouraging results demonstrate that the proposed algorithms obtain significantly lower errors (in terms of both RMSE and MAE) than the state-of-the-art online recommendation methods when receiving the same amount of training data in the online learning process. Zhou Zhao 0001, Hanqing Lu, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2015 | Collaborative Correlation TrackingabstractCorrelation filter based tracking has attracted many researchers’ attention in recent years for high efficiency and robustness. Most existing works focus on exploiting different characteristics with correlation filters for visual tracking, e.g. circulant structure, kernel trick, effective feature representation and context information. However, how to handle the scale variation and the model drift is still an open problem. In this paper, we propose a collaborative correlation tracker to deal with the above problems. Firstly, we extend the correlation tracking filter by embedding the scale factor into the kernelized matrix to handle the scale variation. Then a novel long-term CUR filter for detection is learnt efficiently with random sampling to alleviate model drift by detecting effective object candidates in the collaborative tracker. In this way, the proposed approach could estimate the object state accurately and handle the model drift problem effectively. Extensive experiments show the superiority of the proposed method. Guibo Zhu, Jinqiao Wang, Yi Wu 0001, Hanqing Lu |
BMVC | 4 |
| 2015 | Personalized Recommendation Meets Your Next FavoriteabstractA comprehensive understanding of user's item selection behavior is not only essential to many scientific disciplines, but also has a profound business impact on online recommendation. Recent researches have discovered that user's favorites can be divided into 2 categories: long-term and short-term. User's item selection behavior is a mixed decision of her long and short-term favorites. In this paper, we propose a unified model, namely States Transition pAir-wise Ranking Model (STAR), to address users' favorites mining for sequential-set recommendation. Our method utilizes a transition graph for collaborative filtering that accounts for mining user's short-term favorites, jointed with a generative topic model for expressing user's long-term favorites. Furthermore, a user's specific prior is introduced into our unified model for better modeling personalization. Technically, we develop a pair-wise ranking loss function for parameters learning. Empirically, we measure the effectiveness of our method using two real-world datasets and the results show that our method outperforms state-of-the-art methods. Jian Cheng 0001, Hanqing Lu |
CIKM | 4 |
| 2015 | Online sketching hashingabstractRecently, hashing based approximate nearest neighbor (ANN) search has attracted much attention. Extensive new algorithms have been developed and successfully applied to different applications. However, two critical problems are rarely mentioned. First, in real-world applications, the data often comes in a streaming fashion but most of existing hashing methods are batch based models. Second, when the dataset becomes huge, it is almost impossible to load all the data into memory to train hashing models. In this paper, we propose a novel approach to handle these two problems simultaneously based on the idea of data sketching. A sketch of one dataset preserves its major characters but with significantly smaller size. With a small size sketch, our method can learn hash functions in an online fashion, while needs rather low computational complexity and storage space. Extensive experiments on two large scale benchmarks and one synthetic dataset demonstrate the efficacy of the proposed method. Cong Leng, Jiaxiang Wu 0001, Jian Cheng 0001, Xiao Bai 0001, Hanqing Lu |
CVPR | 5 |
| 2015 | Relaxing from Vocabulary: Robust Weakly-Supervised Deep Learning for Vocabulary-Free Image TaggingabstractThe development of deep learning has empowered machines with comparable capability of recognizing limited image categories to human beings. However, most existing approaches heavily rely on human-curated training data, which hinders the scalability to large and unlabeled vocabularies in image tagging. In this paper, we propose a weakly-supervised deep learning model which can be trained from the readily available Web images to relax the dependence on human labors and scale up to arbitrary tags (categories). Specifically, based on the assumption that features of true samples in a category tend to be similar and noises tend to be variant, we embed the feature map of the last deep layer into a new affinity representation, and further minimize the discrepancy between the affinity representation and its low-rank approximation. The discrepancy is finally transformed into the objective function to give relevance feedback to back propagation. Experiments show that we can achieve a performance gain of 14.0% in terms of a semantic-based relevance metric in image tagging with 63,043 tags from the WordNet, against the typical deep model trained on the ImageNet 1,000 vocabulary set. Jianlong Fu, Tao Mei 0001, Jinqiao Wang, Hanqing Lu, Yong Rui |
ICCV | 5 |
| 2015 | Multiple features based shared models for background subtractionabstractBackground modeling is a fundamental problem in computer vision and usually as the first step for high-level applications. Pixel based approaches usually ignore the spatial coherence, while region based approaches are sensitive to region size and scene complexity. In this paper, we propose a robust background subtraction approach via multiple features based shared models. Each shared model is represented by a sequence of samples based on sample consensus. Each pixel dynamically searches a matched model around the neighborhood. This shared mechanism not only enhances the robustness for background noise and jitter but also significantly reduces the number of models and samples for each model. Besides, we concatenate color and texture features as multiple features according to the discriminability and complementarity, so that each pixel can find a proper model more easily. Finally, the shared models are updated by random selecting a pixel matched the model with an adaptive update rate. Experiments on ChangeDetection benchmark 2014 show that the proposed approach outperforms the state-of-the-art methods. Yingying Chen 0003, Jinqiao Wang, Jianqiang Li 0002, Hanqing Lu |
ICIP | 4 |
| 2015 | Learning deep compact descriptor with bagging auto-encoders for object retrievalabstractContent based object retrieval across large scale surveillance video dataset is a significant and challenging task, in which learning an effective compact object descriptor plays a critical role. In this paper, we propose an efficient deep compact descriptor with bagging auto-encoders. Specifically, we take advantage of discriminative CNN to extract efficient deep features, which not only involve rich semantic information but also can filter background noise. Besides, to boost the retrieval speed, auto-encoders are used to map the high-dimensional real-valued CNN features into short binary codes. Considering the instability of auto-encoder, we adopt a bagging strategy to fuse multiple auto-encoders to reduce the generalization error, thus further improving the retrieval accuracy. In addition, bagging is easy for parallel computing, so retrieval efficiency can be guaranteed. Retrieval experimental results on the dataset of 100k visual objects extracted from multi-camera surveillance videos demonstrate the effectiveness of the proposed deep compact descriptor. Haiyun Guo, Jinqiao Wang, Hanqing Lu |
ICIP | 3 |
| 2015 | Color names learning using convolutional neural networksabstractIn this paper, we propose a two-stage CNN-based framework to learn color names from web images, aiming to predict color names for tiny image patches. To deal with the noisy labels widespread in web images, we propose a self-supervised CNN (SS-CNN) model in the first stage. The SS-CNN model is trained on image patches with their own color histograms as supervision information. Thus its outputs are able to reflect the color characteristics of images without the influence of the noisy labels. In the second stage, we finetune the SS-CNN model to learn the mapping from image patches to color names, where the patch labels are inherited from its father images. Besides, sample selection is imported iteratively in turns with the finetuning process, which helps filtering out some noisy samples and further improves the model accuracy. Our model shows high representation ability to colors and achieves better performance of color naming compared with the state-of-the-art methods. Jing Liu 0001, Jinqiao Wang, Yong Li 0034, Hanqing Lu |
ICIP | 5 |
| 2015 | Dictionary learning based superpixels clustering for weakly-supervised semantic segmentationabstractThe task of weakly-supervised semantic segmentation is solved by assigning image-level labels to over-segmented superpixels. Considering that superpixels are geometrically and semantically ambiguous for label assignment, we propose a joint solution of semantic segmentation to enhance the learnability of superpixels. First, our model includes a spectral clustering item and a discriminative clustering item to obtain some clustering subsets of superpixels (ideally semantic regions), which are more separable semantically than independent superpixels. Second, sparse coding based feature for superpixel is adopted to make the representation robust to noise, and the dictionary for the sparse representation is learned together with the above clustering items. Third, a weakly supervised item for superpixels, transferred from image-level labels, is attached. We jointly formulate the above problems as a non-convex objective function, and optimize it by the constraint concave-convex programming (CCCP) algorithm. Extensive experiments on MSRC-21 and LabelMe datasets prove the effectiveness of our approach. Peng Ying, Jing Liu 0001, Hanqing Lu |
ICIP | 3 |
| 2015 | Multi-modal learning for gesture recognitionabstractWith the development of sensing equipments, data from different modalities is available for gesture recognition. In this paper, we propose a novel multi-modal learning framework. A coupled hidden Markov model (CHMM) is employed to discover the correlation and complementary information across different modalities. In this framework, we use two configurations: one is multi-modal learning and multi-modal testing, where all the modalities used during learning are still available during testing; the other is multi-modal learning and single-modal testing, where only one modality is available during testing. Experiments on two real-world gesture recognition data sets have demonstrated the effectiveness of our multi-modal learning framework. Improvements on both of the multi-modal and single-modal testing have been observed. Congqi Cao, Yifan Zhang 0001, Hanqing Lu |
ICME | 3 |
| 2015 | Learning sharable models for robust background subtractionabstractBackground modeling and subtraction is a classical topic in compute vision. Gaussian mixture modeling (GMM) is a popular choice for its capability of adaptation to background variations. Lots of improvements have been made to enhance the robustness by considering spatial consistency and temporal correlation. In this paper, we propose a sharable GMM based background subtraction approach. Firstly, a sharable mechanism is presented to model the many-to-one relationship between pixels and models. Each pixel dynamically searches the best matched model in the neighborhood. This kind of space-sharing way is robust to camera jitter, dynamic background, etc. Secondly, the sharable models are built for both background and foreground. The noises resulted by local small movements could be effectively eliminated through the background sharable models, while the integrity of moving objects is enhanced by the foreground sharable models, especially for small objects. Finally, each sharable model is updated through randomly selecting a pixel which matches this model. And a flexible mechanism is added for switching between background and foreground models. Experiments on ChangeDetection benchmark dataset demonstrate the effectiveness of our approach. Yingying Chen 0003, Jinqiao Wang, Hanqing Lu |
ICME | 3 |
| 2015 | Hashing for Distributed DataabstractRecently, hashing based approximate nearest neighbors search has attracted much attention. Extensive centralized hashing algorithms have been proposed and achieved promising performance. However, due to the large scale of many applications, the data is often stored or even collected in a distributed manner. Learning hash functions by aggregating all the data into a fusion center is infeasible because of the prohibitively expensive communication and computation overhead. In this paper, we develop a novel hashing model to learn hash functions in a distributed setting. We cast a centralized hashing model as a set of subproblems with consensus constraints. We find these subproblems can be analytically solved in parallel on the distributed compute nodes. Since no training data is transmitted across the nodes in the learning process, the communication cost of our model is independent to the data size. Extensive experiments on several large scale datasets containing up to 100 million samples demonstrate the efficacy of our method. Cong Leng, Jiaxiang Wu 0001, Jian Cheng 0001, Xi Zhang 0018, Hanqing Lu |
ICML | 5 |
| 2015 | Weakly Supervised RBM for Semantic Segmentation
Yong Li 0034, Jing Liu 0001, Hanqing Lu, Songde Ma |
IJCAI | 4 |
| 2015 | Face Clustering in Videos with Proportion Prior
Yifan Zhang 0001, Zechao Li, Hanqing Lu |
IJCAI | 4 |
| 2015 | Mobile Media ThumbnailingabstractWith the development of Multimedia and Internet techniques, massively increasing visual data, such as image and video, need to be shown and browsed as thumbnails in various digital display platforms, like PC, cell phone, etc. This demonstration presents a grid based adaptive media thumb-nailing approach to maximize user experience in mobile image and video browsing. After representative frame extraction by spectral clustering and salient region detection, we obtain thumbnails with three resizing operators: cropping, warping and scaling, and adaptively fuse them into a unified grid based convex programming problem which could be solved simultaneously and efficiently through numerical optimization. Extensive experiments and comparisons on HUAWEI Honor 6 and Samsung S5 demonstrate that the proposed method achieves an excellent information preservation for thumbnails in mobile devices. Yingying Chen 0003, Jinqiao Wang, Jing Liu 0001, Hanqing Lu |
ICMR | 4 |
| 2015 | Spatio-Temporal Triangular-Chain CRF for Activity RecognitionabstractUnderstanding human activities in video is a fundamental problem in computer vision. In real life, human activities are composed of temporal and spatial arrangement of actions. Understanding such complex activities requires recognizing not only each individual action, but more importantly, capturing their spatio-temporal relationships. This paper addresses the problem of complex activity recognition with a unified hierarchical model. We expand triangular-chain CRFs (TriCRFs) to the spatial dimension. The proposed architecture can be perceived as a spatio-temporal version of the TriCRFs, in which the labels of actions and activity are modeled jointly and their complex dependencies are exploited. Experiments show that our model generates promising results, outperforming competing methods significantly. The framework also can be applied to model other structured sequential data. Congqi Cao, Yifan Zhang 0001, Hanqing Lu |
ACM Multimedia | 3 |
| 2015 | Learning Multi-view Deep Features for Small Object Retrieval in Surveillance ScenariosabstractWith the explosive growth of surveillance videos, object retrieval has become a significant task for security monitoring. However, visual objects in surveillance videos are usually of small size with complex light conditions, view changes and partial occlusions, which increases the difficulty level of efficiently retrieving objects of interest in a large-scale dataset. Although deep features have achieved promising results on object classification and retrieval and have been verified to contain rich semantic structure property, they lack of adequate color information, which is as crucial as structure information for effective object representation. In this paper, we propose to leverage discriminative Convolutional Neural Network (CNN) to learn deep structure and color feature to form an efficient multi-view object representation. Specifically, we utilize CNN trained on ImageNet to abstract rich semantic structure information. Meanwhile, we propose a CNN model supervised by 11 color names to extract deep color features. Compared with traditional color descriptors, deep color features can capture the common color property across different illumination conditions. Then, the complementary multi-view deep features are encoded into short binary codes by Locality-Sensitive Hash (LSH) and fused to retrieve objects. Retrieval experiments are performed on a dataset of 100k objects extracted from multi-camera surveillance videos. Comparison results with several popular visual descriptors show the effectiveness of the proposed approach. Haiyun Guo, Jinqiao Wang, Min Xu 0001, Zhengjun Zha, Hanqing Lu |
ACM Multimedia | 5 |
| 2015 | Semi- and Weakly- Supervised Semantic Segmentation with Deep Convolutional Neural NetworksabstractSuccessful semantic segmentation methods typically rely on the training datasets containing a large number of pixel-wise labeled images. To alleviate the dependence on such a fully annotated training dataset, in this paper, we propose a semi- and weakly-supervised learning framework by exploring images most only with image-level labels and very few with pixel-level labels, in which two stages of Convolutional Neural Network (CNN) training are included. First, a pixel-level supervised CNN is trained on very few fully annotated images. Second, given a large number of images with only image-level labels available, a collaborative-supervised CNN is designed to jointly perform the pixel-level and image-level classification tasks, while the pixel-level labels are predicted by the fully-supervised network in the first stage. The collaborative-supervised network can remain the discriminative ability of the fully-supervised model learned with fully labeled images, and further enhance the performance by importing more weakly labeled data. Our experiments on two challenging datasets, i.e, PASCAL VOC 2007 and LabelMe LMO, demonstrate the satisfactory performance of our approach, nearly matching the results achieved when all training images have pixel-level labels. Jing Liu 0001, Yong Li 0034, Hanqing Lu |
ACM Multimedia | 4 |
| 2015 | Exclusive Constrained Discriminative Learning for Weakly-Supervised Semantic SegmentationabstractHow to import image-level labels as weak supervision to direct the region-level labeling task is the core task of weakly-supervised semantic segmentation. In this paper, we focus on designing an effective but simple weakly-supervised constraint, and propose an exclusive constrained discriminative learning model for image semantic segmentation. To be specific, we employ a discriminative linear regression model to assign subsets of superpixels with different labels. During the assignment, we construct an exclusive weakly-supervised constraint term to suppress the labeling responses of each superpixel on the labels outside its parent image-level label set. Besides, a spectral smoothing term is integrated to encourage that both visually and semantically similar superpixels have similar labels. Combining these terms, we formulate the problem as a convex objective function, which can be easily optimized via alternative iterations. Extensive experiments on MSRC-21 and LabelMe datasets demonstrate the effectiveness of the proposed model. Peng Ying, Jing Liu 0001, Hanqing Lu, Songde Ma |
ACM Multimedia | 3 |
| 2015 | A Real-Time People Counting Approach in Indoor Environment
Jinqiao Wang, Huazhong Xu, Hanqing Lu |
MMM (1) | 4 |
| 2015 | Incremental Matrix Factorization via Feature Space Re-learning for Recommender SystemabstractMatrix factorization is widely used in Recommender Systems. Although existing popular incremental matrix factorization methods are effectively in reducing time complexity, they simply assume that the similarity between items or users is invariant. For instance, they keep the item feature matrix unchanged and just update the user matrix without re-training the entire model. However, with the new users growing continuously, the fitting error would be accumulated since the extra distribution information of items has not been utilized. In this paper, we present an alternative and reasonable approach, with a relaxed assumption that the similarity between items (users) is relatively stable after updating. Concretely, utilizing the prediction error of the new data as the auxiliary features, our method updates both feature matrices simultaneously, and thus users' preference can be better modeled than merely adjusting one corresponded feature matrix. Besides, our method maintains the feature dimension in a smaller size through taking advantage of matrix sketching. Experimental results show that our proposal outperforms the existing incremental matrix factorization methods. Jian Cheng 0001, Hanqing Lu |
RecSys | 3 |
| 2015 | When Personalization Meets Conformity: Collective Similarity based Multi-Domain RecommendationabstractExisting recommender systems place emphasis on personalization to achieve promising accuracy. However, in the context of multiple domain, users are likely to seek the same behaviors as domain authorities. This conformity effect provides a wealth of prior knowledge when it comes to multi-domain recommendation, but has not been fully exploited. In particular, users whose behaviors are significant similar with the public tastes can be viewed as domain authorities. To detect these users meanwhile embed conformity into recommendation, a domain-specific similarity matrix is intuitively employed. Therefore, a collective similarity is obtained to leverage the conformity with personalization. In this paper, we establish a Collective Structure Sparse Representation(CSSR) method for multi-domain recommendation. Based on adaptive $k$-Nearest-Neighbor framework, we impose the lasso and group lasso penalties as well as least square loss to jointly optimize the collective similarity. Experimental results on real-world data confirm the effectiveness of the proposed method. Xi Zhang 0018, Jian Cheng 0001, Shuang Qiu 0002, Zhenfeng Zhu, Hanqing Lu |
SIGIR | 5 |
| 2015 | Tagging Personal Photos with Transfer Deep LearningabstractThe advent of mobile devices and media cloud services has led to the unprecedented growing of personal photo collections. One of the fundamental problems in managing the increasing number of photos is automatic image tagging. Existing research has predominantly focused on tagging general Web images with a well-labelled image database, e.g., ImageNet. However, they can only achieve limited success on personal photos due to the domain gaps between personal photos and Web images. These gaps originate from the differences in semantic distribution and visual appearance. To deal with these challenges, in this paper, we present a novel transfer deep learning approach to tag personal photos. Specifically, to solve the semantic distribution gap, we have designed an ontology consisting of a hierarchical vocabulary tailored for personal photos. This ontology is mined from $10,000$ active users in Flickr with 20 million photos and 2.7 million unique tags. To deal with the visual appearance gap, we discover the intermediate image representations and ontology priors by deep learning with bottom-up and top-down transfers across two domains, where Web images are the source domain and personal photos are the target. Moreover, we present two modes (single and batch-modes) in tagging and find that the batch-mode is highly effective to tag photo collections. We conducted personal photo tagging on 7,000 real personal photos and personal photo search on the MIT-Adobe FiveK photo dataset. The proposed tagging approach is able to achieve a performance gain of $12.8\%$ and $4.5\%$ in terms of [email protected], against the state-of-the-art hand-crafted feature-based and deep learning-based methods, respectively. Jianlong Fu, Tao Mei 0001, Kuiyuan Yang, Hanqing Lu, Yong Rui |
WWW | 4 |
| 2015 | Learning representative and discriminative image representation by deep appearance and spatial coding
Bingyuan Liu, Jing Liu 0001, Hanqing Lu |
Comput. Vis. Image Underst. | 3 |
| 2015 | Automatic face annotation in TV series by video/script alignment
Yifan Zhang 0001, Chunjie Zhang 0001, Jing Liu 0001, Hanqing Lu |
Neurocomputing | 5 |
| 2015 | Low rank driven robust facial landmark regression
Jiankang Deng, Yubao Sun, Qingshan Liu 0001, Hanqing Lu |
Neurocomputing | 4 |
| 2015 | How friends affect user behaviors? An exploration of social relation analysis for recommendation
Jian Cheng 0001, Xi Zhang 0018, Qingshan Liu 0001, Hanqing Lu |
Knowl. Based Syst. | 5 |
| 2015 | DualDS: A dual discriminative rating elicitation framework for cold start recommendation
Xi Zhang 0018, Jian Cheng 0001, Shuang Qiu 0002, Guibo Zhu, Hanqing Lu |
Knowl. Based Syst. | 5 |
| 2015 | Finding logos in real-world images with point-context representation-based region search
Jinqiao Wang, Jianlong Fu, Hanqing Lu |
Multim. Syst. | 3 |
| 2015 | Learning latent semantic model with visual consistency for image analysis
Jian Cheng 0001, Ting Rui, Hanqing Lu |
Multim. Tools Appl. | 4 |
| 2015 | Boosted MIML method for weakly-supervised image semantic segmentation
Yang Liu 0021, Zechao Li, Jing Liu 0001, Hanqing Lu |
Multim. Tools Appl. | 4 |
| 2015 | Robust Structured Subspace Learning for Data RepresentationabstractTo uncover an appropriate latent subspace for data representation, in this paper we propose a novel Robust Structured Subspace Learning (RSSL) algorithm by integrating image understanding and feature learning into a joint learning framework. The learned subspace is adopted as an intermediate space to reduce the semantic gap between the low-level visual features and the high-level semantics. To guarantee the subspace to be compact and discriminative, the intrinsic geometric structure of data, and the local and global structural consistencies over labels are exploited simultaneously in the proposed algorithm. Besides, we adopt the l2,1-norm for the formulations of loss function and regularization respectively to make our algorithm robust to the outliers and noise. An efficient algorithm is designed to solve the proposed optimization problem. It is noted that the proposed framework is a general one which can leverage several well-known algorithms as special cases and elucidate their intrinsic relationships. To validate the effectiveness of the proposed method, extensive experiments are conducted on diversity datasets for different image understanding tasks, i.e., image tagging, clustering, and classification, and the more encouraging results are achieved compared with some state-of-the-art approaches. Zechao Li, Jing Liu 0001, Jinhui Tang 0001, Hanqing Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Detection guided deconvolutional network for hierarchical feature learning
Jing Liu 0001, Bingyuan Liu, Hanqing Lu |
Pattern Recognit. | 3 |
| 2015 | Image Tag Refinement With View-Dependent Concept RepresentationsabstractImage tag refinement is the task of refining initial tags of an image such that the refined tags can better reflect the content of the image and, therefore, can help users better access that image. The quality of tag refinement depends on the quality of concept representations that build a mapping from concepts to visual images. While good progress was made in the past decade on tag refinement, the previous approaches only achieved a limited success due to their limited concept representations. In this paper, we show that the visual appearances of a concept consist of both a generic view and a specific view, and therefore we can comprehensively represent a concept by two components. To ensure a clean concept representation, this representation is learned on clean click-through data, where noises are greatly reduced. In the framework, a coarse-to-fine image tag refinement is proposed, which: (1) first generates an efficient star graph to find candidate tags but missing in the initial tag list of an input image and (2) guided by this view-dependent concept representation, formulates a probabilistic objective function to eliminate irrelevant tags. Extensive experiments on two widely used standard data sets (MIRFlickr-25K and NUS-WIDE-270K) demonstrate the effectiveness of our approach. Jianlong Fu, Jinqiao Wang, Yong Rui, Xin-Jing Wang, Tao Mei 0001, Hanqing Lu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2015 | Human Age Estimation Based on Locality and Ordinal InformationabstractIn this paper, we propose a novel feature selection-based method for facial age estimation. The face aging is a typical temporal process, and facial images should have certain ordinal patterns in the aging feature space. From the geometrical perspective, a facial image can be usually seen as sampled from a low-dimensional manifold embedded in the original high-dimensional feature space. Thus, we first measure the energy of each feature in preserving the underlying local structure information and the ordinal information of the facial images, respectively, and then we intend to learn a low-dimensional aging representation that can maximally preserve both kinds of information. To further improve the performance, we try to eliminate the redundant local information and ordinal information as much as possible by minimizing nonlinear correlation and rank correlation among features. Finally, we formulate all these issues into a unified optimization problem, which is similar to linear discriminant analysis in format. Since it is expensive to collect the labeled facial aging images in practice, we extend the proposed supervised method to a semi-supervised learning mode including the semi-supervised feature selection method and the semi-supervised age prediction algorithm. Extensive experiments are conducted on the FACES dataset, the Images of Groups dataset, and the FG-NET aging dataset to show the power of the proposed algorithms, compared to the state-of-the-arts. Qingshan Liu 0001, Weishan Dong, Xiaobin Zhu 0001, Jing Liu 0001, Hanqing Lu |
IEEE Trans. Cybern. | 6 |
| 2015 | Weighted Part Context Learning for Visual TrackingabstractContext information is widely used in computer vision for tracking arbitrary objects. Most of the existing studies focus on how to distinguish the object of interest from background or how to use keypoint-based supporters as their auxiliary information to assist them in tracking. However, in most cases, how to discover and represent both the intrinsic properties inside the object and the surrounding context is still an open problem. In this paper, we propose a unified context learning framework that can effectively capture spatiotemporal relations, prior knowledge, and motion consistency to enhance tracker's performance. The proposed weighted part context tracker (WPCT) consists of an appearance model, an internal relation model, and a context relation model. The appearance model represents the appearances of the object and the parts. The internal relation model utilizes the parts inside the object to directly describe the spatiotemporal structure property, while the context relation model takes advantage of the latent intersection between the object and background regions. Then, the three models are embedded in a max-margin structured learning framework. Furthermore, prior label distribution is added, which can effectively exploit the spatial prior knowledge for learning the classifier and inferring the object state in the tracking process. Meanwhile, we define online update functions to decide when to update WPCT, as well as how to reweight the parts. Extensive experiments and comparisons with the state of the arts demonstrate the effectiveness of the proposed method. Guibo Zhu, Jinqiao Wang, Chaoyang Zhao, Hanqing Lu |
IEEE Trans. Image Process. | 4 |
| 2015 | Ordinal Distance Metric Learning for Image RankingabstractRecently, distance metric learning (DML) has attracted much attention in image retrieval, but most previous methods only work for image classification and clustering tasks. In this brief, we focus on designing ordinal DML algorithms for image ranking tasks, by which the rank levels among the images can be well measured. We first present a linear ordinal Mahalanobis DML model that tries to preserve both the local geometry information and the ordinal relationship of the data. Then, we develop a nonlinear DML method by kernelizing the above model, considering of real-world image data with nonlinear structures. To further improve the ranking performance, we finally derive a multiple kernel DML approach inspired by the idea of multiple-kernel learning that performs different kernel operators on different kinds of image features. Extensive experiments on four benchmarks demonstrate the power of the proposed algorithms against some related state-of-the-art methods. Qingshan Liu 0001, Jing Liu 0001, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2015 | Partially Shared Latent Factor Learning With Multiview DataabstractMultiview representations reveal the fundamental attributes of the studied instances from different perspectives. Some common perspectives are reviewed by multiple views simultaneously, while some specific ones are reflected by individual views. That is, there are two kinds of properties embedded in the multiview data: 1) consistency and 2) complementarity. Different from most multiview learning approaches only focusing on either consistency or complementarity, this paper proposes a novel semisupervised multiview learning algorithm, called partially shared latent factor (PSLF) learning, which jointly exploits both consistent and complementary information among multiple views. In PSLF, a nonnegative matrix factorization (NMF)-based formulation is adopted to learn a compact and comprehensive partially shared latent representation, which is composed of common latent factors shared by multiple views and some specific latent factors to each view. With the learned representations of multiview data, we introduce a robust sparse regression model to predict the cluster labels of labeled data. By integrating the NMF-based model and the regression model, we obtain a unified formulation and propose a multiplicative-based alternative algorithm for optimization. In addition, PSLF can learn the weights of different views adaptively according to the reconstruction precisions of data matrices. Our experimental study indicates different multiview data that contains consistent and complementary information in different degrees. In addition, the encouraging results of the proposed algorithm are achieved in comparison with the state-of-the-art algorithms on real-world data sets. Jing Liu 0001, Zechao Li, Zhi-Hua Zhou, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2014 | Learning Low-Rank Representations with Classwise Block-Diagonal Structure for Robust Face RecognitionabstractFace recognition has been widely studied due to its importance in various applications. However, the case that both training images and testing images are corrupted is not well addressed. Motivated by the success of low-rank matrix recovery, we propose a novel semi-supervised low-rank matrix recovery algorithm for robust face recognition. The proposed method can learn robust discriminative representations for both training images and testing images simultaneously by exploiting the classwise block-diagonal structure. Specifically, low-rank matrix approximation can handle the possible contamination of data. Moreover, the classwise block-diagonal structure is exploited to promote discrimination of representations for robust recognition. The above issues are formulated into a unified objective function and we design an efficient optimization procedure based on augmented Lagrange multiplier method to solve it. Extensive experiments on three public databases are performed to validate the effectiveness of our approach. The strong identification capability of representations with block-diagonal structure is verified. Yong Li 0034, Jing Liu 0001, Zechao Li, Yangmuzi Zhang, Hanqing Lu, Songde Ma |
AAAI | 5 |
| 2014 | Recommendation by Mining Multiple User Behaviors with Group SparsityabstractRecently, some recommendation methods try to improvethe prediction results by integrating informationfrom user’s multiple types of behaviors. How to modelthe dependence and independence between differentbehaviors is critical for them. In this paper, we proposea novel recommendation model, the Group-Sparse MatrixFactorization (GSMF), which factorizes the ratingmatrices for multiple behaviors into the user and itemlatent factor space with group sparsity regularization.It can (1) select out the different subsets of latent factorsfor different behaviors, addressing that users’ decisionson different behaviors are determined by differentsets of factors;(2) model the dependence and independencebetween behaviors by learning the sharedand private factors for multiple behaviors automatically; (3) allow the shared factors between different behaviorsto be different, instead of all the behaviors sharingthe same set of factors. Experiments on the real-world dataset demonstrate that our model can integrate users’multiple types of behaviors into recommendation better,compared with other state-of-the-arts. Jian Cheng 0001, Xi Zhang 0018, Shuang Qiu 0002, Hanqing Lu |
AAAI | 5 |
| 2014 | What Visual Attributes Characterize an Object Class?
Jianlong Fu, Jinqiao Wang, Xin-Jing Wang, Yong Rui, Hanqing Lu |
ACCV (1) | 5 |
| 2014 | Image Representation Learning by Deep Appearance and Spatial Coding
Bingyuan Liu, Jing Liu 0001, Zechao Li, Hanqing Lu |
ACCV (1) | 4 |
| 2014 | Learning a Representative and Discriminative Part Model with Deep Convolutional Features for Scene Recognition
Bingyuan Liu, Jing Liu 0001, Jinqiao Wang, Hanqing Lu |
ACCV (1) | 4 |
| 2014 | Clustering Ensemble Tracking
Guibo Zhu, Jinqiao Wang, Hanqing Lu |
ACCV (5) | 3 |
| 2014 | Part Context Learning for Visual Tracking
Guibo Zhu, Jinqiao Wang, Chaoyang Zhao, Hanqing Lu |
BMVC | 4 |
| 2014 | Supervised Hashing with Soft ConstraintsabstractDue to the ability to preserve semantic similarity in Hamming space, supervised hashing has been extensively studied recently. Most existing approaches encourage two dissimilar samples to have maximum Hamming distance. This may lead to an unexpected consequence that two unnecessarily similar samples would have the same code if they are both dissimilar with another sample. Besides, in existing methods, all labeled pairs are treated with equal importance without considering the semantic gap, which is not conducive to thoroughly leverage the supervised information. We present a general framework for supervised hashing to address the above two limitations. We do not toughly require a dissimilar pair to have maximum Hamming distance. Instead, a soft constraint which can be viewed as a regularization to avoid over-fitting is utilized. Moreover, we impose different weights to different training pairs, and these weights can be automatically adjusted in the learning process. Experiments on two benchmarks show that the proposed method can easily outperform other state-of-the-art methods. Cong Leng, Jian Cheng 0001, Jiaxiang Wu 0001, Xi Zhang 0018, Hanqing Lu |
CIKM | 5 |
| 2014 | Fast and Accurate Image Matching with Cascade Hashing for 3D ReconstructionabstractImage matching is one of the most challenging stages in 3D reconstruction, which usually occupies half of computational cost and inaccurate matching may lead to failure of reconstruction. Therefore, fast and accurate image matching is very crucial for 3D reconstruction. In this paper, we proposed a Cascade Hashing strategy to speed up the image matching. In order to accelerate the image matching, the proposed Cascade Hashing method is designed to be three-layer structure: hashing lookup, hashing remapping, and hashing ranking. Each layer adopts different measures and filtering strategies, which is demonstrated to be less sensitive to noise. Extensive experiments show that image matching can be accelerated by our approach in hundreds times than brute force matching, even achieves ten times or more than Kd-tree based matching while retaining comparable accuracy. Jian Cheng 0001, Cong Leng, Jiaxiang Wu 0001, Hainan Cui, Hanqing Lu |
CVPR | 5 |
| 2014 | Video face naming using global sequence alignmentabstractThis paper explores the problem of automatically naming faces in TV series or films. A novel method is proposed to build association between the faces in the video and the names in the script by a global sequence alignment algorithm. We firstly build two heterogenous sequences: a face sequence and a name sequence. The elements of the two sequences are cluster labels, computed from the clustering process, and speaking names, respectively. Then the alignment of the two sequences is considered as a problem of surjection between the cluster set and the name set. The optimal solution is obtained by minimizing the Levenshtein Distance between the two sequences which is constrained by the temporal order information. Experiments on public videos demonstrate the effectiveness of our method. Yifan Zhang 0001, Shuang Qiu 0002, Hanqing Lu |
ICIP | 4 |
| 2014 | Object tracking with part-based discriminative context modelsabstractObject tracking is a classic problem in computer vision. Part-based appearance model has been applied to object tracking and shown good performance. However, how to initialize the parts is still an open question. In this paper, we believe that the selection of discriminative parts and effectively modeling the structural context information could improve the tracking performance. Therefore, we tackle the tracking problem by discovering discriminative parts through exemplar-SVM in the initialization, and then exploit the structural relationship between discriminative context parts and the object in the process of tracking, which is consensual in the spatio-temporal domain. Experimental results demonstrate that our approach outperforms state-of-the-art trackers on benchmark videos. Guibo Zhu, Jinqiao Wang, Hanqing Lu |
ICIP | 3 |
| 2014 | Community discovering guided cold-start recommendation: A discriminative approachabstractRecommendation for new users is a key challenge due to the lack of prior information from them, which is the well-known cold-start problem. Preference elicitation has been proposed as an efficient strategy for eliciting new users preference through an initial interview where new users are queried by elaborately selected items. In this paper, we propose a novel community discovering guided discriminative selection (CDDS) model for constructing query set. We exploit the community as an effective information which is not fully used in existing approaches. By integrating item selection and community discovery into one framework, our model selects most discriminative items for preference elicitation, with guidance of unsupervised community discovering process. To perform community discovering process, the model utilizes rating similarity graph and social network as a graph regular-ization. Experimental results on real-world datasets Flixster and Douban demonstrate that the proposed method outperforms traditional preference elicitation methods for cold-start recommendation. Shuang Qiu 0002, Jian Cheng 0001, Xi Zhang 0018, Biao Niu, Hanqing Lu |
ICME | 5 |
| 2014 | Regularized Hierarchical Feature Learning with Non-negative Sparsity and Selectivity for Image ClassificationabstractRecently, many deep networks are proposed to learn hierarchical image representation to replace traditional hand-designed features. To enhance the ability of the generative model to tackle discriminative computer vision tasks (e.g. image classification), we propose a hierarchical deconvolutional network with two biologically inspired properties incorporated, i.e., non-negative sparsity and selectivity. First, we propose a single layer deconvolutional model with a raw image as input, attempting to decompose the input as a weighted sum of feature maps convolving with filters. Here, the filters are the model parameters common to all the inputs, while the feature maps and the summing weights are specific to the input. The non-negative sparsity is formulated as the /i-norm regularizer on the feature map, which is used to generate feature representations for image classification. And the selectivity is forced on the filters to make different filters active different inputs, through requiring the sparsity on the summing weights specifically. The two properties are summarized into an overall cost function, which can be solved with an alternatively iterative algorithm. Then, we build multiple layer deconvolutional network by stacking the single models, where the next-layer inputs are the results of a 3D max-pooling operation on the inferred feature maps of the front layer, and train the network in a greedy layer wise scheme. Finally, we explore the feature maps of each layer to generate the image representations and input them to a SVM classifier for the classification task. Experiments on two image benchmark datasets of Caltech-101 and Caltech-256 demonstrate the encouraging performance of our model compared with other deep feature learning models as well as some hand-designed features. Bingyuan Liu, Jing Liu 0001, Xiao Bai 0001, Hanqing Lu |
ICPR | 4 |
| 2014 | Discriminative Context Models for Collective Activity RecognitionabstractContext information has been widely studied for recognizing collective activities. Most existing works assume that all individuals in a single image share the same activity label. However, in many cases, multiple activities can be coexisted and serve as the context for each other in real-world scenarios. Based on this observation, we propose a novel approach to model both the intra-class and inter-class behavior interactions among persons in the scenario. By introducing the intra-class and inter-class context descriptors, we propose a unified discriminative model to jointly capture the individual appearance information and the context patterns around the focal person in a max-margin framework. Finally, a greedy forward search method is utilized to optimally label the activities in the testing scene. Experimental results demonstrate the superiority of our approach in activity recognition. Chaoyang Zhao, Jinqiao Wang, Xiao Bai 0001, Qingshan Liu 0001, Hanqing Lu |
ICPR | 6 |
| 2014 | Mask Assisted Object Coding with Deep Learning for Object Retrieval in Surveillance VideosabstractRetrieving visual object from a large-scale video dataset is one of multimedia research focuses but a challenging task due to imprecise object extraction and partial occlusion. This paper presents a novel approach to efficiently encode and retrieve visual objects, which addresses some practical complications in surveillance videos. Specifically, we take advantage of the mask information to assist object representation, and develop an encoding method by utilizing highly nonlinear mapping with a deep neural network. Furthermore, we add some occluded noise into the learning process to enhance the robustness of dealing with background noise and partial occlusions. A real-life surveillance video data containing over 10 million objects are built to evaluate the proposed approach. Experimental results show our approach significantly outperforms state-of-the-art solutions for object retrieval in large-scale video dataset. Kezhen Teng, Jinqiao Wang, Min Xu 0001, Hanqing Lu |
ACM Multimedia | 4 |
| 2014 | Learning Binary Codes with Bagging PCA
Cong Leng, Jian Cheng 0001, Xiao Bai 0001, Hanqing Lu |
ECML/PKDD (2) | 5 |
| 2014 | Group latent factor model for recommendation with multiple user behaviorsabstractRecently, some recommendation methods try to relieve the data sparsity problem of Collaborative Filtering by exploiting data from users' multiple types of behaviors. However, most of the exist methods mainly consider to model the correlation between different behaviors and ignore the heterogeneity of them, which may make improper information transferred and harm the recommendation results. To address this problem, we propose a novel recommendation model, named Group Latent Factor Model (GLFM), which attempts to learn a factorization of latent factor space into subspaces that are shared across multiple behaviors and subspaces that are specific to each type of behaviors. Thus, the correlation and heterogeneity of multiple behaviors can be modeled by these shared and specific latent factors. Experiments on the real-world dataset demonstrate that our model can integrate users' multiple types of behaviors into recommendation better. Jian Cheng 0001, Jinqiao Wang, Hanqing Lu |
SIGIR | 4 |
| 2014 | Random subspace for binary codes learning in large scale image retrievalabstractDue to the fast query speed and low storage cost, hashing based approximate nearest neighbor search methods have attracted much attention recently. Many state of the art methods are based on eigenvalue decomposition. In these approaches, the information caught in different dimensions is unbalanced and generally most of the information is contained in the top eigenvectors. We demonstrate that this leads to an unexpected phenomenon that longer hashing code does not necessarily yield better performance. In this work, we introduce a random subspace strategy to address this limitation. At first, a small fraction of the whole feature space is randomly sampled to train the hashing algorithms each time and only the top eigenvectors are kept to generate one piece of short code. This process will be repeated several times and then the obtained many pieces of short codes are concatenated into one piece of long code. Theoretical analysis and experiments on two benchmarks confirm the effectiveness of the proposed strategy for hashing. Cong Leng, Jian Cheng 0001, Hanqing Lu |
SIGIR | 3 |
| 2014 | Item group based pairwise preference learning for personalized rankingabstractCollaborative filtering with implicit feedbacks has been steadily receiving more attention, since the abundant implicit feedbacks are more easily collected while explicit feedbacks are not necessarily always available. Several recent work address this problem well utilizing pairwise ranking method with a fundamental assumption that a user prefers items with positive feedbacks to the items without observed feedbacks, which also implies that the items without observed feedbacks are treated equally without distinction. However, users have their own preference on different items with different degrees which can be modeled into a ranking relationship. In this paper, we exploit this prior information of a user's preference from the nearest neighbor set by the neighbors' implicit feedbacks, which can split items into different item groups with specific ranking relations. We propose a novel PRIGP(Personalized Ranking with Item Group based Pairwise preference learning) algorithm to integrate item based pairwise preference and item group based pairwise preference into the same framework. Experimental results on three real-world datasets demonstrate the proposed method outperforms the competitive baselines on several ranking-oriented evaluation metrics. Shuang Qiu 0002, Jian Cheng 0001, Cong Leng, Hanqing Lu |
SIGIR | 5 |
| 2014 | Semi-supervised multi-graph hashing for scalable similarity search
Jian Cheng 0001, Cong Leng, Meng Wang 0001, Hanqing Lu |
Comput. Vis. Image Underst. | 5 |
| 2014 | Projective Matrix Factorization with unified embedding for social image tagging
Zechao Li, Jing Liu 0001, Jinhui Tang 0001, Hanqing Lu |
Comput. Vis. Image Underst. | 4 |
| 2014 | Online video synopsis of structured motion
Jinqiao Wang, Liangke Gui, Hanqing Lu, Songde Ma |
Neurocomputing | 4 |
| 2014 | Adaptive spatial partition learning for image classification
Bingyuan Liu, Jing Liu 0001, Hanqing Lu |
Neurocomputing | 3 |
| 2014 | Beyond semantic attributes: Auxiliary feature discovery for image classification
Biao Niu, Jian Cheng 0001, Yang Liu 0021, Hanqing Lu |
Neurocomputing | 4 |
| 2014 | Sparse semantic metric learning for image retrieval
Jing Liu 0001, Zechao Li, Hanqing Lu |
Multim. Syst. | 3 |
| 2014 | Interactive ads recommendation with contextual search on product topic space
Jinqiao Wang, Bo Wang 0011, Ling-Yu Duan, Qi Tian 0001, Hanqing Lu |
Multim. Tools Appl. | 5 |
| 2014 | A three-level framework for affective content analysis and its case studies
Min Xu 0001, Jinqiao Wang, Xiangjian He, Jesse S. Jin, Suhuai Luo, Hanqing Lu |
Multim. Tools Appl. | 6 |
| 2014 | Semi-supervised Unified Latent Factor learning with multi-view data
Jing Liu 0001, Zechao Li, Hanqing Lu |
Mach. Vis. Appl. | 4 |
| 2014 | Key observation selection-based effective video synopsis for camera network
Xiaobin Zhu 0001, Jing Liu 0001, Jinqiao Wang, Hanqing Lu |
Mach. Vis. Appl. | 4 |
| 2014 | Sparse representation for robust abnormality detection in crowded scenes
Xiaobin Zhu 0001, Jing Liu 0001, Jinqiao Wang, Hanqing Lu |
Pattern Recognit. | 5 |
| 2014 | A hybrid domain enhanced framework for video retargeting with spatial-temporal importance and 3D grid optimization
Jinqiao Wang, Min Xu 0001, Xiangjian He, Hanqing Lu, Doan B. Hoang |
Signal Process. | 4 |
| 2014 | Spatiotemporal Group Context for Pedestrian CountingabstractPedestrian counting has been a challenging topic, especially in video surveillance, for a long time due to the view variations, scale changes, and spatial occlusions. While most of the previous approaches try to count people within one frame, our approach addresses this problem with a group context model, which is to segment individuals into groups and model the spatiotemporal relationships between them. With the basic definitions of the group state, group event, and group relative, a group correspondence matrix is built to model the bidirectional correspondences between the groups in two consecutive frames. Then, a group context is modeled with a sequence of context masks, which encodes not only the spatiotemporal changes within a group, but also the historical relevance and spatial dependency between different groups. Finally, we assemble context masks from multiple frames and formulate the problem of pedestrian counting as a joint maximum a posteriori problem. Markov-chain Monte Carlo is utilized to search for an optimal configuration set to match the group context model. Comprehensive experiments on the PETS2009 data set and UCSD pedestrian data set show the promising performance of the proposed approach. Jinqiao Wang, Hanqing Lu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | Learning Robust Face Representation With Classwise Block-Diagonal StructureabstractFace recognition has been widely studied due to its importance in various applications. However, the case that both training images and testing images are corrupted is not well solved. To address such a problem, this paper proposes a semisupervised learning algorithm for robust face recognition. In particular, we consider three items in the proposed formulation. First, a low-rank and sparse representation for face recognition is required to handle the possible contamination of the whole data. Second, a classwise block-diagonal structure of the learned representation is expected to promote discrimination among different classes. With the structure regularization, we make the samples from different classes be reconstructed with different bases as much as possible. Third, a compact and discriminative dictionary should be learnt to handle the problem of corrupted data. Extensive experiments on three public databases are performed to validate the effectiveness of our approach. The strong identification capability of representation with block-diagonal structure is verified. Yong Li 0034, Jing Liu 0001, Hanqing Lu, Songde Ma |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | Bilayer Sparse Topic Model for Scene Analysis in Imbalanced Surveillance VideosabstractDynamic scene analysis has become a popular research area especially in video surveillance. The goal of this paper is to mine semantic motion patterns and detect abnormalities deviating from normal ones occurring in complex dynamic scenarios. To address this problem, we propose a data-driven and scene-independent approach, namely, Bilayer sparse topic model (BiSTM), where a given surveillance video is represented by a word-document hierarchical generative process. In this BiSTM, motion patterns are treated as latent topics sparsely distributed over low-level motion vectors, whereas a video clip can be sparsely reconstructed by a mixture of topics (motion pattern). In addition to capture the characteristic of extreme imbalance between numerous typical normal activities and few rare abnormalities in surveillance video data, a one-class constraint is directly imposed on the distribution of documents as a discriminant priori. By jointly learning topics and one-class document representation within a discriminative framework, the topic (pattern) space is more specific and explicit. An effective alternative iteration algorithm is presented for the model learning. Experimental results and comparisons on various public data sets demonstrate the promise of the proposed approach. Jinqiao Wang, Hanqing Lu, Songde Ma |
IEEE Trans. Image Process. | 3 |
| 2014 | Snap & Play: Auto-Generated Personalized Find-the-Difference GameabstractIn this article, by taking a popular game, the Find-the-Difference (FiDi) game, as a concrete example, we explore how state-of-the-art image processing techniques can assist in developing a personalized, automatic, and dynamic game. Unlike the traditional FiDi game, where image pairs (source image and target image) with five different patches are manually produced by professional game developers, the proposed Personalized FiDi (P-FiDi) electronic game can be played in a fully automatic Snap & Play mode.Snapmeans that players first take photos with their digital cameras. The newly captured photos are used as source images and fed into the P-FiDi system to autogenerate the counterpart target images for users toplay. Four steps are adopted to autogenerate target images: enhancing the visual quality of source images, extracting some changeable patches from the source image, selecting the most suitable combination of changeable patches and difference styles for the image, and generating the differences on the target image with state-of-the-art image processing techniques. In addition, the P-FiDi game can be easily redesigned for the im-game advertising. Extensive experiments show that the P-FiDi electronic game is satisfying in terms of player experience, seamless advertisement, and technical feasibility. Si Liu 0001, Qiang Chen 0007, Shuicheng Yan, Changsheng Xu, Hanqing Lu |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2014 | Clustering-Guided Sparse Structural Learning for Unsupervised Feature SelectionabstractMany pattern analysis and data mining problems have witnessed high-dimensional data represented by a large number of features, which are often redundant and noisy. Feature selection is one main technique for dimensionality reduction that involves identifying a subset of the most useful features. In this paper, a novel unsupervised feature selection algorithm, named clustering-guided sparse structural learning (CGSSL), is proposed by integrating cluster analysis and sparse structural analysis into a joint framework and experimentally evaluated. Nonnegative spectral clustering is developed to learn more accurate cluster labels of the input samples, which guide feature selection simultaneously. Meanwhile, the cluster labels are also predicted by exploiting the hidden structure shared by different features, which can uncover feature correlations to make the results more reliable. Row-wise sparse models are leveraged to make the proposed model suitable for feature selection. To optimize the proposed formulation, we propose an efficient iterative algorithm. Finally, extensive experiments are conducted on 12 diverse benchmarks, including face data, handwritten digit data, document data, and biomedical data. The encouraging experimental results in comparison with several representative algorithms and the theoretical analysis demonstrate the efficiency and effectiveness of the proposed algorithm for feature selection. Zechao Li, Jing Liu 0001, Yi Yang 0001, Xiaofang Zhou 0001, Hanqing Lu |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2014 | Personalized Geo-Specific Tag Recommendation for Photos on Social WebsitesabstractSocial tagging becomes increasingly important to organize and search large-scale community-contributed photos on social websites. To facilitate generating high-quality social tags, tag recommendation by automatically assigning relevant tags to photos draws particular research interest. In this paper, we focus on the personalized tag recommendation task and try to identify user-preferred, geo-location-specific as well as semantically relevant tags for a photo by leveraging rich contexts of the freely available community-contributed photos. For users and geo-locations, we assume they have different preferred tags assigned to a photo, and propose a subspace learning method to individually uncover the both types of preferences. The goal of our work is to learn a unified subspace shared by the visual and textual domains to make visual features and textual information of photos comparable. Considering the visual feature is a lower level representation on semantics than the textual information, we adopt a progressive learning strategy by additionally introducing an intermediate subspace for the visual domain, and expect it to have consistent local structure with the textual space. Accordingly, the unified subspace is mapped from the intermediate subspace and the textual space respectively. We formulate the above learning problems into a united form, and present an iterative optimization with its convergence proof. Given an untagged photo with its geo-location to a user, the user-preferred and the geo-location-specific tags are found by the nearest neighbor search in the corresponding unified spaces. Then we combine the obtained tags and the visual appearance of the photo to discover the semantically and visually related photos, among which the most frequent tags are used as the recommended tags. Experiments on a large-scale data set collected from Flickr verify the effectivity of the proposed solution. Jing Liu 0001, Zechao Li, Jinhui Tang 0001, Hanqing Lu |
IEEE Trans. Multim. | 5 |
| 2013 | Semi-supervised discriminative preference elicitation for cold-start recommendationabstractRecommendation for cold users is fairly challenging because no prior rating can be used in preference prediction. To tackle this cold-start scenario, rating elicitation is usually employed through an initial interview in which users are queried by some carefully selected items. In this paper, we propose a novel framework to mine the most valuable items to construct query set using a semi-supervised discriminative selection (SSDS) model. To learn a low dimensional representation for users in item space which can reflect their tastes to a large extent, the model incorporates category labels as discriminative information. To ensure the used labels reliable as well as all users considered, the model utilizes a semi-supervised scheme leveraging expert guidance with graph regularization. Experimental results on real-world dataset MovieLens demonstrate that the proposed SSDS model outperforms traditional preference elicitation methods on top-N measures for cold-start recommendation. Xi Zhang 0018, Jian Cheng 0001, Biao Niu, Hanqing Lu |
CIKM | 5 |
| 2013 | Weakly-Supervised Dual Clustering for Image Semantic SegmentationabstractIn this paper, we propose a novel Weakly-Supervised Dual Clustering (WSDC) approach for image semantic segmentation with image-level labels, i.e., collaboratively performing image segmentation and tag alignment with those regions. The proposed approach is motivated from the observation that super pixels belonging to an object class usually exist across multiple images and hence can be gathered via the idea of clustering. In WSDC, spectral clustering is adopted to cluster the super pixels obtained from a set of over-segmented images. At the same time, a linear transformation between features and labels as a kind of discriminative clustering is learned to select the discriminative features among different classes. The both clustering outputs should be consistent as much as possible. Besides, weakly-supervised constraints from image-level labels are imposed to restrict the labeling of super pixels. Finally, the non-convex and non-smooth objective function are efficiently optimized using an iterative CCCP procedure. Extensive experiments conducted on MSRC and Label Me datasets demonstrate the encouraging performance of our method in comparison with some state-of-the-arts. Yang Liu 0021, Jing Liu 0001, Zechao Li, Jinhui Tang 0001, Hanqing Lu |
CVPR | 5 |
| 2013 | Event Detection in Complex Scenes Using Interval Temporal ConstraintsabstractIn complex scenes with multiple atomic events happening sequentially or in parallel, detecting each individual event separately may not always obtain robust and reliable result. It is essential to detect them in a holistic way which incorporates the causality and temporal dependency among them to compensate the limitation of current computer vision techniques. In this paper, we propose an interval temporal constrained dynamic Bayesian network to extend Allen's interval algebra network (IAN) [2] from a deterministic static model to a probabilistic dynamic system, which can not only capture the complex interval temporal relationships, but also model the evolution dynamics and handle the uncertainty from the noisy visual observation. In the model, the topology of the IAN on each time slice and the interlinks between the time slices are discovered by an advanced structure learning method. The duration of the event and the unsynchronized time lags between two correlated event intervals are captured by a duration model, so that we can better determine the temporal boundary of the event. Empirical results on two real world datasets show the power of the proposed interval temporal constrained model. Yifan Zhang 0001, Hanqing Lu |
ICCV | 3 |
| 2013 | Robust Feature Encoding with Neighborhood Information for Image ClassificationabstractThe bag of visual words (BoW) model is one of the most successful model in image classification task. However, the major problem of the BoW model lies in the determination of visual words, which consists of codebook training and feature encoding phases. The traditional K-means and hard-assignment method completely ignore the structure of the local feature space, leading to high loss of information. To alleviate the information loss, we propose to incorporate the neighborhood information of the features into the codebook training and feature encoding process. We firstly propose a model to roughly measure the influence of the distribution of the neighboring features. Then we combine the proposed model with the traditional K-means method in a probability perspective to train the visual codebook. Finally, in the feature encoding phase, both the hard-assignment and soft-assignment method are improved with the proposed neighborhood information term. We investigate our method on two popular datasets: 15-Scenes and Caltech-101. Experimental results demonstrate the effectiveness of our proposed method. Bingyuan Liu, Jing Liu 0001, Chunjie Zhang 0001, Hanqing Lu |
ICIG | 5 |
| 2013 | Improving scene classification with weakly spatial symmetry informationabstractThe bag-of-visual-words (BOW) model has been widely used in the field of scene classification. Since it ignores the spatial information, the spatial-pyramid-matching (SPM) model [1] was presented by partitioning the image into increasingly fine blocks and computing histograms of local features in each block. However, the spatial symmetry has never been considered explicitly in scene classification as we known. In this paper, a novel descriptor named weakly spatial symmetry (WSS) is proposed to boost the performance of image classification. After region segmentation, the spatial symmetry is represented by L1 distances of region histograms. Four kinds of spatial symmetry are extracted in blocks of increasing scales as in SPM [1]. The WSS descriptor can be used independently or combined with BOW or SPM for scene classification. Experiments on scene-15 and caltech 101 dataset demonstrate the effectiveness of the proposed approach. Kezhen Teng, Jinqiao Wang, Qi Tian 0001, Hanqing Lu |
ICIP | 4 |
| 2013 | Fusing multi-modal features for gesture recognitionabstractThis paper proposes a novel multi-modal gesture recognition framework and introduces its application to continuous sign language recognition. A Hidden Markov Model is used to construct the audio feature classifier. A skeleton feature classifier is trained to provided complementary information based on the Dynamic Time Warping model. The confidence scores generated by two classifiers are firstly normalized and then combined to produce a weighted sum for the final recognition. Experimental results have shown that the precision and recall scores for 20 classes of our multi-modal recognition framework can achieve 0.8829 and 0.8890 respectively, which proves that our method is able to correctly reject false detection caused by single classifier. Our approach scored 0.12756 in mean Levenshtein distance and was ranked 1st in the Multi-modal Gesture Recognition Challenge in 2013. Jiaxiang Wu 0001, Jian Cheng 0001, Chaoyang Zhao, Hanqing Lu |
ICMI | 4 |
| 2013 | Object co-segmentation via discriminative low rank matrix recoveryabstractThe goal of this paper is to simultaneously segment the object regions appearing in a set of images of the same object class, known as object co-segmentation. Different from typical methods, simply assuming that the regions common among images are the object regions, we additionally consider the disturbance from consistent backgrounds, and indicate not only common regions but salient ones among images to be the object regions. To this end, we propose a Discriminative Low Rank matrix Recovery (DLRR) algorithm to divide the over-completely segmented regions (i.e.,superpixels) of a given image set into object and non-object ones. In DLRR, a low-rank matrix recovery term is adopted to detect salient regions in an image, while a discriminative learning term is used to distinguish the object regions from all the super-pixels. An additional regularized term is imported to jointly measure the disagreement between the predicted saliency and the objectiveness probability corresponding to each super-pixel of the image set. For the unified learning problem by connecting the above three terms, we design an efficient optimization procedure based on block-coordinate descent. Extensive experiments are conducted on two public datasets, i.e., MSRC and iCoseg, and the comparisons with some state-of-the-arts demonstrate the effectiveness of our work. Yong Li 0034, Jing Liu 0001, Zechao Li, Yang Liu 0021, Hanqing Lu |
ACM Multimedia | 5 |
| 2013 | Object Categorization Using Local Feature Context
Chunjie Zhang 0001, Jing Liu 0001, Hanqing Lu |
MMM (2) | 4 |
| 2013 | A Weighted One Class Collaborative Filtering with Content Topic Features
Jian Cheng 0001, Xi Zhang 0018, Qingshan Liu 0001, Hanqing Lu |
MMM (2) | 5 |
| 2013 | Collaborative Tracking: Dynamically Fusing Short-Term Trackers and Long-Term Detector
Guibo Zhu, Jinqiao Wang, Hanqing Lu |
MMM (2) | 4 |
| 2013 | TopRec: domain-specific recommendation through community topic mining in social networkabstractTraditionally, Collaborative Filtering assumes that similar users have similar responses to similar items. However, human activities exhibit heterogenous features across multiple domains such that users own similar tastes in one domain may behave quite differently in other domains. Moreover, highly sparse data presents crucial challenge in preference prediction. Intuitively, if users' interested domains are captured first, the recommender system is more likely to provide the enjoyed items while filter out those uninterested ones. Therefore, it is necessary to learn preference profiles from the correlated domains instead of the entire user-item matrix. In this paper, we propose a unified framework, TopRec, which detects topical communities to construct interpretable domains for domain-specific collaborative filtering. In order to mine communities as well as the corresponding topics, a semi-supervised probabilistic topic model is utilized by integrating user guidance with social network. Experimental results on real-world data from Epinions and Ciao demonstrate the effectiveness of the proposed framework. Xi Zhang 0018, Jian Cheng 0001, Biao Niu, Hanqing Lu |
WWW | 5 |
| 2013 | Structure preserving non-negative matrix factorization for dimensionality reduction
Zechao Li, Jing Liu 0001, Hanqing Lu |
Comput. Vis. Image Underst. | 3 |
| 2013 | Hashing with dual complementary projection learning for fast image retrieval
Jian Cheng 0001, Hanqing Lu |
Neurocomputing | 3 |
| 2013 | Nonlinear matrix factorization with unified embedding for social tag relevance learning
Zechao Li, Jing Liu 0001, Hanqing Lu |
Neurocomputing | 3 |
| 2013 | Correlation consistency constrained probabilistic matrix factorization for social tag refinement
Jing Liu 0001, Yifan Zhang 0001, Zechao Li, Hanqing Lu |
Neurocomputing | 4 |
| 2013 | Hierarchical Remote Sensing Image Analysis via Graph Laplacian EnergyabstractSegmentation and classification are important tasks in remote sensing image analysis. Recent research shows that images can be described in hierarchical structure or regions. Such hierarchies can produce the state-of-the-art segmentations and can be used in the classification. However, they often contain more levels and regions than required for an efficient image description, which may cause increased computational complexity. In this letter, we propose a new hierarchical segmentation method that applies graph Laplacian energy as a generic measure for segmentation. It reduces the redundancy in the hierarchy by an order of magnitude with little or no loss of performance. In the classification stage, we apply local self-similarity feature to capture the internal geometric layouts of regions in an image. By incorporating advantages from both semantic hierarchical segmentation and local geometric region description, we have achieved better performance than those from the methods being compared. In the experimental section, we validate the effectiveness of our method by showing results on QuickBird and GeoEye-1 image data sets. Huigang Zhang, Xiao Bai 0001, Huaxin Zheng, Huijie Zhao, Jun Zhou 0001, Jian Cheng 0001, Hanqing Lu |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2013 | Dynamic scene understanding by improved sparse topical coding
Jinqiao Wang, Hanqing Lu, Songde Ma |
Pattern Recognit. | 3 |
| 2013 | MLRank: Multi-correlation Learning to Rank for image annotation
Zechao Li, Jing Liu 0001, Changsheng Xu, Hanqing Lu |
Pattern Recognit. | 4 |
| 2013 | M4L: Maximum margin Multi-instance Multi-cluster Learning for scene modeling
Tianzhu Zhang 0001, Si Liu 0001, Changsheng Xu, Hanqing Lu |
Pattern Recognit. | 4 |
| 2013 | Dual local consistency hashing with discriminative projections selection
Jian Cheng 0001, Hanqing Lu |
Signal Process. | 3 |
| 2013 | Ordinal regularized manifold feature extraction for image ranking
Qingshan Liu 0001, Jing Liu 0001, Hanqing Lu |
Signal Process. | 4 |
| 2013 | Asymmetric propagation based batch mode active learning for image retrieval
Biao Niu, Jian Cheng 0001, Xiao Bai 0001, Hanqing Lu |
Signal Process. | 4 |
| 2013 | Mining Semantic Context Information for Intelligent Video Surveillance of Traffic ScenesabstractAutomated visual surveillance systems are attracting extensive interest due to public security. In this paper, we attempt to mine semantic context information including object-specific context information and scene-specific context information (learned from object-specific context information) to build an intelligent system with robust object detection, tracking, and classification and abnormal event detection. By means of object-specific context information, a cotrained classifier, which takes advantage of the multiview information of objects and reduces the number of labeling training samples, is learned to classify objects into pedestrians or vehicles with high object classification performance. For each kind of object, we learn its corresponding semantic scene-specific context information: motion pattern, width distribution, paths, and entry/exist points. Based on this information, it is efficient to improve object detection and tracking and abnormal event detection. Experimental results demonstrate the effectiveness of our semantic context features for multiple real-world traffic scenes. Tianzhu Zhang 0001, Si Liu 0001, Changsheng Xu, Hanqing Lu |
IEEE Trans. Ind. Informatics | 4 |
| 2013 | Spectral Hashing With Semantically Consistent Graph for Image IndexingabstractThe ability of fast similarity search in a large-scale dataset is of great importance to many multimedia applications. Semantic hashing is a promising way to accelerate similarity search, which designs compact binary codes for a large number of images so that semantically similar images are mapped to close codes. Retrieving similar neighbors is then simply accomplished by retrieving images that have codes within a small Hamming distance of the code of the query. Among various hashing approaches, spectral hashing (SH) has shown promising performance by learning the binary codes with a spectral graph partitioning method. However, the Euclidean distance is usually used to construct the graph Laplacian in SH, which may not reflect the inherent distribution of the data. Therefore, in this paper, we propose a method to directly optimize the graph Laplacian. The learned graph, which can better represent similarity between samples, is then applied to SH for effective binary code learning. Meanwhile, our approach, unlike metric learning, can automatically determine the scale factor during the optimization. Extensive experiments are conducted on publicly available datasets and the comparison results demonstrate the effectiveness of our approach. Meng Wang 0001, Jian Cheng 0001, Changsheng Xu, Hanqing Lu |
IEEE Trans. Multim. | 5 |
| 2013 | Script-to-Movie: A Computational Framework for Story Movie CompositionabstractTraditional movie production has always been a highly professional work that needs team collaboration, advanced devices and techniques, and vast time and money investment. These high threshold requirements not only prevent mass amateur enthusiasts entering this field, but also hinder professionals quickly previewing their conceived story plots. In this paper, we raise a novel application, named script-to-movie (S2M) composition, to automatically produce new movies from existing videos in accordance with user created script. Our motivation is to liberate producers from complex filming and editing operations, thereby people's story idea can be instantly converted into the vivid movie video. To support the novel “What You Dream Is What You See” (WYDIWYS) production mode, we first propose a hierarchical alignment method to automatically construct a video material database with detailed semantic description. Considering diverse story plots in user designed script, the database contains abundant video materials about different characters appearing in various time and places conditions. On this basis, the S2M composition is formulated as a constrained optimization problem, where semantic story plot and syntactic visual content are synthetically considered to identify a group of optimal video segments to narrate the user designed script story. Both quantitative and qualitative experimental results are reported to illustrate the effectiveness of the proposed S2M application. Chao Liang 0001, Changsheng Xu, Jian Cheng 0001, Weiqing Min, Hanqing Lu |
IEEE Trans. Multim. | 5 |
| 2013 | Context-Aware Video Retargeting via Graph ModelabstractVideo retargeting is a crowded but challenging research area. In order to maximally comfort the viewers' watching experience, the most challenging issue is how to retain the spatial shape of important objects while ensure temporal smoothness and coherence. Existing retargeting techniques deal with these spatial-temporal requirements individually, which preserve the spatial geometry and temporal coherence for each region. However, the spatial-temporal property of the video content should be context-relevant, i.e., the regions belonging to the same object are supposed to undergo uniform spatial-temporal transformation. Regardless of the contextual information, the divide-and-rule strategy of existing techniques usually incurs various spatial-temporal artifacts. In order to achieve satisfactory spatial-temporal coherent video retargeting, in this paper, a novel context-aware solution is proposed via graph model. First, we employ a grid-based warping framework to preserve the spatial structure and temporal motion trend at the unit of grid cell. Second, we propose a graph-based motion layer partition algorithm to estimate motions of different regions, which simultaneously provides the evaluation of contextual relationship between grid cells while estimating the motions of regions. Third, complementing the salience-based spatial-temporal information preservation, two novel context constraints are encoded for encouraging the grid cells of the same object to undergo uniform spatial and temporal transformation, respectively. Finally, we formulate the objective function as a quadratic programming problem. Our method achieves a satisfactory spatial-temporal coherence while maximally avoiding the influence of artifacts. In addition, the grid-cell-wise motion estimation could be calculated every few frames, which obviously improves the speed. Experimental results and comparisons with state-of-the-art methods demonstrate the effectiveness and efficiency of our approach. Jinqiao Wang, Min Xu 0001, Hanqing Lu |
IEEE Trans. Multim. | 4 |
| 2013 | Enhancing news organization for convenient retrieval and browsingabstractTo facilitate users to access news quickly and comprehensively, we design a news search and browsing system named GeoVisNews, in which the news elements of “Where”, “Who”, “What” and “When” are enhanced via news geo-localization, image enrichment and joint ranking, respectively. For news geo-localization, an Ordinal Correlation Consistent Matrix Factorization (OCCMF) model is proposed to maintain the relevance rankings of locations to a specific news document and simultaneously capture intra-relations among locations and documents. To visualize news, we develop a novel method to enrich news documents with appropriate web images. Specifically, multiple queries are first generated from news documents for image search, and then the appropriate images are selected from the collected web images by an intelligent fusion approach based on multiple features. Obtaining the geo-localized and image enriched news resources, we further employ a joint ranking strategy to provide relevant, timely and popular news items as the answer of user searching queries. Extensive experiments on a large-scale news dataset collected from the web demonstrate the superior performance of the proposed approaches over related methods. Zechao Li, Jing Liu 0001, Meng Wang 0001, Changsheng Xu, Hanqing Lu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2013 | Exploiting content relevance and social relevance for personalized ad recommendation on internet TVabstractThere have been not many interactions between the two dominant forms of mass communication: television and the Internet, while nowadays the appearance of Internet television makes them more closely. Different with traditional TV in a passive mode of transmission, Internet TV makes it more possible to make personalized service recommendation because of the interactivity between users and the Internet. In this article, we introduce a scheme to provide targeted ad recommendation to Internet TV users by exploiting the content relevance and social relevance. First, we annotate TV videos in terms of visual content analysis and textual analysis by aligning visual and textual information. Second, with user-user, video-video and user-video relationships, we employ Multi-Relationship based Probabilistic Matrix Factorization (MRPMF) to learn representative tags for modeling user preference. And then semantic content relevance (between product/ad and TV video) and social relevance (between product/ad and user interest) are calculated by projecting the corresponding tags into our advertising concept space. Finally, with relevancy scores we make ranking for relevant product/ads to effectively provide users personalized recommendation. The experimental results demonstrate attractiveness and effectiveness of our proposed approach. Bo Wang 0011, Jinqiao Wang, Hanqing Lu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2012 | Unsupervised Feature Selection Using Nonnegative Spectral AnalysisabstractIn this paper, a new unsupervised learning algorithm, namely Nonnegative Discriminative Feature Selection (NDFS), is proposed. To exploit the discriminative information in unsupervised scenarios, we perform spectral clustering to learn the cluster labels of the input samples, during which the feature selection is performed simultaneously. The joint learning of the cluster labels and feature selection matrix enables NDFS to select the most discriminative features. To learn more accurate cluster labels, a nonnegative constraint is explicitly imposed to the class indicators. To reduce the redundant or even noisy features, l2,1-norm minimization constraint is added into the objective function, which guarantees the feature selection matrix sparse in rows. Our algorithm exploits the discriminative information and feature correlation simultaneously to select a better feature subset. A simple yet efficient iterative algorithm is designed to optimize the proposed objective function. Experimental results on different real world datasets demonstrate the encouraging performance of our algorithm over the state-of-the-arts. Zechao Li, Yi Yang 0001, Jing Liu 0001, Xiaofang Zhou 0001, Hanqing Lu |
AAAI | 5 |
| 2012 | Efficient Clothing Retrieval with Semantic-Preserving Visual Phrases
Jianlong Fu, Jinqiao Wang, Zechao Li, Min Xu 0001, Hanqing Lu |
ACCV (2) | 5 |
| 2012 | Co-regularized PLSA for Multi-view Clustering
Jing Liu 0001, Zechao Li, Hanqing Lu |
ACCV (2) | 5 |
| 2012 | Modeling Hidden Topics with Dual Local Consistency for Image Analysis
Jian Cheng 0001, Hanqing Lu |
ACCV (1) | 3 |
| 2012 | Fusing Warping, Cropping, and Scaling for Optimal Image Thumbnail Generation
Jinqiao Wang, Min Xu 0001, Hanqing Lu |
ACCV (4) | 4 |
| 2012 | Hierarchical Object Representations for Visual Recognition via Weakly Supervised Learning
Tianzhu Zhang 0001, Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001, Hanqing Lu |
ACCV (1) | 5 |
| 2012 | Weighted Interaction Force Estimation for Abnormality Detection in Crowd Scenes
Xiaobin Zhu 0001, Jing Liu 0001, Jinqiao Wang, Hanqing Lu |
ACCV (3) | 5 |
| 2012 | Learning ordinal discriminative features for age estimationabstractIn this paper, we present a new method for facial age estimation based on ordinal discriminative feature learning. Considering the temporally ordinal and continuous characteristic of aging process, the proposed method not only aims at preserving the local manifold structure of facial images, but also it wants to keep the ordinal information among aging faces. Moreover, we try to remove redundant information from both the locality information and ordinal information as much as possible by minimizing nonlinear correlation and rank correlation. Finally, we formulate these two issues into a unified optimization problem of feature selection and present an efficient solution. The experiments are conducted on the public available Images of Groups dataset and the FG-NET dataset, and the experimental results demonstrate the power of the proposed method against the state-of-the-art methods. Qingshan Liu 0001, Jing Liu 0001, Hanqing Lu |
CVPR | 4 |
| 2012 | Street-to-shop: Cross-scenario clothing retrieval via parts alignment and auxiliary setabstractIn this paper, we address a practical problem of cross-scenario clothing retrieval - given a daily human photo captured in general environment, e.g., on street, finding similar clothing in online shops, where the photos are captured more professionally and with clean background. There are large discrepancies between daily photo scenario and online shopping scenario. We first propose to alleviate the human pose discrepancy by locating 30 human parts detected by a well trained human detector. Then, founded on part features, we propose a two-step calculation to obtain more reliable one-to-many similarities between the query daily photo and online shopping photos: 1) the within-scenario one-to-many similarities between a query daily photo and the auxiliary set are derived by direct sparse reconstruction; and 2) by a cross-scenario many-to-many similarity transfer matrix inferred offline from an extra auxiliary set and the online shopping set, the reliable cross-scenario one-to-many similarities between the query daily photo and all online shopping photos are obtained. We collect a large online shopping dataset and a daily photo dataset, both of which are thoroughly labeled with 15 clothing attributes via Mechanic Turk. The extensive experimental evaluations on the collected datasets well demonstrate the effectiveness of the proposed framework for cross-scenario clothing retrieval. Si Liu 0001, Guangcan Liu, Changsheng Xu, Hanqing Lu, Shuicheng Yan |
CVPR | 5 |
| 2012 | Object-centered narratives for video surveillanceabstractEffective video presentation and summarization techniques are critical for fast browsing of video content. In this paper, we propose a novel presentation approach to vividly depict the moving process of a specific object in a surveillance video, which aims at effectively summarizing video content by a static image named narrative. Firstly, the object of interest is extracted and segmented from the video to form a spatio-temporal object tube. Then three criteria are proposed to select the most representative objects from this tube. We formulate the object selecting process as an energy minimization problem, in which each energy term measures a corresponding criterion cost. We maximally preserve the changes of appearance and behavior while remove other redundant content as much as possible. Finally, the selected representative objects are stitched to the background image by Poisson editing. Experimental results show the promise of the proposed approach. Jinqiao Wang, Chaoyang Zhao, Hanqing Lu, Songde Ma |
ICIP | 4 |
| 2012 | Beyond local image features: Scene calssification using supervised semantic representationabstractThe use of local features for image representation has been proven very effective for a variety of visual tasks such as object localization and scene classification. However, local image features carry little semantic information which is potentially not enough for high level visual tasks. To solve this problem, in this paper, we propose to use a supervised semantic image representation for scene classification, where an image is represented as a response histogram. This response histogram is a combination of the prediction of pre-trained generic object classifiers and classifiers generated by supervised learning. Besides, the use of sparsity constraints makes the proposed representation more efficient and effective to compute. Performances on the UIUC-Sports dataset, the MIT Indoor scene dataset and the Scene-15 dataset demonstrate the effectiveness of the proposed method. Chunjie Zhang 0001, Jing Liu 0001, Chao Liang 0001, Jinhui Tang 0001, Hanqing Lu |
ICIP | 5 |
| 2012 | Anomaly detection in crowded scene via appearance and dynamics joint modelingabstractIn this paper, we propose a novel solution of anomaly detection in crowd scene by jointly modeling appearance and dynamics of motion. First, a novel high-frequency feature based on optical flow (HFOF) is introduced. It can well capture the dynamic information of optical flow. Besides, we adopt the other two types of features, namely multi-scale histogram of optical(MHOF), and dynamic textures (DT). MHOF reserves the motion direction information, while DT captures appearance variant property. The three types of features can complement each other in modeling crowd motions. Finally, multiple kernel learning (MKL) is adopted to train a classifier for anomaly detection. Experiments are conducted on a publicly available dataset of escaping scenarios from University of Minnesota and a challenging dataset from Internet. The results of comparative experiments show the promising performance against other related work. Xiaobin Zhu 0003, Jing Liu 0001, Jinqiao Wang, Yikai Fang, Hanqing Lu |
ICIP | 5 |
| 2012 | Learning Semantic Motion Patterns for Dynamic Scenes by Improved Sparse Topical CodingabstractWith the proliferation of cameras in public areas, it becomes increasingly desirable to develop fully automated surveillance and monitoring systems. In this paper, we propose a novel unsupervised approach to automatically explore motion patterns occurring in dynamic scenes under an improved sparse topical coding (STC) framework. Given an input video with a fixed camera, we first segment the whole video into a sequence of clips (documents) without overlapping. Optical flow features are extracted from each pair of consecutive frames, and quantized into discrete visual words. Then the video is represented by a word-document hierarchical topic model through a generative process. Finally, an improved sparse topical coding approach is proposed for model learning. The semantic motion patterns (latent topics) are learned automatically and each video clip is represented as a weighted summation of these patterns with only a few nonzero coefficients. The proposed approach is purely data-driven and scene independent (not an object-class specific), which make it suitable for very large range of scenarios. Experiments demonstrate that our approach outperforms the state-of-the art technologies in dynamic scene analysis. Jinqiao Wang, Zechao Li, Hanqing Lu, Songde Ma |
ICME | 4 |
| 2012 | Noisy Tag Alignment with Image RegionsabstractWith the permeation of Web 2.0, large-scale user contributed images with tags are easily available on social websites. How to align these social tags with image regions is a challenging task while no additional human intervention is considered, but a valuable one since the alignment can provide more detailed image semantic information and improve the accuracy of image retrieval. To this end, we propose a large margin discriminative model for automatically locating unaligned and possibly noisy image-level tags to the corresponding regions, and the model is optimized using concave-convex procedure (CCCP). In the model, each image is considered as a bag of segmented regions, associated with a set of candidate labeling vectors. Each labeling vector encodes a possible label arrangement for the regions of an image. To make the size of admissible labels tractable, we adopt an effective strategy based on the consistency between visual similarity and semantic correlation to generate a more compact set of labeling vectors. Extensive experiments on MSRC and SAIAPR TC-12 databases have been conducted to demonstrate the encouraging performance of our method comparing with other baseline methods. Yang Liu 0021, Jing Liu 0001, Zechao Li, Hanqing Lu |
ICME | 4 |
| 2012 | Improving Relevance Feedback for Image Retrieval with Asymmetric SamplingabstractRelevance feedback is a quite effective approach to improve performance for image retrieval. Recently, active learning method has attracted much attention due to its capability of alleviating the burden of labeling in relevance feedback. However, most of the traditional studies focus on single sample selection in each feedback which needs heavy computational cost in practice. In this paper, we presents a novel batch mode active learning method for informative sample selection. Inspired by graph propagation, we consider the certainty of labels as asymmetric propagation information on graph, and formulate the correlation between labeled samples and unlabeled samples in an united scheme. Extensive experiments on publicly available data sets show that the proposed method is promising. Biao Niu, Jian Cheng 0001, Hanqing Lu |
ICME | 3 |
| 2012 | Collaborative PLSA for multi-view clustering
Jing Liu 0001, Zechao Li, Hanqing Lu |
ICPR | 4 |
| 2012 | Learning distance metric regression for facial age estimation
Qingshan Liu 0001, Jing Liu 0001, Hanqing Lu |
ICPR | 4 |
| 2012 | Key observation selection for effective video synopsis
Xiaobin Zhu 0003, Jing Liu 0001, Jinqiao Wang, Hanqing Lu |
ICPR | 4 |
| 2012 | Ordinal preserving projection: a novel dimensionality reduction method for image rankingabstractLearning to rank has been demonstrated as a powerful tool for image ranking, but the issue of the "curse of dimensionality" is a key challenge of learning a ranking model from a large image database. This paper proposes a novel dimensionality reduction algorithm named ordinal preserving projection (OPP) for learning to rank. We first define two matrices, which work in the row direction and column direction respectively. The two matrices aim at leveraging the global structure of the data set and ordinal information of the observations. By maximizing the corresponding objective functions, we can obtain two optimal projection matrices mapping original data points into low-dimensional subspace, in which both global structure and ordinal information can be preserved. The experiments are conducted on the public available MSRA-MM image data set and "Web Queries" image data set, and the experimental results demonstrate the effectiveness of the proposed method. Jing Liu 0001, Yan Liu 0004, Changsheng Xu, Qingshan Liu 0001, Hanqing Lu |
ICMR | 6 |
| 2012 | Low rank metric learning for social image retrievalabstractWith the popularity of social media applications, large amounts of social images associated with rich context are available, which is helpful for many applications. In this paper, we propose a Low Rank distance Metric Learning (LRML) algorithm by discovering knowledge from these rich contextual data, to boost the performance of CBIR. Different from traditional approaches that often use the must-links and cannot-links between images, the proposed method exploits information from the visual and textual domains. We assume that the visual similarity estimated by the learned metric is expected to be consistent with the semantic similarity in the textual domain. Since tags are usually noisy, misspelling or meaningless, we also leverage the preservation of visual structure to prevent overfitting those noisy tags. On the other hand, the metric is straightforward constrained to be low rank. We formulate it as a convex optimization problem with nuclear norm minimization and propose an effective optimization algorithm based on proximal gradient method. With the learned metric for image retrieval, some experimental evaluations on a real-world dataset demonstrate the outperformance of our approach over other related work. Zechao Li, Jing Liu 0001, Jinhui Tang 0001, Hanqing Lu |
ACM Multimedia | 5 |
| 2012 | Hi, magic closet, tell me what to wear!abstractIn this paper, we aim at a practical system, magic closet, for automatic occasion-oriented clothing recommendation. Given a user-input occasion, e.g., wedding, shopping or dating, magic closet intelligently suggests the most suitable clothing from the user's own clothing photo album, or automatically pairs the user-specified reference clothing (upper-body or lower-body) with the most suitable one from online shops. Si Liu 0001, Jiashi Feng, Tianzhu Zhang 0001, Hanqing Lu, Changsheng Xu, Shuicheng Yan |
ACM Multimedia | 5 |
| 2012 | Social tag alignment with image regions by sparse reconstructionsabstractHow to align social tags with image regions without additional human intervention is a challenging but a valuable task since it can provide more detailed image semantic information and improve the accuracy of image retrieval. To this end, we propose a novel tag-to-region method with two phases of sparse reconstructions by exploring the large-scale user contributed resources. Given an image with social tags, we first explore the tagging information of large-scale social images to sparsely reconstruct the label vector of the given image, and then use the reconstructing weights as the semantic relevance to the image. With the top $T$ semantically relevant images, we further employ a group sparse coding algorithm to reconstruct each region of the given image, in which the regions from the social images with a common label are deemed as a label group. The group sparsity works on the assumption that one image region corresponds to tags as few as possible. Finally, the region-level tags can be predicted based on the reconstruction error in the corresponding label groups. Extensive experiments on MSRC and SAIAPR TC-12 datasets demonstrate the encouraging performance of our method in comparison with other baselines. Yang Liu 0021, Jing Liu 0001, Zechao Li, Biao Niu, Hanqing Lu |
ACM Multimedia | 5 |
| 2012 | Street-to-shop: cross-scenario clothing retrieval via parts alignment and auxiliary setabstractWe address a cross-scenario clothing retrieval problem- given a daily human photo captured in general environment, e.g., on street, finding similar clothing in online shops, where the photos are captured more professionally and with clean background. There are large discrepancies between daily photo scenario and online shopping scenario. We first propose to alleviate the human pose discrepancy by locating 30 human parts detected by a well trained human detector. Then, founded on part features, we propose a two-step calculation to obtain more reliable one-to-many similarities between the query daily photo and online shopping photos: 1) the within-scenario one-to-many similarities between a query daily photo and an extra auxiliary set are derived by direct sparse reconstruction; 2) by a cross-scenario many-to-many similarity transfer matrix inferred offline from the auxiliary set and the online shopping set, the reliable cross-scenario one-to-many similarities between the query daily photo and all online shopping photos are obtained. Si Liu 0001, Meng Wang 0001, Changsheng Xu, Hanqing Lu, Shuicheng Yan |
ACM Multimedia | 5 |
| 2012 | Real-time multiple object instances detectionabstractIn this paper, we present a novel, real-time multiple object instance detection system via template matching and pairwise classification. Instance detection aims to find and locate exactly the same object instances as specified. Our system is composed of two heterogeneous stages. The first stage adopts instance-specific detection to generate candidates. And the second stage makes use of a pairwise-based classifier across instance categories to test and verify these candidates with respect to templates. Experiments show the superiority of our approach. Chengli Xie, Jinqiao Wang, Yifan Zhang 0001, Hanqing Lu |
ACM Multimedia | 4 |
| 2012 | Real-Time Probabilistic Covariance Tracking With Efficient Model UpdateabstractThe recently proposed covariance region descriptor has been proven robust and versatile for a modest computational cost. The covariance matrix enables efficient fusion of different types of features, where the spatial and statistical properties, as well as their correlation, are characterized. The similarity between two covariance descriptors is measured on Riemannian manifolds. Based on the same metric but with a probabilistic framework, we propose a novel tracking approach on Riemannian manifolds with a novel incremental covariance tensor learning (ICTL). To address the appearance variations, ICTL incrementally learns a low-dimensional covariance tensor representation and efficiently adapts online to appearance changes of the target with only O(1) computational complexity, resulting in a real-time performance. The covariance-based representation and the ICTL are then combined with the particle filter framework to allow better handling of background clutter, as well as the temporary occlusions. We test the proposed probabilistic ICTL tracker on numerous benchmark sequences involving different types of challenges including occlusions and variations in illumination, scale, and pose. The proposed approach demonstrates excellent real-time performance, both qualitatively and quantitatively, in comparison with several previously proposed trackers. Yi Wu 0001, Jian Cheng 0001, Jinqiao Wang, Hanqing Lu, Haibin Ling, Erik Blasch, Li Bai 0002 |
IEEE Trans. Image Process. | 4 |
| 2012 | Weakly Supervised Graph Propagation Towards Collective Image ParsingabstractIn this work, we propose a weakly supervised graph propagation method to automatically assign the annotated labels at image level to those contextually derived semantic regions. The graph is constructed with the over-segmented patches of the image collection as nodes. Image-level labels are imposed on the graph as weak supervision information over subgraphs, each of which corresponds to all patches of one image, and the contextual information across different images at patch level are then mined to assist the process of label propagation from images to their descendent regions. The ultimate optimization problem is efficiently solved by Convex Concave Programming (CCCP). Extensive experiments on four benchmark datasets clearly demonstrate the effectiveness of our proposed method for the task of collective image parsing. Two extensions including image annotation and concept map based image retrieval demonstrate the proposed image parsing algorithm can effectively aid other vision tasks. Si Liu 0001, Shuicheng Yan, Tianzhu Zhang 0001, Changsheng Xu, Jing Liu 0001, Hanqing Lu |
IEEE Trans. Multim. | 6 |
| 2012 | A Generic Framework for Video Annotation via Semi-Supervised LearningabstractLearning-based video annotation is essential for video analysis and understanding, and many various approaches have been proposed to avoid the intensive labor costs of purely manual annotation. However, there lacks a generic framework due to several difficulties, such as dependence of domain knowledge, insufficiency of training data, no precise localization and in efficacy for large-scale video dataset. In this paper, we propose a novel approach based on semi-supervised learning by means of information from the Internet for interesting event annotation in videos. Concretely, a Fast Graph-based Semi-Supervised Multiple Instance Learning (FGSSMIL) algorithm, which aims to simultaneously tackle these difficulties in a generic framework for various video domains (e.g., sports, news, and movies), is proposed to jointly explore small-scale expert labeled videos and large-scale unlabeled videos to train the models. The expert labeled videos are obtained from the analysis and alignment of well-structured video related text (e.g., movie scripts, web-casting text, close caption). The unlabeled data are obtained by querying related events from the video search engine (e.g., YouTube, Google) in order to give more distributive information for event modeling. Two critical issues of FGSSMIL are: (1) how to calculate the weight assignment for a graph construction, where the weight of an edge specifies the similarity between two data points. To tackle this problem, we propose a novel Multiple Instance Learning Induced Similarity (MILIS) measure by learning instance sensitive classifiers; (2) how to solve the algorithm efficiently for large-scale dataset through an optimization approach. To address this issue, Concave-Convex Procedure (CCCP) and nonnegative multiplicative updating rule are adopted. We perform the extensive experiments in three popular video domains: movies, sports, and news. The results compared with the state-of-the-arts are promising and demonstrate the effectiveness and efficiency of our proposed approach. Tianzhu Zhang 0001, Changsheng Xu, Guangyu Zhu 0002, Si Liu 0001, Hanqing Lu |
IEEE Trans. Multim. | 5 |
| 2011 | Size Adaptive Selection of Most Informative FeaturesabstractIn this paper, we propose a novel method to select the most informativesubset of features, which has little redundancy andvery strong discriminating power. Our proposed approach automaticallydetermines the optimal number of features and selectsthe best subset accordingly by maximizing the averagepairwise informativeness, thus has obvious advantage overtraditional filter methods. By relaxing the essential combinatorialoptimization problem into the standard quadratic programmingproblem, the most informative feature subset canbe obtained efficiently, and a strategy to dynamically computethe redundancy between feature pairs further greatly acceleratesour method through avoiding unnecessary computationsof mutual information. As shown by the extensive experiments,the proposed method can successfully select the mostinformative subset of features, and the obtained classificationresults significantly outperform the state-of-the-art results onmost test datasets. Si Liu 0001, Hairong Liu, Longin Jan Latecki, Shuicheng Yan, Changsheng Xu, Hanqing Lu |
AAAI | 6 |
| 2011 | TVParser: An automatic TV video parsing methodabstractIn this paper, we propose an automatic approach to simultaneously name faces and discover scenes in TV shows. We follow the multi-modal idea of utilizing script to assist video content understanding, but without using timestamp (provided by script-subtitles alignment) as the connection. Instead, the temporal relation between faces in the video and names in the script is investigated in our approach, and an global optimal video-script alignment is inferred according to the character correspondence. The contribution of this paper is two-fold: (1) we propose a generative model, named TVParser, to depict the temporal character correspondence between video and script, from which face-name relationship can be automatically learned as a model parameter, and meanwhile, video scene structure can be effectively inferred as a hidden state sequence; (2) we find fast algorithms to accelerate both model parameter learning and state inference, resulting in an efficient and global optimal alignment. We conduct extensive comparative experiments on popular TV series and report comparable and even superior performance over existing methods. Chao Liang 0001, Changsheng Xu, Jian Cheng 0001, Hanqing Lu |
CVPR | 4 |
| 2011 | Image classification by non-negative sparse coding, low-rank and sparse decompositionabstractWe propose an image classification framework by leveraging the non-negative sparse coding, low-rank and sparse matrix decomposition techniques (LR-Sc+SPM). First, we propose a new non-negative sparse coding along with max pooling and spatial pyramid matching method (Sc+SPM) to extract local features' information in order to represent images, where non-negative sparse coding is used to encode local features. Max pooling along with spatial pyramid matching (SPM) is then utilized to get the feature vectors to represent images. Second, motivated by the observation that images of the same class often contain correlated (or common) items and specific (or noisy) items, we propose to leverage the low-rank and sparse matrix recovery technique to decompose the feature vectors of images per class into a low-rank matrix and a sparse error matrix. To incorporate the common and specific attributes into the image representation, we still adopt the idea of sparse coding to recode the Sc+SPM representation of each image. In particular, we collect the columns of the both matrixes as the bases and use the coding parameters as the updated image representation by learning them through the locality-constrained linear coding (LLC). Finally, linear SVM classifier is leveraged for the final classification. Experimental results show that the proposed method achieves or outperforms the state-of-the-art results on several benchmarks. Chunjie Zhang 0001, Jing Liu 0001, Qi Tian 0001, Changsheng Xu, Hanqing Lu, Songde Ma |
CVPR | 5 |
| 2011 | Video Reshuffling with Narratives toward Effective Video BrowsingabstractWith the rapid increasing of video cameras, large amount of video data everyday brings the problem of video storage and browsing. In this paper, we propose a novel approach to video reshuffling with a group of static images to effectively summarize the video content. Each static image called narrative is generated to depict the behavior of a specific object or a special event. Firstly background subtraction and object tracking are employed to extract the segmentations of moving objects and corresponding trajectories. After that, we apply three sampling rules to optimized select representative object samples from the spatial-temporal object tube and stitch them to the background image by Poisson blending. Experimental results show the promise of the proposed approach. Jinqiao Wang, Xiaobin Zhu 0003, Hanqing Lu, Songde Ma |
ICIG | 4 |
| 2011 | Global Trajectory Construction across Multi-cameras via Graph MatchingabstractBehavior analysis across multi-cameras becomes more and more popular with the rapid development of camera network in video surveillance. In this paper, we propose a novel unsupervised graph matching framework to associate trajectories across partially overlapping cameras. Firstly, trajectory extraction is based on object extraction and tracking and is followed by a homographic projection to a mosaic-plane. And we extract appearance and spatio-temporal features for trajectory description. Then a robust graph matching algorithm based on reweighted random walk is adopted for trajectory association. The association is formulated as node ranking and selection on an association graph whose nodes represent candidate correspondences of trajectories. Finally, the pairs of corresponding trajectories in overlapping regions are fused by an adaptive averaging scheme, in which trajectories with more observations and longer length is given higher weight. Experiments and comparison on real scenarios demonstrate the effectiveness of the proposed approach. Xiaobin Zhu 0003, Jing Liu 0001, Jinqiao Wang, Hanqing Lu, Yikai Fang |
ICIG | 5 |
| 2011 | Representative sampling with certainty propagation for image retrievalabstractSelective sampling has been widely used in relevance feedback of image retrieval to alleviate the burden of labeling by selecting the most informative instances for user to label. Traditional sample selection scheme often selects a batch of instances each time and label them simultaneously, which ignores the correlation among instances and results in redundant labeling. In this paper, we propose an improved representative sampling method with certainty propagation to improve the performance of sampling. In our method, two kinds of correlations among instances are explored to reduce the redundancy in sampling. One is the correlation between labeled instances and unlabeled instances. The other is the correlation among unlabeled instances. Extensive experiments show that the proposed method achieve encouraging results. Jian Cheng 0001, Biao Niu, Yikai Fang, Hanqing Lu |
ICIP | 4 |
| 2011 | One step beyond bags of features: Visual categorization using componentsabstractThe bag-of-visual-words (BoW) representation has received wide application and public acceptance for visual categorization. However, the histogram based image representation ignores the spatial information and correlations among visual words. To tackle these problems, in this paper, we propose to use some image regions called `components', as the higher-level visual elements to represent an image associating with the lower-level elements of `visual words'. Then we formulate the task of visual categorization into two progressive relationships among a given concept and the two-level visual elements of images, i.e., visual-words-to-components and components-to-concept. Firstly, component level linear SVM classifiers are learned to model the relationship between visual words and components, then the output of these SVM classifiers are linearly combined to model the relationships between components and concept. Experiments on the Scene-15 dataset and the Oxford Flowers dataset demonstrate the effectiveness of the proposed method. Jing Liu 0001, Chunjie Zhang 0001, Qi Tian 0001, Changsheng Xu, Hanqing Lu, Songde Ma |
ICIP | 5 |
| 2011 | Using context saliency for movie shot classificationabstractMovie shot classification is vital but challenging task due to various movie genres, different movie shooting techniques and much more shot types than other video domain. Variety of shot types are used in movies in order to attract audiences attention and enhance their watching experience. In this pa per, we introduce context saliency to measure visual attention distributed in keyframes for movie shot classification. Different from traditional saliency maps, context saliency map is generated by removing redundancy from contrast saliency and incorporating geometry constrains. Context saliency is later combined with color and texture features to generate feature vectors. Support Vector Machine (SVM) is used to classify keyframes into pre-defined shot classes. Different from the existing works of either performing in a certain movie genre or classifying movie shot into limited directing semantic classes, the proposed method has three unique features: 1) context saliency significantly improves movie shot classification; 2) our method works for all movie genres; 3) our method deals with the most common types of video shots in movies. The experimental results indicate that the proposed method is effective and efficient for movie shot classification. Min Xu 0001, Jinqiao Wang, Muhammad Abul Hasan, Xiangjian He, Changsheng Xu, Hanqing Lu, Jesse S. Jin |
ICIP | 6 |
| 2011 | Learning to detect salient region of image under weak supervisionabstractSalient region of an image usually contains the crucial information for image analysis and understanding. Most conventional approaches learn the saliency by utilizing the low-level features, which ignore the participation of human. In this paper, we propose an effective and robust approach to detect the salient region of an image by combining the bottom-up and top-down cues. The proposed method not only consider the low-level attention features, but also take human into the loop for better understanding of human attention. Furthermore, we build an asymmetrical graph model to integrate these bottom-up and top-down cues into an energy function of saliency. A compact but exact saliency region can be obtained by minimizing posterior energy function. The compact constraint and global minimization manner of the asymmetrical graph cuts guarantee the good performance of saliency extraction. Extensive experiments demonstrate the proposed method is promising. Jian Cheng 0001, Hanqing Lu |
ICME | 3 |
| 2011 | News contextualization with geographic and visual informationabstractIn this paper, we investigate the contextualization of news documents with geographic and visual information. We propose a matrix factorization approach to analyze the location relevance for each news document. We also propose a method to enrich the document with a set of web images. For location relevance analysis, we first perform toponym extraction and expansion to obtain a toponym list from news documents. We then propose a matrix factorization method to estimate the location-document relevance scores while simultaneously capturing the correlation of locations and documents. For image enrichment, we propose a method to generate multiple queries from each news document for image search and then employ an intelligent fusion approach to collect a set of images from the search results. Based on the location relevance analysis and image enrichment, we introduce a news browsing system named NewsMap which can support users in reading news via browsing a map and retrieving news with location queries. The news documents with the corresponding enriched images are presented to help users quickly get information. Extensive experiments demonstrate the effectiveness of our approaches. Zechao Li, Meng Wang 0001, Jing Liu 0001, Changsheng Xu, Hanqing Lu |
ACM Multimedia | 5 |
| 2011 | Snap & play: auto-generate personalized find-the-difference mobile gameabstractAccording to the year 2010 report of the Entertainment Software Association [5], 42% of USA heads of households reported playing games on mobile devices, rising quickly from the 20% in 2002 and bringing huge market for mobile games. In this paper, by taking the popular game, Find-the-Difference (FiDi), as a concrete example, we explore new mobile game design principles and techniques for enhancing player's gaming experience in personalized, automatic, and dynamic aspects. Unlike traditional FiDi game, where image pairs (source image vs. target image) with M different patches are manually produced by game developer and players may feel boring or cheat after practicing all image pairs, our proposed Personalized FiDi (P-FiDi) mobile game may be played under a new Snap & Play mode. The player may first take photos with one s mobile device (or select from one's own albums). Then, these photos serve as source images, and the P-FiDi system automatically generates the counterpart target images by sequential operations of aesthetic image quality enhancement, image patch and differentiating style joint selection, music adaptation, dynamic difficulty level determination, and ultimate automatic image editing with a rich set of popular differentiating styles used in traditional FiDi game. Finally, the player enjoys the unique gaming with one's own (instant) photos and music, and the freedom to have new gaming image pairs any time. The user studies show that the P-FiDi mobile game is satisfying in terms of player experience. Si Liu 0001, Qiang Chen 0007, Jian Dong 0011, Shuicheng Yan, Changsheng Xu, Hanqing Lu |
ACM Multimedia | 6 |
| 2011 | Correlated PLSA for Image Clustering
Jian Cheng 0001, Zechao Li, Hanqing Lu |
MMM (1) | 4 |
| 2011 | Adaptive Model for Robust Pedestrian Counting
Jinqiao Wang, Hanqing Lu |
MMM (1) | 3 |
| 2011 | Boosting part-sense multi-feature learners toward effective object detection
Shi Chen 0008, Jinqiao Wang, Bo Wang 0011, Changsheng Xu, Hanqing Lu |
Comput. Vis. Image Underst. | 6 |
| 2011 | Boosted multi-class semi-supervised learning for human action recognition
Tianzhu Zhang 0001, Si Liu 0001, Changsheng Xu, Hanqing Lu |
Pattern Recognit. | 4 |
| 2011 | Boosted Exemplar Learning for Action Recognition and AnnotationabstractHuman action recognition and annotation is an active research topic in computer vision. How to model various actions, varying with time resolution, visual appearance, and others, is a challenging task. In this paper, we propose a boosted exemplar learning (BEL) approach to model various actions in a weakly supervised manner, i.e., only action bag-level labels are provided but action instance level ones are not. The proposed BEL method can be summarized as three steps. First, for each action category, amount of class-specific candidate exemplars are learned through an optimization formulation considering their discrimination and co-occurrence. Second, each action bag is described as a set of similarities between its instances and candidate exemplars. Instead of simply using a heuristic distance measure, the similarities are decided by the exemplar-based classifiers through the multiple instance learning, in which a positive (or negative) video or image set is deemed as a positive (or negative) action bag and those frames similar to the given exemplar in Euclidean Space as action instances. Third, we formulate the selection of the most discriminative exemplars into a boosted feature selection framework and simultaneously obtain an action bag-based detector. Experimental results on two publicly available datasets: the KTH dataset and Weizmann dataset, demonstrate the validity and effectiveness of the proposed approach for action recognition. We also apply BEL to learn representations of actions by using images collected from the Web and use this knowledge to automatically annotate action in YouTube videos. Results are very impressive, which proves that the proposed algorithm is also practical in unconstraint environments. Tianzhu Zhang 0001, Jing Liu 0001, Si Liu 0001, Changsheng Xu, Hanqing Lu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2011 | Exploiting Visual-Audio-Textual Characteristics for Automatic TV Commercial Block Detection and SegmentationabstractAutomatic TV commercial block detection (CBD) and commercial block segmentation (CBS) are two key components of a smart commercial digesting system. In this paper, we focus our research on CBD and CBS by the means of collaborative exploitation of visual-audio-textual characteristics embedded in commercials. Rather than utilizing exclusively visual-audio characteristics like most previous works, an abundance of textual characteristics associated with commercials are fully exploited. Additionally, Tri-AdaBoost, an interactive ensemble learning manner, is proposed to form a consolidated semantic fusion across visual, audio, and textual characteristics. In order to segment a detected commercial block into multiple individual commercials, additional informative descriptors including textual characteristics are introduced to boost the robustness in the detection of frame marked with product information (FMPI). Together with the characteristics of audio spectral variation pointer and silent position, FMPI can provide a kind of complementary representation architecture to model the similarity of intra-commercial and the dissimilarity of inter-commercial. Experiments are conducted on a large video dataset from both China central television (CCTV) channels and TRECVID'05, and promising experimental results show the effectiveness of the proposed scheme. Nan Liu 0007, Yao Zhao 0001, Zhenfeng Zhu, Hanqing Lu |
IEEE Trans. Multim. | 4 |
| 2010 | Image Classification Using Spatial Pyramid Coding and Visual Word Reweighting
Chunjie Zhang 0001, Jing Liu 0001, Jinqiao Wang, Qi Tian 0001, Changsheng Xu, Hanqing Lu, Songde Ma |
ACCV (3) | 6 |
| 2010 | Multi-modal multi-correlation person-centric news retrievalabstractIn this paper, we propose a framework of multi-modal multi-correlation person-centric news retrieval, which integrates news event correlations, news entity correlations, and event-entity correlations simultaneously by exploring both text and image information. The proposed framework is confined to a person-name query and enables a more vivid and informative person-centric news retrieval by providing two views of result presentation, namely a query-oriented multi-correlation map and a ranking list of news items with necessary descriptions including news image, news title and summary, central entities and relevant news events. First, we pre-process news articles using natural language techniques, and initialize the three correlations by statistical analysis about events and entities in news articles and face images. Second, a Multi-correlation Probabilistic Matrix Factorization (MPMF) algorithm is proposed to complete and refine the three correlations. Different from traditional Probabilistic Matrix Factorization (PMF), the proposed MPFM additionally considers the event correlations and the entity correlations as well as the event-entity correlations during the factor analysis. Third, the result ranking and visualization are conducted to present search results relevant to a target news topic. Experimental results on a news dataset collected from multiple news websites demonstrate the attractive performance of the proposed solution for news retrieval. Zechao Li, Jing Liu 0001, Xiaobin Zhu 0003, Hanqing Lu |
CIKM | 4 |
| 2010 | A co-Gaussian Process based framework for remote sensing image change detectionabstractInspired by the idea of co-training algorithm, in this paper we propose a novel semi-supervised learning algorithm, co-Gaussian Process (co-GP), under a Bayesian framework. Image data are characterized in two distinct views, i.e. two disjoint feature sets. A latent function with a GP prior is employed for each view. In learning process of co-GP, knowledge acquired in each view is transferred by probabilistic labels to the other in turns to enhance learning effect. In this manner, proper parameters are estimated in a bootstrap mode and a satisfying performance can be maintained with only small amount of labeled data. The experiments carried out on multitemporal images validate the proposed algorithm. Zhenglong Li 0001, Jian Cheng 0001, Zhixin Zhou, Hanqing Lu |
ICASSP | 5 |
| 2010 | Multi-level trajectory modeling for video copy detectionabstractThe main issue of video copy detection is to estimate a constant spatial-temporal transformation in object level between the original video and the copies. In this paper, we propose a multi-level trajectory modeling approach for video copy detection. It includes a rich trajectory description and a robust trajectory-to-trajectory matching to preserve and explore the trajectory characteristics in both spatial-temporal space and feature space. In summary, we will describe the trajectories in three levels: feature-level descriptor, spatial-temporal coordinates and high-level dynamic behaviors. After extracting the trajectories of videos, we apply a two-stage trajectory-to-trajectory based parametric matching technique to achieve an optimal spatial-temporal transformation between query video and the database videos. To speed up the detection process, we use Locality Sensitive Hashing (LSH) to index and query trajectories with the dynamic behavior and features. Extensive experiments on 100 hours of videos from the TRECVID 2008 demonstrate the effectiveness of our approach. Shi Chen 0008, Jinqiao Wang, Bo Wang 0011, Qi Tian 0001, Hanqing Lu |
ICASSP | 6 |
| 2010 | A improved silhouette tracking approach integrating particle filter with graph cutsabstractIn this paper, we propose a novel approach that combines particle filter tracking and 3D graph cut based segmentation to achieve silhouette tracking against drastic scale change and occlusion. The segmentation module offers particle filter tracking procedure the target shape information to compensate spatial information loss in the histogram based particle filter tracking process. Meanwhile, particle filter predicts a location for the shape prior in the segmentation module to overcome the global nature of graph cut algorithm, that is outlying regions similar with object are prone to be captured. The above two parts are linked with a guide mask that is initialized at the beginning of tracking and updated in a mask update mechanism. We demonstrate the effectiveness of our approach with experiments in several challenging image sequences. Jing Liu 0001, Jinqiao Wang, Jian Cheng 0001, Hanqing Lu |
ICASSP | 5 |
| 2010 | Sparse constraint nearest neighbour selection in cross-media retrievalabstractWith the rapid increasing multimedia documents including videos, images or text, the cross-media retrieval is being focused on. Currently, most of state-of-art methods belonging to the retrieval methods are developed within the scope of the transductive learning. And as soon as the query samples are outside the database, k-nearest-neighbor method is always adopted. However, under such circumstances the fixed global parameter k is not robust for all queries with diverse semantics. In this paper, we propose an alternative method based on the sparse representation. The query sample is considered as a sparse linear combination of all training samples, and the number of nearest neighbors is determined automatically according to the sparse coefficients to the query. Then we import the selection of nearest neighbors into a cross-media ranking model with Local Regression and Global Alignment (LRGA) to get the relevant documents to the query. We conduct extensive experiments for cross-media retrieval to demonstrate the efficiency and effectiveness of our methods. Zechao Li, Jing Liu 0001, Hanqing Lu |
ICIP | 3 |
| 2010 | Multi-modal characteristics analysis and fusion for TV commercial detectionabstractAutomatic TV commercial detection has become an indispensable part of content-based video analysis technique due to the explosive growth in TV commercial volume. In this paper, a multi-modal (i.e. visual, audio and textual modalities) commercial digesting scheme is proposed to alleviate two challenges in commercial detection, which are the generation of mid-level semantic descriptor and the application of effective discrimination method. Compared with the general program, some unique semantic characteristics are purposely embedded in the commercial to grasp more attention from audience. Aiming at exploring the power of these semantic characteristics, a kind of novel commercial-oriented descriptor from textual modality is proposed, besides taking advantage of those commonly used description means in light of audio and visual modalities. To boost the ability of discrimination of commercial from general program in multi-modal representation space, Tri-AdaBoost, a self-learning method by an interactive way across multiple modalities, is introduced to form a final consolidated decision for discrimination. Moreover, a heuristic post processing strategy based on the temporal consistency is taken to further reduce the false alarms. The promising experimental results show the effectiveness of the proposed scheme with respect to large video data collections. Nan Liu 0007, Yao Zhao 0001, Zhenfeng Zhu, Hanqing Lu |
ICME | 4 |
| 2010 | Interactive Web Video Advertising with Context Analysis and SearchabstractOnline media services and electronic commerce are booming recently. Previous studies have been devoted to contextual advertising, but few work deals with interactive web advertising. In this paper, we propose to put users in the loop of collecting contextual ad information with an interaction process, establishing semantic ad links across media platforms. Given an ad video, the key frames with explicit product information are located, which allow users to click favorite key frames for searching ads interactively. A three-stage contextual search is applied to find relevant products or services from web pages, i.e., searching visually similar product images on shopping websites, ranking product tags by text aggregation, and re-search textual items consisting of semantic meaningful tags to make a recommendation. In addition, users can choose automatically suggested keywords to reflect their intentions. Subjective evaluation has demonstrated the effectiveness of the proposed approach to interactive video advertising over the Web. Bo Wang 0011, Jinqiao Wang, Ling-Yu Duan, Qi Tian 0001, Hanqing Lu, Wen Gao 0001 |
ICPR | 5 |
| 2010 | Extracting Key Sub-trajectory Features for Supervised Tactic Detection in Sports VideoabstractTactic analysis is receiving more attention in sports video analysis for its assistance to coaches and players. This paper proposes an efficient key sub-trajectory feature representation of ball trajectory for tactic analysis. Ball trajectories are modeled with the generalized suffix tree where frequent sub-trajectory patterns are searched for. Key sub-trajectory patterns are extracted by further filtering these frequent sub-trajectory patterns. Instead of directly using individual sub-trajectories as features to train tactic detectors, we take key sub-trajectory patterns as a whole. Key sub-trajectory feature representation effectively removes noise, reduces the dimension of features, and improves the performance of supervised learning to detect tactics. Yi Zhang 0007, Changsheng Xu, Hanqing Lu |
ICPR | 3 |
| 2010 | Fast feature selection and training for AdaBoost-based concept detection with large scale datasetsabstractAdaBoost has been proved a successful statistical learning method for concept detection with high performance of discrimination and generalization. However, it is computationally expensive to train a concept detector using boosting, especially on large scale datasets. The bottleneck of training phase is to select the best learner among massive learners. Traditional approaches for selecting a weak classifier usually run in O(NT), with N examples and T learners. In this paper, we treat the best learner selection as a Nearest Neighbor Search problem in the function space instead of feature space. With the help of Locality Sensitive Hashing (LSH) algorithm, the best learner searching procedure can be speeded up in the time of O(NL), where L is the number of buckets in LSH. Compared with the T (~500,000), the L (~600) is much smaller in our experiments. In addition, through studying the distribution of weak learners and candidate query points, we present an efficient method to try to partition the weak learner points and the feasible region of query points uniformly as much as possible, which can achieve significant improvement in both recall and precision compared with the random projection in traditional LSH algorithm. Experimental results reveal our method can significantly reduce the training time. And still the performance of our method is comparable with the state-of-art methods. Shi Chen 0008, Jinqiao Wang, Yang Liu 0021, Changsheng Xu, Hanqing Lu |
ACM Multimedia | 5 |
| 2010 | Effective logo retrieval with adaptive local feature selectionabstractTowards building a practical large-scale logo retrieval system, we propose a novel approach to extract and combine local features for effective logo retrieval. Instead of global feature extraction by modeling the web logo as a whole, we extract the local feature phrases to form a visual codebook and build an inverted file storing the features to accelerate the indexing process. Then we divide logos into several groups according to local feature type based on which feature can model the logo best and naming as "Point-type", "Shape-type" and "Patch-type". We develop a strategy of adaptive feature selection by a weight updating mechanism. To evaluate the performance, we have built a new challenging dataset which consists of 60 international corporations' logos. Experiments and comparisons demonstrate the superior performance to previous retrieval algorithms. Jianlong Fu, Jinqiao Wang, Hanqing Lu |
ACM Multimedia | 3 |
| 2010 | Image annotation using multi-correlation probabilistic matrix factorizationabstractThe image-word correlation estimation is an essential issue in image annotation. In this paper, we propose a multi-correlation probabilistic matrix factorization (MPMF) algorithm for the correlation estimation. Different from the traditional solutions which treat the image-word correlation, image similarity and word relation independently or sequentially, in the proposed MPMF, these three elements are integrated together simultaneously and seamlessly. Specifically, we have derived two low-dimensional sets by conducting a joint factorization upon the word-to-image relation matrix, the image similarity matrix, and the word relation matrix to derive two low-dimensional sets of latent word factors and latent image factors. Finally, the annotation words of each untagged or noisily tagged image can be predicted by reconstructing the image-word correlations with the both derived latent factors. Experimental results on the Corel dataset and a Flickr image dataset show the superior performance of our proposed algorithm over the state-of-the-arts. Zechao Li, Jing Liu 0001, Xiaobin Zhu 0003, Tinglin Liu, Hanqing Lu |
ACM Multimedia | 5 |
| 2010 | A generic framework for event detection in various video domainsabstractEvent detection is essential for the extensively studied video analysis and understanding area. Although various approaches have been proposed for event detection, there is a lack of a generic event detection framework that can be applied to various video domains (e.g. sports, news, movies, surveillance). In this paper, we present a generic event detection approach based on semi-supervised learning and Internet vision. Concretely, a Graph-based Semi-Supervised Multiple Instance Learning (GSSMIL) algorithm is proposed to jointly explore small-scale expert labeled videos and large-scale unlabeled videos to train the event models to detect video event boundaries. The expert labeled videos are obtained from the analysis and alignment of well-structured video related text (e.g. movie scripts, web-casting text, close caption). The unlabeled data are obtained by querying related events from the video search engine (e.g. YouTube) in order to give more distributive information for event modeling. A critical issue of GSSMIL in constructing a graph is the weight assignment, where the weight of an edge specifies the similarity between two data points. To tackle this problem, we propose a novel Multiple Instance Learning Induced Similarity (MILIS) measure by learning instance sensitive classifiers. We perform the thorough experiments in three popular video domains: movies, sports and news. The results compared with the state-of-the-arts are promising and demonstrate our proposed approach is performance-effective. Tianzhu Zhang 0001, Changsheng Xu, Guangyu Zhu 0002, Si Liu 0001, Hanqing Lu |
ACM Multimedia | 5 |
| 2010 | AdVR: Linking Ad Video with Products or Service
Shi Chen 0008, Jinqiao Wang, Bo Wang 0011, Ling-Yu Duan, Qi Tian 0001, Hanqing Lu |
MMM | 6 |
| 2010 | Extended CBIR via Learning Semantics of Query Image
Chuanghua Gui, Jing Liu 0001, Changsheng Xu, Hanqing Lu |
MMM | 4 |
| 2010 | Personalized Sports Video Customization for Mobile Devices
Chao Liang 0001, Jian Cheng 0001, Changsheng Xu, Jinqiao Wang, Hanqing Lu, Jian Ma 0001 |
MMM | 8 |
| 2010 | Human Action Recognition in Videos Using Hybrid Motion Features
Si Liu 0001, Jing Liu 0001, Tianzhu Zhang 0001, Hanqing Lu |
MMM | 4 |
| 2010 | Building topographic subspace model with transfer learning for sparse representation
Yang Liu 0021, Jian Cheng 0001, Changsheng Xu, Hanqing Lu |
Neurocomputing | 4 |
| 2010 | Fast Object-Level Change Detection for VHR ImagesabstractA novel approach is presented for change detection of very high resolution images, which is accomplished by fast object-level change feature extraction and progressive change feature classification. Object-level change feature is helpful for improving the discriminability between the changed class and the unchanged class. Progressive change feature classification helps improve the accuracy and the degree of automation, which is implemented by dynamically adjusting the training samples and gradually tuning the separating hyperplane. Experiments demonstrate the effectiveness of the proposed approach. Chunlei Huo, Zhixin Zhou, Hanqing Lu, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2010 | Lennard-Jones force field for geometric active contour
Zhenglong Li 0001, Qingshan Liu 0001, Hanqing Lu, Dimitris N. Metaxas |
Signal Process. | 3 |
| 2009 | Learning semantic scene models by object classification and trajectory clusteringabstractActivity analysis is a basic task in video surveillance and has become an active research area. However, due to the diversity of moving objects category and their motion patterns, developing robust semantic scene models for activity analysis remains a challenging problem in traffic scenarios. This paper proposes a novel framework to learn semantic scene models. In this framework, the detected moving objects are first classified as pedestrians or vehicles via a co-trained classifier which takes advantage of the multiview information of objects. As a result, the framework can automatically learn motion patterns respectively for pedestrians and vehicles. Then, a graph is proposed to learn and cluster the motion patterns. To this end, trajectory is parameterized and the image is cut into multiple blocks which are taken as the nodes in the graph. Based on the parameters of trajectories, the primary motion patterns in each node (block) are extracted via Gaussian mixture model (GMM), and supplied to this graph. The graph cut algorithm is finally employed to group the motion patterns together, and trajectories are clustered to learn semantic scene models. Experimental results and applications to real world scenes show the validity of our proposed method. Tianzhu Zhang 0001, Hanqing Lu, Stan Z. Li |
CVPR | 2 |
| 2009 | Robust Bayesian tracking on Riemannian manifolds via fragments-based representationabstractRecently, the covariance region descriptor [1] has been proved robust and versatile for a modest computational cost. It enables efficient fusion of different types of features. Based on the covariance descriptor and the metric on Riemannian manifolds, we develop a robust Bayesian tracking framework via fragments-based representation in this paper. In this framework, the template object is represented by multiple image fragments or patches. Every patch votes on the possible state of the object in the current frame, by comparing its covariance descriptor with the corresponding image patch model. Tracking is then led by the Bayesian state inference framework in which a particle filter is used for propagating sample distributions over time. The weight of each particle is formulated by combining the votes of the patches using a robust statistic. Further, we extend the fast covariance computation to the Bayesian tracking problem, which makes the tracking procedure more efficient. We present extensive experimental results on challenging sequences, which demonstrate the robust tracking achieved by our algorithm. Yi Wu 0001, Jinqiao Wang, Hanqing Lu |
ICASSP | 3 |
| 2009 | A robust boosting tracker with minimum error bound in a co-training frameworkabstractThe varying object appearance and unlabeled data from new frames are always the challenging problem in object tracking. Recently machine learning methods are widely applied to tracking, and some online and semi-supervised algorithms are developed to handle these difficulties. In this paper, we consider tracking as a classification problem and present a novel tracking method based on boosting in a co-training framework. The proposed tracker can be online updated and boosted with multi-view weak hypothesis. The most important contribution of this paper is that we find a boosting error upper bound in a co-training framework to guide the novel tracker construction. In theory, the proposed tracking method is proved to minimize this error bound. In experiments, the accuracy rate of foreground/ background classification and the tracking results are both served as evaluation metrics. Experimental results show good performance of proposed novel tracker on challenging sequences. Jian Cheng 0001, Hanqing Lu |
ICCV | 3 |
| 2009 | Real-time visual tracking via Incremental Covariance Tensor LearningabstractVisual tracking is a challenging problem, as an object may change its appearance due to pose variations, illumination changes, and occlusions. Many algorithms have been proposed to update the target model using the large volume of available information during tracking, but at the cost of high computational complexity. To address this problem, we present a tracking approach that incrementally learns a low-dimensional covariance tensor representation, efficiently adapting online to appearance changes for each mode of the target with only ̃(1) computational complexity. Moreover, a weighting scheme is adopted to ensure less modeling power is expended fitting older observations. Both of these features contribute measurably to improving overall tracking performance. Tracking is then led by the Bayesian inference framework in which a particle filter is used to propagate sample distributions over time. With the help of integral images, our tracker achieves real-time performance. Extensive experiments demonstrate the effectiveness of the proposed tracking algorithm for the targets undergoing appearance variations. Yi Wu 0001, Jian Cheng 0001, Jinqiao Wang, Hanqing Lu |
ICCV | 4 |
| 2009 | Expanded bag of words representation for object classificationabstractCurrently, the bag of visual words (BOW) representation has received wide applications in object categorization. However, the BOW representation ignores the dependency relationship among visual words, which could provide informative knowledge to understand an image. In this paper, we first design a simple method to discover this dependency through computing the spatial correlation between visual words in overlapped local patches. Obtaining the dependency relationship, we further propose a novel update strategy to modify the BOW representation. The modification is motivated by the idea of Query Expansion applied successfully in text retrieval. We implement our approach on challenging PASCAL 2006 database, and the experimental results show its improved performance against the BOW representation. Tinglin Liu, Jing Liu 0001, Qingshan Liu 0001, Hanqing Lu |
ICIP | 4 |
| 2009 | Category sensitive codebook construction for object category recognitionabstractRecently, the bag of visual words based image representation is getting popular in object category recognition. Since the codebook of the bag-of-words (BOW) based image representation approach is typically constructed by only measuring the visual similarity of local image features (e.g., k-means), the resulting codebooks may not capture the desired information for object category recognition. This paper proposes a novel optimization method for discriminative codebook construction that considers the category information of local image features as an additional term in traditional visual-similarity-only based codebook construction methods. The category sensitive codebook is constructed through solving an optimization problem. Therefore, the category sensitive codebook construction method goes one step beyond visual-similarity-only methods. Besides, the proposed category sensitive codebook construction method can be implemented with k-means clustering very efficiently and effectively. Experimental results on PASCAL VOC Challenge 2006 data set demonstrate the effectiveness of our method. Chunjie Zhang 0001, Jing Liu 0001, Qi Tian 0001, Hanqing Lu, Songde Ma |
ICIP | 5 |
| 2009 | Web image retrieval via learning semantics of query imageabstractThe performance of traditional image retrieval approaches remains unsatisfactory, as they are restricted by the wellknown semantic gap and the diversity of textual semantics. To tackle these problems, we propose an improved image retrieval framework when querying with an image. The framework considers not only the discriminative power of various visual properties but also the semantic representation of the query image. Given a query image, we first perform CBIR to obtain some visually similar image sets corresponding to different visual properties separately. Then, a semantic representation to the query image is learnt from each image set. The semantic consistence among the textual indexes of each image set is measured in order to judge the confidence of various visual properties and the obtained semantic representation in search. Obtaining these items, both visually and semantically relevant images are returned to the user by a combined similarity measure. Experiments on a large-scale Web images demonstrate the effectiveness and potential of the proposed framework. Chuanghua Gui, Jing Liu 0001, Changsheng Xu, Hanqing Lu |
ICME | 4 |
| 2009 | A variational multi-view learning framework and its application to image segmentationabstractThe paper presents a novel multi-view learning framework based on variational inference. We formulate the framework as a graph representation in form of graph factorization: the graph comprises of factor graphs, which are used to describe internal states of views. Each view is modeled with a Gaussian mixture model. The proposed framework has three main advantages (1) less constraint assumed on data, (2) effective utilization of unlabeled data, and (3) automatic data structure inferring: proper data structure can be inferred in only one round. The experiments on image segmentation demonstrate its effectiveness. Zhenglong Li 0001, Qingshan Liu 0001, Hanqing Lu |
ICME | 3 |
| 2009 | Learning local features for object categorizationabstractIn this paper, for every local feature, we propose to learn its similar local features across all positive images, instead of using heuristic distance as similarity measure. Specifically, multiple instance learning (MIL) is employed to simultaneously determine the similar points of a local feature and learn its corresponding discriminative function which can be regarded as some kind of similarity measure. For each local feature, a weak learner is constructed based on such similarity measure. Then AdaBoost selects the most discriminative local features and combines them to form a strong classifier. Experimental results show encouraging performance of our method. Ming Tang 0001, Shi Chen 0008, Jinqiao Wang, Hanqing Lu, Songde Ma |
ICME | 5 |
| 2009 | Context saliency based image summarizationabstractImage summarization is to determine a smaller but faithful representation of the original visual content. In this paper, we propose a context saliency based image summarization approach, incorporating statistical saliency and geometric information as the importance measurement instead of visual saliency. To ensure image summaries to be adaptive to target device under perception constraint, we present a grid- based piecewise linear image warping scaleplate, and adopt the sweet spot evaluation to generate a flexible model combining the cropping and warping methods. Additionally, we explore potential extensions on image retargeting, thumbnail generation, digital matting and photo browsing. Experimental results show comparable performance compared to the-state- of-art on common data sets. Jinqiao Wang, Hanqing Lu, Changsheng Xu |
ICME | 4 |
| 2009 | Linking video ADS with product or service information by web searchabstractWith the proliferation of online media services, video ads are pervasive across various platforms involving Internet services and interactive TV services. Existing research efforts such as Google AdSense and MSRA videosense/imagesense have been devoted to the less intrusive insertion of relevant textual or video ads in streams or Web pages through text/image/video content analysis whereas the inherent semantics of video ads is much less exploited. In this paper, we propose to link video ads with relevant product/service information across e-commerce Web sites or portals towards ad recommendation in a cross-media manner. Firstly, we carry out semantic analysis within ad videos in which frames marked with product images (FMPI) are extracted. Secondly, we link ad videos with relevant ads on the Web by utilizing FMPI to search visually similar product images (e.g. appearance or logo) and to collect their accompanying text (brand name, category, description, or other tags) over popular e-commerce Websites or portals such as EBay, Amazon, Taobao, etc. We search visually similar product images with local sensitive hashing (LSH) in a naive Bayes near neighbor classifier. Finally, we may recommend more relevant products/services for ad videos through ranking those matched product images and categorizing useful tags of top ranked ads from the Web. Preliminary experiments have been carried out to demonstrate the idea of linking ad videos with product/service information from the Web. Jinqiao Wang, Ling-Yu Duan, Bo Wang 0011, Shi Chen 0008, Jing Liu 0001, Hanqing Lu, Wen Gao 0001 |
ICME | 7 |
| 2009 | Human activity recognition based on the blob featuresabstractIn this paper, we present a novel approach for human activities recognition in the video. We analyze human activities in the sequential frames because human activities can be considered as a temporal object which contains a series of frames. Firstly, we establish a statistical background model and extract foreground object through background subtraction in the video stream. Then, we use foreground blobs of the current frame and a series of frames before the current frame to form a new feature image in certain rules. Finally, we combine the non-zero pixels in the feature image into blobs using the connected component method. Then each blob corresponds to an activity which is characterized by the blob appearance. By recognizing blob features we can recognize activities. We use Gaussian Mixture to model features for each type of human activities and employ Mahalanobis distance to measure the similarity. Jie Yang 0002, Jian Cheng 0001, Hanqing Lu |
ICME | 3 |
| 2009 | Human-centered picture slideshow personalization for mobile devicesabstractThis paper presents a human-centered picture slideshow system for mobile users. In contrast to conventional ROIs (region-of-interest) detection based systems, we provide mobile users the freedom of personalizing ROIs in a convenient and effective way. Here, we import a simple human interaction, i.e., only a single click, to give a hint for users' ROIs. First, local saliency map (LSM) is generated, which considers not only multi-scale contrast, but also the self-correlation measure and central effect. Then a local fuzzy growing method is adopted to extract ROIs automatically based on LSM. Extensive experiments and user studies show the encouraging performance of the proposed system. Cunxun Zang, Jian Cheng 0001, Hanqing Lu, Jian Ma 0001 |
ICME | 4 |
| 2009 | Multi-view multi-label active learning for image classificationabstractImage classification is an important topic in multimedia analysis, among which multi-label image classification is a very challenging task with respect to the large demand for human annotation of multi-label samples. In this paper, we propose a multi-view multi-label active learning strategy, which integrates the mechanism of active learning and multi-view learning. On one hand we explore the sample and label uncertainties within each view; on the other hand we capture the uncertainty over different views based on multi-view fusion. Then the overall uncertainty along the sample, label and view dimensions are obtained to detect the most informative sample-label pairs. Experimental results demonstrate the effectiveness of the proposed scheme. Xiaoyu Zhang 0002, Jian Cheng 0001, Changsheng Xu, Hanqing Lu, Songde Ma |
ICME | 4 |
| 2009 | Web image mining using concept sensitive Markov stationary featuresabstractWith the explosive growth of Web resources, how to mine semantically relevant images efficiently becomes a challenging and necessary task. In this paper, we propose a concept sensitive Markov stationary feature (C-MSF) to represent images and also present a classifier based scheme for web image mining. First, through analyzing the results of Google Image Searcher, we collect an image set, which are highly relevant to a concept. Then the image set is explored to learn a C-MSF about the concept by the algorithm of random walk with restart (RWR), in which the spatial co-occurrence of the bag-of-words representation and the concept information are integrated. Obtaining the concept sensitive representation, SVM is applied to mine the web images, while the highly relevant set are considered as positive examples and other random images as negative ones. Finally, experiments on a crawled web dataset demonstrate the improved performance of the proposed scheme. Chunjie Zhang 0001, Jing Liu 0001, Hanqing Lu, Songde Ma |
ICME | 3 |
| 2009 | Naming faces in films using hypergraph matchingabstractIn this paper, we aim to address the problem of naming faces in feature-length films using video and film script. Different from the state-of-the-art methods on naming faces in the videos, most of which used a local matching between a visible face and one of the names extracted from the local video transcript, we use a global matching between names and faces as it is not easy to obtain enough local name cues in the films. In the video, we cluster the faces into groups corresponding to characters and build a face network according to face co-occurrence relationship. Similarly in the film script, a name network is also built according to name cooccurrence relationship. The vertices of the two networks are finally matched by a hypergraph matching method. Experiments are conducted on five feature-length films and give encouraging results. Yifan Zhang 0001, Changsheng Xu, Jian Cheng 0001, Hanqing Lu |
ICME | 4 |
| 2009 | Semi-supervised Change Detection via Gaussian ProcessesabstractThis paper introduces a semi-supervised change detection method that exploits both labeled and unlabeled samples via Gaussian Process (GP). The proposed method is based on recent development in Gaussian Process classifier named NCNM [3]. NCNM is a probabilistic approach to learning a GP classifier in the presence of unlabeled data. It involves a novel transductive learning under a probabilistic framework. Experimental results obtained on two sets of multitemporal remote sensing images confirm the effectiveness of the proposed approach. It also proves that NCNM can compete seriously with the state-of-the-art support vector machines (SVM) classifier for remote sensing image change detection. Chunlei Huo, Zhixin Zhou, Hanqing Lu, Jian Cheng 0001 |
IGARSS (2) | 4 |
| 2009 | A Variational Bayesian Approach to Remote Sensing Image Change DetectionabstractIn this paper, we present a variational Bayesian (VB) approach to multitemporal remote sensing image change detection. The content of the so called `difference image' is modeled by finite Gaussians Mixture Model (GMM), then with the factor analysis techniques, underlying structure of image content is inferred automatically. Compared with the Expectation-Maximization (EM) algorithm, the proposed method can adaptively determine the number of components in the mixture model without usual sub- or over-segmentation problem. Moreover, to overcome the local optimization problem, a component split strategy is employed in inference process. Experimental results confirm the effectiveness of the proposed method. Zhenglong Li 0001, Jian Cheng 0001, Zhixin Zhou, Hanqing Lu |
IGARSS (3) | 5 |
| 2009 | A Variational Co-training Framework for Remote Sensing Image SegmentationabstractInspired by the idea of co-training algorithm, in this paper we propose a novel remote sensing image segmentation approach using co-training strategy under variational Bayesian (VB) framework. Image data are characterized in two distinct views, i.e. two disjoint feature sets. A Gaussian mixture model (GMM) is employed for each view. On one hand, underlying structure of image content is inferred automatically with the factor analysis techniques. On the other hand, parameters are estimated in a bootstrap mode with the co-training strategy. In this manner, a satisfying performance can be achieved. Experimental analyses carried out on several different sets of high resolution optical images validate the proposed algorithm. Zhenglong Li 0001, Jian Cheng 0001, Zhixin Zhou, Hanqing Lu |
IGARSS (4) | 5 |
| 2009 | Consumer video retargeting: context assisted spatial-temporal grid optimizationabstractPervasive multimedia devices require accurate video retargeting, especially in connected consumer electronics platforms. In this paper, we present a context assisted spatialtemporal grid scheme for consumer video retargeting. First, we parse consumer videos from low-level features to highlevel visual concepts, combining visual attention into a more accurate importance description. Then, a semantic importance map is built up representing the spatial importance and temporal continuity, which is incorporated with a 3D rectilinear grid scaleplate to map frames to the target display, thereby keeping the aspect ratio of semantically salient objects as well as the perceptual coherency. Extensive evaluations were done on two popular video genres, sports and advertisements. The comparison with state-of-the-art approaches on both images and videos have demonstrated the advantages of the proposed approach. Jinqiao Wang, Ling-Yu Duan, Hanqing Lu |
ACM Multimedia | 4 |
| 2009 | Sports video retargetingabstractWith the proliferation of diverse multimedia terminals, the request for elegantly retargeting videos to different display devices is evident, especially in sports. This demonstration presents a Sports Video Retargeting(SVR) technique, that utilized domain based structure parsing to build a semantic importance map for video retargeting. The system enables flexible and coherent aspect-ratio change of the output sports videos with a spatial-temporal 3D rectilinear grid framework, which are free from significant loss of information or distortion on salient and important regions. Results in various sports type have shown that SVR is promising for content adaptation on mobile media. Jinqiao Wang, Ling-Yu Duan, Hanqing Lu |
ACM Multimedia | 4 |
| 2009 | Visual Tracking Using Particle Filters with Gaussian Process Regression
Yi Wu 0001, Hanqing Lu |
PSIVT | 3 |
| 2009 | Personalized retrieval of sports video based on multi-modal analysis and user preference acquisition
Yifan Zhang 0001, Changsheng Xu, Xiaoyu Zhang 0002, Hanqing Lu |
Multim. Tools Appl. | 4 |
| 2009 | Image annotation via graph learning
Jing Liu 0001, Mingjing Li, Qingshan Liu 0001, Hanqing Lu, Songde Ma |
Pattern Recognit. | 4 |
| 2009 | Effective Annotation and Search for Video Blogs with Integration of Context and Content AnalysisabstractIn recent years, weblogs (or blogs) have received great popularity worldwide, among which video blogs (or vlogs) are playing an increasingly important role. However, research on vlog analysis is still in the early stage, and how to manage vlogs effectively so that they can be more easily accessible is a challenging problem. In this paper, we propose a novel vlog management model which is comprised of automatic vlog annotation and user-oriented vlog search. For vlog annotation, we extract informative keywords from both the target vlog itself and relevant external resources; besides semantic annotation, we perform sentiment analysis on comments to obtain the overall evaluation. For vlog search, we present saliency-based matching to simulate human perception of similarity, and organize the results by personalized ranking and category-based clustering. An evaluation criterion is also proposed for vlog annotation, which assigns a score to an annotation according to its accuracy and completeness in representing the vlog's semantics. Experimental results demonstrate the effectiveness of the proposed management model for vlogs. Xiaoyu Zhang 0002, Changsheng Xu, Jian Cheng 0001, Hanqing Lu, Songde Ma |
IEEE Trans. Multim. | 4 |
| 2009 | Character Identification in Feature-Length Films Using Global Face-Name MatchingabstractIdentification of characters in films, although very intuitive to humans, still poses a significant challenge to computer methods. In this paper, we investigate the problem of identifying characters in feature-length films using video and film script. Different from the state-of-the-art methods on naming faces in the videos, most of which used the local matching between a visible face and one of the names extracted from the temporally local video transcript, we attempt to do a global matching between names and clustered face tracks under the circumstances that there are not enough local name cues that can be found. The contributions of our work include: 1) A graph matching method is utilized to build face-name association between a face affinity network and a name affinity network which are, respectively, derived from their own domains (video and script). 2) An effective measure of face track distance is presented for face track clustering. 3) As an application, the relationship between characters is mined using social network analysis. The proposed framework is able to create a new experience on character-centered film browsing. Experiments are conducted on ten feature-length films and give encouraging results. Yifan Zhang 0001, Changsheng Xu, Hanqing Lu, Yueh-Min Huang |
IEEE Trans. Multim. | 3 |
| 2008 | Image Segmentation Based on Supernodes and Region Size Estimation
Lihong Ma 0002, Hanqing Lu |
ACIVS | 3 |
| 2008 | Adaptive Spread-Transform Dither Modulation for Color Image WatermarkingabstractBased on the improved Watson's model, an adaptive spread-transform dither modulation (ASTDM) algorithm is proposed for efficient color image watermarking. The luminance masking threshold is modified to be consistent with amplitude without affecting the original parameters. Then, by incorporating the improved model with STDM, we achieve a maximum labeling strength and an appropriate quantization step-size shrinking with the local values of a host image. Finally, the Normalized SI Correction method is implemented to correct the possible color distortion. Experiments demonstrate that with high embedding rate and transparency, ASTDM can greatly increase the robustness against amplitude scaling while retaining anti-re-quantization property, which means the two main drawbacks of QIM algorithm are solved simultaneously. Lihong Ma 0002, Hanqing Lu |
GLOBECOM | 3 |
| 2008 | Boosted Interactively Distributed Particle Filter for automatic multi-object trackingabstractIn this paper, we propose a boosted interactively distributed particle filter (BIDPF) to address the problem of automatic multi-object tracking in the application of player tracking in broadcast soccer video. The interactively distributed particle filter technique (IDPF) is adopted to handle the mutual occlusions among targets. The proposal distribution using a mixture model that incorporates information from the dynamic model and the boosting detection is introduced into the IDPF framework. The boosting proposal distribution quickly detects targets, while the IDPF process keeps the identity of targets during mutual occlusions. Moreover, the foreground observation is extracted by using the color model of the playfield to speed up the boosting detection and reduce false alarms. The foreground is also used to develop a data-driven potential model to improve the IDPF performance. We test the proposed approach on several video sequences and the results demonstrate that our system is able to track a variable number of objects in a dynamic scene and correctly maintain their identities regardless of camera motion and frequent mutual occlusions. Yi Wu 0001, Xiaofeng Tong, Yimin Zhang 0002, Hanqing Lu |
ICIP | 4 |
| 2008 | Adaptive spread-transform dither modulation using an improved luminance-masked thresholdabstractBased on the improved Watson’s model, an adaptive spread-transform dither modulation (ASTDM) algorithm is proposed for efficient watermarking to resist the amplitude scaling and coefficient re-quantization attacks. Firstly, the luminance masking is modified to be consistent with amplitude scaling without affecting the original parameters. Then, by incorporating the improved model with STDM algorithm, a maximum labeling strength and an appropriate quantization step-size shrinking with the local values of a host image are achieved. Experiments demonstrate that with high embedding rate and transparence, ASTDM can greatly increase the robustness against amplitude scaling while retaining anti-re-quantization property, which means the two drawbacks of QIM algorithm are solved simultaneously. Lihong Ma 0002, Guoxi Wang, Hanqing Lu |
ICIP | 4 |
| 2008 | Synchronization analysis for synchronized diving videosabstractThe judgements in some sports competitions are subjective tasks, especially in competitions with high requirements on skills, which could lead to unfair results. Synchronized diving is such a skillful competition. Using computer to judge competitions automatically or assist referees to judge can effectively decrease the unfair results. This paper presents a framework for analyzing synchronization for synchronized diving videos. The framework is composed of three steps: (1) silhouette extraction, that can effectively get the diverpsilas silhouette; (2) synchronization feature extraction and representation, that can explicitly achieve synchronization features from synchronized diving videos; (3) synchronization evaluation, that can evaluate synchronization with evaluation function which is constructed by statistic learning method with diving video segments come from different games judged by different referees. The experiment shows that the method can accurately and effectively evaluate the synchronization for synchronized diving videos. Haoyang Ding, Jian Cheng 0001, Hanqing Lu, Zhixin Zhou |
ICME | 3 |
| 2008 | A novel contextual descriptors for category recognitionabstractIn this paper, we propose a novel contextual descriptor which combines the contextual information and local appearance. Based on Gibbs distribution, a local descriptor is designed. By assembling the contextual information and local descriptors, a new partial contextual descriptor (PCD) is finally presented. Combining Pyramid Match Kernel (PMK) and SVM, we test our new descriptor and obtain higher average precision of classification than using local appearance descriptor. Ming Tang 0001, Jian Cheng 0001, Jinqiao Wang, Hanqing Lu, Songde Ma |
ICME | 5 |
| 2008 | Online video advertising based on user's attention relavancy computingabstractInformation overload has become an important problem in the internet, and that all kinds of existing ads flood into people’s eyes causes scarcity of user’s attention. To provide relevant information under user’s control, we propose an online video advertising framework based on user’s attention relevancy computing. Users receive relevant video ads in exchange of their attention consumption. Multimodal concept detectors are trained to annotate the video databases, and a multimodal video ads categorization and related concept-to-ad relevancy and ad-to-concept relevancy ranking algorithm are proposed to compute user’s attention relevancy. Experiments and a subjective evaluation show the feasibility and effectiveness of the proposed approach. Jinqiao Wang, Yikai Fang, Hanqing Lu |
ICME | 3 |
| 2008 | Human-centered image navigation on mobile devicesabstractHow to navigate large images on mobile device is an open problem due to small display screen. Currently region of interest (ROI) image compressing is a popular approach. However, since the semantics of image is diversified by different users, it is hard to extract general ROI. In this paper, we propose a human-centered image navigation method for mobile users with simple human interaction. Different from previous work, we first extract Local Saliency Map (LSM) of an image according to personalized requirement, which can reduce the semantic ambiguity of the image for different users. Based on LSM, we can detect the ROIs of the image, which fully satisfies the different interests of the different users. Experiments and user studies show an encouraging performance of the proposed method. Cunxun Zang, Qingshan Liu 0001, Jian Cheng 0001, Hanqing Lu |
ICME | 4 |
| 2008 | Automatic semantic annotation for video blogsabstractIn recent years, Weblogs (or blogs) have received great popularity worldwide, among which video blogs (or vlogs) are playing an increasingly important role. As vlogs gain in population, how to make them more easily accessible has become a hot research topic. In this paper, we propose a novel automatic annotation model for vlogs. We extract informative keywords from both the target vlog itself and external resources which are semantically and visually relevant to it. We also present a new evaluation criterion, which assigns a score to an annotation according to its accuracy and completeness in representing the vlogpsilas semantics. Experimental results demonstrate the effectiveness of both the annotation model and evaluation criterion. Xiaoyu Zhang 0002, Changsheng Xu, Jian Cheng 0001, Hanqing Lu, Songde Ma |
ICME | 4 |
| 2008 | Automatic character identification in feature-length filmsabstractThis paper presents a novel approach to automatically identify characters in films using audio visual cues and text analysis. The approach consists of three stages: (i) frontal face track detection and clustering, (ii) face track classification, (iii) name assignment. A Finite State Machine (FSM) method is utilized to filter faces detected on each frame and build face tracks. The face tracks are clustered using constrained K-Centers. The tracks located in the center area of each cluster are set as exemplars. The marginal points of each cluster and the newly detected non-frontal face tracks are classified to these exemplars using complementary cues of audio and visual. The names of characters are ranked based on their occurrences in the film script and the face track clusters are ranked based on track counts. The names are assigned to the clusters according to the ranking order. Experiments were conducted on two feature-length films and gave promising results. Yifan Zhang 0001, Changsheng Xu, Hanqing Lu |
ICME | 3 |
| 2008 | Change detection based on adaptive Markov Random FieldsabstractUsually changes in remote sensing images go along with the appearance or disappearance of some edges. In addition, pixels located along the edges are likely to weakly influenced by its neighborhood pixels, while pixels located far from the edges commonly have a tightly correlation among them. In this paper, we propose a novel change detection technique based on adaptive Markov Random Fields (MRFs) for high resolution satellite images with combined color and texture features. The technique is composed of two main steps: (1) the input images are marked with different region indexes by the combined color and edge features; (2) change maps are obtained under the MRF framework with alterable order of neighborhood and variable smooth weight coefficient controlled by the index map. The main contribution of this paper is that the spatial-contextual information included in the remote sensing imagery is correctly and adaptively exploited under an adaptive MRF framework. Experiments results obtained on a set of remote sensing imagery confirm the effectiveness of the proposed approach. Chunlei Huo, Jian Cheng 0001, Zhixin Zhou, Hanqing Lu |
ICPR | 5 |
| 2008 | Hand posture recognition with co-trainingabstractAs an emerging human-computer interaction approach vision based hand interaction is more natural and efficient. However in order to achieve high accuracy, most of the existing hand posture recognition methods need a large number of labeled samples which is expensive or unavailable in practice. In this paper, a co-training based method is proposed to recognize different hand postures with a small quantity of labeled data. Hand postures examples are represented with different features and disparate classifiers are trained simultaneously with labeled data. Then the semi-supervised learning treats each new posture as unlabeled data and updates the classifiers in a co-training framework. Experiments show that the proposed method outperforms the traditional methods with much less labeled examples. Yikai Fang, Jian Cheng 0001, Jinqiao Wang, Kongqiao Wang, Jing Liu 0001, Hanqing Lu |
ICPR | 6 |
| 2008 | Saliency Cuts: An automatic approach to object segmentationabstractInteractive graph cuts are widely used in object segmentation but with some disadvantages: 1) Manual interactions may cause inaccurate or even incorrect segmentation results and involve more interactions especially for novices. 2) In some situations, the manual interactions are infeasible. To overcome these disadvantages, we propose a novel approach, namely Saliency cuts, to segment object from background automatically. By exploring the effects of labels to graph cuts, the so called ldquoprofessional labelsrdquo is introduced to evaluate labels. With the aid of saliency detection, a multiresolution framework is designed to provide ldquoprofessional labelsrdquo automatically and implement object segmentation using graph cuts. The experiments demonstrate the promising performance of Saliency cuts. Jian Cheng 0001, Zhenglong Li 0001, Hanqing Lu |
ICPR | 4 |
| 2008 | A variational inference based approach for image segmentationabstractIn this paper, we present a variational Bayes (VB) approach for image segmentation. First, image is modeled by a mixture model, and then with the techniques of factor analyzer, the underlying structure of image content is inferred automatically. Different from the traditional EM algorithm that seriously suffers from component number selection, the proposed method can accurately infer the underlying image structure including suitable component number without usual sub- or over-segmentation problem. To overcome the problem of local optimization, a component split strategy is adopted in inference optimization process. Extensive experiments on various images validate the proposed method. Zhenglong Li 0001, Qingshan Liu 0001, Jian Cheng 0001, Hanqing Lu |
ICPR | 4 |
| 2008 | Lennard-Jones force field for Geometric Active ContourabstractThis paper presents a new Geometric Active Contour (GAC) model based on Lennard-Jones (L-J) force field, which is inspired by the theory of intermolecular interaction. It is different from gradient based GAC models in that the proposed model does not rely on any pre-computed edge map and is directly computed from image data. Moreover, it can integrate various information including grayscale, color and texture etc. The proposed L-J force field has two different characteristics controlled by a switch parameter c. In the case of c = 0, the force vector flows will bi-directionally converge to boundaries, and it will obtain a morphological dilation-like effect with c ≠ 0. We test the proposed method on various images, and the experimental results are very promising. Zhenglong Li 0001, Qingshan Liu 0001, Hanqing Lu, Dimitris N. Metaxas |
ICPR | 3 |
| 2008 | Multi-cue collaborative kernel tracking with cross ratio invariant constraintabstractIn this paper, a novel multi-cue collaborative kernel tracking algorithm is proposed. A new constraint based on the property of cross ratio invariant enables tracking of objects insensitive to complex motions, including scale changes, rotation and especially views changes, without labeling and training. Meanwhile, invariant moments are introduced into the kernel based tracking method as the shape representation. The integration of shape and color information makes tracking more robust, and avoids the kernels drifting when color information is not sufficient. Experiments show our method is robust to arbitrary motions of articulated objects and other rigid objects in complex environment. Jian Cheng 0001, Hanqing Lu |
ICPR | 3 |
| 2008 | Probabilistic tracking on Riemannian manifoldsabstractThe covariance region descriptor recently proposed in [1] has been proved robust and versatile for a modest computational cost. The covariance matrix enables efficient fusion of different types of features, where the spatial and statistical properties as well as their correlation are characterized. The similarity of two covariance descriptor is measured on Riemannian manifolds. Relying on the same metric, but within a probabilistic framework, we propose a novel tracking approach on Riemannian manifolds. The particle filtering technique allows us to better handle background clutter, as well as the temporary occlusions of the target. Furthermore, we extend the fast covariance computation to the tracking problem, which makes the tracking procedure more efficient. The proposed approach is robust to noises and much faster than the original search-based covariance tracker [2]. Extensive experimental results demonstrate greatly improved performance over classical color-based Bayesian tracker. Yi Wu 0001, Hanqing Lu |
ICPR | 4 |
| 2008 | Collaborate ball and player trajectory extraction in broadcast soccer videoabstractEnormous accessible broadcast soccer videos demand an efficient ball and player trajectory extraction framework to represent the tactic semantics for the automatic analysis. Camera motions, noise and blurs in broadcast videos make it difficult to extract the trajectories with a single existing object tracking algorithm. In this paper, we propose a novel framework for ball and player trajectory extraction in broadcast soccer videos. The framework generates candidate ball trajectories and player trajectory segments, then it searches the optimal ball trajectory with the likelihood ranking and refines player trajectories with MCMC data association. Instead of extracting ball and player trajectories respectively, our framework employs the motion relationship of the ball and players to build a collaborate scheme to improve the tracking and trajectory refinement results. The experimental results show the proposed framework is more effective than previous works. Yi Zhang 0007, Hanqing Lu, Changsheng Xu |
ICPR | 2 |
| 2008 | Unsupervised Change Detection in SAR Image using Graph CutsabstractIn this paper, we present an unsupervised change detection approach in temporal sets of SAR images. The change detection is represented as a task of energy minimization and the energy function is minimized using graph cuts. Neighboring pixels are taken into account in a priority sequence according to their distance from the center pixel, and the energy function is formed based on Markov Random Field (MRF) model. Graph cuts algorithm is employed for computing maximum a-posteriori (MAP) estimates of the MRF. Experiments results obtained on a SAR data set confirm the effectiveness of the proposed approach. The comparisons between graph cuts algorithm and iterated conditional modes (ICM) algorithm about the quality of change map and running time of energy minimization illustrate that graph cuts algorithm is a huge improvement over ICM. Chunlei Huo, Zhixin Zhou, Hanqing Lu |
IGARSS (3) | 4 |
| 2008 | A Multilevel Contextual Approach to Change Detection for very high Resolution ImagesabstractA multilevel contextual approach is proposed in this paper for change detection of VHR images. By representing the change features in a hierarchical contextual manner, the changes are detected level-by-level. By taking advantages of SVMs, the ambiguity of changes is mitigated and the optimal changes are detected peculiar to the specific user. Compared to the traditional methods, the proposed approach is more accurate, more robust and faster. Experiments demonstrate the effectiveness and advantages of the proposed approach. Chunlei Huo, Zhixin Zhou, Hanqing Lu, Jian Cheng 0001, Qingshan Liu 0001 |
IGARSS (4) | 4 |
| 2008 | Urban Change Detection based on Local Features and Multiscale FusionabstractA multiscale approach is presented in this paper for urban change detection of VHR images. The proposed approach detects the changes at different scales by local-region-based approach, which consists of local region extraction, local region description and local region comparison. To combine the changes at different scales and improve the accuracy, multiscale fusion strategy is applied to local-region-based change detection. Experimental results obtained on Quickbird images confirm the effectiveness of the proposed approach. Chunlei Huo, Zhixin Zhou, Qingshan Liu 0001, Jian Cheng 0001, Hanqing Lu |
IGARSS (3) | 5 |
| 2008 | IQ evaluation based adaptive wavelet denoising and enhancement for a VTRAN systemabstractAn Image Quality (IQ) evaluation based wavelet domain denoising and enhancement model for a Vehicle Target Recognition and Assistant Navigation (VTRAN) system is introduced. In order to adapt to a complex atmosphere environment, our model utilizes the IQ parameters as a criterion to search the proper estimation values of the wavelet domain denoising and enhancement model. Firstly, before the wavelet processing, we evaluate the IQ and estimate the expected image gray values from some sub-regions of the original image. Then our algorithm will add some ldquocleanrdquo and ldquoclearrdquo points with those expected gray values in the original image. After that our model will calculate the wavelet based denoising and enhancement model and evaluate the result by the IQ with those added reference points. Finally, if the IQ is not good enough, a feedback will be given and the model parameters will be modified until it gets a better result. Experiment results show our algorithm works well in many outdoor tests and it can be used in a complex atmosphere environment. Haoting Liu, Hanqing Lu |
IROS | 2 |
| 2008 | Hierarchical clustering-based navigation of image search resultsabstractUsually, the image search results contain multiple topics on semantic level and even semantically consistent images have diverse appearances on visual level. How to organize the results into semantically and visually consistent clusters becomes a necessary task to facilitate users' navigation. To attack this, HiCluster, an effective method to organize image search results is designed in this paper, which employs both textual and visual analysis. First, we extract some query-related key phrases to enumerate specific semantics of the given query and cluster them into some semantic clusters using K-lines-based clustering algorithm. Second, the resulting images corresponding to each key phrase are clustered with Bregman Bubble Clustering (BBC) algorithm, which partially groups images in the whole set while discarding some scattered noisy ones. At last, a novel user interface (UI) is designed to provide users with the diverse and helpful information based on the hierarchical clustering structure. Experiments on web images demonstrate the effectiveness and potential of the system. Haoyang Ding, Jing Liu 0001, Hanqing Lu |
ACM Multimedia | 3 |
| 2008 | Boosting relative spaces for categorizing objects with large intra-class variationabstractIn this paper, a novel method for object categorization is proposed. We first analyze the phenomenon of large intra-class variation and attribute it to the "subcategory" problem. To reveal the local and distinct properties of the different subcategories, relative spaces are constructed. Then the weighted FLDs (Fisher Linear Discriminant) as weak learners trained in relative spaces are integrated with the boosting framework to form the final classifier. Experiments on 8 categories from Caltech database show the effectiveness of our algorithm. Ming Tang 0001, Jinqiao Wang, Hanqing Lu, Songde Ma |
ACM Multimedia | 4 |
| 2008 | Selective Sampling Based on Dynamic Certainty Propagation for Image Retrieval
Xiaoyu Zhang 0002, Jian Cheng 0001, Hanqing Lu, Songde Ma |
MMM | 3 |
| 2008 | A new extension of kernel feature and its application for visual recognition
Qingshan Liu 0001, Hongliang Jin, Xiaoou Tang, Hanqing Lu, Songde Ma |
Neurocomputing | 4 |
| 2008 | Automatic composition of broadcast sports video
Jinjun Wang, Changsheng Xu, Chng Eng Siong, Hanqing Lu, Qi Tian 0002 |
Multim. Syst. | 4 |
| 2008 | A graph-based image annotation framework
Jing Liu 0001, Hanqing Lu, Songde Ma |
Pattern Recognit. Lett. | 3 |
| 2008 | A Multimodal Scheme for Program Segmentation and Representation in Broadcast Video StreamsabstractWith the advance of digital video recording and playback systems, the request for efficiently managing recorded TV video programs is evident so that users can readily locate and browse their favorite programs. In this paper, we propose a multimodal scheme to segment and represent TV video streams. The scheme aims to recover the temporal and structural characteristics of TV programs with visual, auditory, and textual information. In terms of visual cues, we develop a novel concept named program-oriented informative images (POIM) to identify the candidate points correlated with the boundaries of individual programs. For audio cues, a multiscale Kullback-Leibler (K-L) distance is proposed to locate audio scene changes (ASC), and accordingly ASC is aligned with video scene changes to represent candidate boundaries of programs. In addition, latent semantic analysis (LSA) is adopted to calculate the textual content similarity (TCS) between shots to model the inter-program similarity and intra-program dissimilarity in terms of speech content. Finally, we fuse the multimodal features of POIM, ASC, and TCS to detect the boundaries of programs including individual commercials (spots). Towards effective program guide and attracting content browsing, we propose a multimodal representation of individual programs by using POIM images, key frames, and textual keywords in a summarization manner. Extensive experiments are carried out over an open benchmarking dataset TRECVID 2005 corpus and promising results have been achieved. Compared with the electronic program guide (EPG), our solution provides a more generic approach to determine the exact boundaries of diverse TV programs even including dramatic spots. Jinqiao Wang, Ling-Yu Duan, Qingshan Liu 0001, Hanqing Lu, Jesse S. Jin |
IEEE Trans. Multim. | 4 |
| 2008 | A Novel Framework for Semantic Annotation and Personalized Retrieval of Sports VideoabstractSports video annotation is important for sports video semantic analysis such as event detection and personalization. In this paper, we propose a novel approach for sports video semantic annotation and personalized retrieval. Different from the state of the art sports video analysis methods which heavily rely on audio/visual features, the proposed approach incorporates web-casting text into sports video analysis. Compared with previous approaches, the contributions of our approach include the following. 1) The event detection accuracy is significantly improved due to the incorporation of web-casting text analysis. 2) The proposed approach is able to detect exact event boundary and extract event semantics that are very difficult or impossible to be handled by previous approaches. 3) The proposed method is able to create personalized summary from both general and specific point of view related to particular game, event, player or team according to user's preference. We present the framework of our approach and details of text analysis, video analysis, text/video alignment, and personalized retrieval. The experimental results on event boundary detection in sports video are encouraging and comparable to the manually selected events. The evaluation on personalized retrieval is effective in helping meet users' expectations. Changsheng Xu, Jinjun Wang, Hanqing Lu, Yifan Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2008 | Using Webcast Text for Semantic Event Detection in Broadcast Sports VideoabstractSports video semantic event detection is essential for sports video summarization and retrieval. Extensive research efforts have been devoted to this area in recent years. However, the existing sports video event detection approaches heavily rely on either video content itself, which face the difficulty of high-level semantic information extraction from video content using computer vision and image processing techniques, or manually generated video ontology, which is domain specific and difficult to be automatically aligned with the video content. In this paper, we present a novel approach for sports video semantic event detection based on analysis and alignment of Webcast text and broadcast video. Webcast text is a text broadcast channel for sports game which is co-produced with the broadcast video and is easily obtained from the Web. We first analyze Webcast text to cluster and detect text events in an unsupervised way using probabilistic latent semantic analysis (pLSA). Based on the detected text event and video structure analysis, we employ a conditional random field model (CRFM) to align text event and video event by detecting event moment and event boundary in the video. Incorporation of Webcast text into sports video analysis significantly facilitates sports video semantic event detection. We conducted experiments on 33 hours of soccer and basketball games for Webcast analysis, broadcast video analysis and text/video semantic alignment. The results are encouraging and compared with the manually labeled ground truth. Changsheng Xu, Yifan Zhang 0001, Guangyu Zhu 0002, Yong Rui, Hanqing Lu, Qingming Huang |
IEEE Trans. Multim. | 5 |
| 2007 | Image Segmentation Using Co-EM Strategy
Zhenglong Li 0001, Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu |
ACCV (2) | 4 |
| 2007 | Topology-Preserved Diffusion Distance for Histogram ComparisonabstractIn most previous works, histograms are simply treated as n-dimensional arrays or even reshaped into vectors when measuring the distances between them. However many histograms have their intrinsic topologies, such as HSV histogram (cone), shape context (polar), orientation histogram (circle). The topologies are important for so-called cross-bin distance, because they determine the similarities between histogram bins, and influence the crossbin distances between histograms. In this paper, we proposed the topologypreserved diffusion distance to take the topology into account. This method extracts the distance by measuring the heat diffusion process defined on the topology of the histogram. Moreover, a fast implementation with time complexity O(N) is developed. Experiments on image retrieval and interest point matching show the effectiveness and efficiency of the proposed method. 1 Wang Yan, Qiqi Wang 0001, Qingshan Liu 0001, Hanqing Lu, Songde Ma |
BMVC | 4 |
| 2007 | Sequential Architecture for Efficient Car DetectionabstractBased on multi-cue integration and hierarchical SVM, we present a sequential architecture for efficient car detection under complex outdoor scene in this paper. On the low level, two novel area templates based on edge and interest-point cues respectively are first constructed, which can be applied to forming the identities of visual perception to some extent and thus utilized to reject rapidly most of the negative non-car objects at the cost of missing few of the true ones. Moreover on the high level, both global structure and local texture cues are exploited to characterize the car objects precisely. To improve the computational efficiency of general SVM, a solution approximating based two-level hierarchical SVM is proposed. The experimental results show that the integration of global structure and local texture properties provides more powerful ability in discrimination of car objects from non-car ones. The final high detection performance also contributes to the utilizing of two novel low level visual cues and the hierarchical SVM. Zhenfeng Zhu, Yao Zhao 0001, Hanqing Lu |
CVPR | 3 |
| 2007 | Robust lip Localization on Multi-View Faces in VideoabstractIn this paper, a fast and robust multi-view lip localization algorithm in video is proposed. We consider lip localization as a binary classification problem, where a classifier is learned to distinguish between the lip and the region surrounding it. The classifier we use here is a histogram-based one which exploits the anthropometrical properties of the human face with the help of face scale normalization. Due to the perceptual uniformity and robustness for lip/skin color variations across different people, we adopt CIELUV color model to represent the color of lip. After classification, we propose a novel projection-cut algorithm by spatial deviation analysis (SDA) to locate the lip, which is effective to deal with the background clutters. Experimental results on teleplay videos demonstrate that the proposed approach is efficient and robust for lip localization. Yi Wu 0001, Wei Hu 0002, Tao Wang 0003, Yimin Zhang 0002, Jian Cheng 0001, Hanqing Lu |
ICIP (4) | 7 |
| 2007 | Weighted Co-SVM for Image Retrieval with MVB StrategyabstractIn relevance feedback, active learning is often used to alleviate the burden of labeling by selecting only the most informative data. Traditional data selection strategies often choose the data closest to the current classification boundary to label, which are in fact not informative enough. In this paper, we propose the moving virtual boundary (MVB) strategy, which is proved to be a more effective way for data selection. The co-SVM algorithm is another powerful method used in relevance feedback. Unfortunately, its basic assumption that each view of the data be sufficient is often untenable in image retrieval. We present our weighted co-SVM as an extension of co-SVM by attaching weight to each view, and thus relax the view sufficiency assumption. The experimental results show that the weighted co-SVM algorithm outperforms co-SVM obviously, especially with the help of MVB data selection strategy. Xiaoyu Zhang 0002, Jian Cheng 0001, Hanqing Lu, Songde Ma |
ICIP (4) | 3 |
| 2007 | A Real-Time Hand Gesture Recognition MethodabstractCompared with the traditional interaction approaches, such as keyboard, mouse, pen, etc, vision based hand interaction is more natural and efficient. In this paper, we proposed a robust real-time hand gesture recognition method. In our method, firstly, a specific gesture is required to trigger the hand detection followed by tracking; then hand is segmented using motion and color cues; finally, in order to break the limitation of aspect ratio encountered in most of learning based hand gesture methods, the scale-space feature detection is integrated into gesture recognition. Applying the proposed method to navigation of image browsing, experimental results show that our method achieves satisfactory performance. Yikai Fang, Kongqiao Wang, Jian Cheng 0001, Hanqing Lu |
ICME | 4 |
| 2007 | Image Annotation Refinement using NSC-Based Word CorrelationabstractImage annotation refinement is crucial to improve the performance of automatic image annotation, in which the estimation of word correlation is a key issue. Typically, the word co-occurrence information may be utilized to estimate the word correlation. However, this approach is not accurate enough because it equally treats any word pair co-occurring in the training data and cannot extract synonymy relationship effectively. In this paper, a novel method is developed to estimate the word correlation based on the improved nearest spanning chains (NSC). It can extract more informative and reasonable relations among keywords. Obtaining the enhanced word correlation, a word-based graph is constructed, which is used to re-rank the candidate annotations for an untagged image. Experiments conducted on the typical Corel dataset demonstrate the effectiveness of the proposed method. Jing Liu 0001, Mingjing Li, Qingshan Liu 0001, Hanqing Lu, Songde Ma |
ICME | 4 |
| 2007 | Lip Localization on Multi-View Faces in VideoabstractIn this paper, a real time and robust multi-view lip localization in video frames is proposed. The objective of this research is to exactly localize lip surrounding box with sub-pixel accuracy in low quality video frames. This is achieved by multi-view face detection in video frames, lip ROI extraction, color space transformation, vertical edge feature detection and lip localization by a feature-spatial profile analysis on a joint color-edge plane. One possible application of talking face detection in video is used to evaluate our work. Experimental results on TV programs and movies demonstrate that the proposed approach is efficient and robust in lip localization. Yi Wu 0001, Wei Hu 0002, Tao Wang 0003, Yimin Zhang 0002, Hanqing Lu |
ICME | 7 |
| 2007 | Robust Commercial Retrieval in Video StreamsabstractTV commercial video is a kind of informative medium. To fast and robustly index and retrieve commercial videos is of interest to commercial monitor, copyright protection, and commercial management, we propose a coarse-to-fine scheme to robustly retrieve commercial videos. Different from previous work using clip or key frames-based matching, our scheme has incorporated the commercial production knowledge to search the candidate commercial positions. Color and ordinal features are extracted for locating the exact commercial positions with dynamic time warping distance. Comparison experiments were carried out over TRECVID 2006 news videos and some videos from Chinese channels. Our scheme has achieved promising simulation results. Jinqiao Wang, Ling-Yu Duan, Qingshan Liu 0001, Hanqing Lu, Jesse S. Jin |
ICME | 4 |
| 2007 | A New Multimedia Message Customizing Framework for mobile DevicesabstractIn this paper, we present a novel framework to customize multimedia messages for mobile users. The goal is to generate a video message from a series of pictures. The framework includes visual attention view detection, image grouping, image ranking, and slideshow generation. Considering the limitation of mobile device, we use a simple color feature based attention model to detect interesting regions of the images. We group the images, and rank them based on the attention view similarities. Finally a human perception based slideshow is designed to keep the mobile users' eye on attention regions efficiently. In addition, a short music is selected to match the video message. Extensive experiments and user studies show the promising performance of the proposed system. Cunxun Zang, Qingshan Liu 0001, Hanqing Lu, Kongqiao Wang |
ICME | 3 |
| 2007 | Semantic Event Extraction from Basketball Games using Multi-Modal AnalysisabstractIn this paper, we present a novel multi-modal framework for semantic event extraction from basketball games based on Webcasting text and broadcast video. We propose novel approaches to text analysis for event detection and semantics extraction, video analysis for event structure modeling and event moment detection, and text/video alignment for event boundary detection in the video. Compared with existing approaches to event detection in sports video which rely heavily on low-level features directly extracted from video itself, our approach aims to bridge the semantic gap between low-level features and high-level events and facilitates personalization of the sports video. Promising results are reported on real-world video clips by using text analysis, video analysis and text/video alignment. Yifan Zhang 0001, Changsheng Xu, Yong Rui, Jinqiao Wang, Hanqing Lu |
ICME | 5 |
| 2007 | Human behaviour consistent relevance feedback model for image retrievalabstractDue to the well known semantic gap, content based image retrieval is a difficult problem. To bridge it, relevance feedback as an effective solution has been extensively studied in literatures. However, existing methods follow a single-line searching philosophy, which may lead to a local optimum in search space. To address the problem, we propose a human behavior consistent relevance feedback model for image retrieval in this paper. Simulating human behaviors, the proposed model enable the user to perform relevance feedback in three manners: Follow up, Go back, and Restart. Each manner is a way for the user to provide the system with his or her opinions about search results. The accumulated feedback information can be used to refine the user query and regulate the similarity metric. We adopt the graph ranking algorithm to model the retrieval process. Experiments conducted on standard Corel dataset and Pascal VOC 2006 dataset demonstrate the effectiveness of the proposed mechanism. Jing Liu 0001, Zhiwei Li 0006, Mingjing Li, Hanqing Lu, Songde Ma |
ACM Multimedia | 4 |
| 2007 | Dual cross-media relevance model for image annotationabstractImage annotation has been an active research topic in recent years due to its potential impact on both image understanding and web image retrieval. Existing relevance-model-based methods perform image annotation by maximizing the joint probability of images and words, which is calculated by the expectation over training images. However, the semantic gap and the dependence on training data restrict their performance and scalability. In this paper, a dual cross-media relevance model (DCMRM) is proposed for automatic image annotation, which estimates the joint probability by the expectation over words in a pre-defined lexicon. DCMRM involves two kinds of critical relations in image annotation. One is the word-to-image relation and the other is the word-to-word relation. Both relations can be estimated by using search techniques on the web data as well as available training data. Experiments conducted on the Corel dataset and a web image dataset demonstrate the effectiveness of the proposed model. Jing Liu 0001, Mingjing Li, Zhiwei Li 0006, Wei-Ying Ma, Hanqing Lu, Songde Ma |
ACM Multimedia | 6 |
| 2007 | Automatic TV Logo Detection, Tracking and Removal in Broadcast Video
Jinqiao Wang, Qingshan Liu 0001, Ling-Yu Duan, Hanqing Lu, Changsheng Xu |
MMM (2) | 4 |
| 2007 | A Fuzzy Segmentation of Salient Region of Interest in Low Depth of Field Image
KeDai Zhang, Hanqing Lu, MiYi Duan |
MMM (1) | 2 |
| 2007 | Differential energy modulation: Novel video labeling scheme for content fidelity and easy rate controlabstractThis paper proposes the differential energy modulation(DEM) algorithm for video watermarking. The DEM algorithm embeds watermark bits by applying Quantization Index Modulation(QIM) to differential energy(DE). The energy difference between horizontal and vertical components in a quantized DCT block is modified by dithering values to make the DE consistent with a watermark bit quantizer. The suggested DEM algorithm is superior to QIM in robustness, since it chooses a region energy to carry a label bit, and spreads embedding distortion over values. On contrast, QIM labels with a coefficient instead. DEM also improves Differential Energy Watermarking(DEW)’s low capacity while preserves image quality, because DEW had to employ many blocks to carry a bit to avoid artifacts and guarantee valid energies. Experiments demonstrate this conclusion. Besides, DEM is a high perceptual and easy rate control scheme. It is robust to attacks of frame dropping, image cropping and additive white Gaussian noise. Lihong Ma 0002, Hanqing Lu |
SMC | 3 |
| 2007 | Generalized optical flow in the scale space
Haifeng Gong, Chunhong Pan, Qing Yang 0002, Hanqing Lu, Songde Ma |
Comput. Vis. Image Underst. | 4 |
| 2007 | Scale multiplication in odd Gabor transform domain for edge detection
Zhenfeng Zhu, Hanqing Lu, Yao Zhao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2007 | An improved variable-size block-matching algorithm
Qingshan Liu 0001, Hanqing Lu |
Multim. Tools Appl. | 3 |
| 2007 | Generation of Personalized Music Sports Video Using Multimodal CuesabstractIn this paper, we propose a novel automatic approach for personalized music sports video generation. Two research challenges are addressed, specifically the semantic sports video content extraction and the automatic music video composition. For the first challenge, we propose to use multimodal (audio, video, and text) feature analysis and alignment to detect the semantics of events in broadcast sports video. For the second challenge, we introduce the video-centric and music-centric music video composition schemes and proposed a dynamic-programming based algorithm to perform fully or semi-automatic generation of personalized music sports video. The experimental results and user evaluations are promising and show that our systems generated music sports video is comparable to professionally generated ones. Our proposed system greatly facilitates the music sports video editing task for both professionals and amateurs Jinjun Wang, Chng Eng Siong, Changsheng Xu, Hanqing Lu, Qi Tian 0002 |
IEEE Trans. Multim. | 4 |
| 2006 | A Geometric Contour Framework with Vector Field Support
Zhenglong Li 0001, Qingshan Liu 0001, Hanqing Lu |
ACCV (2) | 3 |
| 2006 | Boosting Multi-gabor Subspaces for Face Recognition
Qingshan Liu 0001, Hongliang Jin, Xiaoou Tang, Hanqing Lu, Songde Ma |
ACCV (1) | 4 |
| 2006 | Motion Detection in Driving Environment Using U-V-Disparity
Zhencheng Hu, Hanqing Lu, Keiichi Uchimura |
ACCV (1) | 3 |
| 2006 | Automatic Moving Object Segmentation with Accurate Boundaries
Qingshan Liu 0001, Hanqing Lu |
ACCV (1) | 4 |
| 2006 | Fast Global Motion Estimation Via Iterative Least-Square Method
Qingshan Liu 0001, Hanqing Lu |
ACCV (2) | 4 |
| 2006 | Multiple Similarities Based Kernel Subspace Learning for Image Classification
Wang Yan, Qingshan Liu 0001, Hanqing Lu, Songde Ma |
ACCV (2) | 3 |
| 2006 | Fusion Method of Fingerprint Quality Evaluation: From the Local Gabor Feature to the Global Spatial-Frequency Structures
Decong Yu, Lihong Ma 0002, Hanqing Lu, Zhiqing Chen |
ACIVS | 3 |
| 2006 | Neural Network Modeling of Spectral EmbeddingabstractMost of spectral embedding algorithms such as Isomap, LLE and Laplacian Eigenmap only give map on training samples. One main problem of these methods is to find the embedding of new samples, which is known as the outof-sample problem of spectral embedding. In this paper, we propose a neural network based method to solve this problem. Neural network is used to train and perform both the forward map from high dimensional image space to low dimensional embedding space, and the backward map in the reverse direction. Additionally, combining the forward and backward network, this method is able to build auto-association model to retrieve high dimensional data, and cross association model to learn high dimensional correspondences. Experiments are conducted on real images for forward and backward map, auto-association and cross association. Haifeng Gong, Chunhong Pan, Qing Yang 0002, Hanqing Lu, Songde Ma |
BMVC | 4 |
| 2006 | Craniofacial Landmark Detection by Layered Diffusion and Dilated Skeleton MapsabstractThis paper proposed a new method for cephalogram landmark location. It firstly employed diffusions in different scales for layered segmentation, making uses of region-homogeneity and edge-saltation, thus the landmark positions could be clearly reflected on edge-maps, i.e. skeleton, of layered diffusion. Secondly, key landmarks were determined via binarization pixel-cliques which were formed by Euclidean distance maps (EDM) on dilated skeleton. To validate the performance of this method, 30 surgery cases were inspected. The comparison of predicted parameters and application values shows that all 9 angle parameters and 4 of 5 distance parameters were correctly calculated, the predicted profiles were similar to the actual contours except one point due to the intrinsic difficulty in labium prediction. Our method took advantages of layered feature preserving and ease landmark extraction, and the latter was also owing to different EDM-Width of pixel-cliques and skeleton points. It is accurate, lifelike, and superior to many similar methods Lihong Ma 0002, Shengmin Jiang, Hanqing Lu |
ICARCV | 5 |
| 2006 | A Mid-Level Scene Change Representation Via Audiovisual AlignmentabstractScene is a series of semantic correlated video shots. An effective scene detection depends on domain knowledge more or less. Most existing approaches try to directly detect various scene changes by applying clustering or supervised learning methods to low level audiovisual features. However, robustly detecting diverse scene changes derived from complex semantic meanings is still a challenging problem. In this paper we are focused on the association of visual signal changes (e.g. cuts, fade-in, fade-out, etc.) and audio signal changes (e.g. speaker change, background music change, etc.) to propose a mid-level scene change representation, which is meant to locate candidate scene change points by characterizing temporally uncorrelated properties of audio and visual track in the case of scene change happening. By incorporating domain knowledge, enhanced features can be further extracted to complement this representation to bridge semantic gap towards scene change detection. We utilize a camera motion estimation algorithm to detect visual signal changes. Such visual change positions are selected as time-stamp points. An alignment is performed to search for candidate audio signal change positions by multi-scale Kullback-Leibler(K-L) distance computing. Both metric-based K-L distance approach and model-based HMM are applied to determine true audio signal changes. The associated visual and audio signal changes are considered as the mid-level scene change representation. This representation has been successfully applied to detect boundaries of individual commercial in TV broadcast stream with an accuracy of around 95%. Particularly the systematic alignment approach can be utilized in video summarization. Jinqiao Wang, Ling-Yu Duan, Hanqing Lu, Jesse S. Jin, Changsheng Xu |
ICASSP (2) | 3 |
| 2006 | Web Image Mining Based on Modeling Concept-Sensitive Salient RegionsabstractIn this paper, we propose a probabilistic model for Web image mining, which is based on concept-sensitive salient regions without human intervene. Our goal is to achieve a middle-level understanding of image semantics to bridge the semantic gap existing in the field of image mining and retrieval. With the help of a popular search engine, semantically relevant images are collected, and concept-sensitive salient regions are extracted automatically based on an attention model. Then the semantic concept model is learned from the joint distribution of all salient regions with Gaussian mixture model and expectation-maximization algorithm. In addition, by incorporating semantically irrelevant un-salient regions as negative samples, the discriminative power of the solution is further enhanced. Experiments demonstrate the encouraging performance of the proposed method Jing Liu 0001, Qingshan Liu 0001, Jinqiao Wang, Hanqing Lu, Songde Ma |
ICME | 4 |
| 2006 | Identify Sports Video Shots with "Happy" or "Sad" EmotionsabstractSemantic video content extraction and selection are critical steps in sports video analysis and editing. The identification of video segments can be from various semantic perspectives, e.g. certain event, player or emotional state. In this paper, we examined the possibility of automatically identifying shots with "happy" or "sad" emotion from broadcast sports video. Our proposed model first performs the sports highlight extraction to obtain candidate shots that possibly contain emotion information and then classifies these shots into either "happy" or "sad" emotion groups using hidden Markov model based method. The final experimental results are satisfactory Jinjun Wang, Chng Eng Siong, Changsheng Xu, Hanqing Lu, Xiaofeng Tong |
ICME | 4 |
| 2006 | A Robust Method for TV Logo Tracking in Video StreamsabstractMost broadcast stations rely on TV logos to claim video content ownership or visually distinguish the broadcast from the interrupting commercial block. Detecting and tracking a TV logo is of interest to TV commercial skipping applications and logo-based broadcasting surveillance (abnormal signal is accompanied by logo absence). Pixel-wise difference computing within predetermined logo regions cannot address semi-transparent TV logos well for the blending effects of a logo itself and inconstant background images. Edge-based template matching is weak for semi-transparent ones when incomplete edges appear. In this paper we present a more robust approach to detect and track TV logos in video streams on the basis of multispectral images gradient. Instead of single frame based detection, our approach makes use of the temporal correlation of multiple consecutive frames. Since it is difficult to manually delineate logos of irregular shape, an adaptive threshold is applied to the gradient image in subpixel space to extract the logo mask. TV logo tracking is finally carried out by matching the masked region with a known template. An extensive comparison experiment has shown our proposed algorithm outperforms traditional methods such as frame difference, single frame-based edge matching. Our experimental dataset comes from part of TRECVID2005 news corpus and several Chinese TV channels with challenging TV logos Jinqiao Wang, Ling-Yu Duan, Zhenglong Li 0001, Jing Liu 0001, Hanqing Lu, Jesse S. Jin |
ICME | 5 |
| 2006 | Fast Progressive Model Refinement Global Motion Estimation Algorithm with PredictionabstractGlobal motion estimation (GME) is an important part in the object-based applications. In this paper, a fast progressive model refinement (FPMR) GME algorithm is proposed. It can select the appropriate motion model according to the complexity of the camera motion. Two techniques are used to accelerate the procedure of FPMR. The first is an outlier prediction based feature point selection method. It can predict outliers from that of the last frame and therefore can effectively remove the influence of outliers on parameter calculation. The second is an intermediate-level model prediction method, which is used to fast the model selection and the parameter calculation procedure. Experiments show that the proposed algorithm is above two times faster than that of the feature-based fast and robust GME technique Qingshan Liu 0001, Hanqing Lu |
ICME | 4 |
| 2006 | Dynamic Similarity Kernel for Visual Recognition
Wang Yan, Qingshan Liu 0001, Hanqing Lu, Songde Ma |
KES (2) | 3 |
| 2006 | Segmentation, categorization, and identification of commercial clips from TV streams using multimodal analysisabstractTV advertising is ubiquitous, perseverant, and economically vital. Millions of people's living and working habits are affected by TV commercials. In this paper, we present a multimodal ("visual + audio + text") commercial video digest scheme to segment individual commercials and carry out semantic content analysis within a detected commercial segment from TV streams.Two challenging issues are addressed. Firstly, we propose a multimodal approach to robustly detect the boundaries of individual commercials. Secondly, we attempt to classify a commercial with respect to advertised products/services. For the first, the boundary detection of individual commercials is reduced to the problem of binary classification of shot boundaries via the mid-level features derived from two concepts: Image Frames Marked with Product Information (FMPI) and Audio Scene Change Indicator (ASCI). Moreover, the accurate individual boundary enables us to perform commercial identification by clip matching via a spatial-temporal signature. For the second, commercial classification is formulated as the task of text categorization by expanding sparse texts from ASR/OCR with external knowledge. Our boundary detection has achieved a good result of F1 = 93.7% on the dataset comprising 499 individual commercials from TRECVID'05 video corpus. Commercial classification has obtained a promising accuracy of 80.9% on 141 distinct ones. Based on these achievements, various applications such as an intelligent digital TV set-top box can be accomplished to enhance the TV viewer's capabilities in monitoring and managing commercials from TV streams. Ling-Yu Duan, Jinqiao Wang, Yantao Zheng, Jesse S. Jin, Hanqing Lu, Changsheng Xu |
ACM Multimedia | 5 |
| 2006 | A Fast Mean Shift Procedure with New Iteration Strategy and Re-samplingabstractMean-shift analysis is a general nonparametric clustering technique based on density estimation for the analysis of complex feature spaces. It has been successfully applied to many applications such as segmentation and tracking. However, despite its promising performance, there are applications for which the algorithm converges too slowly and is not practical. In this paper, an improved version of mean shift algorithm is proposed and implemented. The fast mean shift procedure uses a new iteration strategy and re-sampling. The new iteration strategy is based on updating cluster centers according to dynamically updated sample set. And the original data set is simplified by re-sampling, which accelerates the algorithm more significantly. Experimental results demonstrate the efficiency of the fast mean shift procedure in clustering problems. Huimin Guo, Ping Guo 0002, Hanqing Lu |
SMC | 3 |
| 2006 | Ensemble learning for independent component analysis
Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001 |
Pattern Recognit. | 3 |
| 2006 | Design and Performance Studies of an Adaptive Scheme for Serving Dynamic Web Content in a Mobile Computing EnvironmentabstractCurrently, people gain easy access to an increasingly diverse range of mobile devices such as personal digital assistants (PDAs), smart phones, and handheld computers. As dynamic content has become dominant on the fast-growing World Wide Web (C. Yuan et al., 2003), it is necessary to provide effective ways for the users to access such prevalent Web content in a mobile computing environment. During a course of browsing dynamic content on mobile devices, the requested content is first dynamically generated by remote Web server, then transmitted over a wireless network, and, finally, adapted for display' on small screens. This leads to considerable latency and processing load on mobile devices. By integrating a novel Web content adaptation algorithm and an enhanced caching strategy, we propose an adaptive scheme called MobiDNA for serving dynamic content in a mobile computing environment. To validate the feasibility and effectiveness of the proposed MobiDNA system, we construct an experimental testbed to investigate its performance. Experimental results demonstrate that this scheme can effectively improve mobile dynamic content browsing, by improving Web content readability on small displays, decreasing mobile browsing latency, and reducing wireless bandwidth consumption Zhigang Hua, Xing Xie 0001, Hao Liu 0007, Hanqing Lu, Wei-Ying Ma |
IEEE Trans. Mob. Comput. | 4 |
| 2006 | Face recognition using kernel scatter-difference-based discriminant analysisabstractThere are two fundamental problems with the Fisher linear discriminant analysis for face recognition. One is the singularity problem of the within-class scatter matrix due to small training sample size. The other is that it cannot efficiently describe complex nonlinear variations of face images because of its linear property. In this letter, a kernel scatter-difference-based discriminant analysis is proposed to overcome these two problems. We first use the nonlinear kernel trick to map the input data into an implicit feature space F. Then a scatter-difference-based discriminant rule is defined to analyze the data in F. The proposed method can not only produce nonlinear discriminant features but also avoid the singularity problem of the within-class scatter matrix. Extensive experiments show encouraging recognition performance of the new algorithm. Qingshan Liu 0001, Xiaoou Tang, Hanqing Lu, Songde Ma |
IEEE Trans. Neural Networks | 3 |
| 2005 | A Nonlinear Approach for Face Sketch Synthesis and RecognitionabstractMost face recognition systems focus on photo-based face recognition. In this paper, we present a face recognition system based on face sketches. The proposed system contains two elements: pseudo-sketch synthesis and sketch recognition. The pseudo-sketch generation method is based on local linear preserving of geometry between photo and sketch images, which is inspired by the idea of locally linear embedding. The nonlinear discriminate analysis is used to recognize the probe sketch from the synthesized pseudo-sketches. Experimental results on over 600 photo-sketch pairs show that the performance of the proposed method is encouraging. Qingshan Liu 0001, Xiaoou Tang, Hongliang Jin, Hanqing Lu, Songde Ma |
CVPR (1) | 4 |
| 2005 | A Semi-Supervised Framework for Mapping Data to the Intrinsic ManifoldabstractThis paper presents a novel scheme for manifold learning. Different from the previous work reducing data to Euclidean space which cannot handle the looped manifold well, we map the scattered data to its intrinsic parameter manifold by semisupervised learning. Given a set of partially labeled points, the map to a specified parameter manifold is computed by an iterative neighborhood average method called anchor points diffusion procedure (APD). We explore this idea on the most frequently used close formed manifolds, Stiefel manifolds whose special cases include hyper sphere and orthogonal group. The experiments show that APD can recover the underlying intrinsic parameters of points on scattered data manifold successfully. Haifeng Gong, Chunhong Pan, Qing Yang 0002, Hanqing Lu, Songde Ma |
ICCV | 4 |
| 2005 | Replay Scene Classification in Soccer Video Using Web Broadcast TextabstractThe automatic extraction of sports video highlights is a typical kind of personalized media production process. Many ways have been studied from the viewpoints of low-level audio/visual processing (e. g. detection of excited commentator speech), event detection (e. g. goal detection), etc. However, the subjectivity of highlights is an unavoidable bottleneck. The replay scene is an effective clue for highlights in broad-cast sports video due to the incorporation of video production knowledge. Most related work deals with the replay detection and/or a simple composition of all detected replays to generate highlights. Different from previous work, our work considers different flavors of different people in terms of highlight content or type through replay scenes classification. The main contributions include: 1) proposing a multi-modal (visual+ textual) approach for refined replay classification; 2) employing the sources of Broadcast Web Text (BWT) to facilitate replay content analysis. An overall accuracy of 79.9% has been achieved on seven soccer matches over seven replay categories Jinhui Dai, Ling-Yu Duan, Xiaofeng Tong, Changsheng Xu, Qi Tian 0002, Hanqing Lu, Jesse S. Jin |
ICME | 6 |
| 2005 | Automatic Annotation of Location Information for WWW ImagesabstractCurrently, a crucial challenge is raised on how to manage a large amount of images on the Web. Due to a real synergy between an image and its location, we propose an automatic solution to annotate contextual location information for WWW images. We construct an image importance model to acquire the dominant images in a page that comprise contextual surrounding text. For each acquired image, we develop an effective algorithm to compute location from its contextual text. We apply our approach to 1,000 pages from various Websites for image location annotation. The experiments demonstrated that more than 30% WWW images are related with geographic location information, and our solution can achieve the satisfactory results. Finally, we present some potential applications involving the utilization of image location information Zhigang Hua, Chuang Wang 0001, Xing Xie 0001, Hanqing Lu, Wei-Ying Ma |
ICME | 4 |
| 2005 | Learning Local Descriptors for Face DetectionabstractIn this paper, we propose a realtime face detection approach based on local structure and texture of the objects in gray-level images. Our strategy is to map the local spatial structures and image textures of face class into binary patterns, and use these binary patterns as local descriptors. Boosting based face detector is constructed using these local descriptors, and cascade scheme is employed to further improve the efficiency of the face detector. Compared to the existing face detection approaches, our proposed method has two advantages: (1) it is robust to illumination changes to some extend, for the features use the information of local relationship instead of the original gray values; (2) the computational cost is very low, both in training procedure and evaluation step. The experimental results show that the proposed method can meet the demand of realtime applications with a satisfied detection performance. Hongliang Jin, Qingshan Liu 0001, Xiaoou Tang, Hanqing Lu |
ICME | 4 |
| 2005 | A Mid-level Visual Concept Generation Framework for Sports AnalysisabstractThe development of mid-level concepts helps to bridge the gap between low-level feature and high-level semantics in video analysis. Most existing work combines the customized mid-level concepts and statistical models to detect particular events. Based on broadcast sports video production knowledge, we extend our previous work to present a unified framework for mid-level concept generation in this paper. A video segment is characterized via three essential aspects: camera shot size, an object appearing in a scene, and video production technology. These three aspects clearly summarize the primary concerns in terms of a generic concept generation. Within this framework, we can flexibly and clearly define meaningful mid-level concepts towards comprehensive video content analysis, such as replay classification and the detection of events (e. g. goal, shoot, attack, foul, offside, and out of bound, etc.). Xiaofeng Tong, Ling-Yu Duan, Hanqing Lu, Changsheng Xu, Qi Tian 0002, Jesse S. Jin |
ICME | 3 |