Kuiyuan Yang

dblp:21/7643 · DBLP profile ↗
← Back
47ranked-venue papers
6as first author
8since 2021 · last 2024
0000-0003-3063-2925ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 25 · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 SurroundSDF: Implicit 3D Scene Understanding Based on Signed Distance Field
abstract
Vision-centric 3D environment understanding is both vi-tal and challenging for autonomous driving systems. Re-cently, object-free methods have attracted considerable at-tention. Such methods perceive the world by predicting the semantics of discrete voxel grids but fail to construct continuous and accurate obstacle surfaces. To this end, in this paper, we propose SurroundSDF to implicitly predict the signed distance field (SDF) and semantic field for the continuous perception from surround images. Specifically, we introduce a query-based approach and utilize SDF con-strained by the Eikonal formulation to accurately describe the surfaces of obstacles. Furthermore, considering the absence of precise SDF ground truth, we propose a novel weakly supervised paradigm for SDF, referred to as the Sandwich Eikonal formulation, which emphasizes applying correct and dense constraints on both sides of the surface, thereby enhancing the perceptual accuracy of the surface. Experiments suggest that our method achieves SOTA for both occupancy prediction and 3D scene reconstruction tasks on the nuScenes dataset.
Lizhe Liu, Bohua Wang, Hongwei Xie, Daqi Liu, Kuiyuan Yang
CVPR7
2024 Sfnet: Faster and Accurate Semantic Segmentation Via Semantic Flow
abstract
Abstract In this paper, we focus on exploring effective methods for faster and accurate semantic segmentation. A common practice to improve the performance is to attain high-resolution feature maps with strong semantic representation. Two strategies are widely used: atrous convolutions and feature pyramid fusion, while both are either computationally intensive or ineffective. Inspired by the Optical Flow for motion alignment between adjacent video frames, we propose a Flow Alignment Module (FAM) to learn Semantic Flow between feature maps of adjacent levels and broadcast high-level features to high-resolution features effectively and efficiently. Furthermore, integrating our FAM to a standard feature pyramid structure exhibits superior performance over other real-time methods, even on lightweight backbone networks, such as ResNet-18 and DFNet. Then to further speed up the inference procedure, we also present a novel Gated Dual Flow Alignment Module to directly align high-resolution feature maps and low-resolution feature maps where we term the improved version network as SFNet-Lite. Extensive experiments are conducted on several challenging datasets, where results show the effectiveness of both SFNet and SFNet-Lite. In particular, when using Cityscapes test set, the SFNet-Lite series achieve 80.1 mIoU while running at 60 FPS using ResNet-18 backbone and 78.8 mIoU while running at 120 FPS using STDC backbone on RTX-3090. Moreover, we unify four challenging driving datasets (i.e., Cityscapes, Mapillary, IDD, and BDD) into one large dataset, which we named Unified Driving Segmentation (UDS) dataset. It contains diverse domain and style information. We benchmark several representative works on UDS. Both SFNet and SFNet-Lite still achieve the best speed and accuracy trade-off on UDS, which serves as a strong baseline in such a challenging setting. The code and models are publicly available at https://github.com/lxtGH/SFSegNets .
Xiangtai Li, Jiangning Zhang, Kuiyuan Yang, Yunhai Tong, Dacheng Tao
Int. J. Comput. Vis.5
2023 Improving Video Instance Segmentation via Temporal Pyramid Routing
abstract
Video Instance Segmentation (VIS) is a new and inherently multi-task problem, which aims to detect, segment, and track each instance in a video sequence. Existing approaches are mainly based on single-frame features or single-scale features of multiple frames, where either temporal information or multi-scale information is ignored. To incorporate both temporal and scale information, we propose a Temporal Pyramid Routing (TPR) strategy to conditionally align and conduct pixel-level aggregation from a feature pyramid pair of two adjacent frames. Specifically, TPR contains two novel components, including Dynamic Aligned Cell Routing (DACR) and Cross Pyramid Routing (CPR), where DACR is designed for aligning and gating pyramid features across temporal dimension, while CPR transfers temporally aggregated features across scale dimension. Moreover, our approach is a light-weight and plug-and-play module and can be easily applied to existing instance segmentation methods. Extensive experiments on three datasets including YouTube-VIS (2019, 2021) and Cityscapes-VPS demonstrate the effectiveness and efficiency of the proposed approach on several state-of-the-art instance and panoptic segmentation methods. Codes will be publicly available at https://github.com/lxtGH/TemporalPyramidRouting.
Xiangtai Li, Hao He 0015, Henghui Ding, Kuiyuan Yang, Yunhai Tong, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 DT-Loc: Monocular Visual Localization on HD Vector Map Using Distance Transforms of 2D Semantic Detections
abstract
Localizing a vehicle on a prebuilt HD vector map is a prerequisite for many autonomous driving applications. Existing visual localization approaches usually require a separate local feature layer to function. The separate localization layer suffers from the robustness issue inherited from the local features. Also, it could be difficult to create a feature layer that aligns perfectly with an existing vector map. In this paper, we propose a monocular visual localization method that exploits the vector map directly as the localization layer. The method detects semantic traffic elements from the images and matches them with the vectors in the map. To deal with the harmful problem of false matches, we propose to align the vector map to the distance transforms of the semantic detections, which enables a non-explicit and differentiable data association process. The system is able to achieve centimeter and sub-meter accuracies in lateral and longitudinal directions, respectively.
Chi Zhang 0069, Hao Liu 0007, Kuiyuan Yang, Rui Cai 0002, Zhiwei Li 0006
IROS5
2021 Robust LiDAR Localization on an HD Vector Map without a Separate Localization Layer
abstract
Many autonomous driving applications nowadays come along with a prebuilt vector map for routing and planning purposes. In order to localize on this map, traditional LiDAR localization methods usually require a separate localization layer to function. On one hand, the separate layer occupies large storage and is not convenient to update. On the other hand, the potential of the vector map itself has not been fully exploited by existing methods. In this paper, we present a LiDAR localization system that leverages the vector map directly as the localization layer. A semantic extraction module is developed to match the heterogeneous data between LiDAR measurements and the 3D vector elements. A local map maintenance module is introduced to keep the system function robustly when there are not enough vector matches. The system adopts an optimization-based framework and infers 6-DOF poses. Experiments show that the proposed system is able to achieve centimeter accuracy robustly in both highway and urban environments, without a separate localization layer.
Chi Zhang 0069, Liwen Liu, Zhoupeng Xue, Kuiyuan Yang, Rui Cai 0002, Zhiwei Li 0006
IROS5
2021 AVP-Loc: Surround View Localization and Relocalization Based on HD Vector Map for Automated Valet Parking
abstract
Localization is a crucial prerequisite for automated valet parking, in which a vehicle is required to navigate itself in a GPS-denied parking lot. Traditional visual localization methods usually build a feature map and use it for future localizations. However, the feature map is not robust to changes in illumination, appearance, and viewing perspective. To deal with this issue, we need a more stable map. In this paper, we propose to use the parking lot’s HD vector map directly for localization. The vector representation is ultimately stable but brings challenges in data association as well. To this end, we present a novel data association method to match the surround-view images with the vector map. In addition, we also propose a closed-form relocalization strategy by exploiting distinctive road mark combinations in the vector map. Experiments show that the proposed method is able to achieve centimeter-level localization accuracy in a multi-floor parking lot.
Chi Zhang 0069, Hao Liu 0007, Zhijun Xie, Kuiyuan Yang, Rui Cai 0002, Zhiwei Li 0006
IROS4
2021 Towards Efficient Scene Understanding via Squeeze Reasoning
abstract
Graph-based convolutional model such as non-local block has shown to be effective for strengthening the context modeling ability in convolutional neural networks (CNNs). However, its pixel-wise computational overhead is prohibitive which renders it unsuitable for high resolution imagery. In this paper, we explore the efficiency of context graph reasoning and propose a novel framework called Squeeze Reasoning. Instead of propagating information on the spatial map, we first learn to squeeze the input feature into a channel-wise global vector and perform reasoning within the single vector where the computation cost can be significantly reduced. Specifically, we build the node graph in the vector where each node represents an abstract semantic concept. The refined feature within the same semantic category results to be consistent, which is thus beneficial for downstream tasks. We show that our approach can be modularized as an end-to-end trained block and can be easily plugged into existing networks. Despite its simplicity and being lightweight, the proposed strategy allows us to establish the considerable results on different semantic segmentation datasets and shows significant improvements with respect to strong baselines on various other scene understanding tasks including object detection, instance segmentation and panoptic segmentation. Code is available at https://github.com/lxtGH/SFSegNets.
Xiangtai Li, Xia Li 0005, Ansheng You, Li Zhang 0040, Kuiyuan Yang, Yunhai Tong, Zhouchen Lin
IEEE Trans. Image Process.6
2021 Global Aggregation Then Local Distribution for Scene Parsing
abstract
Modelling long-range contextual relationships is critical for pixel-wise prediction tasks such as semantic segmentation. However, convolutional neural networks (CNNs) are inherently limited to model such dependencies due to the naive structure in its building modules (e.g., local convolution kernel). While recent global aggregation methods are beneficial for long-range structure information modelling, they would oversmooth and bring noise to the regions contain fine details (e.g., boundaries and small objects), which are very much cared in the semantic segmentation task. To alleviate this problem, we propose to explore the local context for making the aggregated long-range relationship being distributed more accurately in local regions. In particular, we design a novel local distribution module which models the affinity map between global and local relationship for each pixel adaptively. Integrating existing global aggregation modules, we show that our approach can be modularized as an end-to-end trainable block and easily plugged into existing semantic segmentation networks, giving rise to the GALD networks. Despite its simplicity and versatility, our approach allows us to build new state of the art on major semantic segmentation benchmarks including Cityscapes, ADE20K, Pascal Context, Camvid and COCO-stuff. Code and trained models are released at https://github.com/lxtGH/GALD-DGCNet to foster further research.
Xiangtai Li, Li Zhang 0040, Kuiyuan Yang, Yunhai Tong, Xiatian Zhu, Tao Xiang 0002
IEEE Trans. Image Process.4
2020 Adaptive Unimodal Cost Volume Filtering for Deep Stereo Matching
abstract
State-of-the-art deep learning based stereo matching approaches treat disparity estimation as a regression problem, where loss function is directly defined on true disparities and their estimated ones. However, disparity is just a byproduct of a matching process modeled by cost volume, while indirectly learning cost volume driven by disparity regression is prone to overfitting since the cost volume is under constrained. In this paper, we propose to directly add constraints to the cost volume by filtering cost volume with unimodal distribution peaked at true disparities. In addition, variances of the unimodal distributions for each pixel are estimated to explicitly model matching uncertainty under different contexts. The proposed architecture achieves state-of-the-art performance on Scene Flow and two KITTI stereo benchmarks. In particular, our method ranked the 1st place of KITTI 2012 evaluation and the 4th place of KITTI 2015 evaluation (recorded on 2019.8.20). The codes of AcfNet are available at: https://github.com/youmi-zym/AcfNet.
Youmin Zhang 0005, Xiao Bai 0001, Suihanjin Yu, Zhiwei Li 0006, Kuiyuan Yang
AAAI7
2020 Gated Fully Fusion for Semantic Segmentation
abstract
Semantic segmentation generates comprehensive understanding of scenes through densely predicting the category for each pixel. High-level features from Deep Convolutional Neural Networks already demonstrate their effectiveness in semantic segmentation tasks, however the coarse resolution of high-level features often leads to inferior results for small/thin objects where detailed information is important. It is natural to consider importing low level features to compensate for the lost detailed information in high-level features. Unfortunately, simply combining multi-level features suffers from the semantic gap among them. In this paper, we propose a new architecture, named Gated Fully Fusion(GFF), to selectively fuse features from multiple levels using gates in a fully connected way. Specifically, features at each level are enhanced by higher-level features with stronger semantics and lower-level features with more details, and gates are used to control the propagation of useful information which significantly reduces the noises during fusion. We achieve the state of the art results on four challenging scene parsing datasets including Cityscapes, Pascal Context, COCO-stuff and ADE20K.
Xiangtai Li, Houlong Zhao, Yunhai Tong, Shaohua Tan, Kuiyuan Yang
AAAI6
2020 Semantic Flow for Fast and Accurate Scene Parsing
Xiangtai Li, Ansheng You, Zhen Zhu 0006, Houlong Zhao, Maoke Yang, Kuiyuan Yang, Shaohua Tan, Yunhai Tong
ECCV (1)6
2020 Feature-Metric Loss for Self-supervised Learning of Depth and Egomotion
Zhixiang Duan, Kuiyuan Yang
ECCV (19)4
2020 Category-Aware Spatial Constraint for Weakly Supervised Detection
abstract
Weakly supervised object detection has attracted increasing research attention recently. To this end, most existing schemes rely on scoring category-independent region proposals, which is formulated as a multiple instance learning problem. During this process, the proposal scores are aggregated and supervised by only image-level labels, which often fails to locate object boundaries precisely. In this paper, we break through such a restriction by taking a deeper look into the score aggregation stage and propose a Category-aware Spatial Constraint (CSC) scheme for proposals, which is integrated into weakly supervised object detection in an end-to-end learning manner. In particular, we incorporate the global shape information of objects as an unsupervised constraint, which is inferred from build-in foreground-and-background cues, termed Category-specific Pixel Gradient (CPG) maps. Specifically, each region proposal is weighted according to how well it covers the estimated shape of objects. For each category, a multi-center regularization is further introduced to penalize the violations between centers cluster and high-score proposals in a given image. Extensive experiments are done on the most widely-used benchmark Pascal VOC and COCO, which shows that our approach significantly improves weakly supervised object detection without adding new learnable parameters to the existing models nor changing the structures of CNNs.
Yunhang Shen, Rongrong Ji, Kuiyuan Yang, Cheng Deng 0002, Changhu Wang
IEEE Trans. Image Process.3
2019 Global Aggregation then Local Distribution in Fully Convolutional Networks
Xiangtai Li, Li Zhang 0040, Ansheng You, Maoke Yang, Kuiyuan Yang, Yunhai Tong
BMVC5
2019 Dual Graph Convolutional Network for Semantic Segmentation
Li Zhang 0040, Xiangtai Li, Anurag Arnab, Kuiyuan Yang, Yunhai Tong, Philip Torr 0001
BMVC4
2019 Flow2Seg: Motion-Aided Semantic Segmentation
Xiangtai Li, Jiangang Bai, Kuiyuan Yang, Yunhai Tong
ICANN (3)3
2019 Filter-in-Filter: Low Cost CNN Improvement by Sub-filter Parameter Sharing
Guotian Xie, Kuiyuan Yang, Jian-Huang Lai
Pattern Recognit.2
2019 Balanced Decoupled Spatial Convolution for CNNs
abstract
In this paper, we are interested in designing lightweight CNNs by decoupling the convolution along the spatial and channel dimension. Most existing decoupling techniques focus on approximating the filter matrix through decomposition. In contrast, we provide a decoupled view of the standard convolution to separate the spatial information and the channel information. The resulting decoupled process is exactly equivalent to the standard convolution. Inspired from our decoupled view, we propose an effective structure, balanced decoupled spatial convolution (BDSC), to relax the sparsity of the filter in spatial aggregation by learning a spatial configuration and reduce the redundancy by reducing the number of intermediate channels. We also designed an adaptive spatial configuration, which is simply adding a nonlinear activation layer [rectified linear units (ReLU)] after the intermediate output. Our experiments verify that the adaptive spatial configuration can improve the classification performance without extra cost. In addition, our BDSC achieves comparable classification performance with the standard convolution but with a smaller model size on Canadian Institute for Advanced Research (CIFAR)-100, CIFAR-10, and ImageNet. To show the potential of further reducing the redundancy of across channel-domain convolution, we also show experiments of our models with a designed lightweight across channel-domain convolution. Finally, we show in our experiments that our models achieve superior performance than the state-of-the-art models.
Guotian Xie, Kuiyuan Yang, Ting Zhang 0002, Jingdong Wang 0001, Jian-Huang Lai
IEEE Trans. Neural Networks Learn. Syst.2
2018 Decoupled Convolutions for CNNs
abstract
In this paper, we are interested in designing small CNNs by decoupling the convolution along the spatial and channel domains. Most existing decoupling techniques focus on approximating the filter matrix through decomposition. In contrast, we provide a two-step interpretation of the standard convolution from the filter at a single location to all locations, which is exactly equivalent to the standard convolution. Motivated by the observations in our decoupling view, we propose an effective approach to relax the sparsity of the filter in spatial aggregation by learning a spatial configuration, and reduce the redundancy by reducing the number of intermediate channels. Our approach achieves comparable classification performance with the standard uncoupled convolution, but with a smaller model size over CIFAR-100, CIFAR-10 and ImageNet.
Guotian Xie, Ting Zhang 0002, Kuiyuan Yang, Jian-Huang Lai, Jingdong Wang 0001
AAAI3
2018 DenseASPP for Semantic Segmentation in Street Scenes
abstract
Semantic image segmentation is a basic street scene understanding task in autonomous driving, where each pixel in a high resolution image is categorized into a set of semantic labels. Unlike other scenarios, objects in autonomous driving scene exhibit very large scale changes, which poses great challenges for high-level feature representation in a sense that multi-scale information must be correctly encoded. To remedy this problem, atrous convolution[14]was introduced to generate features with larger receptive fields without sacrificing spatial resolution. Built upon atrous convolution, Atrous Spatial Pyramid Pooling (ASPP)[2] was proposed to concatenate multiple atrous-convolved features using different dilation rates into a final feature representation. Although ASPP is able to generate multi-scale features, we argue the feature resolution in the scale-axis is not dense enough for the autonomous driving scenario. To this end, we propose Densely connected Atrous Spatial Pyramid Pooling (DenseASPP), which connects a set of atrous convolutional layers in a dense way, such that it generates multi-scale features that not only cover a larger scale range, but also cover that scale range densely, without significantly increasing the model size. We evaluate DenseASPP on the street scene benchmark Cityscapes[4] and achieve state-of-the-art performance.
Maoke Yang, Chi Zhang 0069, Zhiwei Li 0006, Kuiyuan Yang
CVPR5
2018 Hierarchical semantic image matching using CNN feature pyramid
Wei Yu 0004, Xiaoshuai Sun, Kuiyuan Yang, Yong Rui, Hongxun Yao
Comput. Vis. Image Underst.3
2018 Automatic Data Augmentation from Massive Web Images for Deep Visual Recognition
abstract
Large-scale image datasets and deep convolutional neural networks (DCNNs) are the two primary driving forces for the rapid progress in generic object recognition tasks in recent years. While lots of network architectures have been continuously designed to pursue lower error rates, few efforts are devoted to enlarging existing datasets due to high labeling costs and unfair comparison issues. In this article, we aim to achieve lower error rates by augmenting existing datasets in an automatic manner. Our method leverages both the web and DCNN, where the web provides massive images with rich contextual information, and DCNN replaces humans to automatically label images under the guidance of web contextual information. Experiments show that our method can automatically scale up existing datasets significantly from billions of web pages with high accuracy. The performance on object recognition tasks and transfer learning tasks have been significantly improved by using the automatically augmented datasets, which demonstrates that more supervisory information has been automatically gathered from the web. Both the dataset and models trained on the dataset have been made publicly available.
Yalong Bai, Kuiyuan Yang, Tao Mei 0001, Wei-Ying Ma, Tiejun Zhao
ACM Trans. Multim. Comput. Commun. Appl.2
2017 Hard-Aware Deeply Cascaded Embedding
abstract
Riding on the waves of deep neural networks, deep metric learning has achieved promising results in various tasks by using triplet network or Siamese network. Though the basic goal of making images from the same category closer than the ones from different categories is intuitive, it is hard to optimize the objective directly due to the quadratic or cubic sample size. Hard example mining is widely used to solve the problem, which spends the expensive computation on a subset of samples that are considered hard. However, hard is defined relative to a specific model. Then complex models will treat most samples as easy ones and vice versa for simple models, both of which are not good for training. It is difficult to define a model with the just right complexity and choose hard examples adequately as different samples are of diverse hard levels. This motivates us to propose the novel framework named Hard-Aware Deeply Cascaded Embedding(HDC) to ensemble a set of models with different complexities in cascaded manner to mine hard examples at multiple levels. A sample is judged by a series of models with increasing complexities and only updates models that consider the sample as a hard case. The HDC is evaluated on CARS196, CUB-200-2011, Stanford Online Products, VehicleID and DeepFashion datasets, and outperforms state-of-the-art methods by a large margin.
Yuhui Yuan, Kuiyuan Yang, Chao Zhang 0001
ICCV2
2017 Exploiting the complementary strengths of multi-layer CNN features for image retrieval
Wei Yu 0004, Kuiyuan Yang, Hongxun Yao, Xiaoshuai Sun, Pengfei Xu 0001
Neurocomputing2
2016 You Lead, We Exceed: Labor-Free Video Concept Learning by Jointly Exploiting Web Videos and Images
abstract
Video concept learning often requires a large set oftraining samples. In practice, however, acquiring noise-free training labels with sufficient positive examples is very expensive. A plausible solution for training data collection is by sampling from the vast quantities of images and videos on the Web. Such a solution is motivated by the assumption that the retrieved images or videos are highly correlated with the query. Still, a number ofchallenges remain. First, Web videos are often untrimmed. Thus, only parts of the videos are relevant to the query. Second, the retrieved Web images are always highly relevant to the issued query. However, thoughtlessly utilizing the images in the video domain may even hurt the performance due to the well-known semantic drift and domain gap problems. As a result, a valid question is how Web images and videos interact for video concept learning. In this paper, we propose a Lead-Exceed Neural Network (LENN), which reinforces the training on Web images and videos in a curriculum manner. Specifically, the training proceeds by inputting frames of Web videos to obtain a network. The Web images are then filtered by the learnt network and the selected images are additionally fed into the network to enhance the architecture and further trim the videos. In addition, Long Short-Term Memory (LSTM) can be applied on the trimmed videos to explore temporal information. Encouraging results are reported on UCFIOl, TRECVID 2013 and 2014 MEDTest in the context ofboth action recognition and event detection. Without using human annotated exemplars, our proposed LENN can achieve 74.4% accuracy on UCFIOI dataset.
Chuang Gan 0001, Ting Yao 0003, Kuiyuan Yang, Yi Yang 0001, Tao Mei 0001
CVPR3
2016 From Seed Discovery to Deep Reconstruction: Predicting Saliency in Crowd via Deep Networks
abstract
Although saliency prediction in crowd has been recently recognized as an essential task for video analysis, it is not comprehensively explored yet. The challenges lie in that eye fixations in crowded scenes are inherently "distinct" and "multi-modal", which differs from those in regular scenes. To this end, the existing saliency prediction schemes typically rely on hand designed features with shallow learning paradigm, which neglect the underlying characteristics of crowded scenes. In this paper, we propose a saliency prediction model dedicated for crowd videos with two novelties: 1) Distinct units are discovered using deep representation learned by a Stacked Denoising Auto-Encoder (SDAE), considering perceptual properties of crowd saliency; 2) Contrast-based saliency is measured through deep reconstruction errors in the second SDAE trained on all units excluding distinct units. A unified model is integrated for online processing crowd saliency. Extensive evaluations on two crowd video benchmark datasets demonstrate that our approach can effectively explore crowd saliency mechanism in two-stage SDAEs and achieve significantly better results than state-of-the-art methods, with robustness to parameters.
Yanhao Zhang 0001, Qingming Huang, Kuiyuan Yang, Jun Zhang 0017, Hongxun Yao
ACM Multimedia4
2015 Exploiting Low-rank Structure for Discriminative Sub-categorization
abstract
In visual recognition, sub-categorization has been proposed to deal with large intraclass variance of samples in a category. Instead of learning a single classifier for each category, discriminant sub-categorization approaches divide a category into several subcategories and simultaneously train classifiers for each sub-category. In this paper, we propose a novel approach for discriminative sub-categorization. Our method jointly trains the exemplar classifier for each positive sample to address the intra-variance of a category and exploits the low rank structure to preserve common information while discovering sub-categories. We formulate the problem as a convex objective function and introduce an efficient solver based on alternating direction method of multipliers. Comprehensive experiments on various datasets demonstrate the effectiveness and efficiency of the proposed method in both sub-category discovery and visual recognition.
Zheng Xu 0002, Kuiyuan Yang, Tom Goldstein
BMVC3
2015 The application of two-level attention models in deep convolutional neural network for fine-grained image classification
abstract
Fine-grained classification is challenging because categories can only be discriminated by subtle and local differences. Variances in the pose, scale or rotation usually make the problem more difficult. Most fine-grained classification systems follow the pipeline of finding foreground object or object parts (where) to extract discriminative features (what). In this paper, we propose to apply visual attention to fine-grained classification task using deep neural network. Our pipeline integrates three types of attention: the bottom-up attention that propose candidate patches, the object-level top-down attention that selects relevant patches to a certain object, and the part-level top-down attention that localizes discriminative parts. We combine these attentions to train domain-specific deep nets, then use it to improve both the what and where aspects. Importantly, we avoid using expensive annotations like bounding box or part information from end-to-end. The weak supervision constraint makes our work easier to generalize. We have verified the effectiveness of the method on the subsets of ILSVRC2012 dataset and CUB200 2011 dataset. Our pipeline delivered significant improvements and achieved the best accuracy under the weakest supervision condition. The performance is competitive against other methods that rely on additional annotations.
Tianjun Xiao, Yichong Xu, Kuiyuan Yang, Yuxin Peng 0001, Zheng Zhang 0001
CVPR3
2015 Automatic Image Dataset Construction from Click-through Logs Using Deep Neural Network
abstract
Labelled image datasets are the backbone for high-level image understanding tasks with wide application scenarios, and continuously drive and evaluate the progress of feature designing and supervised learning models. Recently, the million scale labelled image dataset further contributes to the rebirth of deep convolutional neural network and bypass manual designing handcraft features. However, the construction process of image dataset is mainly manual-based and quite labor intensive, which often take years' efforts to construct a million scale dataset with high quality. In this paper, we propose a deep learning based method to construct large scale image dataset in an automatic way. Specifically, word representation and image representation are learned in a deep neural network from large amount of click-through logs, and further used to define word-word similarity and image-word similarity. These two similarities are used to automatize the two labor intensive steps in manual-based image dataset construction: query formation and noisy image removal. With a new proposed cross convolutional filter regularizer, we can construct a million scale image dataset in one week. Finally, two image datasets are constructed to verify the effectiveness of the method. In addition to scale, the automatically constructed dataset has comparable accuracy, diversity and cross-dataset generalization with manually labelled image datasets.
Yalong Bai, Kuiyuan Yang, Wei Yu 0004, Chang Xu 0008, Wei-Ying Ma, Tiejun Zhao
ACM Multimedia2
2015 Tagging Personal Photos with Transfer Deep Learning
abstract
The advent of mobile devices and media cloud services has led to the unprecedented growing of personal photo collections. One of the fundamental problems in managing the increasing number of photos is automatic image tagging. Existing research has predominantly focused on tagging general Web images with a well-labelled image database, e.g., ImageNet. However, they can only achieve limited success on personal photos due to the domain gaps between personal photos and Web images. These gaps originate from the differences in semantic distribution and visual appearance. To deal with these challenges, in this paper, we present a novel transfer deep learning approach to tag personal photos. Specifically, to solve the semantic distribution gap, we have designed an ontology consisting of a hierarchical vocabulary tailored for personal photos. This ontology is mined from $10,000$ active users in Flickr with 20 million photos and 2.7 million unique tags. To deal with the visual appearance gap, we discover the intermediate image representations and ontology priors by deep learning with bottom-up and top-down transfers across two domains, where Web images are the source domain and personal photos are the target. Moreover, we present two modes (single and batch-modes) in tagging and find that the batch-mode is highly effective to tag photo collections. We conducted personal photo tagging on 7,000 real personal photos and personal photo search on the MIT-Adobe FiveK photo dataset. The proposed tagging approach is able to achieve a performance gain of $12.8\%$ and $4.5\%$ in terms of [email protected], against the state-of-the-art hand-crafted feature-based and deep learning-based methods, respectively.
Jianlong Fu, Tao Mei 0001, Kuiyuan Yang, Hanqing Lu, Yong Rui
WWW3
2015 Mining Latent Attributes From Click-Through Logs for Image Recognition
abstract
Attribute-based image representation, which represents an image by projecting it into a space spanned by attributes, has attracted increasing attention from both computer vision and multimedia communities for its compactness and potential to bridge the semantic gap. While many works focus on learning attribute models and utilizing them in image recognition and retrieval, few touch on the problem of how to effectively construct a vocabulary of attributes, which is an essential part of effective attribute-based representation. Most existing approaches define the attribute vocabulary by human experts or through existing ontology, which is often limited in coverage of general concept space. In this paper, we propose automatically constructing the attribute vocabulary by mining latent topics from the click-through log of a commercial image search engine. These attributes are referred to as latent topic attributes (LTA), which take advantage of tens of millions of interactions between user submitted queries and images, thereby providing better coverage for the concept space than existing approaches. The mining of latent topics from the click log is formulated as a matrix factorization problem, and further improved by weighted terms-based matrix factorization to address the extreme sparsity of the click-through matrix. Both qualitative results of the mined LTA and quantitative results on the standard image recognition benchmark demonstrate the mined LTA's effectiveness.
Yi-Jie Lu, Linjun Yang, Kuiyuan Yang, Yong Rui
IEEE Trans. Multim.3
2015 Query-Dependent Aesthetic Model With Deep Learning for Photo Quality Assessment
abstract
The automatic assessment of photo quality from an aesthetic perspective is a very challenging problem. Most existing research has predominantly focused on the learning of a universal aesthetic model based on hand-crafted visual descriptors . However, this research paradigm can achieve only limited success because (1) such hand-crafted descriptors cannot well preserve abstract aesthetic properties , and (2) such a universal model cannot always capture the full diversity of visual content. To address these challenges, we propose in this paper a novel query-dependent aesthetic model with deep learning for photo quality assessment. In our method, deep aesthetic abstractions are discovered from massive images , whereas the aesthetic assessment model is learned in a query- dependent manner. Our work addresses the first problem by learning mid-level aesthetic feature abstractions via powerful deep convolutional neural networks to automatically capture the underlying aesthetic characteristics of the massive training images . Regarding the second problem, because photographers tend to employ different rules of photography for capturing different images , the aesthetic model should also be query- dependent . Specifically, given an image to be assessed, we first identify which aesthetic model should be applied for this particular image. Then, we build a unique aesthetic model of this type to assess its aesthetic quality. We conducted extensive experiments on two large-scale datasets and demonstrated that the proposed query-dependent model equipped with learned deep aesthetic abstractions significantly and consistently outperforms state-of-the-art hand-crafted feature -based and universal model-based methods.
Xinmei Tian 0001, Kuiyuan Yang, Tao Mei 0001
IEEE Trans. Multim.3
2015 Learning Cross Space Mapping via DNN Using Large Scale Click-Through Logs
abstract
The gap between low-level visual signals and high-level semantics has been progressively bridged by continuous development of deep neural network (DNN). With recent progress of DNN, almost all image classification tasks have achieved new records of accuracy. To extend the ability of DNN to image retrieval tasks, we proposed a unified DNN model for image-query similarity calculation by simultaneously modeling image and query in one network. The unified DNN is named the cross space mapping (CSM) model, which contains two parts, a convolutional part and a query-embedding part. The image and query are mapped to a common vector space via these two parts respectively, and image-query similarity is naturally defined as an inner product of their mappings in the space. To ensure good generalization ability of the DNN, we learn weights of the DNN from a large number of click-through logs which consists of 23 million clicked image-query pairs between 1 million images and 11.7 million queries. Both the qualitative results and quantitative results on an image retrieval evaluation task with 1000 queries demonstrate the superiority of the proposed method.
Wei Yu 0004, Kuiyuan Yang, Yalong Bai, Hongxun Yao, Yong Rui
IEEE Trans. Multim.2
2014 DNN Flow: DNN Feature Pyramid based Image Matching
Wei Yu 0004, Kuiyuan Yang, Yalong Bai, Hongxun Yao, Yong Rui
BMVC2
2014 Bag-of-Words Based Deep Neural Network for Image Retrieval
abstract
This work targets image retrieval task hold by MSR-Bing Grand Challenge. Image retrieval is considered as a challenge task because of the gap between low-level image representation and high-level textual query representation. Recently further developed deep neural network sheds light on narrowing the gap by learning high-level image representation from raw pixels. In this paper, we proposed a bag-of-words based deep neural network for image retrieval task, which learns high-level image representation and maps images into bag-of-words space. The DNN model is trained on the large scale clickthrough data, and the relevance between query and image is measured by the cosine similarity of query's bag-of-words representation and image's bag-of-words representation predicted by DNN, the visual similarity of images is computed by high-level image representation extracted via the DNN model too. Finally, PageRank algorithm is used to further improve the ranking list by considering visual similarity of images for each query. The experimental results achieved state-of-the-art performance and verified the effectiveness of our proposed method.
Yalong Bai, Wei Yu 0004, Tianjun Xiao, Chang Xu 0008, Kuiyuan Yang, Wei-Ying Ma, Tiejun Zhao
ACM Multimedia5
2014 Organizing Video Search Results to Adapted Semantic Hierarchies for Topic-based Browsing
abstract
Organizing video search results into semantically structured hierarchies can greatly improve the efficiency of browsing complex query topics. Traditional hierarchical clustering techniques are inadequate since they lack the ability to generate semantically interpretable structures. In this paper, we introduce an approach to organize video search results to an adapted semantic hierarchy. As many hot search topics such as celebrities and famous cities have Wikipedia pages where hierarchical topic structures are available, we start from the Wikipedia hierarchies and adjust the structures according to the characteristics of the returned videos from a search engine. Ordinary clustering based on textual information of the videos is performed to discover the hidden topic structures in the video search results, which are used to adapt the hierarchy extracted from Wikipedia. After that, a simple optimization problem is formulated to assign the videos to each node of the hierarchy considering three important criteria. Experiments conducted on a Youtube video dataset verify the effectiveness of our approach.
Yu-Gang Jiang 0001, Kuiyuan Yang, Chong-Wah Ngo
ACM Multimedia4
2014 Error-Driven Incremental Learning in Deep Convolutional Neural Network for Large-Scale Image Classification
abstract
Supervised learning using deep convolutional neural network has shown its promise in large-scale image classification task. As a building block, it is now well positioned to be part of a larger system that tackles real-life multimedia tasks. An unresolved issue is that such model is trained on a static snapshot of data. Instead, this paper positions the training as a continuous learning process as new classes of data arrive. A system with such capability is useful in practical scenarios, as it gradually expands its capacity to predict increasing number of new classes. It is also our attempt to address the more fundamental issue: a good learning system must deal with new knowledge that it is exposed to, much as how human do.
Tianjun Xiao, Kuiyuan Yang, Yuxin Peng 0001, Zheng Zhang 0001
ACM Multimedia3
2013 The shortest warping path based multiple images alignment
abstract
In this paper, we propose a method to align multiple images of the same category. Images with large variations are aligned via a smooth transition formed by some intermediate images. These intermediate images are found by shortest warping path algorithm on a directed complete graph. Moreover, the common regions in the images are discovered to further improve alignment performance. The experimental results show that our method is effective to map and align images of the same category but with large variations of appearance, shape and view.
Wei Yu 0004, Hongxun Yao, Kuiyuan Yang, Lei Zhang 0001
ICIP3
2013 Image search by graph-based label propagation with image representation from DNN
abstract
Our objective is to estimate the relevance of an image to a query for image search purposes. We address two limitations of the existing image search engines in this paper. First, there is no straightforward way of bridging the gap between semantic textual queries as well as users' search intents and image visual content. Image search engines therefore primarily rely on static and textual features. Visual features are mainly used to identify potentially useful recurrent patterns or relevant training examples for complementing search by image reranking. Second, image rankers are trained on query-image pairs labeled by human experts, making the annotation intellectually expensive and time-consuming. Furthermore, the labels may be subjective when the queries are ambiguous, resulting in difficulty in predicting the search intention. We demonstrate that the aforementioned two problems can be mitigated by exploring the use of click-through data, which can be viewed as the footprints of user searching behavior, as an effective means of understanding query. The correspondences between an image and a query are determined by whether the image was searched and clicked by users under the query in a commercial image search engine. We therefore hypothesize that the image click counts in response to a query are as their relevance indications. For each new image, our proposed graph-based label propagation algorithm employs neighborhood graph search to find the nearest neighbors on an image similarity graph built up with visual representations from deep neural networks and further aggregates their clicked queries/click counts to get the labels of the new image. We conduct experiments on MSR-Bing Grand Challenge and the results show consistent performance gain over various baselines. In addition, the proposed approach is very efficient, completing annotation of each query-image pair within just 15 milliseconds on a regular PC.
Yingwei Pan, Ting Yao 0003, Kuiyuan Yang, Houqiang Li, Chong-Wah Ngo, Jingdong Wang 0001, Tao Mei 0001
ACM Multimedia3
2012 Constructing visual tag dictionary by mining community-contributed media corpus
Meng Wang 0001, Kuiyuan Yang
Neurocomputing2
2011 Semantic point detector
abstract
Local features are the building blocks of many visual systems, and local point detector is usually the first component for local feature extraction. Existing local point detector are designed with target for matching and it may not perform well when applied in image content representation. Actually many existing studies demonstrate that the simple dense sampling strategy can achieve better performance than many local point detection methods in image classification tasks. In this paper, we propose a novel point detector named semantic point detector, which detects a set of semantically meaningful patches from each image and yields more compact and complete image representation. It is learned from an set of images with concepts from a large ontology. We conduct extensive experiments based on the proposed detector, and the experimental results demonstrate the effectiveness of our approach.
Kuiyuan Yang, Lei Zhang 0001, Meng Wang 0001, HongJiang Zhang
ACM Multimedia1
2011 Assemble New Object Detector With Few Examples
abstract
Learning a satisfactory object detector generally requires sufficient training data to cover the most variations of the object. In this paper, we show that the performance of object detector is severely degraded when training examples are limited. We propose an approach to handle this issue by exploring a set of pretrained auxiliary detectors for other categories. By mining the global and local relationships between the target object category and auxiliary objects, a robust detector can be learned with very few training examples. We adopt the deformable part model proposed by Felzenszwalb and simultaneously explore the root and part filters in the auxiliary object detectors under the guidance of the few training examples from the target object category. An iterative solution is introduced for such a process. The extensive experiments on the PASCAL VOC 2007 challenge data set show the encouraging performance of the new detector assembled from those related auxiliary detectors.
Kuiyuan Yang, Meng Wang 0001, Xian-Sheng Hua 0001, Shuicheng Yan, HongJiang Zhang
IEEE Trans. Image Process.1
2011 Tag Tagging: Towards More Descriptive Keywords of Image Content
abstract
Tags have been demonstrated to be effective and efficient for organizing and searching social image content. However, these human-provided keywords are far from a comprehensive description of the image content, which limits their effectiveness in tag-based image search. In this paper, we propose an automatic scheme called tag tagging to supplement semantic image descriptions by associating a group of property tags with each existing tag. For example, an initial tag “tiger” may be further tagged with “white”, “stripes”, and “bottom-right” along three tag properties: color, texture, and location, respectively. In this way, the descriptive ability of the existing tags can be greatly enhanced. In the proposed scheme, a lazy learning approach is first applied to estimate the corresponding image regions of each initial tag, and then a set of property tags that correspond to six properties, including location, color, texture, size, shape, and dominance, are derived for each initial tag. These tag properties enable much more precise image search especially when certain tag properties are included in the query. The results of the empirical evaluation show that tag properties remarkably boost the performance of social image retrieval.
Kuiyuan Yang, Xian-Sheng Hua 0001, Meng Wang 0001, HongJiang Zhang
IEEE Trans. Multim.1
2010 Tagging tags
abstract
Social image sharing websites like Flickr have successfully motivated users around the world to annotate images with tags, which greatly facilitate search and organization of social image content. However, these manually-input tags are far from a comprehensive description of the image content, which limits effectiveness of the tags in content-based image search. In this paper, we propose an automatic scheme called tagging tags to supplement semantic image descriptions by associating a group of property tags with each existing tag. For example, an initial tag "tiger" will be further tagged with "white", "stripes" and "bottom-right" along three tag properties: color, texture and location, respectively. In the proposed scheme, a lazy learning approach is first applied to estimate the corresponding image regions of each initial tag, and then a set of property tags, which involve six exemplary property aspects including location, color, texture, shape, size and dominance, are derived for each tag according to the content of the regions and the entire image. These tag properties enable much more precise image search especially when certain tag properties are included in the query. The results of the empirical evaluation show that tag properties remarkably boost the performance of social image retrieval.
Kuiyuan Yang, Xian-Sheng Hua 0001, Meng Wang 0001, HongJiang Zhang
ACM Multimedia1
2010 Social Image Search with Diverse Relevance Ranking
Kuiyuan Yang, Meng Wang 0001, Xian-Sheng Hua 0001, HongJiang Zhang
MMM1
2010 Towards a Relevant and Diverse Search of Social Images
abstract
Recent years have witnessed the great success of social media websites. Tag-based image search is an important approach to accessing the image content on these websites. However, the existing ranking methods for tag-based image search frequently return results that are irrelevant or not diverse. This paper proposes a diverse relevance ranking scheme that is able to take relevance and diversity into account by exploring the content of images and their associated tags. First, it estimates the relevance scores of images with respect to the query term based on both the visual information of images and the semantic information of associated tags. Then, we estimate the semantic similarities of social images based on their tags. Based on the relevance scores and the similarities, the ranking list is generated by a greedy ordering algorithm which optimizes average diverse precision, a novel measure that is extended from the conventional average precision. Comprehensive experiments and user studies demonstrate the effectiveness of the approach. We also apply the scheme for web image search reranking, and it is shown that the diversity of search results can be enhanced while maintaining a comparable level of relevance.
Meng Wang 0001, Kuiyuan Yang, Xian-Sheng Hua 0001, HongJiang Zhang
IEEE Trans. Multim.2
2009 Active tagging for image indexing
abstract
Concept labeling and ontology-free tagging are the two typical manners of image annotation. Despite extensive research efforts have been dedicated to labeling, currently automatic image labeling algorithms are still far from satisfactory, and meanwhile manual labeling is rather labor-intensive. In contrast with labeling, tagging works in a free way and therefore it has better user experience for annotators. In this paper, we introduce an active tagging scheme that combines human and computer to assign tags to images. The scheme works in an iterative way. In each round, the most informative images are selected for manual tagging, and the remained images can be annotated by a tag prediction component. We have integrated multiple criteria for sample selection, including ambiguity, citation, and diversity. Experiments are conducted on different datasets and empirical results have demonstrated the effectiveness of the proposed approach.
Kuiyuan Yang, Meng Wang 0001, HongJiang Zhang
ICME1