EDBT 2026 Demo / reviewers in the wild / expert
Yuanqing Lin
dblp:97/2644
· DBLP profile ↗
37ranked-venue papers
8as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-authorSystems, architecture and hardware · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
21 papers |
Image recognition and object detection · 36% Representation and self-supervised learning · 18% 3D vision · 14% | |
| Databases, data mining, and information retrieval
6 papers |
Information retrieval · 95% Indexing and storage engines · 3% Machine learning and data management · 2% | |
| Computer graphics and multimedia
3 papers |
Multimedia analysis and retrieval · 56% Image and video processing · 24% Audio and music processing · 20% |
Topics — the 30 heaviest of 56, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection › image classification
fine-grained image classification |
1.8 | 7 | 2017 | Kernel Pooling for Convolutional Neural Networks · CVPR 2017 Localizing by Describing: Attribute-Guided Attention Localization for Fine-Grained Recognition · AAAI 2017 Fine-Grained Image Classification by Exploring Bipartite-Graph Labels · CVPR 2016 |
Information retrieval
image retrieval |
0.9 | 5 | 2016 | Scalable Feature Matching by Dual Cascaded Scalar Quantization for Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2016 Cross Indexing With Grouplets · IEEE Trans. Multim. 2015 Semantic-Aware Co-Indexing for Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Computer vision › Image recognition and object detection
object detection |
0.8 | 5 | 2016 | Exploit All the Layers: Fast and Accurate CNN Object Detector with Scale Dependent Pooling and Cascaded Rejection Classifiers · CVPR 2016 Regionlets for Generic Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2015 Regionlets for Generic Object Detection · ICCV 2013 |
Computer vision › Face, body and person analysis
face recognition |
0.6 | 1 | 2022 | Improving Federated Learning Face Recognition via Privacy-Agnostic Clusters · ICLR 2022 |
Machine learning › Efficient and distributed learning
federated learning |
0.6 | 1 | 2022 | Improving Federated Learning Face Recognition via Privacy-Agnostic Clusters · ICLR 2022 |
Machine learning › Representation and self-supervised learning › representation learning › metric learning
deep metric learning |
0.5 | 2 | 2017 | Deep Metric Learning with Angular Loss · ICCV 2017 Fine-Grained Categorization and Dataset Bootstrapping Using Deep Metric Learning with Humans in the Loop · CVPR 2016 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.5 | 2 | 2017 | Deep Metric Learning with Angular Loss · ICCV 2017 Fine-grained visual categorization via multi-stage metric learning · CVPR 2015 |
Information retrieval › indexing
inverted index |
0.4 | 2 | 2015 | Semantic-Aware Co-Indexing for Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2015 Semantic-Aware Co-indexing for Image Retrieval · ICCV 2013 |
Computer vision › 3D vision
camera pose estimation |
0.3 | 1 | 2018 | DeLS-3D: Deep Localization and Segmentation With a 3D Semantic Map · CVPR 2018 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.3 | 1 | 2018 | DeLS-3D: Deep Localization and Segmentation With a 3D Semantic Map · CVPR 2018 |
Machine learning › Deep learning architectures and training › neural network layer design
pooling |
0.3 | 1 | 2017 | Kernel Pooling for Convolutional Neural Networks · CVPR 2017 |
Computer vision › 3D vision
object representation |
0.3 | 2 | 2015 | Regionlets for Generic Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2015 Regionlets for Generic Object Detection · ICCV 2013 |
Computer vision › Image recognition and object detection › object detection › deep learning object detection
CNN-based detection |
0.2 | 1 | 2016 | Exploit All the Layers: Fast and Accurate CNN Object Detector with Scale Dependent Pooling and Cascaded Rejection Classifiers · CVPR 2016 |
Information retrieval
image matching |
0.2 | 1 | 2016 | Scalable Feature Matching by Dual Cascaded Scalar Quantization for Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2016 |
Information retrieval › image retrieval
large-scale image retrieval |
0.2 | 1 | 2016 | Scalable Feature Matching by Dual Cascaded Scalar Quantization for Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2016 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding |
0.2 | 2 | 2011 | Learning image representations from the pixel level via hierarchical sparse coding · CVPR 2011 Deep Coding Network · NIPS 2010 |
Computer vision › 3D vision
3d object detection |
0.2 | 1 | 2015 | Data-driven 3D Voxel Patterns for object category recognition · CVPR 2015 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.2 | 1 | 2015 | Data-driven 3D Voxel Patterns for object category recognition · CVPR 2015 |
Machine learning › Learning paradigms
multi-task learning |
0.2 | 1 | 2015 | Hyper-class augmented and regularized deep learning for fine-grained image classification · CVPR 2015 |
Computer vision › 3D vision
object pose estimation |
0.2 | 1 | 2015 | Data-driven 3D Voxel Patterns for object category recognition · CVPR 2015 |
Machine learning › Deep learning architectures and training
regularization |
0.2 | 1 | 2015 | Hyper-class augmented and regularized deep learning for fine-grained image classification · CVPR 2015 |
Information retrieval › image retrieval
content-based image retrieval |
0.2 | 1 | 2015 | Semantic-Aware Co-Indexing for Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Information retrieval › image retrieval
image indexing |
0.2 | 1 | 2015 | Cross Indexing With Grouplets · IEEE Trans. Multim. 2015 |
Information retrieval
indexing |
0.2 | 1 | 2015 | Semantic-Aware Co-Indexing for Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Image and video processing › image matching
feature matching |
0.2 | 1 | 2014 | Towards Codebook-Free: Scalable Cascaded Hashing for Mobile Image Search · IEEE Trans. Multim. 2014 |
Multimedia analysis and retrieval
image retrieval |
0.2 | 1 | 2014 | Towards Codebook-Free: Scalable Cascaded Hashing for Mobile Image Search · IEEE Trans. Multim. 2014 |
Multimedia analysis and retrieval › visual search
mobile visual search |
0.2 | 1 | 2014 | Towards Codebook-Free: Scalable Cascaded Hashing for Mobile Image Search · IEEE Trans. Multim. 2014 |
Privacy and data protection
privacy-preserving machine learning |
0.2 | 1 | 2022 | Improving Federated Learning Face Recognition via Privacy-Agnostic Clusters · ICLR 2022 |
Computer vision › 3D vision
3d reconstruction |
0.2 | 1 | 2013 | Dense Object Reconstruction with Semantic Priors · CVPR 2013 |
Computer vision › 3D vision › 3d reconstruction
multi-view stereo |
0.2 | 1 | 2013 | Dense Object Reconstruction with Semantic Priors · CVPR 2013 |
Methods — techniques the papers use, named apart from their topics
convolutional neural network · 1.3federated averaging · 1.1clustering · 1.1triplet loss · 0.8multi-task learning · 0.7semantic attributes · 0.4cascaded boosting classifier · 0.4sensor fusion · 0.3recurrent neural network · 0.3reward strategy · 0.3reinforcement learning · 0.3range-based neighbor search · 0.2label structures · 0.2cascaded scalar quantization · 0.2PCA · 0.2multi-stage optimization · 0.2distance metric learning · 0.2co-indexing · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Improving Federated Learning Face Recognition via Privacy-Agnostic Clusters
Feng Zhou 0002, Hainan Ren, Tianshu Feng, Guochao Liu, Yuanqing Lin |
ICLR | 6 |
| 2018 | DeLS-3D: Deep Localization and Segmentation With a 3D Semantic MapabstractFor applications such as augmented reality, autonomous driving, self-localization/camera pose estimation and scene parsing are crucial technologies. In this paper, we propose a unified framework to tackle these two problems simultaneously. The uniqueness of our design is a sensor fusion scheme which integrates camera videos, motion sensors (GPS/IMU), and a 3D semantic map in order to achieve robustness and efficiency of the system. Specifically, we first have an initial coarse camera pose obtained from consumer-grade GPS/IMU, based on which a label map can be rendered from the 3D semantic map. Then, the rendered label map and the RGB image are jointly fed into a pose CNN, yielding a corrected camera pose. In addition, to incorporate temporal information, a multi-layer recurrent neural network (RNN) is further deployed improve the pose accuracy. Finally, based on the pose from RNN, we render a new label map, which is fed together with the RGB image into a segment CNN which produces perpixel semantic label. In order to validate our approach, we build a dataset with registered 3D point clouds and video camera images. Both the point clouds and the images are semantically-labeled. Each video frame has ground truth pose from highly accurate motion sensors. We show that practically, pose estimation solely relying on images like PoseNet [25] may fail due to street view confusion, and it is important to fuse multiple sensors. Finally, various ablation studies are performed, which demonstrate the effectiveness of the proposed system. In particular, we show that scene parsing and pose estimation are mutually beneficial to achieve a more robust and accurate system. Peng Wang 0001, Ruigang Yang, Binbin Cao, Wei Xu 0017, Yuanqing Lin |
CVPR | 5 |
| 2017 | Localizing by Describing: Attribute-Guided Attention Localization for Fine-Grained RecognitionabstractA key challenge in fine-grained recognition is how to find and represent discriminative local regions.Recent attention models are capable of learning discriminative region localizers only from category labels with reinforcement learning. However, not utilizing any explicit part information, they are not able to accurately find multiple distinctive regions.In this work, we introduce an attribute-guided attention localization scheme where the local region localizers are learned under the guidance of part attribute descriptions.By designing a novel reward strategy, we are able to learn to locate regions that are spatially and semantically distinctive with reinforcement learning algorithm. The attribute labeling requirement of the scheme is more amenable than the accurate part location annotation required by traditional part-based fine-grained recognition methods.Experimental results on the CUB-200-2011 dataset demonstrate the superiority of the proposed scheme on both fine-grained recognition and attribute recognition. Xiao Liu 0022, Shilei Wen, Errui Ding, Yuanqing Lin |
AAAI | 5 |
| 2017 | Kernel Pooling for Convolutional Neural NetworksabstractConvolutional Neural Networks (CNNs) with Bilinear Pooling, initially in their full form and later using compact representations, have yielded impressive performance gains on a wide range of visual tasks, including fine-grained visual categorization, visual question answering, face recognition, and description of texture and style. The key to their success lies in the spatially invariant modeling of pairwise (2ndorder) feature interactions. In this work, we propose a general pooling framework that captures higher order interactions of features in the form of kernels. We demonstrate how to approximate kernels such as Gaussian RBF up to a given order using compact explicit feature maps in a parameter-free manner. Combined with CNNs, the composition of the kernel can be learned from data in an end-to-end fashion via error back-propagation. The proposed kernel pooling scheme is evaluated in terms of both kernel approximation error and visual recognition accuracy. Experimental evaluations demonstrate state-of-the-art performance on commonly used fine-grained recognition datasets. Yin Cui, Feng Zhou 0002, Xiao Liu 0022, Yuanqing Lin, Serge J. Belongie |
CVPR | 5 |
| 2017 | Deep Metric Learning with Angular LossabstractThe modern image search system requires semantic understanding of image, and a key yet under-addressed problem is to learn a good metric for measuring the similarity between images. While deep metric learning has yielded impressive performance gains by extracting high level abstractions from image data, a proper objective loss function becomes the central issue to boost the performance. In this paper, we propose a novel angular loss, which takes angle relationship into account, for learning better similarity metric. Whereas previous metric learning methods focus on optimizing the similarity (contrastive loss) or relative similarity (triplet loss) of image pairs, our proposed method aims at constraining the angle at the negative point of triplet triangles. Several favorable properties are observed when compared with conventional methods. First, scale invariance is introduced, improving the robustness of objective against feature variance. Second, a third-order geometric constraint is inherently imposed, capturing additional local structure of triplet triangles than contrastive loss or triplet loss. Third, better convergence has been demonstrated by experiments on three publicly available datasets. Feng Zhou 0002, Shilei Wen, Xiao Liu 0022, Yuanqing Lin |
ICCV | 5 |
| 2017 | Subcategory-Aware Convolutional Neural Networks for Object Proposals and DetectionabstractIn Convolutional Neural Network (CNN)-based object detection methods, region proposal becomes a bottleneck when objects exhibit significant scale variation, occlusion or truncation. In addition, these methods mainly focus on 2D object detection and cannot estimate detailed properties of objects. In this paper, we propose subcategory-aware CNNs for object detection. We introduce a novel region proposal network that uses subcategory information to guide the proposal generating process, and a new detection network for joint detection and subcategory classification. By using subcategories related to object pose, we achieve state of-the-art performance on both detection and pose estimation on commonly used benchmarks. Wongun Choi, Yuanqing Lin, Silvio Savarese |
WACV | 3 |
| 2016 | Fine-Grained Categorization and Dataset Bootstrapping Using Deep Metric Learning with Humans in the LoopabstractExisting fine-grained visual categorization methods often suffer from three challenges: lack of training data, large number of fine-grained categories, and high intraclass vs. low inter-class variance. In this work we propose a generic iterative framework for fine-grained categorization and dataset bootstrapping that handles these three challenges. Using deep metric learning with humans in the loop, we learn a low dimensional feature embedding with anchor points on manifolds for each category. These anchor points capture intra-class variances and remain discriminative between classes. In each round, images with high confidence scores from our model are sent to humans for labeling. By comparing with exemplar images, labelers mark each candidate image as either a "true positive" or a "false positive." True positives are added into our current dataset and false positives are regarded as "hard negatives" for our metric learning model. Then the model is retrained with an expanded dataset and hard negatives for the next round. To demonstrate the effectiveness of the proposed framework, we bootstrap a fine-grained flower dataset with 620 categories from Instagram images. The proposed deep metric learning scheme is evaluated on both our dataset and the CUB-200-2001 Birds dataset. Experimental evaluations show significant performance gain using dataset bootstrapping and demonstrate state-of-the-art results achieved by the proposed deep metric learning methods. Yin Cui, Feng Zhou 0002, Yuanqing Lin, Serge J. Belongie |
CVPR | 3 |
| 2016 | Exploit All the Layers: Fast and Accurate CNN Object Detector with Scale Dependent Pooling and Cascaded Rejection ClassifiersabstractIn this paper, we investigate two new strategies to detect objects accurately and efficiently using deep convolutional neural network: 1) scale-dependent pooling and 2) layerwise cascaded rejection classifiers. The scale-dependent pooling (SDP) improves detection accuracy by exploiting appropriate convolutional features depending on the scale of candidate object proposals. The cascaded rejection classifiers (CRC) effectively utilize convolutional features and eliminate negative object proposals in a cascaded manner, which greatly speeds up the detection while maintaining high accuracy. In combination of the two, our method achieves significantly better accuracy compared to other state-of-the-arts in three challenging datasets, PASCAL object detection challenge, KITTI object detection benchmark and newly collected Inner-city dataset, while being more efficient. Wongun Choi, Yuanqing Lin |
CVPR | 3 |
| 2016 | Embedding Label Structures for Fine-Grained Feature RepresentationabstractRecent algorithms in convolutional neural networks (CNN) considerably advance the fine-grained image classification, which aims to differentiate subtle differences among subordinate classes. However, previous studies have rarely focused on learning a fined-grained and structured feature representation that is able to locate similar images at different levels of relevance, e.g., discovering cars from the same make or the same model, both of which require high precision. In this paper, we propose two main contributions to tackle this problem. 1) A multitask learning framework is designed to effectively learn fine-grained feature representations by jointly optimizing both classification and similarity constraints. 2) To model the multi-level relevance, label structures such as hierarchy or shared attributes are seamlessly embedded into the framework by generalizing the triplet loss. Extensive and thorough experiments have been conducted on three finegrained datasets, i.e., the Stanford car, the Car-333, and the food datasets, which contain either hierarchical labels or shared attributes. Our proposed method has achieved very competitive performance, i.e., among state-of-the-art classification accuracy when not using parts. More importantly, it significantly outperforms previous fine-grained feature representations for image retrieval at different levels of relevance. Xiaofan Zhang 0002, Feng Zhou 0002, Yuanqing Lin, Shaoting Zhang 0001 |
CVPR | 3 |
| 2016 | Fine-Grained Image Classification by Exploring Bipartite-Graph LabelsabstractGiven a food image, can a fine-grained object recognition engine tell "which restaurant which dish" the food belongs to? Such ultra-fine grained image recognition is the key for many applications like search by images, but it is very challenging because it needs to discern subtle difference between classes while dealing with the scarcity of training data. Fortunately, the ultra-fine granularity naturally brings rich relationships among object classes. This paper proposes a novel approach to exploit the rich relationships through bipartite-graph labels (BGL). We show how to model BGL in an overall convolutional neural networks and the resulting system can be optimized through back-propagation. We also show that it is computationally efficient in inference thanks to the bipartite structure. To facilitate the study, we construct a new food benchmark dataset, which consists of 37,885 food images collected from 6 restaurants and totally 975 menus. Experimental results on this new food and three other datasets demonstrate BGL advances previous works in fine-grained object recognition. An online demo is available at http: //www.f-zhou.com/fg_demo/. Feng Zhou 0002, Yuanqing Lin |
CVPR | 2 |
| 2016 | Scalable Feature Matching by Dual Cascaded Scalar Quantization for Image RetrievalabstractIn this paper, we investigate the problem of scalable visual feature matching in large-scale image search and propose a novel cascaded scalar quantization scheme in dual resolution. We formulate the visual feature matching as a range-based neighbor search problem and approach it by identifying hyper-cubes with a dual-resolution scalar quantization strategy. Specifically, for each dimension of the PCA-transformed feature, scalar quantization is performed at both coarse and fine resolutions. The scalar quantization results at the coarse resolution are cascaded over multiple dimensions to index an image database. The scalar quantization results over multiple dimensions at the fine resolution are concatenated into a binary super-vector and stored into the index list for efficient verification. The proposed cascaded scalar quantization (CSQ) method is free of the costly visual codebook training and thus is independent of any image descriptor training set. The index structure of the CSQ is flexible enough to accommodate new image features and scalable to index large-scale image database. We evaluate our approach on the public benchmark datasets for large-scale image retrieval. Experimental results demonstrate the competitive retrieval performance of the proposed method compared with several recent retrieval algorithms on feature quantization. Wengang Zhou 0001, Ming Yang 0007, Xiaoyu Wang 0002, Houqiang Li, Yuanqing Lin, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2015 | Fine-grained visual categorization via multi-stage metric learningabstractFine-grained visual categorization (FGVC) is to categorize objects into subordinate classes instead of basic classes. One major challenge in FGVC is the co-occurrence of two issues: 1) many subordinate classes are highly correlated and are difficult to distinguish, and 2) there exists the large intra-class variation (e.g., due to object pose). This paper proposes to explicitly address the above two issues via distance metric learning (DML). DML addresses the first issue by learning an embedding so that data points from the same class will be pulled together while those from different classes should be pushed apart from each other; and it addresses the second issue by allowing the flexibility that only a portion of the neighbors (not all data points) from the same class need to be pulled together. However, feature representation of an image is often high dimensional, and DML is known to have difficulty in dealing with high dimensional feature vectors since it would require O(d2) for storage and O(d3) for optimization. To this end, we proposed a multi-stage metric learning framework that divides the large-scale high dimensional learning problem to a series of simple subproblems, achieving O(d) computational complexity. The empirical study with FVGC benchmark datasets verifies that our method is both effective and efficient compared to the state-of-the-art FGVC approaches. Qi Qian 0001, Rong Jin 0001, Shenghuo Zhu, Yuanqing Lin |
CVPR | 4 |
| 2015 | Data-driven 3D Voxel Patterns for object category recognitionabstractDespite the great progress achieved in recognizing objects as 2D bounding boxes in images, it is still very challenging to detect occluded objects and estimate the 3D properties of multiple objects from a single image. In this paper, we propose a novel object representation, 3D Voxel Pattern (3DVP), that jointly encodes the key properties of objects including appearance, 3D shape, viewpoint, occlusion and truncation. We discover 3DVPs in a data-driven way, and train a bank of specialized detectors for a dictionary of 3DVPs. The 3DVP detectors are capable of detecting objects with specific visibility patterns and transferring the meta-data from the 3DVPs to the detected objects, such as 2D segmentation mask, 3D pose as well as occlusion or truncation boundaries. The transferred meta-data allows us to infer the occlusion relationship among objects, which in turn provides improved object recognition results. Experiments are conducted on the KITTI detection benchmark [17] and the outdoor-scene dataset [41]. We improve state-of-the-art results on car detection and pose estimation with notable margins (6% in difficult data of KITTI). We also verify the ability of our method in accurately segmenting objects from the background and localizing them in 3D. Wongun Choi, Yuanqing Lin, Silvio Savarese |
CVPR | 3 |
| 2015 | Hyper-class augmented and regularized deep learning for fine-grained image classificationabstractDeep convolutional neural networks (CNN) have seen tremendous success in large-scale generic object recognition. In comparison with generic object recognition, fine-grained image classification (FGIC) is much more challenging because (i) fine-grained labeled data is much more expensive to acquire (usually requiring domain expertise); (ii) there exists large intra-class and small inter-class variance. Most recent work exploiting deep CNN for image recognition with small training data adopts a simple strategy: pre-train a deep CNN on a large-scale external dataset (e.g., ImageNet) and fine-tune on the small-scale target data to fit the specific classification task. In this paper, beyond the fine-tuning strategy, we propose a systematic framework of learning a deep CNN that addresses the challenges from two new perspectives: (i) identifying easily annotated hyper-classes inherent in the fine-grained data and acquiring a large number of hyper-class-labeled images from readily available external sources (e.g., image search engines), and formulating the problem into multitask learning; (ii) a novel learning model by exploiting a regularization between the fine-grained recognition model and the hyper-class recognition model. We demonstrate the success of the proposed framework on two small-scale fine-grained datasets (Stanford Dogs and Stanford Cars) and on a large-scale car dataset that we collected. Saining Xie, Tianbao Yang, Xiaoyu Wang 0002, Yuanqing Lin |
CVPR | 4 |
| 2015 | Regionlets for Generic Object DetectionabstractGeneric object detection is confronted by dealing with different degrees of variations, caused by viewpoints or deformations in distinct object classes, with tractable computations. This demands for descriptive and flexible object representations which can be efficiently evaluated in many locations. We propose to model an object class with a cascaded boosting classifier which integrates various types of features from competing local regions, each of which may consist of a group of subregions, named as regionlets. A regionlet is a base feature extraction region defined proportionally to a detection window at an arbitrary resolution (i.e., size and aspect ratio). These regionlets are organized in small groups with stable relative positions to be descriptive to delineate fine-grained spatial layouts inside objects. Their features are aggregated into a one-dimensional feature within one group so as to be flexible to tolerate deformations. The most discriminative regionlets for each object class are selected through a boosting learning procedure. Our regionlet approach achieves very competitive performance on popular multi-class detection benchmark datasets with a single method, without any context. It achieves a detection mean average precision of 41.7 percent on the PASCAL VOC 2007 dataset, and 39.7 percent on the VOC 2010 for 20 object categories. We further develop support pixel integral images to efficiently augment regionlet features with the responses learned by deep convolutional neural networks. Our regionlet based method won second place in the ImageNet Large Scale Visual Object Recognition Challenge (ILSVRC 2013). Xiaoyu Wang 0002, Ming Yang 0007, Shenghuo Zhu, Yuanqing Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Semantic-Aware Co-Indexing for Image RetrievalabstractIn content-based image retrieval, inverted indexes allow fast access to database images and summarize all knowledge about the database. Indexing multiple clues of image contents allows retrieval algorithms search for relevant images from different perspectives, which is appealing to deliver satisfactory user experiences. However, when incorporating diverse image features during online retrieval, it is challenging to ensure retrieval efficiency and scalability. In this paper, for large-scale image retrieval, we propose a semantic-aware co-indexing algorithm to jointly embed two strong cues into the inverted indexes: 1) local invariant features that are robust to delineate low-level image contents, and 2) semantic attributes from large-scale object recognition that may reveal image semantic meanings. Specifically, for an initial set of inverted indexes of local features, we utilize semantic attributes to filter out isolated images and insert semantically similar images to this initial set. Encoding these two distinct and complementary cues together effectively enhances the discriminative capability of inverted indexes. Such co-indexing operations are totally off-line and introduce small computation overhead to online retrieval, because only local features but no semantic attributes are employed for the query. Hence, this co-indexing is different from existing image retrieval methods fusing multiple features or retrieval results. Extensive experiments and comparisons with recent retrieval methods manifest the competitive performance of our method. Shiliang Zhang, Ming Yang 0007, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Cross Indexing With GroupletsabstractMost of the current image indexing systems for retrieval view a database as a set of individual images. It limits the flexibility of the retrieval framework to conduct sophisticated cross-image analysis, resulting in higher memory consumption and sub-optimal retrieval accuracy. To conquer this issue, we propose cross indexing with grouplets, where the core idea is to view the database images as a set of grouplets, each of which is defined as a group of highly relevant images. Because a grouplet groups similar images together, the number of grouplets is smaller than the number of images, thus naturally leading to less memory cost. Moreover, the definition of a grouplet could be based on customized relations, allowing for seamless integration of advanced image features and data mining techniques like the deep convolutional neural network (DCNN) in off-line indexing . To validate the proposed framework, we construct three different types of grouplets , which are respectively based on local similarity , regional relation, and global semantic modeling. Extensive experiments on public benchmark datasets demonstrate the efficiency and superior performance of our approach. Shiliang Zhang, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001 |
IEEE Trans. Multim. | 3 |
| 2014 | Accurate Object Detection with Location Relaxation and Regionlets Re-localization
Chengjiang Long, Xiaoyu Wang 0002, Gang Hua 0001, Ming Yang 0007, Yuanqing Lin |
ACCV (1) | 5 |
| 2014 | Generic Object Detection with Dense Neural Patterns and Regionlets
Will Y. Zou, Yuanqing Lin |
BMVC | 4 |
| 2014 | Towards Codebook-Free: Scalable Cascaded Hashing for Mobile Image SearchabstractState-of-the-art image retrieval algorithms using local invariant features mostly rely on a large visual codebook to accelerate the feature quantization and matching. This codebook typically contains millions of visual words, which not only demands for considerable resources to train offline but also consumes large amount of memory at the online retrieval stage. This is hardly affordable in resource limited scenarios such as mobile image search applications. To address this issue, we propose a codebook-free algorithm for large scale mobile image search. In our method, we first employ a novel scalable cascaded hashing scheme to ensure the recall rate of local feature matching. Afterwards, we enhance the matching precision by an efficient verification with the binary signatures of these local features. Consequently, our method achieves fast and accurate feature matching free of a huge visual codebook. Moreover, the quantization and binarizing functions in the proposed scheme are independent of small collections of training images and generalize well for diverse image datasets. Evaluated on two public datasets with a million distractor images, the proposed algorithm demonstrates competitive retrieval accuracy and scalability against four recent retrieval methods in literature. Wengang Zhou 0001, Ming Yang 0007, Houqiang Li, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001 |
IEEE Trans. Multim. | 5 |
| 2013 | Dense Object Reconstruction with Semantic PriorsabstractWe present a dense reconstruction approach that overcomes the drawbacks of traditional multiview stereo by incorporating semantic information in the form of learned category-level shape priors and object detection. Given training data comprised of 3D scans and images of objects from various viewpoints, we learn a prior comprised of a mean shape and a set of weighted anchor points. The former captures the commonality of shapes across the category, while the latter encodes similarities between instances in the form of appearance and spatial consistency. We propose robust algorithms to match anchor points across instances that enable learning a mean shape for the category, even with large shape variations across instances. We model the shape of an object instance as a warped version of the category mean, along with instance-specific details. Given multiple images of an unseen instance, we collate information from 2D object detectors to align the structure from motion point cloud with the mean shape, which is subsequently warped and refined to approach the actual shape. Extensive experiments demonstrate that our model is general enough to learn semantic priors for different object categories, yet powerful enough to reconstruct individual shapes with large variations. Qualitative and quantitative evaluations show that our framework can produce more accurate reconstructions than alternative state-of-the-art multiview stereo systems. Sid Ying-Ze Bao, Manmohan Krishna Chandraker, Yuanqing Lin, Silvio Savarese |
CVPR | 3 |
| 2013 | Regionlets for Generic Object DetectionabstractGeneric object detection is confronted by dealing with different degrees of variations in distinct object classes with tractable computations, which demands for descriptive and flexible object representations that are also efficient to evaluate for many locations. In view of this, we propose to model an object class by a cascaded boosting classifier which integrates various types of features from competing local regions, named as region lets. A region let is a base feature extraction region defined proportionally to a detection window at an arbitrary resolution (i.e. size and aspect ratio). These region lets are organized in small groups with stable relative positions to delineate fine grained spatial layouts inside objects. Their features are aggregated to a one-dimensional feature within one group so as to tolerate deformations. Then we evaluate the object bounding box proposal in selective search from segmentation cues, limiting the evaluation locations to thousands. Our approach significantly outperforms the state-of-the-art on popular multi-class detection benchmark datasets with a single method, without any contexts. It achieves the detection mean average precision of 41.7% on the PASCAL VOC 2007 dataset and 39.7% on the VOC 2010 for 20 object categories. It achieves 14.7% mean average precision on the Image Net dataset for 200 object categories, outperforming the latest deformable part-based model (DPM) by 4.7%. Xiaoyu Wang 0002, Ming Yang 0007, Shenghuo Zhu, Yuanqing Lin |
ICCV | 4 |
| 2013 | Semantic-Aware Co-indexing for Image RetrievalabstractInverted indexes in image retrieval not only allow fast access to database images but also summarize all knowledge about the database, so that their discriminative capacity largely determines the retrieval performance. In this paper, for vocabulary tree based image retrieval, we propose a semantic-aware co-indexing algorithm to jointly embed two strong cues into the inverted indexes: 1) local invariant features that are robust to delineate low-level image contents, and 2) semantic attributes from large-scale object recognition that may reveal image semantic meanings. For an initial set of inverted indexes of local features, we utilize 1000 semantic attributes to filter out isolated images and insert semantically similar images to the initial set. Encoding these two distinct cues together effectively enhances the discriminative capability of inverted indexes. Such co-indexing operations are totally off-line and introduce small computation overhead to online query cause only local features but no semantic attributes are used for query. Experiments and comparisons with recent retrieval methods on 3 datasets, i.e., UKbench, Holidays, Oxford5K, and 1.3 million images from Flickr as distractors, manifest the competitive performance of our method. Shiliang Zhang, Ming Yang 0007, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001 |
ICCV | 4 |
| 2013 | Image segmentation for large-scale subcategory flower recognitionabstractWe propose a segmentation algorithm for the purposes of large-scale flower species recognition. Our approach is based on identifying potential object regions at the time of detection. We then apply a Laplacian-based segmentation, which is guided by these initially detected regions. More specifically, we show that 1) recognizing parts of the potential object helps the segmentation and makes it more robust to variabilities in both the background and the object appearances, 2) segmenting the object of interest at test time is beneficial for the subsequent recognition. Here we consider a large-scale dataset containing 578 flower species and 250,000 images. This dataset is developed by our team for the purposes of providing a flower recognition application for general use and is the largest in its scale and scope. We tested the proposed segmentation algorithm on the well-known 102 Oxford flowers benchmark [11] and on the new challenging large-scale 578 flower dataset, that we have collected. We observed about 4% improvements in the recognition performance on both datasets compared to the baseline. The algorithm also improves all other known results on the Oxford 102 flower benchmark dataset. Furthermore, our method is both simpler and faster than other related approaches, e.g. [3, 14], and can be potentially applicable to other subcategory recognition datasets. Anelia Angelova, Shenghuo Zhu, Yuanqing Lin |
WACV | 3 |
| 2012 | Multi-component Models for Object Detection
Chunhui Gu, Pablo Andrés Arbeláez, Yuanqing Lin, Kai Yu 0001, Jitendra Malik |
ECCV (4) | 3 |
| 2012 | Object-Centric Spatial Pooling for Image Classification
Olga Russakovsky, Yuanqing Lin, Kai Yu 0001, Li Fei-Fei 0001 |
ECCV (2) | 2 |
| 2011 | Large-scale image classification: Fast feature extraction and SVM trainingabstractMost research efforts on image classification so far have been focused on medium-scale datasets, which are often defined as datasets that can fit into the memory of a desktop (typically 4G~48G). There are two main reasons for the limited effort on large-scale image classification. First, until the emergence of ImageNet dataset, there was almost no publicly available large-scale benchmark data for image classification. This is mostly because class labels are expensive to obtain. Second, large-scale classification is hard because it poses more challenges than its medium-scale counterparts. A key challenge is how to achieve efficiency in both feature extraction and classifier training without compromising performance. This paper is to show how we address this challenge using ImageNet dataset as an example. For feature extraction, we develop a Hadoop scheme that performs feature extraction in parallel using hundreds of mappers. This allows us to extract fairly sophisticated features (with dimensions being hundreds of thousands) on 1.2 million images within one day. For SVM training, we develop a parallel averaging stochastic gradient descent (ASGD) algorithm for training one-against-all 1000-class SVM classifiers. The ASGD algorithm is capable of dealing with terabytes of training data and converges very fast-typically 5 epochs are sufficient. As a result, we achieve state-of-the-art performance on the ImageNet 1000-class classification, i.e., 52.9% in classification accuracy and 71.8% in top 5 hit rate. Yuanqing Lin, Fengjun Lv, Shenghuo Zhu, Ming Yang 0007, Timothée Cour, Kai Yu 0001, Liangliang Cao, Thomas S. Huang |
CVPR | 1 |
| 2011 | Learning image representations from the pixel level via hierarchical sparse codingabstractWe present a method for learning image representations using a two-layer sparse coding scheme at the pixel level. The first layer encodes local patches of an image. After pooling within local regions, the first layer codes are then passed to the second layer, which jointly encodes signals from the region. Unlike traditional sparse coding methods that encode local patches independently, this approach accounts for high-order dependency among patterns in a local image neighborhood. We develop algorithms for data encoding and codebook learning, and show in experiments that the method leads to more invariant and discriminative image representations. The algorithm gives excellent results for hand-written digit recognition on MNIST and object recognition on the Caltech101 benchmark. This marks the first time that such accuracies have been achieved using automatically learned features from the pixel level, rather than using hand-designed descriptors. Kai Yu 0001, Yuanqing Lin, John D. Lafferty |
CVPR | 2 |
| 2010 | Deep Coding NetworkabstractThis paper proposes a principled extension of the traditional single-layer flat sparse coding scheme, where a two-layer coding scheme is derived based on theoretical analysis of nonlinear functional approximation that extends recent results for local coordinate coding. The two-layer approach can be easily generalized to deeper structures in a hierarchical multiple-layer manner. Empirically, it is shown that the deep coding approach yields improved performance in benchmark datasets. Yuanqing Lin, Tong Zhang 0001, Shenghuo Zhu, Kai Yu 0001 |
NIPS | 1 |
| 2007 | Blind channel identification for speech dereverberation using l1-norm sparse learningabstractSpeech dereverberation remains an open problem after more than three decades of research. The most challenging step in speech dereverberation is blind chan- nel identification (BCI). Although many BCI approaches have been developed, their performance is still far from satisfactory for practical applications. The main difficulty in BCI lies in finding an appropriate acoustic model, which not only can effectively resolve solution degeneracies due to the lack of knowledge of the source, but also robustly models real acoustic environments. This paper proposes a sparse acoustic room impulse response (RIR) model for BCI, that is, an acous- tic RIR can be modeled by a sparse FIR filter. Under this model, we show how to formulate the BCI of a single-input multiple-output (SIMO) system into a l1- norm regularized least squares (LS) problem, which is convex and can be solved efficiently with guaranteed global convergence. The sparseness of solutions is controlled by l1-norm regularization parameters. We propose a sparse learning scheme that infers the optimal l1-norm regularization parameters directly from microphone observations under a Bayesian framework. Our results show that the proposed approach is effective and robust, and it yields source estimates in real acoustic environments with high fidelity to anechoic chamber measurements. Yuanqing Lin, Jingdong Chen, Youngmoo E. Kim, Daniel D. Lee |
NIPS | 1 |
| 2007 | Multiplicative Updates for Nonnegative Quadratic ProgrammingabstractMany problems in neural computation and statistical learning involve optimizations with nonnegativity constraints. In this article, we study convex problems in quadratic programming where the optimization is confined to an axis-aligned region in the nonnegative orthant. For these problems, we derive multiplicative updates that improve the value of the objective function at each iteration and converge monotonically to the global minimum. The updates have a simple closed form and do not involve any heuristics or free parameters that must be tuned to ensure convergence. Despite their simplicity, they differ strikingly in form from other multiplicative updates used in machine learning. We provide complete proofs of convergence for these updates and describe their application to problems in signal processing and pattern recognition. Fei Sha, Yuanqing Lin, Lawrence K. Saul, Daniel D. Lee |
Neural Comput. | 2 |
| 2006 | Bayesian L1-Norm Sparse LearningabstractWe propose a Bayesian framework for learning the optimal regularization parameter in the L1-norm penalized least-mean-square (LMS) problem, also known as LASSO [1] or basis pursuit [2]. The setting of the regularization parameter is critical for deriving a correct solution. In most existing methods, the scalar regularization parameter is often determined in a heuristic manner; in contrast, our approach infers the optimal regularization setting under a Bayesian framework. Furthermore, Bayesian inference enables an independent regularization scheme where each coefficient (or weight) is associated with an independent regularization parameter. Simulations illustrate the improvement using our method in discovering sparse structure from noisy data. Yuanqing Lin, Daniel D. Lee |
ICASSP (5) | 1 |
| 2005 | Relevant deconvolution for acoustic source estimationabstractWe describe a robust deconvolution algorithm for simultaneously estimating an acoustic source signal and convolutive filters associated with the acoustic room impulse responses from a pair of microphone signals. In contrast to conventional blind deconvolution techniques which rely upon a knowledge of the statistics of the source signal, our algorithm exploits the nonnegativity and sparsity structure of room impulse responses. The algorithm is formulated as a quadratic optimization problem with respect to both the source signal and filter coefficients, and proceeds by iteratively solving the optimization in two alternating steps. In the H-step, the nonnegative filter coefficients are optimally estimated within a Bayesian framework using a relevant set of regularization parameters. In the S-step, the source signal is estimated without any prior assumption on its statistical distribution. The resulting estimates converge to a relevant solution exhibiting appropriate sparseness in the filters. Simulation results indicate that the algorithm is able to precisely recover both the source signal and filter coefficients, even in the presence of large ambient noise. Yuanqing Lin, Daniel D. Lee |
ICASSP (5) | 1 |
| 2005 | Learning nonlinear appearance manifolds for robot localizationabstractWe propose a nonlinear method for learning the low-dimensional pose of a robot from high-dimensional panoramic images. The panoramic images are assumed to lie on a nonlinear low-dimensional appearance manifold that is embedded in a high-dimensional image space. We demonstrate that the local geometry of a point and its nearest neighbors on this manifold can be used to project the point onto a low-dimensional coordinate space. Using this embedding, the unknown camera position can be estimated from a novel panoramic image. We show how the image-based position measurements can be integrated with odometry information in a Bayesian framework to yield an online estimate of a robot's position. Results from simulated data show that the proposed method outperforms other appearance-based models based upon principal components analysis and kernel density estimation. Jihun Hamm, Yuanqing Lin, Daniel D. Lee |
IROS | 2 |
| 2005 | Cooperative relative robot localization with audible acoustic sensingabstractWe describe a method for estimating the relative poses of a team of mobile robots using only acoustic sensing. The relative distances and bearing angles of the robots are estimated using the time of arrival of audible sound signals on stereo microphones. The robots emit specially designed sound waveforms that simultaneously enable robot identification and time of arrival estimation. These acoustic observations are then combined with odometry to update a belief state describing the positions and heading angles of all the robots. To efficiently resolve the ambiguity in the heading angle of the observing robot as well as the back-front ambiguity of the observed robot, we employ a Rao-Blackwellised particle filter (RBPF) where the distribution over heading angles is represented by a discrete set of particles, and the uncertainty in the translational positions conditioned on each of these particles is described by a Gaussian. This approach combines the representational accuracy of conventional particle filters with the efficiency of Kalman filter updates in modeling the pose distribution over a number of robots. We demonstrate how the RBPF can quickly resolve uncertainties in the binaural acoustic measurements and yield a globally consistent pose estimate. Simulations as well as an experimental implementation on robots with generic sound hardware illustrate the accuracy and the convergence of the resulting pose estimates. Yuanqing Lin, Paul Vernaza, Jihun Hamm, Daniel D. Lee |
IROS | 1 |
| 2004 | Nonnegative deconvolution for time of arrival estimationabstractThe interaural time difference (ITD) of arrival is a primary cue for acoustic sound source localization. Traditional estimation techniques for ITD based upon cross-correlation are related to maximum-likelihood estimation of a simple generative model. We generalize the time difference estimation into a deconvolution problem with nonnegativity constraints. The resulting nonnegative least squares optimization can be efficiently solved using a novel iterative algorithm with guaranteed global convergence properties. We illustrate the utility of this algorithm using simulations and experimental results from a robot platform. Yuanqing Lin, Daniel D. Lee, Lawrence K. Saul |
ICASSP (2) | 1 |
| 2004 | Bayesian Regularization and Nonnegative Deconvolution for Time Delay EstimationabstractBayesian Regularization and Nonnegative Deconvolution (BRAND) is proposed for estimating time delays of acoustic signals in reverberant environments. Sparsity of the nonnegative filter coefficients is enforced using an L1-norm regularization. A probabilistic generative model is used to simultaneously estimate the regularization parameters and filter coefficients from the signal data. Iterative update rules are derived under a Bayesian framework using the Expectation-Maximization procedure. The resulting time delay estimation algorithm is demonstrated on noisy acoustic data. Yuanqing Lin, Daniel D. Lee |
NIPS | 1 |