Xiang Ruan

dblp:06/7779 · DBLP profile ↗
← Back
45ranked-venue papers
0as first author
7since 2021 · last 2024
0000-0003-4500-7516ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 5 since 2021Artificial intelligence and machine learning · 33 · 4 since 2021
YearPublicationVenuePosition
2024 Deep Boosting Learning: A Brand-New Cooperative Approach for Image-Text Matching
abstract
Image-text matching remains a challenging task due to heterogeneous semantic diversity across modalities and insufficient distance separability within triplets. Different from previous approaches focusing on enhancing multi-modal representations or exploiting cross-modal correspondence for more accurate retrieval, in this paper we aim to leverage the knowledge transfer between peer branches in a boosting manner to seek a more powerful matching model. Specifically, we propose a brand-new Deep Boosting Learning (DBL) algorithm, where an anchor branch is first trained to provide insights into the data properties, with a target branch gaining more advanced knowledge to develop optimal features and distance metrics. Concretely, an anchor branch initially learns the absolute or relative distance between positive and negative pairs, providing a foundational understanding of the particular network and data distribution. Building upon this knowledge, a target branch is concurrently tasked with more adaptive margin constraints to further enlarge the relative distance between matched and unmatched samples. Extensive experiments validate that our DBL can achieve impressive and consistent improvements based on various recent state-of-the-art models in the image-text matching field, and outperform related popular cooperative strategies, e.g., Conventional Distillation, Mutual Learning, and Contrastive Learning. Beyond the above, we confirm that DBL can be seamlessly integrated into their training scenarios and achieve superior performance under the same computational costs, demonstrating the flexibility and broad applicability of our proposed method.
Haiwen Diao, Ying Zhang 0021, Shang Gao 0012, Xiang Ruan, Huchuan Lu
IEEE Trans. Image Process.4
2023 High-Performance Transformer Tracking
abstract
Correlation has a critical role in the tracking field, especially in recent popular Siamese-based trackers. The correlation operation is a simple fusion method that considers the similarity between the template and the search region. However, the correlation operation is a local linear matching process, losing semantic information and easily falling into a local optimum, which may be the bottleneck in designing high-accuracy tracking algorithms. In this work, to determine whether a better feature fusion method exists than correlation, a novel attention-based feature fusion network, inspired by the transformer, is presented. This network effectively combines the template and search region features using attention mechanism. Specifically, the proposed method includes an ego-context augment module based on self-attention and a cross-feature augment module based on cross-attention. First, we present a transformer tracking (named TransT) method based on the Siamese-like feature extraction backbone, the designed attention-based fusion mechanism, and the classification and regression heads. Based on the TransT baseline, we also design a segmentation branch to generate the accurate mask. Finally, we propose a stronger version of TransT by extending it with a multi-template scheme and an IoU prediction head, named TransT-M. Experiments show that our TransT and TransT-M methods achieve promising results on seven popular benchmarks. Code and models are available at https://github.com/chenxin-dlut/TransT-M.
Xin Chen 0032, Bin Yan 0004, Jiawen Zhu 0003, Huchuan Lu, Xiang Ruan, Dong Wang 0004
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Transformer vision-language tracking via proxy token guided cross-modal fusion
Haojie Zhao, Xiao Wang 0014, Dong Wang 0004, Huchuan Lu, Xiang Ruan
Pattern Recognit. Lett.5
2023 Plug-and-Play Regulators for Image-Text Matching
abstract
Exploiting fine-grained correspondence and visual-semantic alignments has shown great potential in image-text matching. Generally, recent approaches first employ a cross-modal attention unit to capture latent region-word interactions, and then integrate all the alignments to obtain the final similarity. However, most of them adopt one-time forward association or aggregation strategies with complex architectures or additional information, while ignoring the regulation ability of network feedback. In this paper, we develop two simple but quite effective regulators which efficiently encode the message output to automatically contextualize and aggregate cross-modal representations. Specifically, we propose (i) a Recurrent Correspondence Regulator (RCR) which facilitates the cross-modal attention unit progressively with adaptive attention factors to capture more flexible correspondence, and (ii) a Recurrent Aggregation Regulator (RAR) which adjusts the aggregation weights repeatedly to increasingly emphasize important alignments and dilute unimportant ones. Besides, it is interesting that RCR and RAR are "plug-and-play": both of them can be incorporated into many frameworks based on cross-modal interaction to obtain significant benefits, and their cooperation achieves further improvements. Extensive experiments on MSCOCO and Flickr30K datasets validate that they can bring an impressive and consistent R@1 gain on multiple models, confirming the general effectiveness and generalization ability of the proposed methods.
Haiwen Diao, Ying Zhang 0021, Wei Liu 0005, Xiang Ruan, Huchuan Lu
IEEE Trans. Image Process.4
2022 Self-Supervised Pretraining for RGB-D Salient Object Detection
abstract
Existing CNNs-Based RGB-D salient object detection (SOD) networks are all required to be pretrained on the ImageNet to learn the hierarchy features which helps provide a good initialization. However, the collection and annotation of large-scale datasets are time-consuming and expensive. In this paper, we utilize self-supervised representation learning (SSL) to design two pretext tasks: the cross-modal auto-encoder and the depth-contour estimation. Our pretext tasks require only a few and unlabeled RGB-D datasets to perform pretraining, which makes the network capture rich semantic contexts and reduce the gap between two modalities, thereby providing an effective initialization for the downstream task. In addition, for the inherent problem of cross-modal fusion in RGB-D SOD, we propose a consistency-difference aggregation (CDA) module that splits a single feature fusion into multi-path fusion to achieve an adequate perception of consistent and differential information. The CDA module is general and suitable for cross-modal and cross-level feature fusion. Extensive experiments on six benchmark datasets show that our self-supervised pretrained model performs favorably against most state-of-the-art methods pretrained on ImageNet. The source code will be publicly available at https://github.com/Xiaoqi-Zhao-DLUT/SSLSOD.
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu, Xiang Ruan
AAAI5
2022 Visible-Thermal UAV Tracking: A Large-Scale Benchmark and New Baseline
abstract
With the popularity of multi-modal sensors, visible-thermal (RGB-T) object tracking is to achieve robust performance and wider application scenarios with the guidance of objects' temperature information. However, the lack of paired training samples is the main bottleneck for unlocking the power of RGB-T tracking. Since it is laborious to collect high-quality RGB-T sequences, recent benchmarks only provide test sequences. In this paper, we construct a large-scale benchmark with high diversity for visible-thermal UAV tracking (VTUAV), including 500 sequences with 1.7 million high-resolution (1920* 1080 pixels) frame pairs. In addition, comprehensive applications (short-term tracking, long-term tracking and segmentation mask prediction) with diverse categories and scenes are considered for exhaustive evaluation. Moreover, we provide a coarse-to-fine attribute annotation, where frame-level attributes are provided to exploit the potential of challenge-specific trackers. In addition, we design a new RGB-T baseline, named Hierarchical Multi-modal Fusion Tracker (HMFT), which fuses RGB-T data in various levels. Numerous experiments on several datasets are conducted to reveal the effectiveness of HMFT and the complement of different fusion types. The project is available at here.
Jie Zhao 0014, Dong Wang 0004, Huchuan Lu, Xiang Ruan
CVPR5
2021 Weakly-Supervised Temporal Action Localization via Cross-Stream Collaborative Learning
abstract
Weakly supervised temporal action localization (WTAL) is a challenging task as only video-level category labels are available during training stage. Without precise temporal annotations, most approaches rely on complementary RGB and optical flow features to predict the start and end frame of each action category in a video. However, existing approaches simply resort to either concatenation or weighted sum to learn how to take advantages of these two modalities for accurate action localization, which ignore the substantial variance between such two modalities. In this paper, we present Cross-Stream Collaborative Learning (CSCL) to address these issues. The proposed CSCL introduce a cross-stream weighting module to identify which modality is more robust during training and take advantage of the robust modality to guide the weaker one. Furthermore, we suppress the snippets which has high action-ness scores in both modalities to further exploiting the complementary property between two modalities. In addition, we bring the concept of co-training for WTAL and take both modalities into account for pseudo label generation to help training a stronger model. Extensive experiments conducted on THUMOS14 and ActivityNet dataset demonstrate that CSCL achieves a favorable performance against state-of-the-arts methods.
Xu Jia 0012, Huchuan Lu, Xiang Ruan
ACM Multimedia4
2020 CLIFFNet for Monocular Depth Estimation with Hierarchical Embedding Loss
Lijun Wang 0001, Jianming Zhang 0001, Yifan Wang 0004, Huchuan Lu, Xiang Ruan
ECCV (5)5
2019 Salient Object Detection with Recurrent Fully Convolutional Networks
abstract
Deep networks have been proved to encode high-level features with semantic meaning and delivered superior performance in salient object detection. In this paper, we take one step further by developing a new saliency detection method based on recurrent fully convolutional networks (RFCNs). Compared with existing deep network based methods, the proposed network is able to incorpor- ate saliency prior knowledge for more accurate inference. In addition, the recurrent architecture enables our method to automatically learn to refine the saliency map by iteratively correcting its previous errors, yielding more reliable final predictions. To train such a netw- ork with numerous parameters, we propose a pre-training strategy using semantic segmentation data, which simultaneously leverages the strong supervision of segmentation tasks for effective training and enables the network to capture generic representations to chara- cterize category-agnostic objects for saliency detection. Extensive experimental evaluations demonstrate that the proposed method compares favorably against state-of-the-art saliency detection approaches. Additional validations are also performed to study the impact of the recurrent architecture and pre-training strategy on both saliency detection and semantic segmentation, which provides important knowledge for network design and training in the future research.
Linzhao Wang, Lijun Wang 0001, Huchuan Lu, Xiang Ruan
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 Person Reidentification by Joint Local Distance Metric and Feature Transformation
abstract
Person reidentification is of great importance in visual surveillance and multiperson tracking across multiple camera views. Two fundamental problems are critical for person reidentification: 1) how to account for appearance variation or feature transformation caused by viewpoint changes and 2) how to learn a discriminative distance metric for reidentification. In this paper, we propose an algorithm in which both feature transformation and metric learning are exploited and jointly optimized. We learn local models from subsets of training samples with regularization imposed by the global model which is trained among the entire data set. The learned local models enhance the discriminative strength and generalization ability. Experimental results on the Viewpoint Invariant PEdestrian Eecognition, Queen Mary University of London ground reidentification, CUHK01, and CUHK03 benchmark data sets show that the proposed sample-specific view-invariant approach performs favorably against the state-of-the-art person reidentification methods.
Zimo Liu, Huchuan Lu, Xiang Ruan, Ming-Hsuan Yang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2018 Detect Globally, Refine Locally: A Novel Approach to Saliency Detection
abstract
Effective integration of contextual information is crucial for salient object detection. To achieve this, most existing methods based on 'skip' architecture mainly focus on how to integrate hierarchical features of Convolutional Neural Networks (CNNs). They simply apply concatenation or element-wise operation to incorporate high-level semantic cues and low-level detailed information. However, this can degrade the quality of predictions because cluttered and noisy information can also be passed through. To address this problem, we proposes a global Recurrent Localization Network (RLN) which exploits contextual information by the weighted response map in order to localize salient objects more accurately. Particularly, a recurrent module is employed to progressively refine the inner structure of the CNN over multiple time steps. Moreover, to effectively recover object boundaries, we propose a local Boundary Refinement Network (BRN) to adaptively learn the local contextual information for each spatial position. The learned propagation coefficients can be used to optimally capture relations between each pixel and its neighbors. Experiments on five challenging datasets show that our approach performs favorably against all existing methods in terms of the popular evaluation metrics.
Tiantian Wang 0002, Lihe Zhang, Huchuan Lu, Gang Yang 0002, Xiang Ruan, Ali Borji
CVPR6
2017 Learning to Detect Salient Objects with Image-Level Supervision
abstract
Deep Neural Networks (DNNs) have substantially improved the state-of-the-art in salient object detection. However, training DNNs requires costly pixel-level annotations. In this paper, we leverage the observation that image-level tags provide important cues of foreground salient objects, and develop a weakly supervised learning method for saliency detection using image-level tags only. The Foreground Inference Network (FIN) is introduced for this challenging task. In the first stage of our training method, FIN is jointly trained with a fully convolutional network (FCN) for image-level tag prediction. A global smooth pooling layer is proposed, enabling FCN to assign object category tags to corresponding object regions, while FIN is capable of capturing all potential foreground regions with the predicted saliency maps. In the second stage, FIN is fine-tuned with its predicted saliency maps as ground truth. For refinement of ground truth, an iterative Conditional Random Field is developed to enforce spatial label consistency and further boost performance. Our method alleviates annotation efforts and allows the usage of existing large scale training sets with image-level tags. Our model runs at 60 FPS, outperforms unsupervised ones with a large margin, and achieves comparable or even superior performance than fully supervised counterparts.
Lijun Wang 0001, Huchuan Lu, Yifan Wang 0004, Mengyang Feng, Dong Wang 0004, Xiang Ruan
CVPR7
2017 Amulet: Aggregating Multi-level Convolutional Features for Salient Object Detection
abstract
Fully convolutional neural networks (FCNs) have shown outstanding performance in many dense labeling problems. One key pillar of these successes is mining relevant information from features in convolutional layers. However, how to better aggregate multi-level convolutional feature maps for salient object detection is underexplored. In this work, we present Amulet, a generic aggregating multi-level convolutional feature framework for salient object detection. Our framework first integrates multi-level feature maps into multiple resolutions, which simultaneously incorporate coarse semantics and fine details. Then it adaptively learns to combine these feature maps at each resolution and predict saliency maps with the combined features. Finally, the predicted results are efficiently fused to generate the final saliency map. In addition, to achieve accurate boundary inference and semantic enhancement, edge-aware feature maps in low-level layers and the predicted results of low resolution features are recursively embedded into the learning framework. By aggregating multi-level convolutional features in this efficient and flexible manner, the proposed saliency model provides accurate salient object labeling. Comprehensive experiments demonstrate that our method performs favorably against state-of-the-art approaches in terms of near all compared evaluation metrics.
Dong Wang 0004, Huchuan Lu, Hongyu Wang 0001, Xiang Ruan
ICCV5
2017 Ranking Saliency
abstract
Most existing bottom-up algorithms measure the foreground saliency of a pixel or region based on its contrast within a local context or the entire image, whereas a few methods focus on segmenting out background regions and thereby salient objects. Instead of only considering the contrast between salient objects and their surrounding regions, we consider both foreground and background cues in this work. We rank the similarity of image elements with foreground or background cues via graph-based manifold ranking. The saliency of image elements is defined based on their relevances to the given seeds or queries. We represent an image as a multi-scale graph with fine superpixels and coarse regions as nodes. These nodes are ranked based on the similarity to background and foreground queries using affinity matrices. Saliency detection is carried out in a cascade scheme to extract background regions and foreground salient objects efficiently. Experimental results demonstrate the proposed method performs well against the state-of-the-art methods in terms of accuracy and speed. We also propose a new benchmark dataset containing 5,168 images for large-scale performance evaluation of saliency detection methods.
Lihe Zhang, Huchuan Lu, Xiang Ruan, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 Salient Object Detection via Multiple Instance Learning
abstract
Object proposals are a series of candidate segments containing objects of interest, which are taken as preprocessing and widely applied in various vision tasks. However, most of existing saliency approaches only utilize the proposals to compute a location prior. In this paper, we naturally take the proposals as the bags of instances of multiple instance learning (MIL), where the instances are the superpixels contained in the proposals, and formulate saliency detection problem as a MIL task (i.e., predict the labels of instances using the classifier in the MIL framework). This method allows some flexibility in finding a decision boundary based on the bag-level representations and can identify salient superpixels from ambiguous proposals. In addition, we introduce the MIL to an optimization mechanism, which iteratively updates training bags from easy to complex ones to learn a strong model. The significant improvement can be consistently achieved when applying the optimization model to existing saliency approaches. Extensive experiments demonstrate that the proposed algorithms perform favorably against the stateof- art saliency detection methods on several benchmark datasets.
Jinqing Qi, Huchuan Lu, Lihe Zhang, Xiang Ruan
IEEE Trans. Image Process.5
2017 Co-Bootstrapping Saliency
abstract
In this paper, we propose a visual saliency detection algorithm to explore the fusion of various saliency models in a manner of bootstrap learning. First, an original bootstrapping model, which combines both weak and strong saliency models, is constructed. In this model, image priors are exploited to generate an original weak saliency model, which provides training samples for a strong model. Then, a strong classifier is learned based on the samples extracted from the weak model. We use this classifier to classify all the salient and non-salient superpixels in an input image. To further improve the detection performance, multi-scale saliency maps of weak and strong model are integrated, respectively. The final result is the combination of the weak and strong saliency maps. The original model indicates that the overall performance of the proposed algorithm is largely affected by the quality of weak saliency model. Therefore, we propose a co-bootstrapping mechanism, which integrates the advantages of different saliency methods to construct the weak saliency model thus addresses the problem and achieves a better performance. Extensive experiments on benchmark data sets demonstrate that the proposed algorithm outperforms the state-of-the-art methods.
Huchuan Lu, Jinqing Qi, Na Tong, Xiang Ruan, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.5
2016 Deep Multi-task Attribute-driven Ranking for Fine-grained Sketch-based Image Retrieval
Jifei Song, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Xiang Ruan
BMVC5
2016 Sample-Specific SVM Learning for Person Re-identification
abstract
Person re-identification addresses the problem of matching people across disjoint camera views and extensive efforts have been made to seek either the robust feature representation or the discriminative matching metrics. However, most existing approaches focus on learning a fixed distance metric for all instance pairs, while ignoring the individuality of each person. In this paper, we formulate the person re-identification problem as an imbalanced classification problem and learn a classifier specifically for each pedestrian such that the matching model is highly tuned to the individual's appearance. To establish correspondence between feature space and classifier space, we propose a Least Square Semi-Coupled Dictionary Learning (LSSCDL) algorithm to learn a pair of dictionaries and a mapping function efficiently. Extensive experiments on a series of challenging databases demonstrate that the proposed algorithm performs favorably against the state-of-the-art approaches, especially on the rank-1 recognition rate.
Ying Zhang 0021, Baohua Li, Huchuan Lu, Atshushi Irie, Xiang Ruan
CVPR5
2016 Pattern Mining Saliency
Yuqiu Kong, Lijun Wang 0001, Xiuping Liu, Huchuan Lu, Xiang Ruan
ECCV (6)5
2016 Saliency Detection with Recurrent Fully Convolutional Networks
Linzhao Wang, Lijun Wang 0001, Huchuan Lu, Xiang Ruan
ECCV (4)5
2016 Query Adaptive Instance Search using Object Sketches
abstract
Sketch-based object search is a challenging problem mainly due to two difficulties: (1) how to match the binary sketch query with the colorful image, and (2) how to locate the small object in a big image with the sketch query. To address the above challenges, we propose to leverage object proposals for object search and localization. However, instead of purely relying on sketch features, e.g., Sketch-a-Net, to locate the candidate object proposals, we propose to fully utilize the appearance information to resolve the ambiguities among object proposals and refine the search results. Our proposed query adaptive search is formulated as a sub-graph selection problem, which can be solved by maximum flow algorithm. By performing query expansion using a smaller set of more salient matches as the query representatives, it can accurately locate the small target objects in cluttered background or densely drawn deformation intensive cartoon (Manga like) images. Our query adaptive sketch based object search on benchmark datasets exhibits superior performance when compared with existing methods, which validates the advantages of utilizing both the shape and appearance features for sketch-based search.
Sreyasee Das Bhattacharjee, Junsong Yuan 0001, Weixiang Hong 0001, Xiang Ruan
ACM Multimedia4
2016 Sketch retrieval via local dense stroke features
Chao Ma 0004, Xiaokang Yang 0001, Xiang Ruan, Ming-Hsuan Yang 0001
Image Vis. Comput.4
2016 Combining motion and appearance cues for anomaly detection
Ying Zhang 0021, Huchuan Lu, Lihe Zhang, Xiang Ruan
Pattern Recognit.4
2016 Video anomaly detection based on locality sensitive hashing filters
Ying Zhang 0021, Huchuan Lu, Lihe Zhang, Xiang Ruan, Shun Sakai
Pattern Recognit.4
2016 Dense and Sparse Reconstruction Error Based Saliency Descriptor
abstract
In this paper, we propose a visual saliency detection algorithm from the perspective of reconstruction error. The image boundaries are first extracted via superpixels as likely cues for background templates, from which dense and sparse appearance models are constructed. First, we compute dense and sparse reconstruction errors on the background templates for each image region. Second, the reconstruction errors are propagated based on the contexts obtained from K -means clustering. Third, the pixel-level reconstruction error is computed by the integration of multi-scale reconstruction errors. Both the pixel-level dense and sparse reconstruction errors are then weighted by image compactness, which could more accurately detect saliency. In addition, we introduce a novel Bayesian integration method to combine saliency maps, which is applied to integrate the two saliency measures based on dense and sparse reconstruction errors. Experimental results show that the proposed algorithm performs favorably against 24 state-of-the-art methods in terms of precision, recall, and F-measure on three public standard salient object detection databases.
Huchuan Lu, Xiaohui Li 0005, Lihe Zhang, Xiang Ruan, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.4
2015 Salient object detection via bootstrap learning
abstract
We propose a bootstrap learning algorithm for salient object detection in which both weak and strong models are exploited. First, a weak saliency map is constructed based on image priors to generate training samples for a strong model. Second, a strong classifier based on samples directly from an input image is learned to detect salient pixels. Results from multiscale saliency maps are integrated to further improve the detection performance. Extensive experiments on six benchmark datasets demonstrate that the proposed bootstrap learning algorithm performs favorably against the state-of-the-art saliency detection methods. Furthermore, we show that the proposed bootstrap learning approach can be easily applied to other bottom-up saliency models for significant improvement.
Na Tong, Huchuan Lu, Xiang Ruan, Ming-Hsuan Yang 0001
CVPR3
2015 Deep networks for saliency detection via local estimation and global search
abstract
This paper presents a saliency detection algorithm by integrating both local estimation and global search. In the local estimation stage, we detect local saliency by using a deep neural network (DNN-L) which learns local patch features to determine the saliency value of each pixel. The estimated local saliency maps are further refined by exploring the high level object concepts. In the global search stage, the local saliency map together with global contrast and geometric information are used as global features to describe a set of object candidate regions. Another deep neural network (DNN-G) is trained to predict the saliency score of each object region based on the global features. The final saliency map is generated by a weighted sum of salient object regions. Our method presents two interesting insights. First, local features learned by a supervised scheme can effectively capture local contrast, texture and shape information for saliency detection. Second, the complex relationship between different global saliency cues can be captured by deep networks and exploited principally rather than heuristically. Quantitative and qualitative experiments on several benchmark data sets demonstrate that our algorithm performs favorably against the state-of-the-art methods.
Lijun Wang 0001, Huchuan Lu, Xiang Ruan, Ming-Hsuan Yang 0001
CVPR3
2015 Salient object detection via global and local cues
Na Tong, Huchuan Lu, Ying Zhang 0021, Xiang Ruan
Pattern Recognit.4
2014 Saliency Detection with Multi-Scale Superpixels
abstract
We propose a salient object detection algorithm via multi-scale analysis on superpixels. First, multi-scale segmentations of an input image are computed and represented by superpixels. In contrast to prior work, we utilize various Gaussian smoothing parameters to generate coarse or fine results, thereby facilitating the analysis of salient regions. At each scale, three essential cues from local contrast, integrity and center bias are considered within the Bayesian framework. Next, we compute saliency maps by weighted summation and normalization. The final saliency map is optimized by a guided filter which further improves the detection results. Extensive experiments on two large benchmark datasets demonstrate the proposed algorithm performs favorably against state-of-the-art methods. The proposed method achieves the highest precision value of 97.39% when evaluated on one of the most popular datasets, the ASD dataset.
Na Tong, Huchuan Lu, Lihe Zhang, Xiang Ruan
IEEE Signal Process. Lett.4
2013 Sketch Retrieval via Dense Stroke Features
abstract
Sketch retrieval aims at retrieving most similar sketches from a large database based on one hand-drawn query. Successful retrieval hinges on an effective representation of sketch images and an efficient search method. In this paper, we propose a representation scheme which takes sketch strokes into account with local features, thereby facilitating efficient retrieval with codebooks. Stroke features are detected via densely sampled points on stroke lines from which local gradients are further enhanced and described by a quantized histogram of gradients. A codebook is organized in a hierarchical vocabulary tree, which maintains structural information of visual words and enables efficient retrieval in sub-linear time. Experimental results on three data sets demonstrate the merits of the proposed algorithm for effective and efficient sketch retrieval.
Chao Ma 0004, Xiaokang Yang 0001, Xiang Ruan, Ming-Hsuan Yang 0001
BMVC4
2013 Discriminative Generative Contour Detection
abstract
Contour detection is an important and fundamental problem in computer vision which finds numerous applications.Despite significant progress has been made in the past decades, contour detection from natural images remains a challenging task due to the difficulty of clearly distinguishing between edges of objects and surrounding backgrounds.To address this problem, we first capture multi-scale features from pixel-level to segmentlevel using local and global information.These features are mapped to a space where discriminative information is captured by computing posterior divergence of Gaussian mixture models and then used to train a random forest classifier for contour detection.We evaluate the proposed algorithm against leading methods in the literature on the Berkeley segmentation and Weizmann horse data sets.Experimental results demonstrate that the proposed contour detection algorithm performs favorably against state-of-the-art methods in terms of speed and accuracy.
Chao Zhang 0010, Xiong Li 0004, Xiang Ruan, Ming-Hsuan Yang 0001
BMVC3
2013 Saliency Detection via Graph-Based Manifold Ranking
abstract
Most existing bottom-up methods measure the foreground saliency of a pixel or region based on its contrast within a local context or the entire image, whereas a few methods focus on segmenting out background regions and thereby salient objects. Instead of considering the contrast between the salient objects and their surrounding regions, we consider both foreground and background cues in a different way. We rank the similarity of the image elements (pixels or regions) with foreground cues or background cues via graph-based manifold ranking. The saliency of the image elements is defined based on their relevances to the given seeds or queries. We represent the image as a close-loop graph with super pixels as nodes. These nodes are ranked based on the similarity to background and foreground queries, based on affinity matrices. Saliency detection is carried out in a two-stage scheme to extract background regions and foreground salient objects efficiently. Experimental results on two large benchmark databases demonstrate the proposed method performs well when against the state-of-the-art methods in terms of accuracy and speed. We also create a more difficult benchmark database containing 5,172 images to test the proposed saliency model and make this database publicly available with this paper for further studies in the saliency field.
Lihe Zhang, Huchuan Lu, Xiang Ruan, Ming-Hsuan Yang 0001
CVPR4
2013 Saliency Detection via Dense and Sparse Reconstruction
abstract
In this paper, we propose a visual saliency detection algorithm from the perspective of reconstruction errors. The image boundaries are first extracted via super pixels as likely cues for background templates, from which dense and sparse appearance models are constructed. For each image region, we first compute dense and sparse reconstruction errors. Second, the reconstruction errors are propagated based on the contexts obtained from K-means clustering. Third, pixel-level saliency is computed by an integration of multi-scale reconstruction errors and refined by an object-biased Gaussian model. We apply the Bayes formula to integrate saliency measures based on dense and sparse reconstruction errors. Experimental results show that the proposed algorithm performs favorably against seventeen state-of-the-art methods in terms of precision and recall. In addition, the proposed algorithm is demonstrated to be more effective in highlighting salient objects uniformly and robust to background noise.
Xiaohui Li 0005, Huchuan Lu, Lihe Zhang, Xiang Ruan, Ming-Hsuan Yang 0001
ICCV4
2012 Hand posture recognition using finger geometric feature
Liwei Liu 0006, Junliang Xing, Haizhou Ai, Xiang Ruan
ICPR4
2012 Contour detection via random forest
Chao Zhang 0010, Xiang Ruan, Ming-Hsuan Yang 0001
ICPR2
2012 Visual saliency: A manifold way of perception
Xiang Ruan
ICPR3
2012 Human body segmentation based on deformable models and two-scale superpixel
Shifeng Li, Huchuan Lu, Xiang Ruan, Yen-Wei Chen 0001
Pattern Anal. Appl.3
2012 Multilinear Supervised Neighborhood Embedding of a Local Descriptor Tensor for Scene/Object Recognition
abstract
In this paper, we propose to represent an image as a local descriptor tensor and use a multilinear supervised neighborhood embedding (MSNE) for discriminant feature extraction, which is able to be used for subject or scene recognition. The contributions of this paper include: 1) a novel feature extraction approach denoted as the histogram of orientation weighted with a normalized gradient (NHOG) for local region representation, which is robust to large illumination variation in an image; 2) an image representation framework denoted as the local descriptor tensor, which can effectively combine a moderate amount of local features together for image representation and be more efficient than the popular existing bag-of-feature model; and 3) an MSNE analysis algorithm, which can directly deal with the local descriptor tensor for extracting discriminant and compact features and, at the same time, preserve neighborhood structure in tensor-feature space for subject/scene recognition. We demonstrate the performance advantages of our proposed approach over existing techniques on different types of benchmark database such as a scene data set (i.e., OT8), face data sets (i.e., YALE and PIE), and view-based object data sets (COIL-100 and ETH-80).
Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan
IEEE Trans. Image Process.3
2011 A co-training framework for visual tracking with multiple instance learning
abstract
This paper proposes a Co-training Multiple Instance Learning algorithm (CoMIL). Our framework is based on the co-training approach which labels incoming data continuously, and then uses the prediction from each classifier to enlarge the training set of the other. The discriminative classifier is implemented using online multiple instance learning (MIL), which can deal with inaccurate positive samples in the updating process and allow some flexibility while finding a decision boundary. Firstly, two classifiers are improved mutually in our CoMIL tracking system. Secondly, our update mechanism uses multiple potential positives according to the MIL which handles the update error due to the risk of extracting only one positive example. Experiments show that our CoMIL tracking algorithm performs better than several state-of-the-art tracking algorithms on challenging sequences.
Huchuan Lu, Qiuhong Zhou, Dong Wang 0004, Xiang Ruan
FG4
2011 Canonical correlation analysis of local feature set for view-based object recognition
abstract
In this paper, we propose to use local feature set for image representation, which can represent variations in an object's appearance due to changing viewpoint or camera pose. It was evidenced that usually only a part of the object are appeared in common when taking a photo of an object in different view points. With comparison of local features set extracted from different positions of images, an object can be recognized when common part is appeared in two images, which take photos of one object in different view points. In this paper, we use Canonical Correlation (also known as principle or canonical angles), which can be thought of as the angles between two d-dimensional subspace, as similarity measure of local feature sets. The proposed approach is evaluated in various view-based object datasets (Coil-100 and ETH80) for object and object category recognition. Experiments show that the performance advantages of our proposed approach can be achieved over existing techniques.
Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan
ICIP3
2011 Pose estimation and body segmentation based on hierarchical searching tree
abstract
In this paper, we propose a novel method for pose estimation and body segmentation. We estimate the partial configuration of adjacent parts instead of detecting each single part, which makes our method more robust and accurate. Further, we develop a general model to calculate the partial configuration. Besides, we present a tree-based hierarchical probabilistic method to derive the global optimal pose. Additionally, the coarse-to-fine strategy is employed to speed up the pose estimation in the whole procedure. After finishing pose estimation, the estimated pose is used to guide body segmentation. Experiments suggest that our method is efficient and effective for pose estimation and body segmentation simultaneously.
Shifeng Li, Huchuan Lu, Xiang Ruan, Yen-Wei Chen 0001
ICIP3
2010 Image recognition by learned linear subspace of combined bag-of-features and low-level features
abstract
Image category recognition is important to access visual information on the level of objects and scene types. This paper combines different feature representations of images and learn a compact subspace of different features for the automatic recognition of object and scene classes. Compact visual-words and low-level-features object class subspaces are automatically learned from a set of training images by a Regularized Linear Discriminant analysis (RLDA) algorithm, and the extracted RLDA-domain features are used for Support Vector Machine (SVM) classifier. The main contribution of this paper is two folds: i) Different features (bag-of-features and low-level features)is fused for image representation. ii) The compact feature subspaces (low-dimension features) of different features are learned for rendering to SVM classifier, which is computationally efficient for image category. High classification accuracy is demonstrated on object recognition database (Caltech). We confirm that the proposed strategy cam improve accuracy rate compared with state-of-the-art methods for object recognition databases.
Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan
ICIP3
2010 Adaptive Color Independent Components Based SIFT Descriptors for Image Classification
abstract
This paper proposes an adaptive color independent components based SIFT descriptor (termed CIC-SIFT) for image classification. Our motivation is to seek an adaptive and efficient color space for color SIFT feature extraction. Our work has two key contributions. First, based on independent component analysis (ICA), an adaptive and efficient color space is proposed for color image representation. Second, in this ICA-based color space, a discriminative CIC-SIFT descriptor is calculated for image classification. The experiment results indicate that (1) contrast between objects and background can be enhanced on the ICA-based color space and (2) the CIC-SIFT descriptor outperforms other conventional color SIFT descriptors on image classification.
Danni Ai, Xianhua Han, Xiang Ruan, Yen-Wei Chen 0001
ICPR3
2010 Image Categorization by Learned Nonlinear Subspace of Combined Visual-Words and Low-Level Features
abstract
Image category recognition is important to access visual information on the level of objects and scene types. This paper presents a new algorithm for the automatic recognition of object and scene classes. Compact and yet discriminative visual-words and low-level-features object class subspaces are automatically learned from a set of training images by a Supervised Nonlinear Neighborhood Embedding (SNNE) algorithm, which can learn an adaptive nonlinear subspace by preserving the neighborhood structure of the visual feature space. The main contribution of this paper is two fold: i) an optimally compact and discriminative feature subspace is learned by the proposed SNNE algorithm for different feature space (visual-word and low-level features). ii) An effective merge of different feature subspace can be implemented simply. High classification accuracy is demonstrated on different database including the scene database (Simplicity) and object recognition database (Caltech). We confirm that the proposed strategy is much better than state-of-the-art methods for different databases.
Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan
ICPR3
2010 Semi-supervised and Interactive Semantic Concept Learning for Scene Recognition
abstract
In this paper, we present a novel semi-supervised and interactive concept learning algorithm for scene recognition by local semantic description. Our work is motivated by the continuing effort in content-based image retrieval to extract and to model the semantic content of images. The basic idea of the semantic modeling is to classify local image regions into semantic concept classes such as water, sunset, or sky. However, labeling concept sampling manually for training semantic model is fairly expensive, and the labeling results is, to some extent, subjective to the operators. In this paper, by using the proposed semi-supervised and interactive learning algorithm, training samples and new concepts can be obtained accurately and efficiently. Through extensive experiments, we demonstrate that the image concept representation is well suited for modeling the semantic content of heterogenous scene categories, and thus for recognition and retrieval. Furthermore, higher recognition accuracy can be achieved by updating new training samples and concepts, which are obtained by the novel proposed algorithm.
Xianhua Han, Yen-Wei Chen 0001, Xiang Ruan
ICPR3