Xiaofeng Ren

dblp:84/3585 · DBLP profile ↗
← Back
49ranked-venue papers
16as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 16 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 13 first-author · 3 since 2021Systems, architecture and hardware · 7Human-computer interaction and ubiquitous computing · 3Computer networks · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
30 papers
Image recognition and object detection · 30% 3D vision · 22% Segmentation and scene understanding · 16%
Computer graphics and multimedia
12 papers
Image and video processing · 71% Multimedia analysis and retrieval · 22% Computational photography and imaging · 7%
Human-computer interaction and pervasive computing
3 papers
Interaction techniques and input · 51% Wearable and physiological sensing · 26% Ubiquitous computing and smart environments · 23%

Topics — the 30 heaviest of 75, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
0.722022
SCMT: Self-Correction Mean Teacher for Semi-supervised Object Detection · IJCAI 2022
Histograms of Sparse Codes for Object Detection · CVPR 2013
Computer vision › Segmentation and scene understanding
semantic segmentation
0.722022
Pyramid Architecture for Multi-Scale Processing in Point Cloud Segmentation · CVPR 2022
RGB-(D) scene labeling: Features and algorithms · CVPR 2012
Machine learning › Deep learning architectures and training
multi-scale feature fusion
0.612022
Pyramid Architecture for Multi-Scale Processing in Point Cloud Segmentation · CVPR 2022
Computer vision › 3D vision
point cloud segmentation
0.612022
Pyramid Architecture for Multi-Scale Processing in Point Cloud Segmentation · CVPR 2022
Computer vision › Image recognition and object detection › object detection
semi-supervised object detection
0.612022
SCMT: Self-Correction Mean Teacher for Semi-supervised Object Detection · IJCAI 2022
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.612022
Doubly-Fused ViT: Fuse Information from Vision Transformer Doubly with Local Representation · ECCV (23) 2022
Computer vision › Image recognition and object detection
image classification
0.432013
Multipath Sparse Coding Using Hierarchical Matching Pursuit · CVPR 2013
Hierarchical Matching Pursuit for Image Classification: Architecture and Fast Algorithms · NIPS 2011
Kernel Descriptors for Visual Recognition · NIPS 2010
Computer vision › Image recognition and object detection
object recognition
0.432011
Object recognition with hierarchical kernel descriptors · CVPR 2011
A Scalable Tree-Based Approach for Joint Object and Pose Recognition · AAAI 2011
Figure-ground segmentation improves handled object recognition in egocentric video · CVPR 2010
Machine learning › Representation and self-supervised learning › representation learning › neural network representation learning › deep representation learning
deep feature representation
0.312018
Visual Search at Alibaba · KDD 2018
Multimedia analysis and retrieval
visual search
0.312018
Visual Search at Alibaba · KDD 2018
Computer vision › Image recognition and object detection › object recognition › multimodal object recognition
RGB-D object recognition
0.332011
Sparse distance learning for object recognition combining RGB and depth information · ICRA 2011
A large-scale hierarchical multi-view RGB-D object dataset · ICRA 2011
Object recognition with hierarchical kernel descriptors · CVPR 2011
Image and video processing › image segmentation
contour detection
0.332012
Discriminatively Trained Sparse Code Gradients for Contour Detection · NIPS 2012
Multi-scale Improves Boundary Detection in Natural Images · ECCV (3) 2008
Scale-Invariant Contour Completion Using Conditional Random Fields · ICCV 2005
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.322012
Discriminatively Trained Sparse Code Gradients for Contour Detection · NIPS 2012
Hierarchical Matching Pursuit for Image Classification: Architecture and Fast Algorithms · NIPS 2011
Computer vision › 3D vision
3d object recognition
0.222011
Sparse distance learning for object recognition combining RGB and depth information · ICRA 2011
A large-scale hierarchical multi-view RGB-D object dataset · ICRA 2011
Computer vision › Image recognition and object detection
object discovery
0.222011
Toward object discovery and modeling via 3-D scene comparison · ICRA 2011
Learning to recognize objects in egocentric activities · CVPR 2011
Machine learning › Representation and self-supervised learning › visual representation › image representation › handcrafted descriptor
kernel descriptors
0.222011
Object recognition with hierarchical kernel descriptors · CVPR 2011
Kernel Descriptors for Visual Recognition · NIPS 2010
Machine learning › Representation and self-supervised learning › visual representation
patch-based representation
0.222011
Object recognition with hierarchical kernel descriptors · CVPR 2011
Kernel Descriptors for Visual Recognition · NIPS 2010
Computer vision › Segmentation and scene understanding
image segmentation
0.232011
Learning to recognize objects in egocentric activities · CVPR 2011
Cue Integration for Figure/Ground Labeling · NIPS 2005
Recovering Human Body Configurations: Combining Segmentation and Recognition · CVPR (2) 2004
Computer vision › 3D vision
depth estimation
0.212014
Depth Enhancement via Low-Rank Matrix Completion · CVPR 2014
Computer vision › 3D vision › depth estimation › depth map refinement
RGB-D depth refinement
0.212014
Depth Enhancement via Low-Rank Matrix Completion · CVPR 2014
Image and video processing › image reconstruction
depth completion
0.212014
Depth Enhancement via Low-Rank Matrix Completion · CVPR 2014
Image and video processing › image enhancement
depth map enhancement
0.212014
Depth Enhancement via Low-Rank Matrix Completion · CVPR 2014
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
foreground-background segmentation
0.222010
Figure-ground segmentation improves handled object recognition in egocentric video · CVPR 2010
Tracking as Repeated Figure/Ground Segmentation · CVPR 2007
Image and video processing › image segmentation
contour completion
0.232008
Learning Probabilistic Models for Contour Completion in Natural Images · Int. J. Comput. Vis. 2008
Scale-Invariant Contour Completion Using Conditional Random Fields · ICCV 2005
A Probabilistic Multi-scale Model for Contour Completion Based on Image Statistics · ECCV (1) 2002
Machine learning › Deep learning architectures and training
teacher-student framework
0.212022
SCMT: Self-Correction Mean Teacher for Semi-supervised Object Detection · IJCAI 2022
Computer vision › 3D vision › motion estimation
3d motion estimation
0.212013
RGB-D flow: Dense 3-D motion estimation using color and depth · ICRA 2013
Computer vision › 3D vision
scene flow estimation
0.212013
RGB-D flow: Dense 3-D motion estimation using color and depth · ICRA 2013
Computer vision › 3D vision
3d object detection
0.112012
Detection-based object labeling in 3D scenes · ICRA 2012
Computer vision › 3D vision › 3d scene understanding
3d scene labeling
0.112012
Detection-based object labeling in 3D scenes · ICRA 2012
Computer vision › 3D vision
3d scene understanding
0.112012
Detection-based object labeling in 3D scenes · ICRA 2012

Methods — techniques the papers use, named apart from their topics

deep CNN · 1.0click behavior mining · 1.0model and search-based fusion · 0.7vision transformer · 0.6self-correction · 0.6pyramid architecture · 0.6pseudo-label re-weighting · 0.6mean teacher · 0.6cross-scale attention · 0.6orthogonal matching pursuit · 0.4sparse coding · 0.3subspace constraint · 0.2low-rank matrix completion · 0.2variational flow · 0.2RGB-D · 0.2shape and appearance analysis · 0.1object tracking · 0.1multi-scale pooling · 0.1
YearPublicationVenuePosition
2026 Identifying the Focus Word in Natural Language Questions Based on Association Rules
abstract
Knowledge base‐based intelligent question‐answering systems have insufficient understanding of the questions. In the early stages of research, it is effective in most cases that the existing natural language question‐understanding methods can answer questions by connecting entities and relationships when ignoring the identification of focus words. However, as research deepens, ignoring focus words has become a shortcoming. To address this, we propose identifying focus words, enabling more precise understanding of user focus. We define focus itemset, frequent focus itemset, focus association rule, and strong focus association rule to express focus‐related information better. Given the unique nature of focus association rules, we propose a prefix tree structure and an algorithm for mining association rules aimed at identifying focus words. We also introduce an inverted index specifically designed for focus association rules and propose an efficient algorithm for identifying focus words based on this index. Experiments verify the effectiveness of our algorithm and the efficiency of the inverted index, with a focus word identification rate exceeding 90%.
Xin Hu 0008, Xiaofeng Ren, Jiangli Duan, Sulan Zhang
Int. J. Intell. Syst.2
2023 Fault Sensing of the Distribution Cable Feeders by Time-Domain Measurements
abstract
Power distribution cable feeders are prone to faults due to the poor working environment and internal defects. Considering the structure and electrical parameters of the typical distribution three-core cable, a fault sensing method is proposed. The fault sensing in this article includes three aspects: fault cable detection, fault phase identification, and fault resistance estimation. The grounding line currents of cable feeders are first analyzed in different neutral grounding modes. Based on the amplitudes and directions of the grounding line currents, a fault feeder detection criterion is then presented. Finally, based on the equivalent circuit of the fault cable, an algorithm for identifying the fault phase is developed by estimating the fault resistance. The experimental model of the distribution cable feeders is created by real time digital simulation (RTDS). Various fault experiments are carried out. The results show that the neutral grounding modes and fault conditions have little effect on the proposed method.
Nan Peng, Zhengyi Zhang, Chengrui Jiang, Peng Zhang 0081, Xiaofeng Ren, Xiuru Wang
IEEE Trans. Ind. Informatics6
2022 Pyramid Architecture for Multi-Scale Processing in Point Cloud Segmentation
abstract
Semantic segmentation of point cloud data is a critical task for autonomous driving and other applications. Recent advances of point cloud segmentation are mainly driven by new designs of local aggregation operators and point sampling methods. Unlike image segmentation, few efforts have been made to understand the fundamental issue of scale and how scales should interact and be fused. In this work, we investigate how to efficiently and effectively integrate features at varying scales and varying stages in a point cloud segmentation network. In particular, we open up the commonly used encoder-decoder architecture, and design scale pyramid architectures that allow information to flow more freely and systematically, both laterally and upward/downward in scale. Moreover, a cross-scale attention feature learning block has been designed to enhance the multi-scale feature fusion which occurs everywhere in the network. Such a design of multi-scale processing and fusion gains large improvements in accuracy without adding much additional computation. When built on top of the popular KPConv network, we see consistent improvements on a wide range of datasets, including achieving state-of-the-art performance on NPM3D and S3DIS. Moreover, the pyramid architecture is generic and can be applied to other network designs: we show an example of similar improvements over RandLANet.
Dong Nie, Rui Lan, Xiaofeng Ren
CVPR4
2022 Doubly-Fused ViT: Fuse Information from Vision Transformer Doubly with Local Representation
Dong Nie, Xiaofeng Ren
ECCV (23)4
2022 SCMT: Self-Correction Mean Teacher for Semi-supervised Object Detection
abstract
Semi-Supervised Object Detection (SSOD) aims to improve performance by leveraging a large amount of unlabeled data. Existing works usually adopt the teacher-student framework to enforce student to learn consistent predictions over the pseudo-labels generated by teacher. However, the performance of the student model is limited since the noise inherently exists in pseudo-labels. In this paper, we investigate the causes and effects of noisy pseudo-labels and propose a simple yet effective approach denoted as Self-Correction Mean Teacher(SCMT) to reduce the adverse effects. Specifically, we propose to dynamically re-weight the unsupervised loss of each student's proposal with additional supervision information from the teacher model, and assign smaller loss weights to possible noisy proposals. Extensive experiments on MS-COCO benchmark have shown the superiority of our proposed SCMT, which can significantly improve the supervised baseline by more than 11% mAP under all 1%, 5% and 10% COCO-standard settings, and surpasses state-of-the-art methods by about 1.5% mAP. Even under the challenging COCO-additional setting, SCMT still improves the supervised baseline by 4.9% mAP, and significantly outperforms previous methods by 1.2% mAP, achieving a new state-of-the-art performance.
Zhihui Hao, Yu-Lin He, Xiaofeng Ren
IJCAI5
2020 Bidirectional Pyramid Networks for Semantic Segmentation
Dong Nie, Jia Xue, Xiaofeng Ren
ACCV (1)3
2018 Visual Search at Alibaba
abstract
This paper introduces the large scale visual search algorithm and system infrastructure at Alibaba. The following challenges are discussed under the E-commercial circumstance at Alibaba (a) how to handle heterogeneous image data and bridge the gap between real-shot images from user query and the online images. (b) how to deal with large scale indexing for massive updating data. (c) how to train deep models for effective feature representation without huge human annotations. (d) how to improve the user engagement by considering the quality of the content. We take advantage of large image collection of Alibaba and state-of-the-art deep learning techniques to perform visual search at scale. We present solutions and implementation details to overcome those problems and also share our learnings from building such a large scale commercial visual search engine. Specifically, model and search-based fusion approach is introduced to effectively predict categories. Also, we propose a deep CNN model for joint detection and feature learning by mining user click behavior. The binary index engine is designed to scale up indexing without compromising recall and precision. Finally, we apply all the stages into an end-to-end system architecture, which can simultaneously achieve highly efficient and scalable performance adapting to real-shot images. Extensive experiments demonstrate the advancement of each module in our system. We hope visual search at Alibaba becomes more widely incorporated into today's commercial applications.
Yanhao Zhang 0002, Yingya Zhang, Xiaofeng Ren, Rong Jin 0001
KDD6
2014 Depth Enhancement via Low-Rank Matrix Completion
abstract
Depth captured by consumer RGB-D cameras is often noisy and misses values at some pixels, especially around object boundaries. Most existing methods complete the missing depth values guided by the corresponding color image. When the color image is noisy or the correlation between color and depth is weak, the depth map cannot be properly enhanced. In this paper, we present a depth map enhancement algorithm that performs depth map completion and de-noising simultaneously. Our method is based on the observation that similar RGB-D patches lie in a very low-dimensional subspace. We can then assemble the similar patches into a matrix and enforce this low-rank subspace constraint. This low-rank subspace constraint essentially captures the underlying structure in the RGB-D patches and enables robust depth enhancement against the noise or weak correlation between color and depth. Based on this subspace constraint, our method formulates depth map enhancement as a low-rank matrix completion problem. Since the rank of a matrix changes over matrices, we develop a data-driven method to automatically determine the rank number for each matrix. The experiments on both public benchmarks and our own captured RGB-D images show that our method can effectively enhance depth maps.
Si Lu, Xiaofeng Ren, Feng Liu 0015
CVPR2
2013 Multipath Sparse Coding Using Hierarchical Matching Pursuit
abstract
Complex real-world signals, such as images, contain discriminative structures that differ in many aspects including scale, invariance, and data channel. While progress in deep learning shows the importance of learning features through multiple layers, it is equally important to learn features through multiple paths. We propose Multipath Hierarchical Matching Pursuit (M-HMP), a novel feature learning architecture that combines a collection of hierarchical sparse features for image classification to capture multiple aspects of discriminative structures. Our building blocks are MI-KSVD, a codebook learning algorithm that balances the reconstruction error and the mutual incoherence of the codebook, and batch orthogonal matching pursuit (OMP), we apply them recursively at varying layers and scales. The result is a highly discriminative image representation that leads to large improvements to the state-of-the-art on many standard benchmarks, e.g., Caltech-101, Caltech-256, MITScenes, Oxford-IIIT Pet and Caltech-UCSD Bird-200.
Liefeng Bo, Xiaofeng Ren, Dieter Fox
CVPR2
2013 Histograms of Sparse Codes for Object Detection
abstract
Object detection has seen huge progress in recent years, much thanks to the heavily-engineered Histograms of Oriented Gradients (HOG) features. Can we go beyond gradients and do better than HOG? We provide an affirmative answer by proposing and investigating a sparse representation for object detection, Histograms of Sparse Codes (HSC). We compute sparse codes with dictionaries learned from data using K-SVD, and aggregate per-pixel sparse codes to form local histograms. We intentionally keep true to the sliding window framework (with mixtures and parts) and only change the underlying features. To keep training (and testing) efficient, we apply dimension reduction by computing SVD on learned models, and adopt supervised training where latent positions of roots and parts are given externally e.g. from a HOG-based detector. By learning and using local representations that are much more expressive than gradients, we demonstrate large improvements over the state of the art on the PASCAL benchmark for both root-only and part-based models.
Xiaofeng Ren, Deva Ramanan
CVPR1
2013 RGB-D flow: Dense 3-D motion estimation using color and depth
abstract
3-D motion estimation is a fundamental problem that has far-reaching implications in robotics. A scene flow formulation is attractive as it makes no assumptions about scene complexity, object rigidity, or camera motion. RGB-D cameras provide new information useful for computing dense 3-D flow in challenging scenes. In this work we show how to generalize two-frame variational 2-D flow algorithms to 3-D. We show that scene flow can be reliably computed using RGB-D data, overcoming depth noise and outperforming previous results on a variety of scenes. We apply dense 3-D flow to rigid motion segmentation.
Evan Herbst, Xiaofeng Ren, Dieter Fox
ICRA2
2013 Research and applications: Ontology-guided organ detection to retrieve web images of disease manifestation: towards the construction of a consumer-based health image library
abstract
BACKGROUND: Visual information is a crucial aspect of medical knowledge. Building a comprehensive medical image base, in the spirit of the Unified Medical Language System (UMLS), would greatly benefit patient education and self-care. However, collection and annotation of such a large-scale image base is challenging. OBJECTIVE: To combine visual object detection techniques with medical ontology to automatically mine web photos and retrieve a large number of disease manifestation images with minimal manual labeling effort. METHODS: As a proof of concept, we first learnt five organ detectors on three detection scales for eyes, ears, lips, hands, and feet. Given a disease, we used information from the UMLS to select affected body parts, ran the pretrained organ detectors on web images, and combined the detection outputs to retrieve disease images. RESULTS: Compared with a supervised image retrieval approach that requires training images for every disease, our ontology-guided approach exploits shared visual information of body parts across diseases. In retrieving 2220 web images of 32 diseases, we reduced manual labeling effort to 15.6% while improving the average precision by 3.9% from 77.7% to 81.6%. For 40.6% of the diseases, we improved the precision by 10%. CONCLUSIONS: The results confirm the concept that the web is a feasible source for automatic disease image retrieval for health image database construction. Our approach requires a small amount of manual effort to collect complex disease images, and to annotate them by standard medical ontology terms.
Yang Chen 0022, Xiaofeng Ren, Guo-Qiang Zhang 0001
J. Am. Medical Informatics Assoc.2
2013 Scheduling Exploiting Frequency and Multi-User Diversity in LTE Downlink Systems
abstract
Scheduling can obtain multi-user diversity if channel state information (CSI) is known, such as for low-mobility users and can exploit frequency diversity if CSI is not available at the transmitter, such as for high-mobility users. In this paper, we investigate resource allocation exploiting frequency and multiuser diversity for LTE downlink systems with users of different mobilities. To facilitate resource allocation, we first develop a user classification algorithm to identify high- and low-mobility users. Based on user mobility classification, we then propose a scheduling algorithm to simultaneously obtain multi-user diversity for those low-mobility users and frequency diversity for those high-mobility users. It is demonstrated by computer simulation that the performance of the proposed scheduling algorithm provides 6% and 23% gain of overall cell throughput, and 5.6% and 18% gain of 10th percentile throughput over proportional fairness based frequency-selective and frequency-diversity scheduling algorithms, respectively. Furthermore, the proposed scheduling algorithm has the same order of computational complexity as the frequency-selective scheduling algorithm.
Jinping Niu, Dae-Won Lee, Xiaofeng Ren, Geoffrey Ye Li
IEEE Trans. Wirel. Commun.3
2013 User Classification and Scheduling in LTE Downlink Systems with Heterogeneous User Mobilities
abstract
In LTE systems with heterogeneous user mobilities, low-mobility users favor frequency selective scheduling while high-mobility users benefit from frequency diversity scheduling. To benefit both low- and high-mobility users simultaneously, scheduling exploiting frequency selectivity and diversity is desired. To enable the scheduling, low-complexity user mobility classification to distinguish these two types of users is required. In this paper, we first propose a user mobility classification algorithm, which is robust to different channel delay profiles (CDPs), for single-transmit-antenna systems. Then, we extend it to multiple-input multiple-output (MIMO) systems. A low-complexity scheduling algorithm, exploiting both frequency-selectivity and diversity for low- and high-mobility users simultaneously, is also developed. As demonstrated by the simulation results, the proposed user classification algorithm is robust to different CDPs and the proposed scheduling algorithm is effective.
Jinping Niu, Dae-Won Lee, Geoffrey Ye Li, Xiaofeng Ren
IEEE Trans. Wirel. Commun.5
2012 SensorSift: balancing sensor data privacy and utility in automated face understanding
abstract
We introduce SensorSift, a new theoretical scheme for balancing utility and privacy in smart sensor applications. At the heart of our contribution is an algorithm which transforms raw sensor data into a 'sifted' representation which minimizes exposure of user defined private attributes while maximally exposing application-requested public attributes. We envision multiple applications using the same platform, and requesting access to public attributes explicitly not known at the time of the platform creation. Support for future-defined public attributes, while still preserving the defined privacy of the private attributes, is a central challenge that we tackle.
Miro Enev, Jaeyeon Jung, Liefeng Bo, Xiaofeng Ren, Tadayoshi Kohno
ACSAC4
2012 RGB-(D) scene labeling: Features and algorithms
abstract
Scene labeling research has mostly focused on outdoor scenes, leaving the harder case of indoor scenes poorly understood. Microsoft Kinect dramatically changed the landscape, showing great potentials for RGB-D perception (color+depth). Our main objective is to empirically understand the promises and challenges of scene labeling with RGB-D. We use the NYU Depth Dataset as collected and analyzed by Silberman and Fergus [30]. For RGB-D features, we adapt the framework of kernel descriptors that converts local similarities (kernels) to patch descriptors. For contextual modeling, we combine two lines of approaches, one using a superpixel MRF, and the other using a segmentation tree. We find that (1) kernel descriptors are very effective in capturing appearance (RGB) and shape (D) similarities; (2) both superpixel MRF and segmentation tree are useful in modeling context; and (3) the key to labeling accuracy is the ability to efficiently train and test with large-scale data. We improve labeling accuracy on the NYU Dataset from 56.6% to 76.1%. We also apply our approach to image-only scene labeling and improve the accuracy on the Stanford Background Dataset from 79.4% to 82.9%.
Xiaofeng Ren, Liefeng Bo, Dieter Fox
CVPR1
2012 Fine-grained kitchen activity recognition using RGB-D
abstract
We present a first study of using RGB-D (Kinect-style) cameras for fine-grained recognition of kitchen activities. Our prototype system combines depth (shape) and color (appearance) to solve a number of perception problems crucial for smart space applications: locating hands, identifying objects and their functionalities, recognizing actions and tracking object state changes through actions. Our proof-of-concept results demonstrate great potentials of RGB-D perception: without need for instrumentation, our system can robustly track and accurately recognize detailed steps through cooking activities, for instance how many spoons of sugar are in a cake mix, or how long it has been mixing. A robust RGB-D based solution to fine-grained activity recognition in real-world conditions will bring the intelligence of pervasive and interactive systems to the next level.
Jinna Lei, Xiaofeng Ren, Dieter Fox
UbiComp2
2012 Detection-based object labeling in 3D scenes
abstract
We propose a view-based approach for labeling objects in 3D scenes reconstructed from RGB-D (color+depth) videos. We utilize sliding window detectors trained from object views to assign class probabilities to pixels in every RGB-D frame. These probabilities are projected into the reconstructed 3D scene and integrated using a voxel representation. We perform efficient inference on a Markov Random Field over the voxels, combining cues from view-based detection and 3D shape, to label the scene. Our detection-based approach produces accurate scene labeling on the RGB-D Scenes Dataset and improves the robustness of object detection.
Kevin Lai 0001, Liefeng Bo, Xiaofeng Ren, Dieter Fox
ICRA3
2012 Discriminatively Trained Sparse Code Gradients for Contour Detection
abstract
Finding contours in natural images is a fundamental problem that serves as the basis of many tasks such as image segmentation and object recognition. At the core of contour detection technologies are a set of hand-designed gradient features, used by most existing approaches including the state-of-the-art Global Pb (gPb) operator. In this work, we show that contour detection accuracy can be significantly improved by computing Sparse Code Gradients (SCG), which measure contrast using patch representations automatically learned through sparse coding. We use K-SVD and Orthogonal Matching Pursuit for efficient dictionary learning and encoding, and use multi-scale pooling and power transforms to code oriented local neighborhoods before computing gradients and applying linear SVM. By extracting rich representations from pixels and avoiding collapsing them prematurely, Sparse Code Gradients effectively learn how to measure local contrasts and find contours. We improve the F-measure metric on the BSDS500 benchmark to 0.74 (up from 0.71 of gPb contours). Moreover, our learning approach can easily adapt to novel sensor data such as Kinect-style RGB-D cameras: Sparse Code Gradients on depth images and surface normals lead to promising contour detection using depth and depth+color, as verified on the NYU Depth Dataset. Our work combines the concept of oriented gradients with sparse representation and opens up future possibilities for learning contour detection and segmentation.
Xiaofeng Ren, Liefeng Bo
NIPS1
2012 Scheduling exploiting frequency and multi-user diversity in LTE downlink systems
abstract
In this paper, we develop a scheduling algorithm to obtain multi-user diversity for those low-mobility users and frequency diversity for those high-mobility users. Computer simulation demonstrates that the proposed scheduling algorithm provides 10% and 16% overall cell throughput gain over proportional fairness based frequency-selective and frequency-diversity scheduling algorithm, respectively. In addition, the proposed scheduling algorithm is shown to have the same order of computational complexity as the frequency-selective scheduling algorithm, and can be easily implemented in the LTE downlink systems.
Jinping Niu, Dae-Won Lee, Xiaofeng Ren, Geoffrey Ye Li
PIMRC3
2011 A Scalable Tree-Based Approach for Joint Object and Pose Recognition
abstract
Recognizing possibly thousands of objects is a crucial capability for an autonomous agent to understand and interact with everyday environments. Practical object recognition comes in multiple forms: Is this a coffee mug (category recognition). Is this Alice's coffee mug? (instance recognition). Is the mug with the handle facing left or right? (pose recognition). We present a scalable framework, Object-Pose Tree, which efficiently organizes data into a semantically structured tree. The tree structure enables both scalable training and testing, allowing us to solve recognition over thousands of object poses in near real-time. Moreover, by simultaneously optimizing all three tasks, our approach outperforms standard nearest neighbor and 1-vs-all classifications, with large improvements on pose recognition. We evaluate the proposed technique on a dataset of 300 household objects collected using a Kinect-style 3D camera. Experiments demonstrate that our system achieves robust and efficient object category, instance, and pose recognition on challenging everyday objects.
Kevin Lai 0001, Liefeng Bo, Xiaofeng Ren, Dieter Fox
AAAI3
2011 Combining Self Training and Active Learning for Video Segmentation
abstract
Presented at the 22nd British Machine Vision Conference (BMVC 2011), 29 August-2 September 2011, University of Dundee, Scotland, UK.
Alireza Fathi, Maria-Florina Balcan, Xiaofeng Ren, James M. Rehg
BMVC3
2011 Toward Robust Material Recognition for Everyday Objects
abstract
Material recognition is a fundamental problem in perception that is receiving increasing attention. Following the recent work using Flickr [16, 23], we empirically study material recognition of real-world objects using a rich set of local features. We use the Kernel Descriptor framework [5] and extend the set of descriptors to include material-motivated attributes using variances of gradient orientation and magnitude. Large-Margin Nearest Neighbor learning is used for a 30-fold dimension reduction. We improve the state-of-the-art accuracy on the Flickr dataset [16] from 45 % to 54%. We also introduce two new datasets using ImageNet and macro photos, extensively evaluating our set of features and showing promising connections between material and object recognition.
Diane Hu, Liefeng Bo, Xiaofeng Ren
BMVC3
2011 HeatWave: thermal imaging for surface user interaction
abstract
We present HeatWave, a system that uses digital thermal imaging cameras to detect, track, and support user interaction on arbitrary surfaces. Thermal sensing has had limited examination in the HCI research community and is generally under-explored outside of law enforcement and energy auditing applications. We examine the role of thermal imaging as a new sensing solution for enhancing user surface interaction. In particular, we demonstrate how thermal imaging in combination with existing computer vision techniques can make segmentation and detection of routine interaction techniques possible in real-time, and can be used to complement or simplify algorithms for traditional RGB and depth cameras. Example interactions include (1) distinguishing hovering above a surface from touch events, (2) shape-based gestures similar to ink strokes, (3) pressure based gestures, and (4) multi-finger gestures. We close by discussing the practicality of thermal sensing for naturalistic user interaction and opportunities for future work.
Eric C. Larson, Gabe Cohn, Sidhant Gupta, Xiaofeng Ren, Beverly L. Harrison, Dieter Fox, Shwetak N. Patel
CHI4
2011 Object recognition with hierarchical kernel descriptors
abstract
Kernel descriptors provide a unified way to generate rich visual feature sets by turning pixel attributes into patch-level features, and yield impressive results on many object recognition tasks. However, best results with kernel descriptors are achieved using efficient match kernels in conjunction with nonlinear SVMs, which makes it impractical for large-scale problems. In this paper, we propose hierarchical kernel descriptors that apply kernel descriptors recursively to form image-level features and thus provide a conceptually simple and consistent way to generate image-level features from pixel attributes. More importantly, hierarchical kernel descriptors allow linear SVMs to yield state-of-the-art accuracy while being scalable to large datasets. They can also be naturally extended to extract features over depth images. We evaluate hierarchical kernel descriptors both on the CIFAR10 dataset and the new RGB-D Object Dataset consisting of segmented RGB and depth images of 300 everyday objects.
Liefeng Bo, Kevin Lai 0001, Xiaofeng Ren, Dieter Fox
CVPR3
2011 Learning to recognize objects in egocentric activities
abstract
This paper addresses the problem of learning object models from egocentric video of household activities, using extremely weak supervision. For each activity sequence, we know only the names of the objects which are present within it, and have no other knowledge regarding the appearance or location of objects. The key to our approach is a robust, unsupervised bottom up segmentation method, which exploits the structure of the egocentric domain to partition each frame into hand, object, and background categories. By using Multiple Instance Learning to match object instances across sequences, we discover and localize object occurrences. Object representations are refined through transduction and object-level classifiers are trained. We demonstrate encouraging results in detecting novel object instances using models produced by weakly-supervised learning.
Alireza Fathi, Xiaofeng Ren, James M. Rehg
CVPR2
2011 Interactive 3D modeling of indoor environments with a consumer depth camera
abstract
Detailed 3D visual models of indoor spaces, from walls and floors to objects and their configurations, can provide extensive knowledge about the environments as well as rich contextual information of people living therein. Vision-based 3D modeling has only seen limited success in applications, as it faces many technical challenges that only a few experts understand, let alone solve. In this work we utilize (Kinect style) consumer depth cameras to enable non-expert users to scan their personal spaces into 3D models. We build a prototype mobile system for 3D modeling that runs in real-time on a laptop, assisting and interacting with the user on-the-fly. Color and depth are jointly used to achieve robust 3D registration. The system offers online feedback and hints, tolerates human errors and alignment failures, and helps to obtain complete scene coverage. We show that our prototype system can both scan large environments (50 meters across) and at the same time preserve fine details (centimeter accuracy). The capability of detailed 3D modeling leads to many promising applications such as accurate 3D localization, measuring dimensions, and interactive visualization.
Hao Du 0004, Peter Henry, Xiaofeng Ren, Marvin Cheng, Dan B. Goldman, Steven M. Seitz, Dieter Fox
UbiComp3
2011 Toward object discovery and modeling via 3-D scene comparison
abstract
The performance of indoor robots that stay in a single environment can be enhanced by gathering detailed knowledge of objects that frequently occur in that environment. We use an inexpensive sensor providing dense color and depth, and fuse information from multiple sensing modalities to detect changes between two 3-D maps. We adapt a recent SLAM technique to align maps. A probabilistic model of sensor readings lets us reason about movement of surfaces. Our method handles arbitrary shapes and motions, and is robust to lack of texture. We demonstrate the ability to find whole objects in complex scenes by regularizing over surface patches.
Evan Herbst, Peter Henry, Xiaofeng Ren, Dieter Fox
ICRA3
2011 A large-scale hierarchical multi-view RGB-D object dataset
abstract
Over the last decade, the availability of public image repositories and recognition benchmarks has enabled rapid progress in visual object category and instance detection. Today we are witnessing the birth of a new generation of sensing technologies capable of providing high quality synchronized videos of both color and depth, the RGB-D (Kinect-style) camera. With its advanced sensing capabilities and the potential for mass adoption, this technology represents an opportunity to dramatically increase robotic object recognition, manipulation, navigation, and interaction capabilities. In this paper, we introduce a large-scale, hierarchical multi-view object dataset collected using an RGB-D camera. The dataset contains 300 objects organized into 51 categories and has been made publicly available to the research community so as to enable rapid progress based on this promising technology. This paper describes the dataset collection procedure and introduces techniques for RGB-D based object recognition and detection, demonstrating that combining color and depth information substantially improves quality of results.
Kevin Lai 0001, Liefeng Bo, Xiaofeng Ren, Dieter Fox
ICRA3
2011 Sparse distance learning for object recognition combining RGB and depth information
abstract
In this work we address joint object category and instance recognition in the context of RGB-D (depth) cameras. Motivated by local distance learning, where a novel view of an object is compared to individual views of previously seen objects, we define a view-to-object distance where a novel view is compared simultaneously to all views of a previous object. This novel distance is based on a weighted combination of feature differences between views. We show, through jointly learning per-view weights, that this measure leads to superior classification performance on object category and instance recognition. More importantly, the proposed distance allows us to find a sparse solution via Group-Lasso regularization, where a small subset of representative views of an object is identified and used, with the rest discarded. This significantly reduces computational cost without compromising recognition accuracy. We evaluate the proposed technique, Instance Distance Learning (IDL), on the RGB-D Object Dataset, which consists of 300 object instances in 51 everyday categories and about 250,000 views of objects with both RGB color and depth. We empirically compare IDL to several alternative state-of-the-art approaches and also validate the use of visual and shape cues and their combination.
Kevin Lai 0001, Liefeng Bo, Xiaofeng Ren, Dieter Fox
ICRA3
2011 Depth kernel descriptors for object recognition
abstract
Consumer depth cameras, such as the Microsoft Kinect, are capable of providing frames of dense depth values at real time. One fundamental question in utilizing depth cameras is how to best extract features from depth frames. Motivated by local descriptors on images, in particular kernel descriptors, we develop a set of kernel features on depth images that model size, 3D shape, and depth edges in a single framework. Through extensive experiments on object recognition, we show that (1) our local features capture different aspects of cues from a depth frame/view that complement one another; (2) our kernel features significantly outperform traditional 3D features (e.g. Spin images); and (3) we significantly improve the capabilities of depth and RGB-D (color+depth) recognition, achieving 10–15% improvement in accuracy over the state of the art.
Liefeng Bo, Xiaofeng Ren, Dieter Fox
IROS2
2011 RGB-D object discovery via multi-scene analysis
abstract
We introduce an algorithm for object discovery from RGB-D (color plus depth) data, building on recent progress in using RGB-D cameras for 3-D reconstruction. A set of 3-D maps are built from multiple visits to the same scene. We introduce a multi-scene MRF model to detect objects that moved between visits, combining shape, visibility, and color cues. We measure similarities between candidate objects using both 2-D and 3-D matching, and apply spectral clustering to infer object clusters from noisy links. Our approach can robustly detect objects and their motion between scenes even when objects are textureless or have the same shape as other objects.
Evan Herbst, Xiaofeng Ren, Dieter Fox
IROS2
2011 Hierarchical Matching Pursuit for Image Classification: Architecture and Fast Algorithms
abstract
Extracting good representations from images is essential for many computer vision tasks. In this paper, we propose hierarchical matching pursuit (HMP), which builds a feature hierarchy layer-by-layer using an efficient matching pursuit encoder. It includes three modules: batch (tree) orthogonal matching pursuit, spatial pyramid max pooling, and contrast normalization. We investigate the architecture of HMP, and show that all three components are critical for good performance. To speed up the orthogonal matching pursuit, we propose a batch tree orthogonal matching pursuit that is particularly suitable to encode a large number of observations that share the same large dictionary. HMP is scalable and can efficiently handle full-size images. In addition, HMP enables linear support vector machines (SVM) to match the performance of nonlinear SVM while being scalable to large datasets. We compare HMP with many state-of-the-art algorithms including convolutional deep belief networks, SIFT based single layer sparse coding, and kernel based feature learning. HMP consistently yields superior accuracy on three types of image classification problems: object recognition (Caltech-101), scene recognition (MIT-Scene), and static event recognition (UIUC-Sports).
Liefeng Bo, Xiaofeng Ren, Dieter Fox
NIPS2
2010 Figure-ground segmentation improves handled object recognition in egocentric video
abstract
Identifying handled objects, i.e. objects being manipulated by a user, is essential for recognizing the person's activities. An egocentric camera as worn on the body enjoys many advantages such as having a natural first-person view and not needing to instrument the environment. It is also a challenging setting, where background clutter is known to be a major source of problems and is difficult to handle with the camera constantly and arbitrarily moving. In this work we develop a bottom-up motion-based approach to robustly segment out foreground objects in egocentric video and show that it greatly improves object recognition accuracy. Our key insight is that egocentric video of object manipulation is a special domain and many domain-specific cues can readily help. We compute dense optical flow and fit it into multiple affine layers. We then use a max-margin classifier to combine motion with empirical knowledge of object location and background movement as well as temporal cues of support region and color appearance. We evaluate our segmentation algorithm on the large Intel Egocentric Object Recognition dataset with 42 objects and 100K frames. We show that, when combined with temporal integration, figure-ground segmentation improves the accuracy of a SIFT-based recognition system from 33% to 60%, and that of a latent-HOG system from 64% to 86%.
Xiaofeng Ren, Chunhui Gu
CVPR1
2010 Discriminative Mixture-of-Templates for Viewpoint Classification
Chunhui Gu, Xiaofeng Ren
ECCV (5)2
2010 Kernel Descriptors for Visual Recognition
abstract
The design of low-level image features is critical for computer vision algorithms. Orientation histograms, such as those in SIFT~\cite{Lowe2004Distinctive} and HOG~\cite{Dalal2005Histograms}, are the most successful and popular features for visual object and scene recognition. We highlight the kernel view of orientation histograms, and show that they are equivalent to a certain type of match kernels over image patches. This novel view allows us to design a family of kernel descriptors which provide a unified and principled framework to turn pixel attributes (gradient, color, local binary pattern, \etc) into compact patch-level features. In particular, we introduce three types of match kernels to measure similarities between image patches, and construct compact low-dimensional kernel descriptors from these match kernels using kernel principal component analysis (KPCA)~\cite{Scholkopf1998Nonlinear}. Kernel descriptors are easy to design and can turn any type of pixel attribute into patch-level features. They outperform carefully tuned and sophisticated features including SIFT and deep belief networks. We report superior performance on standard image classification benchmarks: Scene-15, Caltech-101, CIFAR10 and CIFAR10-ImageNet.
Liefeng Bo, Xiaofeng Ren, Dieter Fox
NIPS2
2008 Finding people in archive films through tracking
abstract
The goal of this work is to find all people in archive films. Challenges include low image quality, motion blur, partial occlusion, non-standard poses and crowded scenes. We base our approach on face detection and take a tracking/temporal approach to detection. Our tracker operates in two modes, following face detections whenever possible, switching to low-level tracking if face detection fails. With temporal correspondences established by tracking, we formulate detection as an inference problem in one-dimensional chains/tracks. We use a conditional random field model to integrate information across frames and to re-score tentative detections in tracks. Quantitative evaluations on full-length films show that the CRF-based temporal detector greatly improves face detection, increasing precision for about 30% (suppressing isolated false positives) and at the same time boosting recall for over 10% (recovering difficult cases where face detectors fail).
Xiaofeng Ren
CVPR1
2008 Local grouping for optical flow
abstract
Optical flow estimation requires spatial integration, which essentially poses a grouping question: what points belong to the same motion and what do not. Classical local approaches to optical flow, such as Lucas-Kanade, use isotropic neighborhoods and have considerable difficulty near motion boundaries. In this work we utilize image-based grouping to facilitate spatial- and scale-adaptive integration. We define soft spatial support using pairwise affinities computed through intervening contour. We sample images at edges and corners, and iteratively estimate affine motion at sample points. Figure-ground organization further improves grouping and flow estimation near boundaries. We show that affinity-based spatial integration enables reliable flow estimation and avoids erroneous motion propagation from and/or across object boundaries. We demonstrate our approach on the Middlebury flow dataset.
Xiaofeng Ren
CVPR1
2008 Multi-scale Improves Boundary Detection in Natural Images
Xiaofeng Ren
ECCV (3)1
2008 Learning Probabilistic Models for Contour Completion in Natural Images
Xiaofeng Ren, Charless C. Fowlkes, Jitendra Malik
Int. J. Comput. Vis.1
2007 Learning and Matching Line Aspects for Articulated Objects
abstract
Traditional aspect graphs are topology-based and are impractical for articulated objects. In this work we learn a small number of aspects, or prototypical views, from video data. Groundtruth segmentations in video sequences are utilized for both training and testing aspect models that operate on static images. We represent aspects of an articulated object as collections of line segments. In learning aspects, where object centers are known, a linear matching based on line location and orientation is used to measure similarity between views. We use K-medoid to find cluster centers. When using line aspects in recognition, matching is based on pairwise cues of relative location, relative orientation as well adjacency and parallelism. Matching with pairwise cues leads to a quadratic optimization that we solve with a spectral approximation. We show that our line aspect matching is capable of locating people in a variety of poses. Line aspect matching performs significantly better than an alternative approach using Hausdorff distance, showing merits of the line representation.
Xiaofeng Ren
CVPR1
2007 Tracking as Repeated Figure/Ground Segmentation
abstract
Tracking over a long period of time is challenging as the appearance, shape and scale of the object in question may vary. We propose a paradigm of tracking by repeatedly segmenting figure from background. Accurate spatial support obtained in segmentation provides rich information about the track and enables reliable tracking of non-rigid objects without drifting. Figure/ground segmentation operates sequentially in each frame by utilizing both static image cues and temporal coherence cues, which include an appearance model of brightness (or color) and a spatial model propagating figure/ground masks through low-level region correspondence. A superpixel-based conditional random field linearly combines cues and loopy belief propagation is used to estimate marginal posteriors of figure vs background. We demonstrate our approach on long sequences of sports video, including figure skating and football.
Xiaofeng Ren, Jitendra Malik
CVPR1
2006 Figure/Ground Assignment in Natural Images
Xiaofeng Ren, Charless C. Fowlkes, Jitendra Malik
ECCV (2)1
2005 Recovering Human Body Configurations Using Pairwise Constraints between Parts
abstract
The goal of this work is to recover human body configurations from static images. Without assuming a priori knowledge of scale, pose or appearance, this problem is extremely challenging and demands the use of all possible sources of information. We develop a framework which can incorporate arbitrary pairwise constraints between body parts, such as scale compatibility, relative position, symmetry of clothing and smooth contour connections between parts. We detect candidate body parts from bottom-up using parallelism, and use various pairwise configuration constraints to assemble them together into body configurations. To find the most probable configuration, we solve an integer quadratic programming problem with a standard technique using linear approximations. Approximate IQP allows us to incorporate much more information than the traditional dynamic programming and remains computationally efficient. 15 hand-labeled images are used to train the low-level part detector and learn the pairwise constraints. We show test results on a variety of images.
Xiaofeng Ren, Alexander C. Berg, Jitendra Malik
ICCV1
2005 Scale-Invariant Contour Completion Using Conditional Random Fields
abstract
We present a model of curvilinear grouping using piece-wise linear representations of contours and a conditional random field to capture continuity and the frequency of different junction types. Potential completions are generated by building a constrained Delaunay triangulation (CDT) over the set of contours found by a local edge detector. Maximum likelihood parameters for the model are learned from human labeled ground truth. Using held out test data, we measure how the model, by incorporating continuity structure, improves boundary detection over the local edge detector. We also compare performance with a baseline local classifier that operates on pairs of edgels. Both algorithms consistently dominate the low-level boundary detector at all thresholds. To our knowledge, this is the first time that curvilinear continuity has been shown quantitatively useful for a large variety of natural images. Better boundary detection has immediate application in the problem of object detection and recognition.
Xiaofeng Ren, Charless C. Fowlkes, Jitendra Malik
ICCV1
2005 Cue Integration for Figure/Ground Labeling
abstract
We present a model of edge and region grouping using a conditional random field built over a scale-invariant representation of images to integrate multiple cues. Our model includes potentials that capture low-level similarity, mid-level curvilinear continuity and high-level object shape. Maximum likelihood parameters for the model are learned from human labeled groundtruth on a large collection of horse images using belief propagation. Using held out test data, we quantify the information gained by incorporating generic mid-level cues and high-level shape.
Xiaofeng Ren, Charless C. Fowlkes, Jitendra Malik
NIPS1
2004 Recovering Human Body Configurations: Combining Segmentation and Recognition
Greg Mori, Xiaofeng Ren, Alexei A. Efros, Jitendra Malik
CVPR (2)2
2003 Learning a Classification Model for Segmentation
abstract
We propose a two-class classification model for grouping. Human segmented natural images are used as positive examples. Negative examples of grouping are constructed by randomly matching human segmentations and images. In a preprocessing stage an image is over-segmented into super-pixels. We define a variety of features derived from the classical Gestalt cues, including contour, texture, brightness and good continuation. Information-theoretic analysis is applied to evaluate the power of these grouping cues. We train a linear classifier to combine these features. To demonstrate the power of the classification model, a simple algorithm is used to randomly search for good segmentations. Results are shown on a wide range of images.
Xiaofeng Ren, Jitendra Malik
ICCV1
2002 A Probabilistic Multi-scale Model for Contour Completion Based on Image Statistics
Xiaofeng Ren, Jitendra Malik
ECCV (1)1