EDBT 2026 Demo / reviewers in the wild / expert
Xiaofeng Ren
dblp:84/3585
· DBLP profile ↗
49ranked-venue papers
16as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 16 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 13 first-author · 3 since 2021Systems, architecture and hardware · 7Human-computer interaction and ubiquitous computing · 3Computer networks · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
30 papers |
Image recognition and object detection · 30% 3D vision · 22% Segmentation and scene understanding · 16% | |
| Computer graphics and multimedia
12 papers |
Image and video processing · 71% Multimedia analysis and retrieval · 22% Computational photography and imaging · 7% | |
| Human-computer interaction and pervasive computing
3 papers |
Interaction techniques and input · 51% Wearable and physiological sensing · 26% Ubiquitous computing and smart environments · 23% |
Topics — the 30 heaviest of 75, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
object detection |
0.7 | 2 | 2022 | SCMT: Self-Correction Mean Teacher for Semi-supervised Object Detection · IJCAI 2022 Histograms of Sparse Codes for Object Detection · CVPR 2013 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.7 | 2 | 2022 | Pyramid Architecture for Multi-Scale Processing in Point Cloud Segmentation · CVPR 2022 RGB-(D) scene labeling: Features and algorithms · CVPR 2012 |
Machine learning › Deep learning architectures and training
multi-scale feature fusion |
0.6 | 1 | 2022 | Pyramid Architecture for Multi-Scale Processing in Point Cloud Segmentation · CVPR 2022 |
Computer vision › 3D vision
point cloud segmentation |
0.6 | 1 | 2022 | Pyramid Architecture for Multi-Scale Processing in Point Cloud Segmentation · CVPR 2022 |
Computer vision › Image recognition and object detection › object detection
semi-supervised object detection |
0.6 | 1 | 2022 | SCMT: Self-Correction Mean Teacher for Semi-supervised Object Detection · IJCAI 2022 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.6 | 1 | 2022 | Doubly-Fused ViT: Fuse Information from Vision Transformer Doubly with Local Representation · ECCV (23) 2022 |
Computer vision › Image recognition and object detection
image classification |
0.4 | 3 | 2013 | Multipath Sparse Coding Using Hierarchical Matching Pursuit · CVPR 2013 Hierarchical Matching Pursuit for Image Classification: Architecture and Fast Algorithms · NIPS 2011 Kernel Descriptors for Visual Recognition · NIPS 2010 |
Computer vision › Image recognition and object detection
object recognition |
0.4 | 3 | 2011 | Object recognition with hierarchical kernel descriptors · CVPR 2011 A Scalable Tree-Based Approach for Joint Object and Pose Recognition · AAAI 2011 Figure-ground segmentation improves handled object recognition in egocentric video · CVPR 2010 |
Machine learning › Representation and self-supervised learning › representation learning › neural network representation learning › deep representation learning
deep feature representation |
0.3 | 1 | 2018 | Visual Search at Alibaba · KDD 2018 |
Multimedia analysis and retrieval
visual search |
0.3 | 1 | 2018 | Visual Search at Alibaba · KDD 2018 |
Computer vision › Image recognition and object detection › object recognition › multimodal object recognition
RGB-D object recognition |
0.3 | 3 | 2011 | Sparse distance learning for object recognition combining RGB and depth information · ICRA 2011 A large-scale hierarchical multi-view RGB-D object dataset · ICRA 2011 Object recognition with hierarchical kernel descriptors · CVPR 2011 |
Image and video processing › image segmentation
contour detection |
0.3 | 3 | 2012 | Discriminatively Trained Sparse Code Gradients for Contour Detection · NIPS 2012 Multi-scale Improves Boundary Detection in Natural Images · ECCV (3) 2008 Scale-Invariant Contour Completion Using Conditional Random Fields · ICCV 2005 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding |
0.3 | 2 | 2012 | Discriminatively Trained Sparse Code Gradients for Contour Detection · NIPS 2012 Hierarchical Matching Pursuit for Image Classification: Architecture and Fast Algorithms · NIPS 2011 |
Computer vision › 3D vision
3d object recognition |
0.2 | 2 | 2011 | Sparse distance learning for object recognition combining RGB and depth information · ICRA 2011 A large-scale hierarchical multi-view RGB-D object dataset · ICRA 2011 |
Computer vision › Image recognition and object detection
object discovery |
0.2 | 2 | 2011 | Toward object discovery and modeling via 3-D scene comparison · ICRA 2011 Learning to recognize objects in egocentric activities · CVPR 2011 |
Machine learning › Representation and self-supervised learning › visual representation › image representation › handcrafted descriptor
kernel descriptors |
0.2 | 2 | 2011 | Object recognition with hierarchical kernel descriptors · CVPR 2011 Kernel Descriptors for Visual Recognition · NIPS 2010 |
Machine learning › Representation and self-supervised learning › visual representation
patch-based representation |
0.2 | 2 | 2011 | Object recognition with hierarchical kernel descriptors · CVPR 2011 Kernel Descriptors for Visual Recognition · NIPS 2010 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.2 | 3 | 2011 | Learning to recognize objects in egocentric activities · CVPR 2011 Cue Integration for Figure/Ground Labeling · NIPS 2005 Recovering Human Body Configurations: Combining Segmentation and Recognition · CVPR (2) 2004 |
Computer vision › 3D vision
depth estimation |
0.2 | 1 | 2014 | Depth Enhancement via Low-Rank Matrix Completion · CVPR 2014 |
Computer vision › 3D vision › depth estimation › depth map refinement
RGB-D depth refinement |
0.2 | 1 | 2014 | Depth Enhancement via Low-Rank Matrix Completion · CVPR 2014 |
Image and video processing › image reconstruction
depth completion |
0.2 | 1 | 2014 | Depth Enhancement via Low-Rank Matrix Completion · CVPR 2014 |
Image and video processing › image enhancement
depth map enhancement |
0.2 | 1 | 2014 | Depth Enhancement via Low-Rank Matrix Completion · CVPR 2014 |
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
foreground-background segmentation |
0.2 | 2 | 2010 | Figure-ground segmentation improves handled object recognition in egocentric video · CVPR 2010 Tracking as Repeated Figure/Ground Segmentation · CVPR 2007 |
Image and video processing › image segmentation
contour completion |
0.2 | 3 | 2008 | Learning Probabilistic Models for Contour Completion in Natural Images · Int. J. Comput. Vis. 2008 Scale-Invariant Contour Completion Using Conditional Random Fields · ICCV 2005 A Probabilistic Multi-scale Model for Contour Completion Based on Image Statistics · ECCV (1) 2002 |
Machine learning › Deep learning architectures and training
teacher-student framework |
0.2 | 1 | 2022 | SCMT: Self-Correction Mean Teacher for Semi-supervised Object Detection · IJCAI 2022 |
Computer vision › 3D vision › motion estimation
3d motion estimation |
0.2 | 1 | 2013 | RGB-D flow: Dense 3-D motion estimation using color and depth · ICRA 2013 |
Computer vision › 3D vision
scene flow estimation |
0.2 | 1 | 2013 | RGB-D flow: Dense 3-D motion estimation using color and depth · ICRA 2013 |
Computer vision › 3D vision
3d object detection |
0.1 | 1 | 2012 | Detection-based object labeling in 3D scenes · ICRA 2012 |
Computer vision › 3D vision › 3d scene understanding
3d scene labeling |
0.1 | 1 | 2012 | Detection-based object labeling in 3D scenes · ICRA 2012 |
Computer vision › 3D vision
3d scene understanding |
0.1 | 1 | 2012 | Detection-based object labeling in 3D scenes · ICRA 2012 |
Methods — techniques the papers use, named apart from their topics
deep CNN · 1.0click behavior mining · 1.0model and search-based fusion · 0.7vision transformer · 0.6self-correction · 0.6pyramid architecture · 0.6pseudo-label re-weighting · 0.6mean teacher · 0.6cross-scale attention · 0.6orthogonal matching pursuit · 0.4sparse coding · 0.3subspace constraint · 0.2low-rank matrix completion · 0.2variational flow · 0.2RGB-D · 0.2shape and appearance analysis · 0.1object tracking · 0.1multi-scale pooling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identifying the Focus Word in Natural Language Questions Based on Association RulesabstractKnowledge base‐based intelligent question‐answering systems have insufficient understanding of the questions. In the early stages of research, it is effective in most cases that the existing natural language question‐understanding methods can answer questions by connecting entities and relationships when ignoring the identification of focus words. However, as research deepens, ignoring focus words has become a shortcoming. To address this, we propose identifying focus words, enabling more precise understanding of user focus. We define focus itemset, frequent focus itemset, focus association rule, and strong focus association rule to express focus‐related information better. Given the unique nature of focus association rules, we propose a prefix tree structure and an algorithm for mining association rules aimed at identifying focus words. We also introduce an inverted index specifically designed for focus association rules and propose an efficient algorithm for identifying focus words based on this index. Experiments verify the effectiveness of our algorithm and the efficiency of the inverted index, with a focus word identification rate exceeding 90%. Xin Hu 0008, Xiaofeng Ren, Jiangli Duan, Sulan Zhang |
Int. J. Intell. Syst. | 2 |
| 2023 | Fault Sensing of the Distribution Cable Feeders by Time-Domain MeasurementsabstractPower distribution cable feeders are prone to faults due to the poor working environment and internal defects. Considering the structure and electrical parameters of the typical distribution three-core cable, a fault sensing method is proposed. The fault sensing in this article includes three aspects: fault cable detection, fault phase identification, and fault resistance estimation. The grounding line currents of cable feeders are first analyzed in different neutral grounding modes. Based on the amplitudes and directions of the grounding line currents, a fault feeder detection criterion is then presented. Finally, based on the equivalent circuit of the fault cable, an algorithm for identifying the fault phase is developed by estimating the fault resistance. The experimental model of the distribution cable feeders is created by real time digital simulation (RTDS). Various fault experiments are carried out. The results show that the neutral grounding modes and fault conditions have little effect on the proposed method. Nan Peng, Zhengyi Zhang, Chengrui Jiang, Peng Zhang 0081, Xiaofeng Ren, Xiuru Wang |
IEEE Trans. Ind. Informatics | 6 |
| 2022 | Pyramid Architecture for Multi-Scale Processing in Point Cloud SegmentationabstractSemantic segmentation of point cloud data is a critical task for autonomous driving and other applications. Recent advances of point cloud segmentation are mainly driven by new designs of local aggregation operators and point sampling methods. Unlike image segmentation, few efforts have been made to understand the fundamental issue of scale and how scales should interact and be fused. In this work, we investigate how to efficiently and effectively integrate features at varying scales and varying stages in a point cloud segmentation network. In particular, we open up the commonly used encoder-decoder architecture, and design scale pyramid architectures that allow information to flow more freely and systematically, both laterally and upward/downward in scale. Moreover, a cross-scale attention feature learning block has been designed to enhance the multi-scale feature fusion which occurs everywhere in the network. Such a design of multi-scale processing and fusion gains large improvements in accuracy without adding much additional computation. When built on top of the popular KPConv network, we see consistent improvements on a wide range of datasets, including achieving state-of-the-art performance on NPM3D and S3DIS. Moreover, the pyramid architecture is generic and can be applied to other network designs: we show an example of similar improvements over RandLANet. Dong Nie, Rui Lan, Xiaofeng Ren |
CVPR | 4 |
| 2022 | Doubly-Fused ViT: Fuse Information from Vision Transformer Doubly with Local Representation
Dong Nie, Xiaofeng Ren |
ECCV (23) | 4 |
| 2022 | SCMT: Self-Correction Mean Teacher for Semi-supervised Object DetectionabstractSemi-Supervised Object Detection (SSOD) aims to improve performance by leveraging a large amount of unlabeled data. Existing works usually adopt the teacher-student framework to enforce student to learn consistent predictions over the pseudo-labels generated by teacher. However, the performance of the student model is limited since the noise inherently exists in pseudo-labels. In this paper, we investigate the causes and effects of noisy pseudo-labels and propose a simple yet effective approach denoted as Self-Correction Mean Teacher(SCMT) to reduce the adverse effects. Specifically, we propose to dynamically re-weight the unsupervised loss of each student's proposal with additional supervision information from the teacher model, and assign smaller loss weights to possible noisy proposals. Extensive experiments on MS-COCO benchmark have shown the superiority of our proposed SCMT, which can significantly improve the supervised baseline by more than 11% mAP under all 1%, 5% and 10% COCO-standard settings, and surpasses state-of-the-art methods by about 1.5% mAP. Even under the challenging COCO-additional setting, SCMT still improves the supervised baseline by 4.9% mAP, and significantly outperforms previous methods by 1.2% mAP, achieving a new state-of-the-art performance. Zhihui Hao, Yu-Lin He, Xiaofeng Ren |
IJCAI | 5 |
| 2020 | Bidirectional Pyramid Networks for Semantic Segmentation
Dong Nie, Jia Xue, Xiaofeng Ren |
ACCV (1) | 3 |
| 2018 | Visual Search at AlibabaabstractThis paper introduces the large scale visual search algorithm and system infrastructure at Alibaba. The following challenges are discussed under the E-commercial circumstance at Alibaba (a) how to handle heterogeneous image data and bridge the gap between real-shot images from user query and the online images. (b) how to deal with large scale indexing for massive updating data. (c) how to train deep models for effective feature representation without huge human annotations. (d) how to improve the user engagement by considering the quality of the content. We take advantage of large image collection of Alibaba and state-of-the-art deep learning techniques to perform visual search at scale. We present solutions and implementation details to overcome those problems and also share our learnings from building such a large scale commercial visual search engine. Specifically, model and search-based fusion approach is introduced to effectively predict categories. Also, we propose a deep CNN model for joint detection and feature learning by mining user click behavior. The binary index engine is designed to scale up indexing without compromising recall and precision. Finally, we apply all the stages into an end-to-end system architecture, which can simultaneously achieve highly efficient and scalable performance adapting to real-shot images. Extensive experiments demonstrate the advancement of each module in our system. We hope visual search at Alibaba becomes more widely incorporated into today's commercial applications. Yanhao Zhang 0002, Yingya Zhang, Xiaofeng Ren, Rong Jin 0001 |
KDD | 6 |
| 2014 | Depth Enhancement via Low-Rank Matrix CompletionabstractDepth captured by consumer RGB-D cameras is often noisy and misses values at some pixels, especially around object boundaries. Most existing methods complete the missing depth values guided by the corresponding color image. When the color image is noisy or the correlation between color and depth is weak, the depth map cannot be properly enhanced. In this paper, we present a depth map enhancement algorithm that performs depth map completion and de-noising simultaneously. Our method is based on the observation that similar RGB-D patches lie in a very low-dimensional subspace. We can then assemble the similar patches into a matrix and enforce this low-rank subspace constraint. This low-rank subspace constraint essentially captures the underlying structure in the RGB-D patches and enables robust depth enhancement against the noise or weak correlation between color and depth. Based on this subspace constraint, our method formulates depth map enhancement as a low-rank matrix completion problem. Since the rank of a matrix changes over matrices, we develop a data-driven method to automatically determine the rank number for each matrix. The experiments on both public benchmarks and our own captured RGB-D images show that our method can effectively enhance depth maps. Si Lu, Xiaofeng Ren, Feng Liu 0015 |
CVPR | 2 |
| 2013 | Multipath Sparse Coding Using Hierarchical Matching PursuitabstractComplex real-world signals, such as images, contain discriminative structures that differ in many aspects including scale, invariance, and data channel. While progress in deep learning shows the importance of learning features through multiple layers, it is equally important to learn features through multiple paths. We propose Multipath Hierarchical Matching Pursuit (M-HMP), a novel feature learning architecture that combines a collection of hierarchical sparse features for image classification to capture multiple aspects of discriminative structures. Our building blocks are MI-KSVD, a codebook learning algorithm that balances the reconstruction error and the mutual incoherence of the codebook, and batch orthogonal matching pursuit (OMP), we apply them recursively at varying layers and scales. The result is a highly discriminative image representation that leads to large improvements to the state-of-the-art on many standard benchmarks, e.g., Caltech-101, Caltech-256, MITScenes, Oxford-IIIT Pet and Caltech-UCSD Bird-200. Liefeng Bo, Xiaofeng Ren, Dieter Fox |
CVPR | 2 |
| 2013 | Histograms of Sparse Codes for Object DetectionabstractObject detection has seen huge progress in recent years, much thanks to the heavily-engineered Histograms of Oriented Gradients (HOG) features. Can we go beyond gradients and do better than HOG? We provide an affirmative answer by proposing and investigating a sparse representation for object detection, Histograms of Sparse Codes (HSC). We compute sparse codes with dictionaries learned from data using K-SVD, and aggregate per-pixel sparse codes to form local histograms. We intentionally keep true to the sliding window framework (with mixtures and parts) and only change the underlying features. To keep training (and testing) efficient, we apply dimension reduction by computing SVD on learned models, and adopt supervised training where latent positions of roots and parts are given externally e.g. from a HOG-based detector. By learning and using local representations that are much more expressive than gradients, we demonstrate large improvements over the state of the art on the PASCAL benchmark for both root-only and part-based models. Xiaofeng Ren, Deva Ramanan |
CVPR | 1 |
| 2013 | RGB-D flow: Dense 3-D motion estimation using color and depthabstract3-D motion estimation is a fundamental problem that has far-reaching implications in robotics. A scene flow formulation is attractive as it makes no assumptions about scene complexity, object rigidity, or camera motion. RGB-D cameras provide new information useful for computing dense 3-D flow in challenging scenes. In this work we show how to generalize two-frame variational 2-D flow algorithms to 3-D. We show that scene flow can be reliably computed using RGB-D data, overcoming depth noise and outperforming previous results on a variety of scenes. We apply dense 3-D flow to rigid motion segmentation. Evan Herbst, Xiaofeng Ren, Dieter Fox |
ICRA | 2 |
| 2013 | Research and applications: Ontology-guided organ detection to retrieve web images of disease manifestation: towards the construction of a consumer-based health image libraryabstractBACKGROUND: Visual information is a crucial aspect of medical knowledge. Building a comprehensive medical image base, in the spirit of the Unified Medical Language System (UMLS), would greatly benefit patient education and self-care. However, collection and annotation of such a large-scale image base is challenging. OBJECTIVE: To combine visual object detection techniques with medical ontology to automatically mine web photos and retrieve a large number of disease manifestation images with minimal manual labeling effort. METHODS: As a proof of concept, we first learnt five organ detectors on three detection scales for eyes, ears, lips, hands, and feet. Given a disease, we used information from the UMLS to select affected body parts, ran the pretrained organ detectors on web images, and combined the detection outputs to retrieve disease images. RESULTS: Compared with a supervised image retrieval approach that requires training images for every disease, our ontology-guided approach exploits shared visual information of body parts across diseases. In retrieving 2220 web images of 32 diseases, we reduced manual labeling effort to 15.6% while improving the average precision by 3.9% from 77.7% to 81.6%. For 40.6% of the diseases, we improved the precision by 10%. CONCLUSIONS: The results confirm the concept that the web is a feasible source for automatic disease image retrieval for health image database construction. Our approach requires a small amount of manual effort to collect complex disease images, and to annotate them by standard medical ontology terms. Yang Chen 0022, Xiaofeng Ren, Guo-Qiang Zhang 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | Scheduling Exploiting Frequency and Multi-User Diversity in LTE Downlink SystemsabstractScheduling can obtain multi-user diversity if channel state information (CSI) is known, such as for low-mobility users and can exploit frequency diversity if CSI is not available at the transmitter, such as for high-mobility users. In this paper, we investigate resource allocation exploiting frequency and multiuser diversity for LTE downlink systems with users of different mobilities. To facilitate resource allocation, we first develop a user classification algorithm to identify high- and low-mobility users. Based on user mobility classification, we then propose a scheduling algorithm to simultaneously obtain multi-user diversity for those low-mobility users and frequency diversity for those high-mobility users. It is demonstrated by computer simulation that the performance of the proposed scheduling algorithm provides 6% and 23% gain of overall cell throughput, and 5.6% and 18% gain of 10th percentile throughput over proportional fairness based frequency-selective and frequency-diversity scheduling algorithms, respectively. Furthermore, the proposed scheduling algorithm has the same order of computational complexity as the frequency-selective scheduling algorithm. Jinping Niu, Dae-Won Lee, Xiaofeng Ren, Geoffrey Ye Li |
IEEE Trans. Wirel. Commun. | 3 |
| 2013 | User Classification and Scheduling in LTE Downlink Systems with Heterogeneous User MobilitiesabstractIn LTE systems with heterogeneous user mobilities, low-mobility users favor frequency selective scheduling while high-mobility users benefit from frequency diversity scheduling. To benefit both low- and high-mobility users simultaneously, scheduling exploiting frequency selectivity and diversity is desired. To enable the scheduling, low-complexity user mobility classification to distinguish these two types of users is required. In this paper, we first propose a user mobility classification algorithm, which is robust to different channel delay profiles (CDPs), for single-transmit-antenna systems. Then, we extend it to multiple-input multiple-output (MIMO) systems. A low-complexity scheduling algorithm, exploiting both frequency-selectivity and diversity for low- and high-mobility users simultaneously, is also developed. As demonstrated by the simulation results, the proposed user classification algorithm is robust to different CDPs and the proposed scheduling algorithm is effective. Jinping Niu, Dae-Won Lee, Geoffrey Ye Li, Xiaofeng Ren |
IEEE Trans. Wirel. Commun. | 5 |
| 2012 | SensorSift: balancing sensor data privacy and utility in automated face understandingabstractWe introduce SensorSift, a new theoretical scheme for balancing utility and privacy in smart sensor applications. At the heart of our contribution is an algorithm which transforms raw sensor data into a 'sifted' representation which minimizes exposure of user defined private attributes while maximally exposing application-requested public attributes. We envision multiple applications using the same platform, and requesting access to public attributes explicitly not known at the time of the platform creation. Support for future-defined public attributes, while still preserving the defined privacy of the private attributes, is a central challenge that we tackle. Miro Enev, Jaeyeon Jung, Liefeng Bo, Xiaofeng Ren, Tadayoshi Kohno |
ACSAC | 4 |
| 2012 | RGB-(D) scene labeling: Features and algorithmsabstractScene labeling research has mostly focused on outdoor scenes, leaving the harder case of indoor scenes poorly understood. Microsoft Kinect dramatically changed the landscape, showing great potentials for RGB-D perception (color+depth). Our main objective is to empirically understand the promises and challenges of scene labeling with RGB-D. We use the NYU Depth Dataset as collected and analyzed by Silberman and Fergus [30]. For RGB-D features, we adapt the framework of kernel descriptors that converts local similarities (kernels) to patch descriptors. For contextual modeling, we combine two lines of approaches, one using a superpixel MRF, and the other using a segmentation tree. We find that (1) kernel descriptors are very effective in capturing appearance (RGB) and shape (D) similarities; (2) both superpixel MRF and segmentation tree are useful in modeling context; and (3) the key to labeling accuracy is the ability to efficiently train and test with large-scale data. We improve labeling accuracy on the NYU Dataset from 56.6% to 76.1%. We also apply our approach to image-only scene labeling and improve the accuracy on the Stanford Background Dataset from 79.4% to 82.9%. Xiaofeng Ren, Liefeng Bo, Dieter Fox |
CVPR | 1 |
| 2012 | Fine-grained kitchen activity recognition using RGB-DabstractWe present a first study of using RGB-D (Kinect-style) cameras for fine-grained recognition of kitchen activities. Our prototype system combines depth (shape) and color (appearance) to solve a number of perception problems crucial for smart space applications: locating hands, identifying objects and their functionalities, recognizing actions and tracking object state changes through actions. Our proof-of-concept results demonstrate great potentials of RGB-D perception: without need for instrumentation, our system can robustly track and accurately recognize detailed steps through cooking activities, for instance how many spoons of sugar are in a cake mix, or how long it has been mixing. A robust RGB-D based solution to fine-grained activity recognition in real-world conditions will bring the intelligence of pervasive and interactive systems to the next level. Jinna Lei, Xiaofeng Ren, Dieter Fox |
UbiComp | 2 |
| 2012 | Detection-based object labeling in 3D scenesabstractWe propose a view-based approach for labeling objects in 3D scenes reconstructed from RGB-D (color+depth) videos. We utilize sliding window detectors trained from object views to assign class probabilities to pixels in every RGB-D frame. These probabilities are projected into the reconstructed 3D scene and integrated using a voxel representation. We perform efficient inference on a Markov Random Field over the voxels, combining cues from view-based detection and 3D shape, to label the scene. Our detection-based approach produces accurate scene labeling on the RGB-D Scenes Dataset and improves the robustness of object detection. Kevin Lai 0001, Liefeng Bo, Xiaofeng Ren, Dieter Fox |
ICRA | 3 |
| 2012 | Discriminatively Trained Sparse Code Gradients for Contour DetectionabstractFinding contours in natural images is a fundamental problem that serves as the basis of many tasks such as image segmentation and object recognition. At the core of contour detection technologies are a set of hand-designed gradient features, used by most existing approaches including the state-of-the-art Global Pb (gPb) operator. In this work, we show that contour detection accuracy can be significantly improved by computing Sparse Code Gradients (SCG), which measure contrast using patch representations automatically learned through sparse coding. We use K-SVD and Orthogonal Matching Pursuit for efficient dictionary learning and encoding, and use multi-scale pooling and power transforms to code oriented local neighborhoods before computing gradients and applying linear SVM. By extracting rich representations from pixels and avoiding collapsing them prematurely, Sparse Code Gradients effectively learn how to measure local contrasts and find contours. We improve the F-measure metric on the BSDS500 benchmark to 0.74 (up from 0.71 of gPb contours). Moreover, our learning approach can easily adapt to novel sensor data such as Kinect-style RGB-D cameras: Sparse Code Gradients on depth images and surface normals lead to promising contour detection using depth and depth+color, as verified on the NYU Depth Dataset. Our work combines the concept of oriented gradients with sparse representation and opens up future possibilities for learning contour detection and segmentation. Xiaofeng Ren, Liefeng Bo |
NIPS | 1 |
| 2012 | Scheduling exploiting frequency and multi-user diversity in LTE downlink systemsabstractIn this paper, we develop a scheduling algorithm to obtain multi-user diversity for those low-mobility users and frequency diversity for those high-mobility users. Computer simulation demonstrates that the proposed scheduling algorithm provides 10% and 16% overall cell throughput gain over proportional fairness based frequency-selective and frequency-diversity scheduling algorithm, respectively. In addition, the proposed scheduling algorithm is shown to have the same order of computational complexity as the frequency-selective scheduling algorithm, and can be easily implemented in the LTE downlink systems. Jinping Niu, Dae-Won Lee, Xiaofeng Ren, Geoffrey Ye Li |
PIMRC | 3 |
| 2011 | A Scalable Tree-Based Approach for Joint Object and Pose RecognitionabstractRecognizing possibly thousands of objects is a crucial capability for an autonomous agent to understand and interact with everyday environments. Practical object recognition comes in multiple forms: Is this a coffee mug (category recognition). Is this Alice's coffee mug? (instance recognition). Is the mug with the handle facing left or right? (pose recognition). We present a scalable framework, Object-Pose Tree, which efficiently organizes data into a semantically structured tree. The tree structure enables both scalable training and testing, allowing us to solve recognition over thousands of object poses in near real-time. Moreover, by simultaneously optimizing all three tasks, our approach outperforms standard nearest neighbor and 1-vs-all classifications, with large improvements on pose recognition. We evaluate the proposed technique on a dataset of 300 household objects collected using a Kinect-style 3D camera. Experiments demonstrate that our system achieves robust and efficient object category, instance, and pose recognition on challenging everyday objects. Kevin Lai 0001, Liefeng Bo, Xiaofeng Ren, Dieter Fox |
AAAI | 3 |
| 2011 | Combining Self Training and Active Learning for Video SegmentationabstractPresented at the 22nd British Machine Vision Conference (BMVC 2011), 29 August-2 September 2011, University of Dundee, Scotland, UK. Alireza Fathi, Maria-Florina Balcan, Xiaofeng Ren, James M. Rehg |
BMVC | 3 |
| 2011 | Toward Robust Material Recognition for Everyday ObjectsabstractMaterial recognition is a fundamental problem in perception that is receiving increasing attention. Following the recent work using Flickr [16, 23], we empirically study material recognition of real-world objects using a rich set of local features. We use the Kernel Descriptor framework [5] and extend the set of descriptors to include material-motivated attributes using variances of gradient orientation and magnitude. Large-Margin Nearest Neighbor learning is used for a 30-fold dimension reduction. We improve the state-of-the-art accuracy on the Flickr dataset [16] from 45 % to 54%. We also introduce two new datasets using ImageNet and macro photos, extensively evaluating our set of features and showing promising connections between material and object recognition. Diane Hu, Liefeng Bo, Xiaofeng Ren |
BMVC | 3 |
| 2011 | HeatWave: thermal imaging for surface user interactionabstractWe present HeatWave, a system that uses digital thermal imaging cameras to detect, track, and support user interaction on arbitrary surfaces. Thermal sensing has had limited examination in the HCI research community and is generally under-explored outside of law enforcement and energy auditing applications. We examine the role of thermal imaging as a new sensing solution for enhancing user surface interaction. In particular, we demonstrate how thermal imaging in combination with existing computer vision techniques can make segmentation and detection of routine interaction techniques possible in real-time, and can be used to complement or simplify algorithms for traditional RGB and depth cameras. Example interactions include (1) distinguishing hovering above a surface from touch events, (2) shape-based gestures similar to ink strokes, (3) pressure based gestures, and (4) multi-finger gestures. We close by discussing the practicality of thermal sensing for naturalistic user interaction and opportunities for future work. Eric C. Larson, Gabe Cohn, Sidhant Gupta, Xiaofeng Ren, Beverly L. Harrison, Dieter Fox, Shwetak N. Patel |
CHI | 4 |
| 2011 | Object recognition with hierarchical kernel descriptorsabstractKernel descriptors provide a unified way to generate rich visual feature sets by turning pixel attributes into patch-level features, and yield impressive results on many object recognition tasks. However, best results with kernel descriptors are achieved using efficient match kernels in conjunction with nonlinear SVMs, which makes it impractical for large-scale problems. In this paper, we propose hierarchical kernel descriptors that apply kernel descriptors recursively to form image-level features and thus provide a conceptually simple and consistent way to generate image-level features from pixel attributes. More importantly, hierarchical kernel descriptors allow linear SVMs to yield state-of-the-art accuracy while being scalable to large datasets. They can also be naturally extended to extract features over depth images. We evaluate hierarchical kernel descriptors both on the CIFAR10 dataset and the new RGB-D Object Dataset consisting of segmented RGB and depth images of 300 everyday objects. Liefeng Bo, Kevin Lai 0001, Xiaofeng Ren, Dieter Fox |
CVPR | 3 |
| 2011 | Learning to recognize objects in egocentric activitiesabstractThis paper addresses the problem of learning object models from egocentric video of household activities, using extremely weak supervision. For each activity sequence, we know only the names of the objects which are present within it, and have no other knowledge regarding the appearance or location of objects. The key to our approach is a robust, unsupervised bottom up segmentation method, which exploits the structure of the egocentric domain to partition each frame into hand, object, and background categories. By using Multiple Instance Learning to match object instances across sequences, we discover and localize object occurrences. Object representations are refined through transduction and object-level classifiers are trained. We demonstrate encouraging results in detecting novel object instances using models produced by weakly-supervised learning. Alireza Fathi, Xiaofeng Ren, James M. Rehg |
CVPR | 2 |
| 2011 | Interactive 3D modeling of indoor environments with a consumer depth cameraabstractDetailed 3D visual models of indoor spaces, from walls and floors to objects and their configurations, can provide extensive knowledge about the environments as well as rich contextual information of people living therein. Vision-based 3D modeling has only seen limited success in applications, as it faces many technical challenges that only a few experts understand, let alone solve. In this work we utilize (Kinect style) consumer depth cameras to enable non-expert users to scan their personal spaces into 3D models. We build a prototype mobile system for 3D modeling that runs in real-time on a laptop, assisting and interacting with the user on-the-fly. Color and depth are jointly used to achieve robust 3D registration. The system offers online feedback and hints, tolerates human errors and alignment failures, and helps to obtain complete scene coverage. We show that our prototype system can both scan large environments (50 meters across) and at the same time preserve fine details (centimeter accuracy). The capability of detailed 3D modeling leads to many promising applications such as accurate 3D localization, measuring dimensions, and interactive visualization. Hao Du 0004, Peter Henry, Xiaofeng Ren, Marvin Cheng, Dan B. Goldman, Steven M. Seitz, Dieter Fox |
UbiComp | 3 |
| 2011 | Toward object discovery and modeling via 3-D scene comparisonabstractThe performance of indoor robots that stay in a single environment can be enhanced by gathering detailed knowledge of objects that frequently occur in that environment. We use an inexpensive sensor providing dense color and depth, and fuse information from multiple sensing modalities to detect changes between two 3-D maps. We adapt a recent SLAM technique to align maps. A probabilistic model of sensor readings lets us reason about movement of surfaces. Our method handles arbitrary shapes and motions, and is robust to lack of texture. We demonstrate the ability to find whole objects in complex scenes by regularizing over surface patches. Evan Herbst, Peter Henry, Xiaofeng Ren, Dieter Fox |
ICRA | 3 |
| 2011 | A large-scale hierarchical multi-view RGB-D object datasetabstractOver the last decade, the availability of public image repositories and recognition benchmarks has enabled rapid progress in visual object category and instance detection. Today we are witnessing the birth of a new generation of sensing technologies capable of providing high quality synchronized videos of both color and depth, the RGB-D (Kinect-style) camera. With its advanced sensing capabilities and the potential for mass adoption, this technology represents an opportunity to dramatically increase robotic object recognition, manipulation, navigation, and interaction capabilities. In this paper, we introduce a large-scale, hierarchical multi-view object dataset collected using an RGB-D camera. The dataset contains 300 objects organized into 51 categories and has been made publicly available to the research community so as to enable rapid progress based on this promising technology. This paper describes the dataset collection procedure and introduces techniques for RGB-D based object recognition and detection, demonstrating that combining color and depth information substantially improves quality of results. Kevin Lai 0001, Liefeng Bo, Xiaofeng Ren, Dieter Fox |
ICRA | 3 |
| 2011 | Sparse distance learning for object recognition combining RGB and depth informationabstractIn this work we address joint object category and instance recognition in the context of RGB-D (depth) cameras. Motivated by local distance learning, where a novel view of an object is compared to individual views of previously seen objects, we define a view-to-object distance where a novel view is compared simultaneously to all views of a previous object. This novel distance is based on a weighted combination of feature differences between views. We show, through jointly learning per-view weights, that this measure leads to superior classification performance on object category and instance recognition. More importantly, the proposed distance allows us to find a sparse solution via Group-Lasso regularization, where a small subset of representative views of an object is identified and used, with the rest discarded. This significantly reduces computational cost without compromising recognition accuracy. We evaluate the proposed technique, Instance Distance Learning (IDL), on the RGB-D Object Dataset, which consists of 300 object instances in 51 everyday categories and about 250,000 views of objects with both RGB color and depth. We empirically compare IDL to several alternative state-of-the-art approaches and also validate the use of visual and shape cues and their combination. Kevin Lai 0001, Liefeng Bo, Xiaofeng Ren, Dieter Fox |
ICRA | 3 |
| 2011 | Depth kernel descriptors for object recognitionabstractConsumer depth cameras, such as the Microsoft Kinect, are capable of providing frames of dense depth values at real time. One fundamental question in utilizing depth cameras is how to best extract features from depth frames. Motivated by local descriptors on images, in particular kernel descriptors, we develop a set of kernel features on depth images that model size, 3D shape, and depth edges in a single framework. Through extensive experiments on object recognition, we show that (1) our local features capture different aspects of cues from a depth frame/view that complement one another; (2) our kernel features significantly outperform traditional 3D features (e.g. Spin images); and (3) we significantly improve the capabilities of depth and RGB-D (color+depth) recognition, achieving 10–15% improvement in accuracy over the state of the art. Liefeng Bo, Xiaofeng Ren, Dieter Fox |
IROS | 2 |
| 2011 | RGB-D object discovery via multi-scene analysisabstractWe introduce an algorithm for object discovery from RGB-D (color plus depth) data, building on recent progress in using RGB-D cameras for 3-D reconstruction. A set of 3-D maps are built from multiple visits to the same scene. We introduce a multi-scene MRF model to detect objects that moved between visits, combining shape, visibility, and color cues. We measure similarities between candidate objects using both 2-D and 3-D matching, and apply spectral clustering to infer object clusters from noisy links. Our approach can robustly detect objects and their motion between scenes even when objects are textureless or have the same shape as other objects. Evan Herbst, Xiaofeng Ren, Dieter Fox |
IROS | 2 |
| 2011 | Hierarchical Matching Pursuit for Image Classification: Architecture and Fast AlgorithmsabstractExtracting good representations from images is essential for many computer vision tasks. In this paper, we propose hierarchical matching pursuit (HMP), which builds a feature hierarchy layer-by-layer using an efficient matching pursuit encoder. It includes three modules: batch (tree) orthogonal matching pursuit, spatial pyramid max pooling, and contrast normalization. We investigate the architecture of HMP, and show that all three components are critical for good performance. To speed up the orthogonal matching pursuit, we propose a batch tree orthogonal matching pursuit that is particularly suitable to encode a large number of observations that share the same large dictionary. HMP is scalable and can efficiently handle full-size images. In addition, HMP enables linear support vector machines (SVM) to match the performance of nonlinear SVM while being scalable to large datasets. We compare HMP with many state-of-the-art algorithms including convolutional deep belief networks, SIFT based single layer sparse coding, and kernel based feature learning. HMP consistently yields superior accuracy on three types of image classification problems: object recognition (Caltech-101), scene recognition (MIT-Scene), and static event recognition (UIUC-Sports). Liefeng Bo, Xiaofeng Ren, Dieter Fox |
NIPS | 2 |
| 2010 | Figure-ground segmentation improves handled object recognition in egocentric videoabstractIdentifying handled objects, i.e. objects being manipulated by a user, is essential for recognizing the person's activities. An egocentric camera as worn on the body enjoys many advantages such as having a natural first-person view and not needing to instrument the environment. It is also a challenging setting, where background clutter is known to be a major source of problems and is difficult to handle with the camera constantly and arbitrarily moving. In this work we develop a bottom-up motion-based approach to robustly segment out foreground objects in egocentric video and show that it greatly improves object recognition accuracy. Our key insight is that egocentric video of object manipulation is a special domain and many domain-specific cues can readily help. We compute dense optical flow and fit it into multiple affine layers. We then use a max-margin classifier to combine motion with empirical knowledge of object location and background movement as well as temporal cues of support region and color appearance. We evaluate our segmentation algorithm on the large Intel Egocentric Object Recognition dataset with 42 objects and 100K frames. We show that, when combined with temporal integration, figure-ground segmentation improves the accuracy of a SIFT-based recognition system from 33% to 60%, and that of a latent-HOG system from 64% to 86%. Xiaofeng Ren, Chunhui Gu |
CVPR | 1 |
| 2010 | Discriminative Mixture-of-Templates for Viewpoint Classification
Chunhui Gu, Xiaofeng Ren |
ECCV (5) | 2 |
| 2010 | Kernel Descriptors for Visual RecognitionabstractThe design of low-level image features is critical for computer vision algorithms. Orientation histograms, such as those in SIFT~\cite{Lowe2004Distinctive} and HOG~\cite{Dalal2005Histograms}, are the most successful and popular features for visual object and scene recognition. We highlight the kernel view of orientation histograms, and show that they are equivalent to a certain type of match kernels over image patches. This novel view allows us to design a family of kernel descriptors which provide a unified and principled framework to turn pixel attributes (gradient, color, local binary pattern, \etc) into compact patch-level features. In particular, we introduce three types of match kernels to measure similarities between image patches, and construct compact low-dimensional kernel descriptors from these match kernels using kernel principal component analysis (KPCA)~\cite{Scholkopf1998Nonlinear}. Kernel descriptors are easy to design and can turn any type of pixel attribute into patch-level features. They outperform carefully tuned and sophisticated features including SIFT and deep belief networks. We report superior performance on standard image classification benchmarks: Scene-15, Caltech-101, CIFAR10 and CIFAR10-ImageNet. Liefeng Bo, Xiaofeng Ren, Dieter Fox |
NIPS | 2 |
| 2008 | Finding people in archive films through trackingabstractThe goal of this work is to find all people in archive films. Challenges include low image quality, motion blur, partial occlusion, non-standard poses and crowded scenes. We base our approach on face detection and take a tracking/temporal approach to detection. Our tracker operates in two modes, following face detections whenever possible, switching to low-level tracking if face detection fails. With temporal correspondences established by tracking, we formulate detection as an inference problem in one-dimensional chains/tracks. We use a conditional random field model to integrate information across frames and to re-score tentative detections in tracks. Quantitative evaluations on full-length films show that the CRF-based temporal detector greatly improves face detection, increasing precision for about 30% (suppressing isolated false positives) and at the same time boosting recall for over 10% (recovering difficult cases where face detectors fail). Xiaofeng Ren |
CVPR | 1 |
| 2008 | Local grouping for optical flowabstractOptical flow estimation requires spatial integration, which essentially poses a grouping question: what points belong to the same motion and what do not. Classical local approaches to optical flow, such as Lucas-Kanade, use isotropic neighborhoods and have considerable difficulty near motion boundaries. In this work we utilize image-based grouping to facilitate spatial- and scale-adaptive integration. We define soft spatial support using pairwise affinities computed through intervening contour. We sample images at edges and corners, and iteratively estimate affine motion at sample points. Figure-ground organization further improves grouping and flow estimation near boundaries. We show that affinity-based spatial integration enables reliable flow estimation and avoids erroneous motion propagation from and/or across object boundaries. We demonstrate our approach on the Middlebury flow dataset. Xiaofeng Ren |
CVPR | 1 |
| 2008 | Multi-scale Improves Boundary Detection in Natural Images
Xiaofeng Ren |
ECCV (3) | 1 |
| 2008 | Learning Probabilistic Models for Contour Completion in Natural Images
Xiaofeng Ren, Charless C. Fowlkes, Jitendra Malik |
Int. J. Comput. Vis. | 1 |
| 2007 | Learning and Matching Line Aspects for Articulated ObjectsabstractTraditional aspect graphs are topology-based and are impractical for articulated objects. In this work we learn a small number of aspects, or prototypical views, from video data. Groundtruth segmentations in video sequences are utilized for both training and testing aspect models that operate on static images. We represent aspects of an articulated object as collections of line segments. In learning aspects, where object centers are known, a linear matching based on line location and orientation is used to measure similarity between views. We use K-medoid to find cluster centers. When using line aspects in recognition, matching is based on pairwise cues of relative location, relative orientation as well adjacency and parallelism. Matching with pairwise cues leads to a quadratic optimization that we solve with a spectral approximation. We show that our line aspect matching is capable of locating people in a variety of poses. Line aspect matching performs significantly better than an alternative approach using Hausdorff distance, showing merits of the line representation. Xiaofeng Ren |
CVPR | 1 |
| 2007 | Tracking as Repeated Figure/Ground SegmentationabstractTracking over a long period of time is challenging as the appearance, shape and scale of the object in question may vary. We propose a paradigm of tracking by repeatedly segmenting figure from background. Accurate spatial support obtained in segmentation provides rich information about the track and enables reliable tracking of non-rigid objects without drifting. Figure/ground segmentation operates sequentially in each frame by utilizing both static image cues and temporal coherence cues, which include an appearance model of brightness (or color) and a spatial model propagating figure/ground masks through low-level region correspondence. A superpixel-based conditional random field linearly combines cues and loopy belief propagation is used to estimate marginal posteriors of figure vs background. We demonstrate our approach on long sequences of sports video, including figure skating and football. Xiaofeng Ren, Jitendra Malik |
CVPR | 1 |
| 2006 | Figure/Ground Assignment in Natural Images
Xiaofeng Ren, Charless C. Fowlkes, Jitendra Malik |
ECCV (2) | 1 |
| 2005 | Recovering Human Body Configurations Using Pairwise Constraints between PartsabstractThe goal of this work is to recover human body configurations from static images. Without assuming a priori knowledge of scale, pose or appearance, this problem is extremely challenging and demands the use of all possible sources of information. We develop a framework which can incorporate arbitrary pairwise constraints between body parts, such as scale compatibility, relative position, symmetry of clothing and smooth contour connections between parts. We detect candidate body parts from bottom-up using parallelism, and use various pairwise configuration constraints to assemble them together into body configurations. To find the most probable configuration, we solve an integer quadratic programming problem with a standard technique using linear approximations. Approximate IQP allows us to incorporate much more information than the traditional dynamic programming and remains computationally efficient. 15 hand-labeled images are used to train the low-level part detector and learn the pairwise constraints. We show test results on a variety of images. Xiaofeng Ren, Alexander C. Berg, Jitendra Malik |
ICCV | 1 |
| 2005 | Scale-Invariant Contour Completion Using Conditional Random FieldsabstractWe present a model of curvilinear grouping using piece-wise linear representations of contours and a conditional random field to capture continuity and the frequency of different junction types. Potential completions are generated by building a constrained Delaunay triangulation (CDT) over the set of contours found by a local edge detector. Maximum likelihood parameters for the model are learned from human labeled ground truth. Using held out test data, we measure how the model, by incorporating continuity structure, improves boundary detection over the local edge detector. We also compare performance with a baseline local classifier that operates on pairs of edgels. Both algorithms consistently dominate the low-level boundary detector at all thresholds. To our knowledge, this is the first time that curvilinear continuity has been shown quantitatively useful for a large variety of natural images. Better boundary detection has immediate application in the problem of object detection and recognition. Xiaofeng Ren, Charless C. Fowlkes, Jitendra Malik |
ICCV | 1 |
| 2005 | Cue Integration for Figure/Ground LabelingabstractWe present a model of edge and region grouping using a conditional random field built over a scale-invariant representation of images to integrate multiple cues. Our model includes potentials that capture low-level similarity, mid-level curvilinear continuity and high-level object shape. Maximum likelihood parameters for the model are learned from human labeled groundtruth on a large collection of horse images using belief propagation. Using held out test data, we quantify the information gained by incorporating generic mid-level cues and high-level shape. Xiaofeng Ren, Charless C. Fowlkes, Jitendra Malik |
NIPS | 1 |
| 2004 | Recovering Human Body Configurations: Combining Segmentation and Recognition
Greg Mori, Xiaofeng Ren, Alexei A. Efros, Jitendra Malik |
CVPR (2) | 2 |
| 2003 | Learning a Classification Model for SegmentationabstractWe propose a two-class classification model for grouping. Human segmented natural images are used as positive examples. Negative examples of grouping are constructed by randomly matching human segmentations and images. In a preprocessing stage an image is over-segmented into super-pixels. We define a variety of features derived from the classical Gestalt cues, including contour, texture, brightness and good continuation. Information-theoretic analysis is applied to evaluate the power of these grouping cues. We train a linear classifier to combine these features. To demonstrate the power of the classification model, a simple algorithm is used to randomly search for good segmentations. Results are shown on a wide range of images. Xiaofeng Ren, Jitendra Malik |
ICCV | 1 |
| 2002 | A Probabilistic Multi-scale Model for Contour Completion Based on Image Statistics
Xiaofeng Ren, Jitendra Malik |
ECCV (1) | 1 |