Haibing Ren

dblp:16/6435 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
10since 2021 · last 2024
0009-0002-5751-7062ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021Systems, architecture and hardware · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Integrating instance-level knowledge to see the unseen: A two-stream network for video object segmentation
Hannan Lu, Zhi Tian, Pengxu Wei, Haibing Ren, Wangmeng Zuo
Neurocomputing4
2023 Adaptive Zone-aware Hierarchical Planner for Vision-Language Navigation
abstract
The task of Vision-Language Navigation (VLN) is for an embodied agent to reach the global goal according to the instruction. Essentially, during navigation, a series of sub-goals need to be adaptively set and achieved, which is naturally a hierarchical navigation process. However, previous methods leverage a single-step planning scheme, i.e., directly performing navigation action at each step, which is unsuitable for such a hierarchical navigation process. In this paper, we propose an Adaptive Zone-aware Hierarchical Planner (AZHP) to explicitly divides the navigation process into two heterogeneous phases, i.e., sub-goal setting via zone partition/selection (high-level action) and sub-goal executing (low-level action), for hierarchical planning. Specifically, AZHP asynchronously performs two levels of action via the designed State-Switcher Module (SSM). For high-level action, we devise a Scene-aware adaptive Zone Partition (SZP) method to adaptively divide the whole navigation area into different zones on-the-fly. Then the Goal-oriented Zone Selection (GZS) method is proposed to select a proper zone for the current sub-goal. For low-level action, the agent conducts navigation-decision multi-steps in the selected zone. Moreover, we design a Hierarchical RL (HRL) strategy and auxiliary losses with curriculum learning to train the AZHP framework, which provides effective supervision signals for each stage. Extensive experiments demonstrate the superiority of our proposed method, which achieves state-of-the-art performance on three VLN benchmarks (REVERIE, SOON, R2R).
Chen Gao 0005, Xingyu Peng, Mi Yan, He Wang 0010, Lirong Yang, Haibing Ren, Hongsheng Li 0001, Si Liu 0001
CVPR6
2022 3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection
abstract
3D visual grounding aims to locate the referred target object in 3D point cloud scenes according to a free-form language description. Previous methods mostly follow a two-stage paradigm, i.e., language-irrelevant detection and cross-modal matching, which is limited by the isolated architecture. In such a paradigm, the detector needs to sample keypoints from raw point clouds due to the inherent properties of 3D point clouds (irregular and large-scale), to generate the corresponding object proposal for each keypoint. However, sparse proposals may leave out the target in detection, while dense proposals may confuse the matching model. Moreover, the language-irrelevant detection stage can only sample a small proportion of keypoints on the target, deteriorating the target prediction. In this paper, we propose a 3D Single-Stage Referred Point Progressive Selection (3D-SPS) method, which progressively selects keypoints with the guidance of language and directly locates the target. Specifically, we propose a Description-aware Keypoint Sampling (DKS) module to coarsely focus on the points of language-relevant objects, which are significant clues for grounding. Besides, we devise a Target-oriented Progressive Mining (TPM) module to finely concentrate on the points of the target, which is enabled by progressive intra-modal relation modeling and inter-modal target mining. 3D-SPS bridges the gap between detection and matching in the 3D visual grounding task, localizing the target at a single stage. Experiments demonstrate that 3D-SPS achieves state-of-the-art performance on both ScanRe-fer and Nr3D/Sr3D datasets.
Junyu Luo 0002, Jiahui Fu 0003, Xianghao Kong, Chen Gao 0005, Haibing Ren, Huaxia Xia, Si Liu 0001
CVPR5
2022 PromptDet: Towards Open-Vocabulary Detection Using Uncurated Images
Chengjian Feng, Zequn Jie, Xiangxiang Chu, Haibing Ren, Xiaolin Wei, Weidi Xie, Lin Ma 0002
ECCV (9)5
2022 IPS300+: a Challenging multi-modal data sets for Intersection Perception System
abstract
Due to high complexity and occlusion, insufficient perception in the crowded urban intersection can be a serious safety risk for both human drivers and autonomous algorithms, whereas CVIS (Cooperative Vehicle Infrastructure System) is a proposed solution for full-participants perception under this scenario. However, the research on roadside multi-modal perception is still in its infancy, and there is no open-source data sets for such scene. Accordingly, this paper fills the gap. Through an IPS (Intersection Perception System) installed at the diagonal of the intersection, this paper proposes a high-quality multi-modal data sets for the intersection perception task. The center of the experimental intersection covers an area of 3000m2, and the extended distance reaches 300m, which is typical for CVIS. The first batch of open-source data includes 14198 frames, and each frame has an average of 319.84 labels, which is 9.6 times larger than the most crowded data sets (H3D data sets in 2019) by now. Our data sets is available at: http://www.openmpd.com/column/IPS300.
Huanan Wang, Xinyu Zhang 0001, Zhiwei Li 0011, Jun Li 0082, Zhu Lei, Haibing Ren
ICRA7
2022 InterFusion: Interaction-based 4D Radar and LiDAR Fusion for 3D Object Detection
abstract
Many recent works detect 3D objects by several sensor modalities for autonomous driving, where high-resolution cameras and high-line LiDARs are mostly used but relatively expensive. To achieve a balance between overall cost and detection accuracy, many multi-modal fusion techniques have been suggested. In recent years, the fusion of LiDAR and Radar has gained ever-increasing attention, especially 4D Radar, which can adapt to bad weather conditions due to its penetrability. Although features have been fused from multiple sensing modalities, most methods cannot learn interactions from different modalities, which does not make for their best use. Inspired by the self-attention mechanism, we present InterFusion, an interaction-based fusion framework, to fuse 16-line LiDAR with 4D Radar. It aggregates features from two modalities and identifies cross-modal relations between Radar and LiDAR features. In experimental evaluations on the Astyx HiRes 2019 dataset, our method outperformed the baseline by 4.20% mAP in 3D and 10.76% BEV mAP for the car class at the moderate level.
Li Wang 0092, Xinyu Zhang 0001, Baowei Xv, Jinzhao Zhang, Haibing Ren, Pingping Lu, Jun Li 0082, Huaping Liu 0001
IROS8
2022 Target-Driven Structured Transformer Planner for Vision-Language Navigation
abstract
Vision-language navigation is the task of directing an embodied agent to navigate in 3D scenes with natural language instructions. For the agent, inferring the long-term navigation target from visual-linguistic clues is crucial for reliable path planning, which, however, has rarely been studied before in literature. In this article, we propose a Target-Driven Structured Transformer Planner (TD-STP) for long-horizon goal-guided and room layout-aware navigation. Specifically, we devise an Imaginary Scene Tokenization mechanism for explicit estimation of the long-term target (even located in unexplored environments). In addition, we design a Structured Transformer Planner which elegantly incorporates the explored room layout into a neural attention architecture for structured and global planning. Experimental results demonstrate that our TD-STP substantially improves previous best methods' success rate by 2% and 5% on the test set of R2R and REVERIE benchmarks, respectively. Our code is available at https://github.com/YushengZhao/TD-STP.
Yusheng Zhao, Chen Gao 0005, Wenguan Wang, Lirong Yang, Haibing Ren, Huaxia Xia, Si Liu 0001
ACM Multimedia6
2022 Design change propagation routing in the modular product
Haibing Ren, Qinming Liu
Adv. Eng. Informatics3
2021 SwiftNet: Real-Time Video Object Segmentation
abstract
In this work we present SwiftNet for real-time semisupervised video object segmentation (one-shot VOS), which reports 77.8% $\mathcal{J}\& \mathcal{F}$ and 70 FPS on DAVIS 2017 validation dataset, leading all present solutions in overall accuracy and speed performance. We achieve this by elaborately compressing spatiotemporal redundancy in matching-based VOS via Pixel-Adaptive Memory (PAM). Temporally, PAM adaptively triggers memory updates on frames where objects display noteworthy inter-frame variations. Spatially, PAM selectively performs memory update and match on dynamic pixels while ignoring the static ones, significantly reducing redundant computations wasted on segmentation-irrelevant pixels. To promote efficient reference encoding, light-aggregation encoder is also introduced in SwiftNet deploying reversed sub-pixel. We hope SwiftNet could set a strong and efficient baseline for real-time VOS and facilitate its application in mobile vision. The source code of SwiftNet can be found at https://github.com/haochenheheda/SwiftNet.
Haibing Ren, Yao Hu 0002, Song Bai 0001
CVPR3
2021 Twins: Revisiting the Design of Spatial Attention in Vision Transformers
abstract
Very recently, a variety of vision transformer architectures for dense prediction tasks have been proposed and they show that the design of spatial attention is critical to their success in these tasks. In this work, we revisit the design of the spatial attention and demonstrate that a carefully devised yet simple spatial attention mechanism performs favorably against the state-of-the-art schemes. As a result, we propose two vision transformer architectures, namely, Twins- PCPVT and Twins-SVT. Our proposed architectures are highly efficient and easy to implement, only involving matrix multiplications that are highly optimized in modern deep learning frameworks. More importantly, the proposed architectures achieve excellent performance on a wide range of visual tasks including image-level classification as well as dense detection and segmentation. The simplicity and strong performance suggest that our proposed architectures may serve as stronger backbones for many vision tasks.
Xiangxiang Chu, Zhi Tian, Bo Zhang 0046, Haibing Ren, Xiaolin Wei, Huaxia Xia, Chunhua Shen
NeurIPS5
2019 Customized Object Recognition and Segmentation by One Shot Learning with Human Robot Interaction
abstract
There are two difficulties to utilize state-of-the-art object recognition/detection/segmentation methods to robotic applications. First, most of the deep learning models heavily depend on large amounts of labeled training data, which are expensive to obtain for each individual application. Second, the object categories must be pre-defined in the dataset, thus not practical to scenarios with varying object categories. To alleviate the reliance on pre-defined big data, this paper proposes a customized object recognition and segmentation method. It aims to recognize and segment any object defined by the user, given only one annotation. There are three steps in the proposed method. First, the user takes an exemplar video of the target object with the robot, defines its name, and mask its boundary on only one frame. Then the robot automatically propagates the annotation through the exemplar video based on a proposed data generation method. In the meantime, a segmentation model continuously updates itself on the generated data. Finally, only a lightweight segmentation net is required at testing stage, to recognize and segment the user-defined object in any scenes.
Lidan Zhang, Yingzhe Shen, Xuesong Shi, Haibing Ren, Yimin Zhang 0002
ICRA6
2018 PCAOT: A Manhattan Point Cloud Registration Method Towards Large Rotation and Small Overlap
abstract
Point cloud registration is a popular research topic and has been widely used in many tasks, such as robot mapping and localization. It is a challenging problem when the overlap is small, or the rotation is large. The problem has not been well solved by existing methods such as the iterative closest point (ICP) and its variants. In this paper, a novel method named principal coordinate alignment with overlap tuning (PCAOT) is proposed based on the Manhattan world assumption. It solves two key problems together, the transformation estimation and the overlap estimation. The overlap is represented by a 3D cuboid and the transformation is computed only within the overlap region. Instead of finding point correspondence as in traditional methods, we estimate the rotation by principal coordinates alignment, which is faster and less sensitive than ICP and its variants to small overlaps and large rotations. Evaluations demonstrate that our method achieves much better results than the ICP and its variants when the overlap ratio is smaller than 50%, or the rotation angle is larger than 60°. Especially, it is effective when the overlap ratio is less than 30%, or the rotation angle is larger than 90°.
Wei Hu 0002, Haibing Ren, Yimin Zhang 0002
IROS3
2016 Fast human detection in RGB-D images based on color-depth joint feature learning
abstract
Human detection in RGB-D images is an important yet very challenging task in computer vision. In this paper, we propose a novel human detection approach in RGB-D images, which integrates ROI (region-of-interest) generation, depth-size relationship estimation and a human detector. Our approach has the following advantages: 1) ROI generation and depth-size relationship estimation take full advantage of color and depth information to fast reject about 70% negative samples while maintaining a high recall rate; 2) the cascade-structured human detector can seamlessly concatenate features extracted from both color and depth images; and 3) our method can detect human at a speed of more than 30 fps on 640 χ 480 images on a single laptop CPU without any GPU acceleration. Experiments on challenging public datasets demonstrate the effectiveness of our method.
Zhan Hu, Haizhou Ai, Haibing Ren, Yimin Zhang 0002
ICIP3
2012 Multiscale superpixel classification for tumor segmentation in breast ultrasound images
abstract
Tumor localization and segmentation in breast ultrasound (BUS) images is an important as well as intractable problem for computer-aided diagnosis (CAD) due to the high variation in shape and appearance. We propose a novel algorithm in this paper without making any assumption on tumor, compared to most previous works. Heterogeneous features are collected via a hierarchical over-segmentation framework, which we have shown has the multiscale property. The superpixels are then classified with their confidences nested into the bottom layer. The ultimate segmentation is made by using an efficient conditional random field model. Experiments on challenging data set show that our algorithm is able to handle almost all kinds of benign and malignant tumors, and also confirm the superiority of our work through a comparison with other two different approaches.
Zhihui Hao, Qiang Wang 0023, Haibing Ren, Kuanhong Xu, Yeong Kyeong Seong, Ji-yeun Kim
ICIP3
2012 Combining CRF and Multi-hypothesis Detection for Accurate Lesion Segmentation in Breast Sonograms
Zhihui Hao, Qiang Wang 0023, Yeong Kyeong Seong, Jong-Ha Lee 0001, Haibing Ren, Ji-yeun Kim
MICCAI (1)5
2009 Face recognition using gender information
abstract
In this paper, we propose a novel method using gender information for achieving better performances of face recognition systems. Gender is one of the important factors for recognizing appearance of human faces and there are many studies on gender classifications such. However, the gender information is not actively applied in vision-based face recognition tasks, because we cannot find out human identity using only gender information. Therefore, we design the face recognition system based on the gender-based facial features with global facial features, and moreover, gender-based score normalization method for verification task. For fair evaluations, we use FRGC database known as a large size face image database.
Wonjun Hwang, Haibing Ren, Seok-Cheol Kee, Junmo Kim 0002
ICIP2
2000 Toward Real-Time Human-Computer Interaction with Continuous Dynamic Hand Gestures
abstract
This paper, aiming at real-time gesture-controlled interaction, describes visual modeling, analysis, and recognition of continuous dynamic hand gestures. By hierarchically integrating multiple cues, a spatio-temporal appearance model and novel approaches are proposed for modeling and analysis of dynamic gestures respectively. At low level, fusion of flesh chrominance analysis and coarse image motion detection is employed to detect and segment hand gestures; at high level, parameters of the spatio-temporal appearance model are recovered by combining robust parameterized image motion estimation and hand shape analysis. The approach, therefore, fulfils real-time processing as well as high recognition rates. Without resorting to any special marks, twelve kinds of hand gestures can be recognized with average accuracy over 89%. A prototype system, gesture-controlled panoramic map browser is designed and implemented to demonstrate the usability of gesture-controlled interaction.
Yuanxin Zhu, Haibing Ren, Guangyou Xu, Xueyin Lin
FG2
2000 Hand Shape Extraction and Understanding by Virtue of Multiple Cues Fusion Technology
Xueyin Lin, Haibing Ren
ICMI3