Ruochen Fan

dblp:202/6198 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
0since 2021 · last 2020
0000-0003-1991-0146ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-authorArtificial intelligence and machine learning · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Segmentation and scene understanding · 92% Image recognition and object detection · 8%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
instance segmentation
0.412019
S4Net: Single Stage Salient-Instance Segmentation · CVPR 2019
Computer vision › Segmentation and scene understanding › instance segmentation
salient instance segmentation
0.412019
S4Net: Single Stage Salient-Instance Segmentation · CVPR 2019
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation
0.312018
Associating Inter-image Salient Instances for Weakly Supervised Semantic Segmentation · ECCV (9) 2018

Methods — techniques the papers use, named apart from their topics

single-stage detection · 0.4context modeling · 0.4weakly supervised learning · 0.3
YearPublicationVenuePosition
2020 S4Net: Single stage salient-instance segmentation
abstract
In this paper, we consider salient instance segmentation. As well as producing bounding boxes, our network also outputs high-quality instance-level segments as initial selections to indicate the regions of interest. Taking into account the category-independent property of each target, we design a single stage salient instance segmentation framework, with a novel segmentation branch. Our new branch regards not only local context inside each detection window but also the surrounding context, enabling us to distinguish instances in the same scope even with partial occlusion. Our network is end-to-end trainable and is fast (running at 40 fps for images with resolution 320 × 320). We evaluate our approach on a publicly available benchmark and show that it outperforms alternative solutions. We also provide a thorough analysis of our design choices to help readers better understand the function of each part of our network. Source code can be found at https://github.com/RuochenFan/S4Net.
Ruochen Fan, Ming-Ming Cheng, Qibin Hou, Tai-Jiang Mu, Jingdong Wang 0001, Shi-Min Hu 0001
Comput. Vis. Media1
2019 S4Net: Single Stage Salient-Instance Segmentation
abstract
We consider an interesting problem---salient instance segmentation. Other than producing approximate bounding boxes, our network also outputs high-quality instance-level segments. Taking into account the category-independent property of each target, we design a single stage salient instance segmentation framework, with a novel segmentation branch. Our new branch regards not only local context inside each detection window but also its surrounding context, enabling us to distinguish the instances in the same scope even with obstruction. Our network is end-to-end trainable and runs at a fast speed (40 fps when processing an image with resolution 320 x 320). We evaluate our approach on a public available benchmark and show that it outperforms other alternative solutions. We also provide a thorough analysis of the design choices to help readers better understand the functions of each part of our network. The source code can be found at https://github.com/RuochenFan/S4Net.
Ruochen Fan, Ming-Ming Cheng, Qibin Hou, Tai-Jiang Mu, Jingdong Wang 0001, Shi-Min Hu 0001
CVPR1
2019 SpinNet: Spinning convolutional network for lane boundary detection
abstract
In this paper, we propose a simple but effective framework for lane boundary detection, called SpinNet. Considering that cars or pedestrians often occlude lane boundaries and that the local features of lane boundaries are not distinctive, therefore, analyzing and collecting global context information is crucial for lane boundary detection. To this end, we design a novel spinning convolution layer and a brand-new lane parameterization branch in our network to detect lane boundaries from a global perspective. To extract features in narrow strip-shaped fields, we adopt strip-shaped convolutions with kernels which have 1 × n or n × 1 shape in the spinning convolution layer. To tackle the problem of that straight strip-shaped convolutions are only able to extract features in vertical or horizontal directions, we introduce the concept of feature map rotation to allow the convolutions to be applied in multiple directions so that more information can be collected concerning a whole lane boundary. Moreover, unlike most existing lane boundary detectors, which extract lane boundaries from segmentation masks, our lane boundary parameterization branch predicts a curve expression for the lane boundary for each pixel in the output feature map. And the network utilizes this information to predict the weights of the curve, to better form the final lane boundaries. Our framework is easy to implement and end-to-end trainable. Experiments show that our proposed SpinNet outperforms state-of-the-art methods.
Ruochen Fan, Xuanrun Wang, Qibin Hou, Tai-Jiang Mu
Comput. Vis. Media1
2019 A three-stage real-time detector for traffic signs in large panoramas
abstract
Traffic sign detection is one of the key components in autonomous driving. Advanced autonomous vehicles armed with high quality sensors capture high definition images for further analysis. Detecting traffic signs, moving vehicles, and lanes is important for localization and decision making. Traffic signs, especially those that are far from the camera, are small, and so are challenging to traditional object detection methods. In this work, in order to reduce computational cost and improve detection performance, we split the large input images into small blocks and then recognize traffic signs in the blocks using another detection module. Therefore, this paper proposes a three-stage traffic sign detector, which connects a BlockNet with an RPN–RCNN detection network. BlockNet, which is composed of a set of CNN layers, is capable of performing block-level foreground detection, making inferences in less than 1 ms. Then, the RPN–RCNN two-stage detector is used to identify traffic sign objects in each block; it is trained on a derived dataset named TT100KPatch. Experiments show that our framework can achieve both state-of-the-art accuracy and recall; its fastest detection speed is 102 fps.
Ruochen Fan, Sharon X. Huang, Zhe Zhu, Ruofeng Tong 0001
Comput. Vis. Media2
2019 Reliable Line Segment Matching for Multispectral Images Guided by Intersection Matches
abstract
Accurate and robust feature matching is a critical issue in the preprocessing of multispectral image data sets. Higher order features, such as lines can provide useful matching information but are heavily affected by the unreliable detection of lines. Existing methods typically make the unrealistic assumption that end points of lines can be accurately detected across the reference and test images. To address the unreliable detection of line end points, this paper proposes mapping line intersections and then employing tentatively mapped intersections as “anchor” points to compute line descriptors. The computed line descriptors are utilized to determine whether the two lines forming an intersection are matched with the two lines forming its mapped intersection. This eliminates the reliance on the accurate detection of line end points and results in improved matching accuracy. The proposed method is tested on a large number of multispectral images containing various scenes. Experimental results show that it can effectively deal with the detection inaccuracy of end points for line matching.
Yong Li 0025, Robert L. Stevenson, Ruochen Fan, Huachun Tan
IEEE Trans. Circuits Syst. Video Technol.4
2018 Associating Inter-image Salient Instances for Weakly Supervised Semantic Segmentation
Ruochen Fan, Qibin Hou, Ming-Ming Cheng, Gang Yu 0002, Ralph R. Martin, Shi-Min Hu 0001
ECCV (9)1
2017 Robust tracking-by-detection using a selection and completion mechanism
abstract
It is challenging to track a target continuously in videos with long-term occlusion, or objects which leave then re-enter a scene. Existing tracking algorithms combined with onlinetrained object detectors perform unreliably in complex conditions, and can only provide discontinuous trajectories with jumps in position when the object is occluded. This paper proposes a novel framework of tracking-by-detection using selection and completion to solve the abovementioned problems. It has two components, tracking and trajectory completion. An offline-trained object detector can localize objects in the same category as the object being tracked. The object detector is based on a highly accurate deep learning model. The object selector determines which object should be used to re-initialize a traditional tracker. As the object selector is trained online, it allows the framework to be adaptable. During completion, a predictive non-linear autoregressive neural network completes any discontinuous trajectory. The tracking component is an online real-time algorithm, and the completion part is an after-theevent mechanism. Quantitative experiments show a significant improvement in robustness over prior state-of- the-art methods.
Ruochen Fan, Min Zhang 0069, Ralph R. Martin
Comput. Vis. Media1