Ruotian Luo

dblp:158/4770 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-authorArtificial intelligence and machine learning · 5 · 3 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 34% Image recognition and object detection · 22% Segmentation and scene understanding · 19%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
image captioning
0.722018
A Multi-task Learning Approach for Image Captioning · IJCAI 2018
Discriminability Objective for Training Descriptive Captions · CVPR 2018
Computer vision › Segmentation and scene understanding
instance segmentation
0.412020
Pixel Consensus Voting for Panoptic Segmentation · CVPR 2020
Computer vision › Image recognition and object detection
object detection
0.412020
Context-Aware Zero-Shot Recognition · AAAI 2020
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.412020
Pixel Consensus Voting for Panoptic Segmentation · CVPR 2020
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot classification
0.412020
Context-Aware Zero-Shot Recognition · AAAI 2020
Computer vision › Image recognition and object detection › object detection › open-vocabulary object detection
zero-shot object detection
0.412020
Context-Aware Zero-Shot Recognition · AAAI 2020
Computer vision › Vision and language › image captioning
discriminative caption generation
0.312018
Discriminability Objective for Training Descriptive Captions · CVPR 2018
Machine learning › Learning paradigms
multi-task learning
0.312018
A Multi-task Learning Approach for Image Captioning · IJCAI 2018
Computer vision › Vision and language › visual grounding
referring expression comprehension
0.312017
Comprehension-Guided Referring Expressions · CVPR 2017
Natural language and speech › Language models and text generation › text generation › sentence planning
referring expression generation
0.312017
Comprehension-Guided Referring Expressions · CVPR 2017
Computer vision › Vision and language
visual context modeling
0.112020
Context-Aware Zero-Shot Recognition · AAAI 2020
Computer vision › Vision and language › image captioning
image caption evaluation
0.112018
Discriminability Objective for Training Descriptive Captions · CVPR 2018
Computer vision › Image recognition and object detection
image classification
0.112018
A Multi-task Learning Approach for Image Captioning · IJCAI 2018
Natural language and speech › Language models and text generation › text generation › constrained text generation
syntax-guided text generation
0.112018
A Multi-task Learning Approach for Image Captioning · IJCAI 2018

Methods — techniques the papers use, named apart from their topics

pixel consensus voting · 0.4inter-object relation prior · 0.4generalized hough transform · 0.4conditional random field · 0.4back-projection · 0.4discriminability loss · 0.3convolutional neural network · 0.3SPICE · 0.3LSTM · 0.3BLEU · 0.3
YearPublicationVenuePosition
2020 Context-Aware Zero-Shot Recognition
abstract
We present a novel problem setting in zero-shot learning, zero-shot object recognition and detection in the context. Contrary to the traditional zero-shot learning methods, which simply infers unseen categories by transferring knowledge from the objects belonging to semantically similar seen categories, we aim to understand the identity of the novel objects in an image surrounded by the known objects using the inter-object relation prior. Specifically, we leverage the visual context and the geometric relationships between all pairs of objects in a single image, and capture the information useful to infer unseen categories. We integrate our context-aware zero-shot learning framework into the traditional zero-shot learning techniques seamlessly using a Conditional Random Field (CRF). The proposed algorithm is evaluated on both zero-shot region classification and zero-shot detection tasks. The results on Visual Genome (VG) dataset show that our model significantly boosts performance with the additional visual context compared to traditional methods.
Ruotian Luo, Bohyung Han
AAAI1
2020 Pixel Consensus Voting for Panoptic Segmentation
abstract
The core of our approach, Pixel Consensus Voting, is a framework for instance segmentation based on the generalized Hough transform. Pixels cast discretized, probabilistic votes for the likely regions that contain instance centroids. At the detected peaks that emerge in the voting heatmap, backprojection is applied to collect pixels and produce instance masks. Unlike a sliding window detector that densely enumerates object proposals, our method detects instances as a result of the consensus among pixel-wise votes. We implement vote aggregation and backprojection using native operators of a convolutional neural network. The discretization of centroid voting reduces the training of instance segmentation to pixel labeling, analogous and complementary to FCN-style semantic segmentation, leading to an efficient and unified architecture that jointly models things and stuff. We demonstrate the effectiveness of our pipeline on COCO and Cityscapes Panoptic Segmentation and obtain competitive results. Code will be open-sourced.
Ruotian Luo, Michael Maire, Gregory Shakhnarovich
CVPR2
2018 Discriminability Objective for Training Descriptive Captions
abstract
One property that remains lacking in image captions generated by contemporary methods is discriminability: being able to tell two images apart given the caption for one of them. We propose a way to improve this aspect of caption generation. By incorporating into the captioning training objective a loss component directly related to ability (by a machine) to disambiguate image/caption matches, we obtain systems that produce much more discriminative caption, according to human evaluation. Remarkably, our approach leads to improvement in other aspects of generated captions, reflected by a battery of standard scores such as BLEU, SPICE etc. Our approach is modular and can be applied to a variety of model/loss combinations commonly proposed for image captioning.
Ruotian Luo, Brian L. Price, Scott Cohen, Gregory Shakhnarovich
CVPR1
2018 A Multi-task Learning Approach for Image Captioning
abstract
In this paper, we propose a Multi-task Learning Approach for Image Captioning (MLAIC ), motivated by the fact that humans have no difficulty performing such task because they possess capabilities of multiple domains. Specifically, MLAIC consists of three key components: (i) A multi-object classification model that learns rich category-aware image representations using a CNN image encoder; (ii) A syntax generation model that learns better syntax-aware LSTM based decoder; (iii) An image captioning model that generates image descriptions in text, sharing its CNN encoder and LSTM decoder with the object classification task and the syntax generation task, respectively. In particular, the image captioning model can benefit from the additional object categorization and syntax knowledge. To verify the effectiveness of our approach, we conduct extensive experiments on MS-COCO dataset. The experimental results demonstrate that our model achieves impressive results compared to other strong competitors.
Wei Zhao 0033, Benyou Wang, Jianbo Ye, Min Yang 0007, Zhou Zhao 0001, Ruotian Luo, Yu Qiao 0001
IJCAI6
2017 Comprehension-Guided Referring Expressions
abstract
We consider generation and comprehension of natural language referring expression for objects in an image. Unlike generic image captioning which lacks natural standard evaluation criteria, quality of a referring expression may be measured by the receivers ability to correctly infer which object is being described. Following this intuition, we propose two approaches to utilize models trained for comprehension task to generate better expressions. First, we use a comprehension module trained on human-generated expressions, as a critic of referring expression generator. The comprehension module serves as a differentiable proxy of human evaluation, providing training signal to the generation module. Second, we use the comprehension model in a generate-and-rerank pipeline, which chooses from candidate expressions generated by a model according to their performance on the comprehension task. We show that both approaches lead to improved referring expression generation on multiple benchmark datasets.
Ruotian Luo, Gregory Shakhnarovich
CVPR1
2016 Person Re-identification by encoding free energy feature maps
Yanna Zhao, Xu Zhao 0001, Ruotian Luo, Yuncai Liu
Multim. Tools Appl.3
2014 Are we still friends: Kernel multivariate survival analysis
abstract
Online Social Network becomes the most prevalent platform for exchanging information between users, maintaining friendships online. As is well-known to us, however, some friendships even those intimate ones might vanish. Therefore, precisely modeling and predicting state of each online relationship is worthwhile in many respects. For social communication services such modeling permits new and novel online services. In addition, constructing this model might enlighten us in exploiting information spreading pattern in online social network. In this paper, we propose a model in determining a probability distribution which describes the ‘surviving time’ of each friendships by applying one commonly used method in sociology, survival analysis. We discuss a series of social explanatory variables that highly affect this probability distribution. Moreover, methods in the moving average process are devoted to determining the appropriate parameter in survival model. Furthermore, to avoid the high computational complexity in kernel learning we impose sparsity in our model. Finally, with the experiments on real data, the proposed survival model is proven to be of high accuracy, and thus of great potential for further applications.
Shiyu Liang, Ruotian Luo, Songjun Ma, Weijie Wu, Li Song 0001, Xiaohua Tian, Xinbing Wang
GLOBECOM2