Nicolas Gauthier

dblp:274/4052 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
3since 2021 · last 2021
0000-0002-2225-5827ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Learning paradigms · 29% Image recognition and object detection · 29% Robot navigation and mapping · 29%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
active perception
0.512021
Towards Efficient Multiview Object Detection with Adaptive Action Prediction · ICRA 2021
Machine learning › Learning paradigms › continual learning
class-incremental learning
0.512021
TAILOR: Teaching with Active and Incremental Learning for Object Registration · AAAI 2021
Machine learning › Learning paradigms
incremental learning
0.512021
TAILOR: Teaching with Active and Incremental Learning for Object Registration · AAAI 2021
Computer vision › Image recognition and object detection › object detection
multi-view object detection
0.512021
Towards Efficient Multiview Object Detection with Adaptive Action Prediction · ICRA 2021
Computer vision › Image recognition and object detection
object detection
0.512021
Towards Efficient Multiview Object Detection with Adaptive Action Prediction · ICRA 2021
Robotics › Robot navigation and mapping
view planning
0.512021
Towards Efficient Multiview Object Detection with Adaptive Action Prediction · ICRA 2021

Methods — techniques the papers use, named apart from their topics

viewpoint selection · 1.0incremental learning · 1.0active learning · 1.0dueling architecture · 0.5deep q-learning · 0.5
YearPublicationVenuePosition
2021 TAILOR: Teaching with Active and Incremental Learning for Object Registration
abstract
When deploying a robot to a new task, one often has to train it to detect novel objects, which is time-consuming and labor- intensive. We present TAILOR - a method and system for ob- ject registration with active and incremental learning. When instructed by a human teacher to register an object, TAILOR is able to automatically select viewpoints to capture informa- tive images by actively exploring viewpoints, and employs a fast incremental learning algorithm to learn new objects without potential forgetting of previously learned objects. We demonstrate the effectiveness of our method with a KUKA robot to learn novel objects used in a real-world gearbox as- sembly task through natural interactions.
Qianli Xu, Nicolas Gauthier, Wenyu Liang, Fen Fang, Hui Li Tan, Ying Sun 0001, Yan Wu 0002, Liyuan Li, Joo-Hwee Lim
AAAI2
2021 Enhancing Multi-Step Action Prediction for Active Object Detection
abstract
Active vision for robots is one promising solution to open world visual detection problems. A fundamental issue is view planning, i.e., predicting next best views to capture images of interest to reduce uncertainty. While multi-step action in a reinforcement learning (RL) setup can boost the efficiency of view planning, existing methods suffer from unstable detection outcome when the Q-values of multiple branches of action advantages (i.e., action range and action type) are combined naively. To tackle this issue, we propose a novel mechanism to disentangle action range from action type through a two-stage training strategy on a deep Q-network. It combines well-crafted loss functions with respect to action range and action type to enforce separated training of these two branches. We evaluate our method on two public datasets and show that it facilitates substantial gain in view planning efficiency, while enhancing detection accuracy.
Fen Fang, Qianli Xu, Nicolas Gauthier, Liyuan Li, Joo-Hwee Lim
ICIP3
2021 Towards Efficient Multiview Object Detection with Adaptive Action Prediction
abstract
Active vision is a desirable perceptual feature for robots. Existing approaches usually make strong assumptions about the task and environment, thus are less robust and efficient. This study proposes an adaptive view planning approach to boost the efficiency and robustness of active object detection. We formulate the multi-object detection task as an active multiview object detection problem given the initial location of the objects. Next, we propose a novel adaptive action prediction (A2P) method built on a deep Q-learning network with a dueling architecture. The A2P method is able to perform view planning based on visual information of multiple objects; and adjust action ranges according to the task status. Evaluated on the AVD dataset, A2P leads to 21.9% increase in detection accuracy in unfamiliar environments, while improving efficiency by 22.7%. On the T-LESS dataset, multi-object detection boosts efficiency by more than 30% while achieving equivalent detection accuracy.
Qianli Xu, Fen Fang, Nicolas Gauthier, Wenyu Liang, Yan Wu 0002, Liyuan Li, Joo-Hwee Lim
ICRA3
2020 Task-Oriented Multi-Modal Question Answering For Collaborative Applications
abstract
Cobots that can work in human workspaces and adapt to human need to understand and respond to human’s inquiry and instruction. In this paper, we propose new question answering (QA) task and dataset for human-robot collaboration on task-oriented operation, i.e., task-oriented collaborative QA (TCQA). Differing from conventional video QA for answering questions about what happened in video clips constrained by scripts and subtitles, TC-QA aims to share common ground for task-oriented operation through question answering. We propose an open-end (OE) format of answer with text reply, image with annotated related objects, and video with operation duration to guide operation execution. Designed for grounding, the TC-QA dataset comprises query videos and questions to seek acknowledgement, correction, attention to task-related objects, and information on objects or operation. Due to the flexibility of real-world task with limited training sample, we propose and evaluate a baseline method based on a hybrid approach. The hybrid approach employs deep learning methods for object detection, hand detection and gesture recognition, and symbolic reasoning to ground question on observation for providing the answer. Our experiments show that the hybrid method is effective for the TC-QA task.
Hui Li Tan, Mei Chee Leong, Qianli Xu, Liyuan Li, Fen Fang, Nicolas Gauthier, Ying Sun 0001, Joo-Hwee Lim
ICIP7
2020 Active Image Sampling on Canonical Views for Novel Object Detection
abstract
To alleviate the costly data annotation problem in deep learning-based object detection, we leverage the canonical view model for active sample selection to improve the effectiveness of learning. Inspired by the view-approximation model, we hypothesize that visual features learned from canonical views denote better representations of objects, thus boosting the effectiveness of object learning. We validate the hypothesis empirically in the context of robot learning for novel object detection. Based on this, we propose a novel on-line viewpoint exploration (OLIVE) method that (1) defines goodness-of-view by combining informativeness of visual features and consistency of model-based object detection, and (2) systematically explores and selects viewpoints to boost learning efficiency. Furthermore, we train a legacy Faster R-CNN model with a data augmentation method while leveraging data samples generated by the OLIVE pipeline. We test our method on the T-LESS dataset and show that the proposed method outperforms competitive benchmarking methods, especially when the samples are few.
Qianli Xu, Fen Fang, Nicolas Gauthier, Liyuan Li, Joo-Hwee Lim
ICIP3