Georgios Kapidis

dblp:227/5354 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0002-4424-9591ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Video understanding and tracking · 77% Transfer learning and domain adaptation · 23%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking › action recognition › human action recognition
egocentric action recognition
0.712023
Multi-Dataset, Multitask Learning of Egocentric Vision Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Transfer learning and domain adaptation › multi-source learning
multi-dataset learning
0.212023
Multi-Dataset, Multitask Learning of Egocentric Vision Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2023

Methods — techniques the papers use, named apart from their topics

multi-task learning · 0.7convolutional neural network · 0.7
YearPublicationVenuePosition
2023 Multi-Dataset, Multitask Learning of Egocentric Vision Tasks
abstract
For egocentric vision tasks such as action recognition, there is a relative scarcity of labeled data. This increases the risk of overfitting during training. In this paper, we address this issue by introducing a multitask learning scheme that employs related tasks as well as related datasets in the training process. Related tasks are indicative of the performed action, such as the presence of objects and the position of the hands. By including related tasks as additional outputs to be optimized, action recognition performance typically increases because the network focuses on relevant aspects in the video. Still, the training data is limited to a single dataset because the set of action labels usually differs across datasets. To mitigate this issue, we extend the multitask paradigm to include datasets with different label sets. During training, we effectively mix batches with samples from multiple datasets. Our experiments on egocentric action recognition in the EPIC-Kitchens, EGTEA Gaze+, ADL and Charades-EGO datasets demonstrate the improvements of our approach over single-dataset baselines. On EGTEA we surpass the current state-of-the-art by 2.47 percent. We further illustrate the cross-dataset task correlations that emerge automatically with our novel training scheme.
Georgios Kapidis, Ronald Poppe, Remco C. Veltkamp
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Saliency Tubes: Visual Explanations for Spatio-Temporal Convolutions
abstract
Deep learning approaches have been established as the main methodology for video classification and recognition. Recently, 3-dimensional convolutions have been used to achieve state-of-the-art performance in many challenging video datasets. Because of the high level of complexity of these methods, as the convolution operations are also extended to an additional dimension in order to extract features from it as well, providing a visualization for the signals that the network interpret as informative, is a challenging task. An effective notion of understanding the network's innerworkings would be to isolate the spatio-temporal regions on the video that the network finds most informative. We propose a method called Saliency Tubes which demonstrate the foremost points and regions in both frame level and over time that are found to be the main focus points of the network. We demonstrate our findings on widely used datasets for thirdperson and egocentric action classification and enhance the set of methods and visualizations that improve 3D Convolutional Neural Networks (CNNs) intelligibility. Our code1and a demo video2are also available.
Alexandros Stergiou, Georgios Kapidis, Grigorios Kalliatakis, Christos Chrysoulas, Remco C. Veltkamp, Ronald Poppe
ICIP2