Naoya Sogi

dblp:220/5672 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0002-2560-4818ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 4 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Robot navigation and mapping · 28% 3D vision · 24% Representation and self-supervised learning · 18%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
search result diversification
0.912025
MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval · IJCAI 2025
Multimedia analysis and retrieval
cross-modal retrieval
0.812024
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval · ECCV (79) 2024
Multimedia analysis and retrieval › cross-modal retrieval
image-text retrieval
0.812024
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval · ECCV (79) 2024
Machine learning › Representation and self-supervised learning › representation learning › feature extraction
discriminant feature extraction
0.712023
Discriminant Feature Extraction by Generalized Difference Subspace · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › Image recognition and object detection › object recognition › appearance-based object recognition
subspace-based object recognition
0.712023
Discriminant Feature Extraction by Generalized Difference Subspace · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Robotics › Robot navigation and mapping
localization
0.412020
Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints* · ICRA 2020
Robotics › Robot navigation and mapping › localization › landmark-based localization
marker-based localization
0.412020
Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints* · ICRA 2020
Computer vision › 3D vision › pose estimation › rigid body pose estimation
planar pose estimation
0.412020
Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints* · ICRA 2020
Computer vision › 3D vision
pose estimation
0.412020
Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints* · ICRA 2020
Computer vision › Vision and language
multimodal representation
0.212024
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval · ECCV (79) 2024
Computer vision › Face, body and person analysis
face recognition
0.212023
Discriminant Feature Extraction by Generalized Difference Subspace · IEEE Trans. Pattern Anal. Mach. Intell. 2023

Methods — techniques the papers use, named apart from their topics

query perturbation · 1.5manifold representation · 0.9determinantal point process · 0.9kernel trick · 0.7geometrical fisher discriminant analysis · 0.7generalized difference subspace · 0.7CNN features · 0.7rotation averaging · 0.4lifted algorithm · 0.4clique constraints · 0.4
YearPublicationVenuePosition
2025 MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval
abstract
Result diversification (RD) is a crucial technique in Text-to-Image Retrieval for enhancing the efficiency of a practical application. Conventional methods focus solely on increasing the diversity metric of image appearances. However, the diversity metric and its desired value vary depending on the application, which limits the applications of RD. This paper proposes a novel task called CDR-CA (Contextual Diversity Refinement of Composite Attributes). CDR-CA aims to refine the diversities of multiple attributes, according to the application's context. To address this task, we propose Multi-Source DPPs, a simple yet strong baseline that extends the Determinantal Point Process (DPP) to multi-sources. We model MS-DPP as a single DPP model with a unified similarity matrix based on a manifold representation. We also introduce Tangent Normalization to reflect contexts. Extensive experiments demonstrate the effectiveness of the proposed method.
Naoya Sogi, Takashi Shibata 0001, Makoto Terao, Masanori Suganuma, Takayuki Okatani
IJCAI1
2024 Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
Naoya Sogi, Takashi Shibata 0001, Makoto Terao
ECCV (79)1
2024 Task Success Classification with Final State of Future Prediction for Robot Control Planning
Taku Fujitomi, Naoya Sogi, Takashi Shibata 0001, Makoto Terao
ICPR (2)2
2024 Disaster Damage Visualization by VLM-Based Interactive Image Retrieval and Cross-View Image Geo-Localization
abstract
We propose a framework for quickly selecting images that show the disaster situation from many images, estimating their locations with high accuracy, and displaying them on a map. The proposed framework introduces interactive image retrieval based on the Vision and Language Model (VLM), which can retrieve images from many images that show the disaster situation according to the user’s intention. Using the correlation between language and images based on VLM and the similarity between images selected interactively enables more accurate retrieval. Next, for selected images for which the location of the affected area is unknown, the location of the image is estimated with street address-level accuracy by matching it with an overhead image covering a large area of the city and map data and then displayed on a map. We confirmed the effectiveness of the proposed method on publicly available datasets such as CrisisNLP.
Naoya Sogi, Takashi Shibata 0001, Makoto Terao, Kenta Senzaki, Masahiro Tani, Royston Rodrigues
IGARSS1
2024 Future Predictive Success-or-Failure Classification for Long-Horizon Robotic Tasks
abstract
Automating long-horizon tasks with a robotic arm has been a central research topic in robotics. Optimization-based action planning is an efficient approach for creating an action plan to complete a given task. Construction of a reliable planning method requires a design process of conditions, e.g., to avoid collision between objects. The design process, however, has two critical issues: 1) iterative trials–the design process is time-consuming due to the trial-and-error process of modifying conditions, and 2) manual redesign–it is difficult to cover all the necessary conditions manually. To tackle these issues, this paper proposes a future-predictive success-or-failure-classification method to obtain conditions automatically. The key idea behind the proposed method is an end-to-end approach for determining whether the action plan can complete a given task instead of manually redesigning the conditions. The proposed method uses a long-horizon future-prediction method to enable success-or-failure classification without the execution of an action plan. This paper also proposes a regularization term called transition consistency regularization to provide easy-to-predict feature distribution. The regularization term improves future prediction and classification performance. The effectiveness of our method is demonstrated through classification and robotic-manipulation experiments.
Naoya Sogi, Hiroyuki Oyama, Takashi Shibata 0001, Makoto Terao
IJCNN1
2023 Visually explaining 3D-CNN predictions for video classification with an adaptive occlusion sensitivity analysis
abstract
This paper proposes a method for visually explaining the decision-making process of 3D convolutional neural networks (CNN) with a temporal extension of occlusion sensitivity analysis. The key idea here is to occlude a specific volume of data by a 3D mask in an input 3D temporalspatial data space and then measure the change degree in the output score. The occluded volume data that produces a larger change degree is regarded as a more critical element for classification. However, while the occlusion sensitivity analysis is commonly used to analyze single image classification, it is not so straightforward to apply this idea to video classification as a simple fixed cuboid cannot deal with the motions. To this end, we adapt the shape of a 3D occlusion mask to complicated motions of target objects. Our flexible mask adaptation is performed by considering the temporal continuity and spatial co-occurrence of the optical flows extracted from the input video data. We further propose to approximate our method by using the first-order partial derivative of the score with respect to an input image to reduce its computational cost. We demonstrate the effectiveness of our method through various and extensive comparisons with the conventional methods in terms of the deletion/insertion metric and the pointing metric on the UCF101. The code is available at: https://github.com/uchiyama33/AOSA.
Tomoki Uchiyama, Naoya Sogi, Koichiro Niinuma, Kazuhiro Fukui
WACV2
2023 Grassmannian learning mutual subspace method for image set recognition
Lincon Sales de Souza, Naoya Sogi, Bernardo Bentes Gatto, Takumi Kobayashi 0001, Kazuhiro Fukui
Neurocomputing2
2023 Discriminant Feature Extraction by Generalized Difference Subspace
abstract
In this paper, we reveal the discriminant capacity of orthogonal data projection onto the generalized difference subspace (GDS), both theoretically and experimentally. In our previous work, we demonstrated that the GDS projection works as a quasi-orthogonalization of class subspaces, which is an effective feature extraction for subspace based classifiers. Here, we further show that GDS projection also works as a discriminant feature extraction through a similar mechanism to the Fisher discriminant analysis (FDA). A direct proof of the connection between GDS projection and FDA is difficult due to the significant difference in their formulations. To circumvent the complication, we first introduce geometrical Fisher discriminant analysis (gFDA) based on a simplified Fisher criterion. It is derived from a heuristic yet practically plausible assumption: the direction of the sample mean vector of a class is largely aligned to the first principal component vector of the class, given that the principal component analysis (PCA) is applied without data centering. gFDA works stably even under few samples, bypassing the small sample size (SSS) problem of FDA. We then prove that gFDA is equivalent to GDS projection with a small correction term. This equivalence ensures GDS projection to inherit the discriminant ability from FDA via gFDA. Furthermore, we discuss two useful extensions of these methods, 1) a nonlinear extension by kernel trick, 2) a combination with CNN features. The equivalence and the effectiveness of the extensions have been verified through extensive experiments on the extended Yale B+, CMU face database, ALOI, ETH80, MNIST, and CIFAR10, mainly focusing on image recognition under small samples.
Kazuhiro Fukui, Naoya Sogi, Takumi Kobayashi 0001, Jing-Hao Xue, Atsuto Maki
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Constrained mutual convex cone method for image set based recognition
Naoya Sogi, Rui Zhu 0006, Jing-Hao Xue, Kazuhiro Fukui
Pattern Recognit.1
2020 Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints*
abstract
Planar markers are useful in robotics and computer vision for mapping and localisation. Given a detected marker in an image, a frequent task is to estimate the 6DOF pose of the marker relative to the camera, which is an instance of planar pose estimation (PPE). Although there are mature techniques, PPE suffers from a fundamental ambiguity problem, in that there can be more than one plausible pose solutions for a PPE instance. Especially when localisation of the marker corners is noisy, it is often difficult to disambiguate the pose solutions based on reprojection error alone. Previous methods choose between the possible solutions using a heuristic criterion, or simply ignore ambiguous markers.We propose to resolve the ambiguities by examining the consistencies of a set of markers across multiple views. Our specific contributions include a novel rotation averaging formulation that incorporates long-range dependencies between possible marker orientation solutions that arise from PPE ambiguities. We analyse the combinatorial complexity of the problem, and develop a novel lifted algorithm to effectively resolve marker pose ambiguities, without discarding any marker observations. Results on real and synthetic data show that our method is able to handle highly ambiguous inputs, and provides more accurate and/or complete marker-based mapping and localisation.
Shin-Fang Ch'ng, Naoya Sogi, Pulak Purkait, Tat-Jun Chin, Kazuhiro Fukui
ICRA2
2020 Discriminative Singular Spectrum Analysis for Bioacoustic Classification
Bernardo Bentes Gatto, Eulanda M. dos Santos, Juan Gabriel Colonna, Naoya Sogi, Lincon Sales de Souza, Kazuhiro Fukui
INTERSPEECH4
2020 A Novel Separating Hyperplane Classification Framework to Unify Nearest-Class-Model Methods for High-Dimensional Data
abstract
In this article, we establish a novel separating hyperplane classification (SHC) framework to unify three nearest-class-model methods for high-dimensional data: the nearest subspace method (NSM), the nearest convex hull method (NCHM), and the nearest convex cone method (NCCM). Nearest-class-model methods are an important paradigm for the classification of high-dimensional data. We first introduce the three nearest-class-model methods and then conduct dual analysis for theoretically investigating them, to understand deeply their underlying classification mechanisms. A new theorem for the dual analysis of NCCM is proposed in this article by discovering the relationship between a convex cone and its polar cone. We then establish the new SHC framework to unify the nearest-class-model methods based on the theoretical results. One important application of this new SHC framework is to help explain empirical classification results: why one class model has a better performance than others on certain data sets. Finally, we propose a new nearest-class-model method, the soft NCCM, under the novel SHC framework to solve the overlapping class model problem. For illustrative purposes, we empirically demonstrate the significance of our SHC framework and the soft NCCM through two types of typical real-world high-dimensional data: the spectroscopic data and the face image data.
Rui Zhu 0006, Ziyu Wang 0003, Naoya Sogi, Kazuhiro Fukui, Jing-Hao Xue
IEEE Trans. Neural Networks Learn. Syst.3
2018 Action Recognition Method Based on Sets of Time Warped ARMA Models
abstract
In this paper, we propose a novel method for recognizing human actions from sequential body skeleton data. Our method is based on ARMA (Autoregressive Mean Average) model, which is constructed from the matrix of 3D joint positions time-series. The intrinsic structure of an action can be compactly summarized by the observability matrix of the ARMA model. Since the column vectors of an observability matrix span a subspace, given two ARMA models, we can measure the similarity between them by the canonical angles between the corresponding subspaces. This framework based on subspace representation is useful for action recognition. However, it does not work well when handling various actions with different action speeds, since optimal row size of each observability matrix depends on the action speed. To address this limitation, we perform a random sampling operation to the row elements in each observability matrix, while preserving the order of the elements. By repeating this operation, we generate a set of various time-warped ARMA models with various local motion speeds. The essence of this idea is that a whole set of such time-warped ARMA models is invariant to the changes in action speed. Furthermore, to construct an effective classification framework, we applied Grassmann discriminant analysis to the time-warped ARMA models. The effectiveness of the proposed method is demonstrated through comparison experiments with state-of-the-art methods on two public datasets: MSR 3D action dataset and UT-Kinect dataset.
Naoya Sogi, Kazuhiro Fukui
ICPR1
2018 A Method Based on Convex Cone Model for Image-Set Classification With CNN Features
abstract
In this paper, we propose a method for image-set classification based on convex cone models, focusing on the effectiveness of convolutional neural network (CNN) features as its input. CNN feature has non-negative values when using the rectified linear unit as an activation function. This naturally leads us to model a set of CNN features by a convex cone and measure the geometrical similarity of convex cones in classification. To achieve this framework, we define sequentially multiple angles between two convex cones by repeating the alternating least square method, and then define the geometrical similarity between the cones by using the obtained angles. Moreover, to enhance our method, we introduce a discriminant space, which maximizes the between-class variance (gaps) and minimizes the within-class variance of the projected convex cones onto the discriminant space, like Fisher discriminant analysis. Finally, the classification is conducted by measuring the similarity between projected convex cones. The effectiveness of the proposed method is demonstrated through evaluation experiments on a private database of a multi-view hand shape dataset, and two public databases.
Naoya Sogi, Taku Nakayama, Kazuhiro Fukui
IJCNN1