Naoyuki Onoe

dblp:87/6219 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
14since 2021 · last 2025
0000-0002-8709-7241ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 AdaPrefix++: Integrating Adapters, Prefixes and Hypernetwork for Continual Learning
Sayanta Adhikari, Dupati Srikar Chandra, P. K. Srijith, Pankaj Wasnik, Naoyuki Onoe
WACV5
2024 Open-Set Object Detection By Aligning Known Class Representations
abstract
Open-Set Object Detection (OSOD) has emerged as a contemporary research direction to address the detection of unknown objects. Recently, few works have achieved remarkable performance in the OSOD task by employing contrastive clustering to separate unknown classes. In contrast, we propose a new semantic clustering-based approach to facilitate a meaningful alignment of clusters in semantic space and introduce a class decorrelation module to enhance inter-cluster separation. Our approach further incorporates an object focus module to predict objectness scores, which enhances the detection of unknown objects. Further, we employ i) an evaluation technique that penalizes low-confidence outputs to mitigate the risk of misclassification of the unknown objects and ii) a new metric called HMP that combines known and unknown precision using harmonic mean. Our extensive experiments demonstrate that the proposed model achieves significant improvement on the MS-COCO & PASCAL VOC dataset for the OSOD task.
Hiran Sarkar, Vishal M. Chudasama, Naoyuki Onoe, Pankaj Wasnik, Vineeth N. Balasubramanian
WACV3
2023 Fiducial Focus Augmentation for Facial Landmark Detection
Purbayan Kar, Vishal M. Chudasama, Naoyuki Onoe, Pankaj Wasnik, Vineeth N. Balasubramanian
BMVC3
2023 Nonparallel Emotional Voice Conversion for Unseen Speaker-Emotion Pairs Using Dual Domain Adversarial Network & Virtual Domain Pairing
abstract
Primary goal of an emotional voice conversion (EVC) system is to convert the emotion of a given speech signal from one style to another style without modifying the linguistic content of the signal. Most of the state-of-the-art approaches convert emotions for seen speaker-emotion combinations only. In this paper, we tackle the problem of converting the emotion of speakers whose only neutral data are present during the time of training and testing (i.e., unseen speaker-emotion combinations). To this end, we extend a recently proposed StartGANv2-VC architecture by utilizing dual encoders for learning the speaker and emotion style embeddings separately along with dual domain source classifiers. For achieving the conversion to unseen speaker-emotion combinations, we propose a Virtual Domain Pairing (VDP) training strategy, which virtually incorporates the speaker-emotion pairs that are not present in the real data without compromising the min-max game of a discriminator and generator in adversarial training. We evaluate the proposed method using a Hindi emotional database.
Nirmesh J. Shah, Mayank Kumar Singh, Naoya Takahashi, Naoyuki Onoe
ICASSP4
2023 Enhancing Social Recommendation with Multi-View BERT Network
abstract
With the emergence of online social platforms enabling users to share their opinions with others, there has been a critical need for developing a recommendation system incorporating users’ social connections to learn their preferences. Social relations between users offer potential information about users’ preferences, which alleviates the data sparsity issue and boosts recommendation performance. Several efforts have been made to develop an efficient social recommendation system. Nevertheless, existing approaches have two significant drawbacks: First, they haven’t efficiently explored the complex correlations between the diverse influence of neighbours on users’ item preferences. Second, contemporary systems rely on the unidirectional context of items and fail to embrace bidirectional contexts and long range dependencies while predicting the next user-item interaction. To address the above issues, we propose a novel framework for Social Recommendation called Multi-View Bert Network (MVBN). This model incorporated Bidirectional Encoder Representations from Transformer(BERT) with Multi-Task learning to learn bidirectional context-aware user and item embeddings with neighbourhood sampling. Neighbourhood Sampling technique samples the most influential neighbours for all the items, and our proposed Sequence Header correlates sampled neighbours with users. The main objective of our model is to predict the next item that a user would interact with based on its interaction behaviour. Experiments on three real-world datasets demonstrate that the proposed MVBN model outperforms the state-of-the-art recommendation methods consistently and significantly. Finally, an extensive ablation study is carried out to validate the importance of each component in the system.
Tushar Prakash, Raksha Jalan, Naoyuki Onoe
ICDM3
2023 Impulsion of Movie's Content-Based Factors in Multi-modal Movie Recommendation System
Prabir Mondal, Pulkit Kapoor, Sriparna Saha 0001, Naoyuki Onoe, Brijraj Singh
ICONIP (15)5
2023 Cd-HRNN: Content-Driven HRNN to Improve Session-Based Recommendation System
abstract
The increasing popularity of digital entertainment systems has made personalization a key factor for success in the industry. Recommendation systems, particularly for videos and movies, are crucial in this regard. However, many existing systems are implicit feedback recommendation system that uses indirect signals to infer user preferences, such as user actions (e.g. clicks, views, purchases) or interactions with items (e.g. listening to a song, watching a movie). The challenge lies in the limited information and uncertainty present in user behavior, making it difficult to predict their interests and preferences. In Previous research, Recurrent Neural Networks (RNNs) have shown to be efficient in predicting the next item in a session, based on past item click sequences, but their effectiveness is limited when only relying on click sequences as input data. In this paper, we extend the Hierarchical RNN architecture (HRNN) for generating recommendations by combining session clicks and item content information, such as item ids and item description respectively. The Bidirectional Encoder Representations from Transformers (BERT) architecture is applied for generating feature vectors from text descriptions of the items. Our model has been extensively tested on the benchmark dataset MovieLens 1m has demonstrated superiority over state-of-the-art (SOTA) session-based recommendation systems (SBRS) models. Experimental results establish the efficacy of using content information along with item ids for recommendation.
Sonal Dabral, Brijraj Singh, Naoyuki Onoe
IJCNN3
2023 A Multi-modal Multi-task based Approach for Movie Recommendation
abstract
An online recommendation system is one of the desires of digital e-commerce sectors and the OTT platforms like Amazon Prime, Netflix, SonyLiv, etc. In recent times, with an increase in the interaction of users with the different e-commerce platforms and then analyzing their liking-disliking essence, the recommendation system tries to predict the preference of the user for recommending new items that may capture his attention. In the current study, a multi-task-based architecture is designed to solve the multi-modal movie recommendation problem. Here our hypothesis is that solving two related tasks, namely (a) genre classification of movies and (b) rating identification for a user-movie pair, helps in generating good quality movie embeddings in an end-to-end setting without using a rating vector. For generating the representation of movies, unlike the state-of-the-art techniques, feature vectors extracted from multiple modalities like textual summary, audio and video information present in the movie trailers, and meta-data information are fused together. For representing the user, average representations of movies that are liked by the user are considered. Different multitasking models, fully shared (FS), shared-private (SP), and adversarial shared-private (ASP) feature models are developed for solving the above-mentioned two tasks simultaneously, genre classification, and user-movie rating prediction. For experimental purposes, MMTF-14K: a multifaceted movie trailer feature dataset was extended by incorporating textual features and meta-data information, and a multi-modal version of the MovieLens-100K dataset is used. Results of different multitasking models are shown in terms of RMSE and different rank-based metrics. The proposed multi-task model along with the adversarial training outperforms the state-of-the-art models when applied to the MMTF-14K and multi-modal version of MovieLens-100K datasets.
Subham Raj, Prabir Mondal, Daipayan Chakder, Sriparna Saha 0001, Naoyuki Onoe
IJCNN5
2023 Iteratively Improving Speech Recognition and Voice Conversion
Mayank Kumar Singh, Naoya Takahashi, Naoyuki Onoe
INTERSPEECH3
2023 LLM Based Generation of Item-Description for Recommendation System
abstract
The description of an item plays a pivotal role in providing concise and informative summaries to captivate potential viewers and is essential for recommendation systems. Traditionally, such descriptions were obtained through manual web scraping techniques, which are time-consuming and susceptible to data inconsistencies. In recent years, Large Language Models (LLMs), such as GPT-3.5, and open source LLMs like Alpaca have emerged as powerful tools for natural language processing tasks. In this paper, we have explored how we can use LLMs to generate detailed descriptions of the items. To conduct the study, we have used the MovieLens 1M dataset comprising movie titles and the Goodreads Dataset consisting of names of books and subsequently, an open-sourced LLM, Alpaca, was prompted with few-shot prompting on this dataset to generate detailed movie descriptions considering multiple features like the names of the cast and directors for the ML dataset and the names of the author and publisher for the Goodreads dataset. The generated description was then compared with the scraped descriptions using a combination of Top Hits, MRR, and NDCG as evaluation metrics. The results demonstrated that LLM-based movie description generation exhibits significant promise, with results comparable to the ones obtained by web-scraped descriptions.
Arkadeep Acharya, Brijraj Singh, Naoyuki Onoe
RecSys3
2023 CR-SoRec: BERT driven Consistency Regularization for Social Recommendation
abstract
In the real world, when we seek our friends’ opinions on various items or events, we request verbal social recommendations. It has been observed that we often turn to our friends for recommendations on a daily basis. The emergence of online social platforms has enabled users to share their opinion with their social connections. Therefore, we should consider users’ social connections to enhance online recommendation performance. The social recommendation aims to fuse social links with user-item interactions to offer more relevant recommendations. Several efforts have been made to develop an effective social recommendation system. However, there are two significant limitations to current methods: First, they haven’t thoroughly explored the intricate relationships between the diverse influences of neighbours on users’ preferences. Second, existing models are vulnerable to overfitting due to the relatively low number of user-item interaction records in the interaction space. For the aforementioned problems, this paper offers a novel framework called CR-SoRec, an effective recommendation model based on BERT and consistency regularization. This model incorporates Bidirectional Encoder Representations from Transformer(BERT) to learn bidirectional context-aware user and item embeddings with neighbourhood sampling. The neighbourhood Sampling technique samples the most influential neighbours for all the users/ items. Further, to effectively use the available user-item interaction data and social ties, we leverage diverse perspectives via consistency regularization to harness the underlying information. The main objective of our model is to predict the next item that a user would interact with based on its interaction behaviour and social connections. Experimental results show that our model defines a new state-of-the-art on various datasets and outperforms previous work by a significant margin. Extensive experiments are also conducted to analyze the proposed method.
Tushar Prakash, Raksha Jalan, Brijraj Singh, Naoyuki Onoe
RecSys4
2022 Semi-supervised Acoustic and Language Modeling for Hindi ASR
Tarun Sai Bandarupalli, Shakti Rath, Nirmesh J. Shah, Naoyuki Onoe, Sriram Ganapathy
INTERSPEECH4
2022 Graph Network based Approaches for Multi-modal Movie Recommendation System
abstract
The three times increase of SonyLiv viewers during the Tokyo Olympic, the 10% hike of YouTube users during the isolation era of covid-pandemic, and the 19% growth in Netflix user count due to the fastest growth of OTT, etc. have made the digital platform’s mode all-time active and specific. The hourly increase of users’ interactions and the e-commerce platform’s desire of letting users engage on their sites are pushing researchers to shape the virtual digital web as user specific and revenue-oriented. This paper develops a deep learning-based approach for building a movie recommendation system with three main aspects: (a) using a knowledge graph to embed text and meta information of movies, (b) using multi-modal information of movies like audio, visual frames, text summary, meta data information to generate movie/user representations without directly using rating information; this multi-modal representation can help in coping up with cold-start problem of recommendation system (c) a graph attention network based approach for developing regression system. For meta encoding, we have built knowledge graph from the meta information of the movies directly. For movie-summary embedding, we extracted nouns, verbs, and object to build a knowledge graph with head-relation-tail relationships. A deep neural network, as well as Graph attention networks, are utilized for measuring performance in terms of RMSE score. The proposed system is tested on an extended MovieLens-100K data-set having multi-modal information. Experimental results establish that only rating-based embeddings in the current setup outperform the state-of-the-art techniques but usage of multi-modal information in embedding generation performs better than its single-modal counterparts.1.
Daipayan Chakder, Prabir Mondal, Subham Raj, Sriparna Saha 0001, Angshuman Ghosh, Naoyuki Onoe
SMC6
2021 A Unified Model for Fingerprint Authentication and Presentation Attack Detection
abstract
Typical fingerprint recognition systems are comprised of a spoof detection module and a subsequent recognition module, running one after the other. In this paper, we reformulate the workings of a typical fingerprint recognition system. In particular, we posit that both spoof detection and fingerprint recognition are correlated tasks. Therefore, rather than performing the two tasks separately, we propose a joint model for spoof detection and matching1to simultaneously perform both tasks without compromising the accuracy of either task. We demonstrate the capability of our joint model to obtain an authentication accuracy (1:1 matching) of TAR = 100% @ FAR = 0.1% on the FVC 2006 DB2A dataset while achieving a spoof detection ACE of 1.44% on the LiveDet 2015 dataset, both maintaining the performance of stand-alone methods. In practice, this reduces the time and memory requirements of the fingerprint recognition system by 50% and 40%, respectively; a significant advantage for recognition systems running on resource-constrained devices and communication channels.
Additya Popli, Saraansh Tandon, Joshua J. Engelsma, Naoyuki Onoe, Atsushi Okubo, Anoop M. Namboodiri
IJCB4