EDBT 2026 Demo / reviewers in the wild / expert
Kyosuke Nishida
dblp:01/5962
· DBLP profile ↗
39ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0002-8443-0651ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 7 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-authorHuman-computer interaction and ubiquitous computing · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Debiasing Reward Models via Causally Motivated Inference-Time InterventionabstractReward models (RMs) play a central role in aligning large language models (LLMs) with human preferences.However, RMs are often sensitive to spurious features such as response length.Existing inference-time approaches to mitigating these biases typically focus exclusively on response length, resulting in performance trade-offs.In this paper, we propose causally motivated intervention for mitigating multiple types of biases in RMs at inference time.Our method first identifies neurons whose activations are strongly correlated with predefined bias attributes, and applies neuron-level intervention that suppresses these signals.We evaluate our method on RM benchmarks and observe reductions in sensitivity to spurious features across diverse bias types, without inducing performance trade-offs.Moreover, when used for preference annotation, small RMs (2B and 7B) with our method, which edits less than 2% of all the neurons in RMs, enable LLMs to improve alignment, achieving performance comparable to that of a state-of-the-art 70B RM on AlpacaEval and MT-Bench.Further analysis reveals that bias signals are primarily encoded by neurons in early layers, shedding light on the internal mechanisms of bias exploitation in RMs. Kazutoshi Shinoda, Kosuke Nishida, Kyosuke Nishida |
ACL (1) | 3 |
| 2026 | InstructSum: A Benchmark to Evaluate Instruction-Following Capability of Large Language Models in Summarization
Kosuke Nishida, Kyosuke Nishida, Itsumi Saito |
LREC | 2 |
| 2025 | ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of MindabstractExisting Theory of Mind (ToM) benchmarks diverge from real-world scenarios in three aspects: 1) they assess a limited range of mental states such as beliefs, 2) false beliefs are not comprehensively explored, and 3) the diverse personality traits of characters are overlooked. To address these challenges, we introduce ToMATO, a new ToM benchmark formulated as multiple-choice QA over conversations. ToMATO is generated via LLM-LLM conversations featuring information asymmetry. By employing a prompting method that requires role-playing LLMs to verbalize their thoughts before each utterance, we capture both first- and second-order mental states across five categories: belief, intention, desire, emotion, and knowledge. These verbalized thoughts serve as answers to questions designed to assess the mental states of characters within conversations. Furthermore, the information asymmetry introduced by hiding thoughts from others induces the generation of false beliefs about various mental states. Assigning distinct personality traits to LLMs further diversifies both utterances and thoughts. ToMATO consists of 5.4k questions, 753 conversations, and 15 personality trait patterns. Our analysis shows that this dataset construction approach frequently generates false beliefs due to the information asymmetry between role-playing LLMs, and effectively reflects diverse personalities. We evaluate nine LLMs on ToMATO and find that even GPT-4o mini lags behind human performance, especially in understanding false beliefs, and lacks robustness to various personality traits. Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno, Keita Suzuki, Ryo Masumura, Hiroaki Sugiyama, Kuniko Saito |
AAAI | 3 |
| 2025 | VDocRAG: Retrieval-Augmented Generation over Visually-Rich DocumentsabstractWe aim to develop a retrieval-augmented generation (RAG) framework that answers questions over a corpus of visually- rich documents presented in mixed modalities (e.g., charts, tables) and diverse formats (e.g., PDF, PPTX). In this paper, we introduce a new RAG framework, VDocRAG, which can directly understand varied documents and modalities in a unified image format to prevent missing information that occurs by parsing documents to obtain text. To improve the performance, we propose novel self-supervised pre-training tasks that adapt large vision-language models for retrieval by compressing visual information into dense token representations while aligning them with textual content in documents. Furthermore, we introduce OpenDocVQA, the first unified collection of open-domain document visual question answering datasets, encompassing diverse document types and formats. OpenDocVQA provides a comprehensive resource for training and evaluating retrieval and question answering models on visually-rich documents in an open-domain setting. Experiments show that VDocRAG substantially outperforms conventional text-based RAG and has strong generalization capability, highlighting the potential of an effective RAG paradigm for real-world documents. Ryota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida, Kuniko Saito, Jun Suzuki 0001 |
CVPR | 4 |
| 2025 | Wavelet-based Positional Representation for Long ContextabstractIn the realm of large-scale language models, a significant challenge arises when extrapolating sequences beyond the maximum allowable length.
This is because the model's position embedding mechanisms are limited to positions encountered during training, thus preventing effective representation of positions in longer sequences.
We analyzed conventional position encoding methods for long contexts and found the following characteristics.
(1) When the representation dimension is regarded as the time axis, Rotary Position Embedding (RoPE) can be interpreted as a restricted wavelet transform using Haar-like wavelets.
However, because it uses only a fixed scale parameter, it does not fully exploit the advantages of wavelet transforms, which capture the fine movements of non-stationary signals using multiple scales (window sizes).
This limitation could explain why RoPE performs poorly in extrapolation.
(2)
Previous research as well as our own analysis indicates that Attention with Linear Biases (ALiBi) functions similarly to windowed attention, using windows of varying sizes.
However, it has limitations in capturing deep dependencies because it restricts the receptive field of the model.
From these insights, we propose a new position representation method that captures multiple scales (i.e., window sizes) by leveraging wavelet transforms without limiting the model's attention field.
Experimental results show that this new method improves the performance of the model in both short and long contexts.
In particular, our method allows extrapolation of position information without limiting the model's attention field. Yui Oka, Taku Hasegawa, Kyosuke Nishida, Kuniko Saito |
ICLR | 3 |
| 2025 | Portable Reward Tuning: Towards Reusable Fine-Tuning across Different Pretrained ModelsabstractWhile foundation models have been exploited for various expert tasks with their fine-tuned parameters, any foundation model will be eventually outdated due to its old knowledge or limited capability, and thus should be replaced by a new foundation model. Subsequently, to benefit from its latest knowledge or improved capability, the new foundation model should be fine-tuned on each task again, which incurs not only the additional training cost but also the maintenance cost of the task-specific data. Existing work address this problem by inference-time tuning, i.e., modifying the output probability from the new foundation model by the outputs from the old foundation model and its fine-tuned model, which involves an additional inference cost by the latter two models. In this paper, we explore a new fine-tuning principle (which we call portable reward tuning; PRT) that reduces the inference cost by its nature, based on the reformulation of fine-tuning as the reward maximization with Kullback-Leibler regularization. Specifically, instead of fine-tuning parameters of the foundation models, PRT trains the reward model explicitly through the same loss as in fine-tuning. During inference, the reward model can be used with any foundation model (with the same set of vocabularies or labels) through the formulation of reward maximization. Experimental results, including both vision and language models, show that the PRT-trained model can achieve comparable accuracy with less inference cost, in comparison to the existing work of inference-time tuning. Daiki Chijiwa, Taku Hasegawa, Kyosuke Nishida, Kuniko Saito, Susumu Takeuchi |
ICML | 3 |
| 2024 | InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with InstructionsabstractWe study the problem of completing various visual document understanding (VDU) tasks, e.g., question answering and information extraction, on real-world documents through human-written instructions. To this end, we propose InstructDoc, the first large-scale collection of 30 publicly available VDU datasets, each with diverse instructions in a unified format, which covers a wide range of 12 tasks and includes open document types/formats. Furthermore, to enhance the generalization performance on VDU tasks, we design a new instruction-based document reading and understanding model, InstructDr, that connects document images, image encoders, and large language models (LLMs) through a trainable bridging module. Experiments demonstrate that InstructDr can effectively adapt to new VDU datasets, tasks, and domains via given instructions and outperforms existing multimodal LLMs and ChatGPT without specific training. Ryota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito, Jun Suzuki 0001 |
AAAI | 3 |
| 2024 | Initialization of Large Language Models via Reparameterization to Mitigate Loss SpikesabstractLoss spikes, a phenomenon in which the loss value diverges suddenly, is a fundamental issue in the pre-training of large language models.This paper supposes that the non-uniformity of the norm of the parameters is one of the causes of loss spikes.Here, in training of neural networks, the scale of the gradients is required to be kept constant throughout the layers to avoid the vanishing and exploding gradients problem.However, to meet these requirements in the Transformer model, the norm of the model parameters must be non-uniform, and thus, parameters whose norm is smaller are more sensitive to the parameter update.To address this issue, we propose a novel technique, weight scaling as reparameterization (WeSaR).WeSaR introduces a gate parameter per parameter matrix and adjusts it to the value satisfying the requirements.Because of the gate parameter, WeSaR sets the norm of the original parameters uniformly, which results in stable training.Experimental results with the Transformer decoders consisting of 130 million, 1.3 billion, and 13 billion parameters showed that WeSaR stabilizes and accelerates training and that it outperformed compared methods including popular initialization methods. Kosuke Nishida, Kyosuke Nishida, Kuniko Saito |
EMNLP | 2 |
| 2023 | SlideVQA: A Dataset for Document Visual Question Answering on Multiple ImagesabstractVisual question answering on document images that contain textual, visual, and layout information, called document VQA, has received much attention recently. Although many datasets have been proposed for developing document VQA systems, most of the existing datasets focus on understanding the content relationships within a single image and not across multiple images. In this study, we propose a new multi-image document VQA dataset, SlideVQA, containing 2.6k+ slide decks composed of 52k+ slide images and 14.5k questions about a slide deck. SlideVQA requires complex reasoning, including single-hop, multi-hop, and numerical reasoning, and also provides annotated arithmetic expressions of numerical answers for enhancing the ability of numerical reasoning. Moreover, we developed a new end-to-end document VQA model that treats evidence selection and question answering as a unified sequence-to-sequence format. Experiments on SlideVQA show that our model outperformed existing state-of-the-art QA models, but that it still has a large gap behind human performance. We believe that our dataset will facilitate research on document VQA. Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa, Itsumi Saito, Kuniko Saito |
AAAI | 2 |
| 2023 | Self-Adaptive Named Entity Recognition by Retrieving Unstructured KnowledgeabstractAlthough named entity recognition (NER) helps us to extract domain-specific entities from text (e.g., artists in the music domain), it is costly to create a large amount of training data or a structured knowledge base to perform accurate NER in the target domain.Here, we propose selfadaptive NER, which retrieves external knowledge from unstructured text to learn the usages of entities that have not been learned well.To retrieve useful knowledge for NER, we design an effective two-stage model that retrieves unstructured knowledge using uncertain entities as queries.Our model predicts the entities in the input and then finds those of which the prediction is not confident.Then, it retrieves knowledge by using these uncertain entities as queries and concatenates the retrieved text to the original input to revise the prediction.Experiments on CrossNER datasets demonstrated that our model outperforms strong baselines by 2.35 points in F 1 metric. Kosuke Nishida, Naoki Yoshinaga 0001, Kyosuke Nishida |
EACL | 3 |
| 2023 | DueT: Image-Text Contrastive Transfer Learning with Dual-adapter TuningabstractThis paper presents DueT, a novel transfer learning method for vision and language models built by contrastive learning.In DueT, adapters are inserted into the image and text encoders, which have been initialized using models pre-trained on uni-modal corpora and then frozen.By training only these adapters, DueT enables efficient learning with a reduced number of trainable parameters.Moreover, unlike traditional adapters, those in DueT are equipped with a gating mechanism, enabling effective transfer and connection of knowledge acquired from pre-trained uni-modal encoders while preventing catastrophic forgetting.We report that DueT outperformed simple finetuning, the conventional method fixing only the image encoder and training only the text encoder, and the LoRA-based adapter method in accuracy and parameter efficiency for 0-shot image and text retrieval in both English and Japanese domains. Taku Hasegawa, Kyosuke Nishida, Koki Maeda, Kuniko Saito |
EMNLP | 2 |
| 2023 | Sparse Neural Retrieval Model for Efficient Cross-Domain Retrieval-Based Question Answering
Kosuke Nishida, Naoki Yoshinaga 0001, Kyosuke Nishida |
PACLIC | 3 |
| 2022 | Dialogue Acts Aided Important Utterance Detection Based on Multiparty and Multimodal Information
Fumio Nihei, Ryo Ishii, Yukiko I. Nakano, Kyosuke Nishida, Ryo Masumura, Atsushi Fukayama, Takao Nakamura |
INTERSPEECH | 4 |
| 2022 | Japanese ASR-Robust Pre-trained Language Model with Pseudo-Error Sentences Generated by Grapheme-Phoneme Conversion
Yasuhito Ohsugi, Itsumi Saito, Kyosuke Nishida, Sen Yoshida |
INTERSPEECH | 3 |
| 2021 | VisualMRC: Machine Reading Comprehension on Document ImagesabstractRecent studies on machine reading comprehension have focused on text-level understanding but have not yet reached the level of human understanding of the visual layout and content of real-world documents. In this study, we introduce a new visual machine reading comprehension dataset, named VisualMRC, wherein given a question and a document image, a machine reads and comprehends texts in the image to answer the question in natural language. Compared with existing visual question answering datasets that contain texts in images, VisualMRC focuses more on developing natural language understanding and generation abilities. It contains 30,000+ pairs of a question and an abstractive answer for 10,000+ document images sourced from multiple domains of webpages. We also introduce a new model that extends existing sequence-to-sequence models, pre-trained with large-scale text corpora, to take into account the visual layout and content of documents. Experiments with VisualMRC show that this model outperformed the base sequence-to-sequence models and a state-of-the-art VQA model. However, its performance is still below that of humans on most automatic evaluation metrics. The dataset will facilitate research aimed at connecting vision and language understanding. Ryota Tanaka, Kyosuke Nishida, Sen Yoshida |
AAAI | 2 |
| 2021 | Towards Interpretable and Reliable Reading Comprehension: A Pipeline Model with Unanswerability PredictionabstractMulti-hop QA with annotated supporting facts, which is the task of reading comprehension (RC) considering the interpretability of the answer, has been extensively studied. In this study, we define an interpretable reading comprehension (IRC) model as a pipeline model with the capability of predicting unanswerable queries. The IRC model justifies the answer prediction by establishing consistency between the predicted supporting facts and the actual rationale for interpretability. The IRC model detects unanswerable questions, instead of outputting the answer forcibly based on the insufficient information, to ensure the reliability of the answer. We also propose an end-to-end training method for the pipeline RC model. To evaluate the interpretability and the reliability, we conducted the experiments considering unanswerability in a multi-hop question for a given passage. We show that our end-to-end trainable pipeline model outperformed a non-interpretable model on our modified HotpotQA dataset. Experimental results also show that the IRC model achieves comparable results to the previous non-interpretable models in spite of the trade-off between prediction performance and interpretability. Kosuke Nishida, Kyosuke Nishida, Itsumi Saito, Sen Yoshida |
IJCNN | 2 |
| 2020 | A Transformer-Based Audio Captioning Model with Keyword EstimationabstractOne of the problems with automated audio captioning (AAC) is the indeterminacy in word selection corresponding to the audio event/scene.Since one acoustic event/scene can be described with several words, it results in a combinatorial explosion of possible captions and difficulty in training.To solve this problem, we propose a Transformer-based audio-captioning model with keyword estimation called TRACKE.It simultaneously solves the word-selection indeterminacy problem with the main task of AAC while executing the sub-task of acoustic event detection/acoustic scene classification (i.e., keyword estimation).TRACKE estimates keywords, which comprise a word set corresponding to audio events/scenes in the input audio, and generates the caption while referring to the estimated keywords to reduce word-selection indeterminacy.Experimental results on a public AAC dataset indicate that TRACKE achieved state-ofthe-art performance and successfully estimated both the caption and its keywords. Yuma Koizumi, Ryo Masumura, Kyosuke Nishida, Masahiro Yasuda, Shoichiro Saito |
INTERSPEECH | 3 |
| 2020 | Unsupervised Domain Adaptation of Language Models for Reading ComprehensionabstractThis study tackles unsupervised domain adaptation of reading comprehension (UDARC). Reading comprehension (RC) is a task to learn the capability for question answering with textual sources. State-of-the-art models on RC still do not have general linguistic intelligence; i.e., their accuracy worsens for out-domain datasets that are not used in the training. We hypothesize that this discrepancy is caused by a lack of the language modeling (LM) capability for the out-domain. The UDARC task allows models to use supervised RC training data in the source domain and only unlabeled passages in the target domain. To solve the UDARC problem, we provide two domain adaptation models. The first one learns the out-domain LM and in-domain RC task sequentially. The second one is the proposed model that uses a multi-task learning approach of LM and RC. The models can retain both the RC capability acquired from the supervised data in the source domain and the LM capability from the unlabeled data in the target domain. We evaluated the models on UDARC with five datasets in different domains. The models outperformed the model without domain adaptation. In particular, the proposed model yielded an improvement of 4.3/4.2 points in EM/F1 in an unseen biomedical domain. Kosuke Nishida, Kyosuke Nishida, Itsumi Saito, Hisako Asano, Junji Tomita |
LREC | 2 |
| 2019 | Answering while Summarizing: Multi-task Learning for Multi-hop QA with Evidence ExtractionabstractKosuke Nishida, Kyosuke Nishida, Masaaki Nagata, Atsushi Otsuka, Itsumi Saito, Hisako Asano, Junji Tomita. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Kosuke Nishida, Kyosuke Nishida, Masaaki Nagata, Atsushi Otsuka, Itsumi Saito, Hisako Asano, Junji Tomita |
ACL (1) | 2 |
| 2019 | Multi-style Generative Reading ComprehensionabstractKyosuke Nishida, Itsumi Saito, Kosuke Nishida, Kazutoshi Shinoda, Atsushi Otsuka, Hisako Asano, Junji Tomita. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Kyosuke Nishida, Itsumi Saito, Kosuke Nishida, Kazutoshi Shinoda, Atsushi Otsuka, Hisako Asano, Junji Tomita |
ACL (1) | 1 |
| 2019 | Generalized Large-Context Language Models Based on Forward-Backward Hierarchical Recurrent Encoder-Decoder ModelsabstractThis paper presents a generalized form of large-context language models (LCLMs) that can take linguistic contexts beyond utterance boundaries into consideration. In discourse-level and conversation-level automatic speech recognition (ASR) tasks, which have to handle a series of utterances, it is essential to capture long-range linguistic contexts beyond utterance boundaries. The LCLMs of previous studies mainly focused on utilizing past contexts, and none fully utilized future contexts because LMs typically process words in a time-ordered manner. Our key idea is to introduce the LCLMs into the situation where ASR results of the whole series of utterances are given by a first decoding pass. This situation makes it possible for the LCLMs to leverage future contexts. In this paper, we propose generalized LCLMs (GLCLMs) based on forward-backward hierarchical recurrent encoder-decoder models in which generative probabilities of individual utterances are computed by leveraging not only past contexts but also future contexts beyond utterance boundaries. In order to efficiently introduce GLCLMs to ASR, we also propose a global-context iterative rescoring method that repeatedly rescores the ASR hypotheses of an individual utterance by using surrounding ASR hypotheses. Experiments on discourse-level ASR tasks demonstrate the effectiveness of our GLCLM approach. Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Itsumi Saito, Kyosuke Nishida, Takanobu Oba |
ASRU | 5 |
| 2018 | Retrieve-and-Read: Multi-task Learning of Information Retrieval and Reading ComprehensionabstractThis study considers the task of machine reading at scale (MRS) wherein, given a question, a system first performs the information retrieval (IR) task of finding relevant passages in a knowledge source and then carries out the reading comprehension (RC) task of extracting an answer span from the passages. Previous MRS studies, in which the IR component was trained without considering answer spans, struggled to accurately find a small number of relevant passages from a large set of passages. In this paper, we propose a simple and effective approach that incorporates the IR and RC tasks by using supervised multi-task learning in order that the IR component can be trained by considering answer spans. Experimental results on the standard benchmark, answering SQuAD questions using the full Wikipedia as the knowledge source, showed that our model achieved state-of-the-art performance. Moreover, we thoroughly evaluated the individual contributions of our model components with our new Japanese dataset and SQuAD. The results showed significant improvements in the IR task and provided a new perspective on IR for RC: it is effective to teach which part of the passage answers the question rather than to give only a relevance score to the whole passage. Kyosuke Nishida, Itsumi Saito, Atsushi Otsuka, Hisako Asano, Junji Tomita |
CIKM | 1 |
| 2018 | Commonsense Knowledge Base Completion and GenerationabstractThis study focuses on acquisition of commonsense knowledge.A previous study proposed a commonsense knowledge base completion (CKB completion) method that predicts a confidence score of triplet-style knowledge for improving the coverage of CKBs.To improve the accuracy of CKB completion and expand the size of CKBs, we formulate a new commonsense knowledge base generation task (CKB generation) and propose a joint learning method that incorporates both CKB completion and CKB generation.Experimental results show that the joint learning method improved completion accuracy and the generation model created reasonable knowledge.Our generation model could also be used to augment data and improve the accuracy of completion. Itsumi Saito, Kyosuke Nishida, Hisako Asano, Junji Tomita |
CoNLL | 2 |
| 2018 | Non-Parallel Voice Conversion Using Variational Autoencoders Conditioned by Phonetic Posteriorgrams and D-VectorsabstractThis paper proposes novel frameworks for non-parallel voice conversion (VC) using variational autoencoders (VAEs). Although conventional VAE-based VC models can be trained using non-parallel speech corpora with given speaker representations, phonetic contents of the converted speech tend to vanish because of an over-regularization issue often observed in latent variables of the VAEs. To overcome the issue, this paper proposes a VAE-based non-parallel VC conditioned by not only the speaker representations but also phonetic contents of speech represented as phonetic posteriorgrams (PPGs). Since the phonetic contents are given during the training, we can expect that the VC models effectively learn speaker-independent latent features of speech. Focusing on the point, this paper also extends the conventional VAE-based non-parallel VC to many-to-many VC that can convert arbitrary speakers' characteristics into another arbitrary speakers' ones. We investigate two methods to estimate speaker representations for speakers not included in speech corpora used for training VC models: 1) adapting conventional speaker codes, and 2) using d-vectors for the speaker representations. Experimental results demonstrate that 1) PPGs successfully improve both naturalness and speaker similarity of the converted speech, and 2) both speaker codes and d-vectors can be adopted to the VAE-based many-to-many non-parallel VC. Yuki Saito 0001, Yusuke Ijima, Kyosuke Nishida, Shinnosuke Takamichi |
ICASSP | 3 |
| 2017 | Understanding the Semantic Structures of Tables with a Hybrid Deep Neural Network ArchitectureabstractWe propose a new deep neural network architecture, TabNet, for table type classification. Table type is essential information for exploring the power of Web tables, and it is important to understand the semantic structures of tables in order to classify them correctly. A table is a matrix of texts, analogous to an image, which is a matrix of pixels, and each text consists of a sequence of tokens. Our hybrid architecture mirrors the structure of tables: its recurrent neural network (RNN) encodes a sequence of tokens for each cell to create a 3d table volume like image data, and its convolutional neural network (CNN) captures semantic features, e.g., the existence of rows describing properties, to classify tables. Experiments using Web tables with various structures and topics demonstrated that TabNet achieved considerable improvements over state-of-the-art methods specialized for table classification and other deep neural network architectures. Kyosuke Nishida, Kugatsu Sadamitsu, Ryuichiro Higashinaka, Yoshihiro Matsuo |
AAAI | 1 |
| 2017 | Automatically Extracting Variant-Normalization Pairs for Japanese Text NormalizationabstractSocial media texts, such as tweets from Twitter, contain many types of non-standard tokens, and the number of normalization approaches for handling such noisy text has been increasing. We present a method for automatically extracting pairs of a variant word and its normal form from unsegmented text on the basis of a pair-wise similarity approach. We incorporated the acquired variant-normalization pairs into Japanese morphological analysis. The experimental results show that our method can extract widely covered variants from large Twitter data and improve the recall of normalization without degrading the overall accuracy of Japanese morphological analysis. Itsumi Saito, Kyosuke Nishida, Kugatsu Sadamitsu, Kuniko Saito, Junji Tomita |
IJCNLP(1) | 2 |
| 2017 | Predicting Destinations from Partial Trajectories Using Recurrent Neural Network
Yuki Endo 0003, Kyosuke Nishida, Hiroyuki Toda, Hiroshi Sawada |
PAKDD (1) | 2 |
| 2016 | How fashionable is each street?: Quantifying road characteristics using social mediaabstractDetermining routes that provide opportunities to satisfy the various demands of users is still an open problem. This is because it is virtually impossible to manually quantify the characteristics of each road and there are few resources describing roads directly such that we meet any demand that may arise. The goal of this study is to automatically quantify the characteristics of roads for demands that can be described using keywords such as “fashionable”. To achieve this goal, we propose a two-stage method that analyzes social media and road networks. First, our method estimates the topic distribution (i.e., the characteristics) of each point-of-interest (POI) by analyzing geotagged texts with the Latent Dirichlet Allocation model. Next, it uses a Markov random field model to estimate the characteristics of each road on the basis of those of POIs and the road networks associated with the POIs. Experiments on real datasets demonstrate that our method achieves statistically significant improvements over baseline methods in terms of ranking quality in the information retrieval for roads in three areas given 25 keywords. Takuya Nishimura 0002, Kyosuke Nishida, Hiroyuki Toda, Hiroshi Sawada |
ASONAM | 2 |
| 2016 | Deep Feature Extraction from Trajectories for Transportation Mode Estimation
Yuki Endo 0003, Hiroyuki Toda, Kyosuke Nishida, Akihisa Kawanobe |
PAKDD (2) | 3 |
| 2014 | Probabilistic identification of visited point-of-interest for personalized automatic check-inabstractAutomatic check-in, which is to identify a user's visited points of interest (POIs) from his or her trajectories, is still an open problem because of positioning errors and the high POI density in small areas. In this study, we propose a probabilistic visited-POI identification method. The method uses a new hierarchical Bayesian model for identifying the latent visited-POI label of stay points, which are automatically extracted from trajectories. This model learns from labeled and unlabeled stay point data (i.e., semi-supervised learning) and takes into account personal preferences, stay locations including positioning errors, stay times for each category, and prior knowledge about typical user preferences and stay times. Experimental results with real user trajectories and POIs of Foursquare demonstrated that our method achieved statistically significant improvements in precision at 1 and recall at 3 over the nearest neighbor method and a conventional method that uses a supervised learning-to-rank algorithm. Kyosuke Nishida, Hiroyuki Toda, Takeshi Kurashima, Yoshihiko Suhara |
UbiComp | 1 |
| 2013 | What is he/she like?: estimating Twitter user attributes from contents and social neighborsabstractWe propose a new method for estimating user attributes (gender, age, occupation, and interests) of a Twitter user from the user's contents (profile document and tweets) and social neighbors, i.e. those whom the user has mentioned. Our labeling method is able to collect a large amount of training data automatically by using Twitter users associated with a blog account. Furthermore, we experiment estimation methods using social neighbors with three adjustable levels of its information and show that our method, which uses the target user's profile document and tweets and the neighbors' profile documents (not including tweets), achieves the best accuracy. Jun Ito, Takahide Hoshide, Hiroyuki Toda, Tadasu Uchiyama, Kyosuke Nishida |
ASONAM | 5 |
| 2012 | Improving tweet stream classification by detecting changes in word probabilityabstractWe propose a classification model of tweet streams in Twitter, which are representative of document streams whose statistical properties will change over time. Our model solves several problems that hinder the classification of tweets; in particular, the problem that the probabilities of word occurrence change at different rates for different words. Our model switches between two probability estimates based on full and recent data for each word when detecting changes in word probability. This switching enables our model to achieve both accurate learning of stationary words and quick response to bursty words. We then explain how to implement our model by using a word suffix array, which is a full-text search index. Using the word suffix array allows our model to handle the temporal attributes of word n-grams effectively. Experiments on three tweet data sets demonstrate that our model offers statistically significant higher topic-classification accuracy than conventional temporally-aware classification models. Kyosuke Nishida, Takahide Hoshide, Ko Fujimura |
SIGIR | 1 |
| 2011 | Dynamic Network Motifs: Evolutionary Patterns of Substructures in Complex Networks
Yutaka Kabutoya, Kyosuke Nishida, Ko Fujimura |
APWeb | 2 |
| 2010 | Hierarchical auto-tagging: organizing Q&A knowledge for everyoneabstractWe propose a hierarchical auto-tagging system, TagHats, to improve users' knowledge sharing. Our system assigns three different levels of tags to Q&A documents: category, theme, and keyword. Multiple category tags can organize a document according to multiple viewpoints, and multiple theme and keyword tags can identify what the document is about clearly. Moreover, these hierarchical tags will be helpful in organizing documents to support everyone because different users have different demands in terms of tag specificity. Our system consists of a hierarchical classification method for assigning category and theme tags, a new keyword extraction method that considers the structure of Q&A documents, and a new method for selecting theme tag candidates from each category. Experiments with the documents of Oshiete! goo demonstrate that our system is able to assign hierarchical tags to the documents appropriately and is capable of outperforming baseline methods significantly. Kyosuke Nishida, Ko Fujimura |
CIKM | 1 |
| 2009 | Learning, detecting, understanding, and predicting concept changesabstractThe demand for learning machines that can adapt to concept change, the change over time of the statistical properties of a target variable, has become more urgent. We, therefore, propose a system in which multiple online and offline classifiers are used for learning changing concepts. Our system is able to: respond to both sudden and gradual changes, handle recurring concepts, detect the occurrence of change, understand the hidden contexts of past concepts, and predict the next concept. We evaluate the effectiveness of our system's elements and demonstrate that our system performed well with synthetic concept-drifting and concept-shifting datasets. Kyosuke Nishida, Koichiro Yamauchi 0001 |
IJCNN | 1 |
| 2008 | Detecting sudden concept drift with knowledge of human behaviorabstractConcept drift, the change over time of the statistical properties of the target variable, is a serious problem for online learning systems. To overcome this problem, we propose a method inspired by human behavior for detecting sudden concept drift. We first conducted a human behavior experiment to investigate our working hypothesis that humans can detect changes quickly when their confident classifications are rejected despite the fact that their recent classification accuracy is high. The human behavior experiments supported our working hypothesis. We then have proposed the leaky integrate-and-detect (LID) model based on our working hypothesis. Our computer experiments showed LID was able to detect sudden changes quickly and accurately in an environment that includes noise and gradual changes. Kyosuke Nishida, Shohei Shimada, Satoru Ishikawa, Koichiro Yamauchi 0001 |
SMC | 1 |
| 2007 | Detecting Concept Drift Using Statistical Testing
Kyosuke Nishida, Koichiro Yamauchi 0001 |
Discovery Science | 1 |
| 2007 | Quick online feature selection method for regression -A feature selection method inspired by human behavior-abstractThe task of variable selection is essential to improving the ability of machine learning systems to generalize. Although there are many conventional variable selection methods, almost all of them need to prepare and learn a large number of samples in advance because they are based on offline learning. This property is not suitable for online learning systems. To overcome this inconvenience, we propose a quick online variable selection method inspired by human problem solving behaviors. The proposed method tries to generate several variable set candidates in a speculative manner using a filter method and evaluates them using a wrapper method. The method can also function in concept-drifting environments, where relevant variable sets are changing. The experimental results show that the new method yields appropriate variable sets from a small number of samples. Youhei Tadeuchi, Ryuji Oshima, Kyosuke Nishida, Koichiro Yamauchi 0001, Takashi Omori |
SMC | 3 |
| 2004 | An On-Line Learning Algorithm with Dimension Selection Using Minimal Hyper Basis Function Networks
Kyosuke Nishida, Koichiro Yamauchi 0001, Takashi Omori |
ICONIP | 1 |