EDBT 2026 Demo / reviewers in the wild / expert
Fumiyo Fukumoto
dblp:57/1318
· DBLP profile ↗
68ranked-venue papers
27as first author
22since 2021 · last 2025
0000-0001-7858-6206ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 19 first-author · 16 since 2021Databases, data management, data science and information retrieval · 22 · 10 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Claim veracity assessment for explainable fake news detectionabstractWith the rapid growth of social network services, misinformation has spread uncontrollably. Most recent approaches to fake news detection use neural network models to predict whether the input text is fake or real. Some of them even provide explanations, in addition to veracity, generated by Large Language Models (LLMs). However, they do not utilize factual evidence, nor do they allude to it or provide evidence/justification, thereby making their predictions less credible. This paper proposes a new fake news detection method that predicts the truth or false-hood of a claim based on relevant factual evidence (if exists) or LLM’s inference mechanisms (such as common-sense reasoning) otherwise. Our method produces the final synthesized prediction, along with well-founded facts or reasoning. Experimental results on several large COVID-19 fake news datasets show that our method achieves state-of-the-art (SOTA) detection and evidence explanation performance. Our source codes are available online. Bassamtiano Renaufalgi Irnawan, Noriko Tomuro, Fumiyo Fukumoto, Yoshimi Suzuki |
COLING | 4 |
| 2024 | Enhanced Coherence-Aware Network with Hierarchical Disentanglement for Aspect-Category Sentiment AnalysisabstractAspect-category-based sentiment analysis (ACSA), which aims to identify aspect categories and predict their sentiments has been intensively studied due to its wide range of NLP applications. Most approaches mainly utilize intrasentential features. However, a review often includes multiple different aspect categories, and some of them do not explicitly appear in the review. Even in a sentence, there is more than one aspect category with its sentiments, and they are entangled intra-sentence, which makes the model fail to discriminately preserve all sentiment characteristics. In this paper, we propose an enhanced coherence-aware network with hierarchical disentanglement (ECAN) for ACSA tasks. Specifically, we explore coherence modeling to capture the contexts across the whole review and to help the implicit aspect and sentiment identification. To address the issue of multiple aspect categories and sentiment entanglement, we propose a hierarchical disentanglement module to extract distinct categories and sentiment features. Extensive experimental and visualization results show that our ECAN effectively decouples multiple categories and sentiments entangled in the coherence representations and achieves state-of-the-art (SOTA) performance. Our codes and data are available online: https://github.com/cuijin-23/ECAN. Jin Cui 0005, Fumiyo Fukumoto, Xinfeng Wang, Yoshimi Suzuki, Jiyi Li, Noriko Tomuro, Wanzeng Kong |
LREC/COLING | 2 |
| 2024 | Enhancing High-order Interaction Awareness in LLM-based Recommender ModelabstractLarge language models (LLMs) have demonstrated prominent reasoning capabilities in recommendation tasks by transforming them into text-generation tasks.However, existing approaches either disregard or ineffectively model the user-item high-order interactions.To this end, this paper presents an enhanced LLM-based recommender (ELMRec).We enhance whole-word embeddings to substantially enhance LLMs' interpretation of graphconstructed interactions for recommendations, without requiring graph pre-training.This finding may inspire endeavors to incorporate rich knowledge graphs into LLM-based recommenders via whole-word embedding.We also found that LLMs often recommend items based on users' earlier interactions rather than recent ones, and present a reranking solution.Our ELMRec outperforms state-of-the-art (SOTA) methods in both direct and sequential recommendations.Our code is available online 1 . Xinfeng Wang, Fumiyo Fukumoto, Yoshimi Suzuki |
EMNLP | 3 |
| 2024 | Reduction-Synthesis: Plug-and-Play for Sentiment Style TransferabstractSentiment style transfer (SST), a variant of text style transfer (TST), has recently attracted extensive interest.Some disentangling-based approaches have improved performance, while most still struggle to properly transfer the input as the sentiment style is intertwined with the content of the text.To alleviate the issue, we propose a plug-and-play method that leverages an iterative self-refinement algorithm with a large language model (LLM).Our approach separates the straightforward Seq2Seq generation into two phases: (1) Reduction phase which generates a style-free sequence for a given text, and (2) Synthesis phase which generates the target text by leveraging the sequence output from the first phase.The experimental results on two datasets demonstrate that our transfer method is effective for challenging SST cases where the baseline methods perform poorly.Our code is available online 1 . Fumiyo Fukumoto, Yoshimi Suzuki |
INLG | 2 |
| 2024 | CaDRec: Contextualized and Debiased Recommender ModelabstractRecommender models aimed at mining users' behavioral patterns have raised great attention as one of the essential applications in daily life. Recent work on graph neural networks (GNNs) or debiasing methods has attained remarkable gains. However, they still suffer from (1) over-smoothing node embeddings caused by recursive convolutions with GNNs, and (2) the skewed distribution of interactions due to popularity and user-individual biases. This paper proposes a contextualized and debiased recommender model (CaDRec). To overcome the over-smoothing issue, we explore a novel hypergraph convolution operator that can select effective neighbors during convolution by introducing both structural context and sequential context. To tackle the skewed distribution, we propose two strategies for disentangling interactions: (1) modeling individual biases to learn unbiased item embeddings, and (2) incorporating item popularity with positional encoding. Moreover, we mathematically show that the imbalance of the gradients to update item embeddings exacerbates the popularity bias, thus adopting regularization and weighting schemes as solutions. Extensive experiments on four datasets demonstrate the superiority of the CaDRec against state-of-the-art (SOTA) methods. Our source code and data are released at https://github.com/WangXFng/CaDRec. Xinfeng Wang, Fumiyo Fukumoto, Jin Cui 0005, Yoshimi Suzuki, Jiyi Li, Dongjin Yu |
SIGIR | 2 |
| 2024 | NFARec: A Negative Feedback-Aware Recommender ModelabstractGraph neural network (GNN)-based models have been extensively studied for recommendations, as they can extract high-order collaborative signals accurately which is required for high-quality recommender systems. However, they neglect the valuable information gained through negative feedback in two aspects: (1) different users might hold opposite feedback on the same item, which hampers optimal information propagation in GNNs, and (2) even when an item vastly deviates from users' preferences, they might still choose it and provide a negative rating. In this paper, we propose a negative feedback-aware recommender model (NFARec) that maximizes the leverage of negative feedback. To transfer information to multi-hop neighbors along an optimal path effectively, NFARec adopts a feedback-aware correlation that guides hypergraph convolutions (HGCs) to learn users' structural representations. Moreover, NFARec incorporates an auxiliary task - predicting the feedback sentiment polarity (i.e., positive or negative) of the next interaction - based on the Transformer Hawkes Process. The task is beneficial for understanding users by learning the sentiment expressed in their previous sequential feedback patterns and predicting future interactions. Extensive experiments demonstrate that NFARec outperforms competitive baselines. Our source code and data are released at https://github.com/WangXFng/NFARec. Xinfeng Wang, Fumiyo Fukumoto, Jin Cui 0005, Yoshimi Suzuki, Dongjin Yu |
SIGIR | 2 |
| 2023 | Multi-Feature and Multi-Channel GCNs for Aspect Based Sentiment Analysis
Wenlong Xi, Xiaoxi Huang, Fumiyo Fukumoto, Yoshimi Suzuki |
DEXA (2) | 3 |
| 2023 | Hierarchy-Aware Bilateral-Branch Network for Imbalanced Hierarchical Text Classification
Jiangjiang Zhao, Jiyi Li, Fumiyo Fukumoto |
DEXA (2) | 3 |
| 2023 | An Efficient Approach for Improving the Recall of Rough Abstract Retrieval in Scientific Claim Verification
Zhiwei Zhang 0011, Jiyi Li, Fumiyo Fukumoto |
ICANN (8) | 3 |
| 2023 | Identifying Self-admitted Technical Debt with Context-Based Ladder Network
Aiyue Gong, Fumiyo Fukumoto, Panitan Muangkammuen, Jiyi Li, Dongjin Yu |
ICONIP (15) | 2 |
| 2023 | Decoupling Style from Contents for Positive Text Reframing
Yoshimi Suzuki, Jiyi Li, Fumiyo Fukumoto |
ICONIP (13) | 4 |
| 2023 | Learning Disentangled Meaning and Style Representations for Positive Text ReframingabstractThe positive text reframing (PTR) task which generates a text giving a positive perspective with preserving the sense of the input text, has attracted considerable attention as one of the NLP applications.Due to the significant representation capability of the pre-trained language model (PLM), a beneficial baseline can be easily obtained by just fine-tuning the PLM.However, how to interpret a diversity of contexts to give a positive perspective is still an open problem.Especially, it is more serious when the size of the training data is limited.In this paper, we present a PTR framework, that learns representations where the meaning and style of text are disentangled.The method utilizes pseudopositive reframing datasets which are generated with two augmentation strategies.A simple but effective multi-task learning-based model is applied to fuse the generation capabilities from these datasets.Experimental results on Positive Psychology Frames (PPF) dataset, show that our approach outperforms the baselines, BART by five and T5 by six evaluation metrics.Our source codes and data are available online. DecoderThis is a challenging work.This work is difficult.This work is difficult. Decoder Layer nLayer 2 Layer 1••• Fumiyo Fukumoto, Jiyi Li, Kentaro Go, Yoshimi Suzuki |
INLG | 2 |
| 2023 | Vision-Language Navigation for Quadcopters with Conditional Transformer and Prompt-based Text RephraserabstractControlling drones with natural language instructions is an important topic in Vision-and-Language Navigation (VLN). However, previous models can not effectively guide drones with the integration of multimodal features, as few of them exploit the correlations between instructions and the environmental contexts and consider the model’s capacity to understand natural languages. Therefore, we propose a novel language-enhanced cross-modal model that has a conditional Transformer to effectively integrate the multimodal features, i.e., the textual instructions and visual contexts. To enhance the ability of language representation, we also employ SentenceBERT. In addition, to address the issue that users could provide various textual instructions even for the same navigation task, we propose a prompt-based approach by introducing an LLM-based intermediary component (LLMIR) for rephrasing users’ instructions. We evaluate our approaches with a quadcopter simulator. Our model improves the absolute task completion rate by 1.39%. To evaluate LLMIR, we create a new test set by extracting the essential and minimal instructions from the original test set. By using the LLM, the task completion rate improves by 1.51%. And it narrows the performance gap between new and original test set by 34.83%. Jiyi Li, Fumiyo Fukumoto, Peng Liu 0027, Yoshimi Suzuki |
MMAsia | 3 |
| 2023 | EEDN: Enhanced Encoder-Decoder Network with Local and Global Context Learning for POI RecommendationabstractThe point-of-interest (POI) recommendation predicts users' destinations, which might be of interest to users and has attracted considerable attention as one of the major applications in location-based social networks (LBSNs). Recent work on graph-based neural networks (GNN) or matrix factorization-based (MF) approaches has resulted in better representations of users and POIs to forecast users' latent preferences. However, they still suffer from the implicit feedback and cold-start problems of check-in data, as they cannot capture both local and global graph-based relations among users (or POIs) simultaneously, and the cold-start neighbors are not handled properly during graph convolution in GNN. In this paper, we propose an enhanced encoder-decoder network (EEDN) to exploit rich latent features between users, POIs, and interactions between users and POIs for POI recommendation. The encoder of EEDN utilizes a hybrid hypergraph convolution to enhance the aggregation ability of each graph convolution step and learns to derive more robust cold-start-aware user representations. In contrast, the decoder mines local and global interactions by both graph- and sequential-based patterns for modeling implicit feedback, especially to alleviate exposure bias. Extensive experiments in three public real-world datasets demonstrate that EEDN outperforms state-of-the-art methods. Our source codes and data are released at https://github.com/WangXFng/EEDN Xinfeng Wang, Fumiyo Fukumoto, Jin Cui 0005, Yoshimi Suzuki, Jiyi Li, Dongjin Yu |
SIGIR | 2 |
| 2023 | STaTRL: Spatial-temporal and text representation learning for POI recommendation
Xinfeng Wang, Fumiyo Fukumoto, Jiyi Li, Dongjin Yu, Xiaoxiao Sun 0001 |
Appl. Intell. | 2 |
| 2022 | BBSN: Bilateral-Branch Siamese Network for Imbalanced Multi-label Text Classification
Jiangjiang Zhao, Jiyi Li, Fumiyo Fukumoto |
ICONIP (3) | 3 |
| 2021 | Semi-Supervised Learning for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis is a rapidly growing domain in natural language processing which is a fine-grained study. Within this broad field, most existing studies use large amounts of labeled data by deep learning methods. However, obtaining massive quantities of labeled data to train a deep neural network model is frequently time-consuming and laborious. In this paper, we focus on semi-supervised learning based on ACSA with few labeled data in restaurant reviews and scholarly paper reviews. In order to leverage information from unlabeled data, the semi-supervised learning method-Ladder network is proposed to fix the problem. Furthermore, the pre-trained language models BERT, ALBERT and Longformer are used for text pre-processing and feature extraction. Extensive experiments on both datasets demonstrate the superiority of the Longformer based Ladder Network compared with supervised learning methods and other semi-supervised learning methods including$\Gamma$-Model and VAT. Yoshimi Suzuki, Fumiyo Fukumoto, Hiromitsu Nishizaki |
CW | 4 |
| 2021 | Multi-task Neural Shared Structure Search: A Study Based on Text Mining
Jiyi Li, Fumiyo Fukumoto |
DASFAA (2) | 2 |
| 2021 | Abstract, Rationale, Stance: A Joint Model for Scientific Claim VerificationabstractScientific claim verification can help the researchers to easily find the target scientific papers with the sentence evidence from a large corpus for the given claim.Some existing works propose pipeline models on the three tasks of abstract retrieval, rationale selection and stance prediction.Such works have the problems of error propagation among the modules in the pipeline and lack of sharing valuable information among modules.We thus propose an approach, named as ARSJOINT, that jointly learns the modules for the three tasks with a machine reading comprehension framework by including claim information.In addition, we enhance the information exchanges and constraints among tasks by proposing a regularization term between the sentence attention scores of abstract retrieval and the estimated outputs of rational selection.The experimental results on the benchmark dataset SCI-FACT show that our approach outperforms the existing works. Zhiwei Zhang 0011, Jiyi Li, Fumiyo Fukumoto, Yanming Ye |
EMNLP (1) | 3 |
| 2021 | Neural Local and Global Contexts Learning for Word Sense Disambiguation
Fumiyo Fukumoto, Taishin Mishima, Jiyi Li, Yoshimi Suzuki |
ICONIP (4) | 1 |
| 2021 | Predominant Sense Acquisition with a Neural Random Walk Model
Attaporn Wangpoonsarp, Fumiyo Fukumoto |
ICONIP (3) | 2 |
| 2021 | Paraphrase Identification with Neural Elaboration Relation Learning
Fumiyo Fukumoto, Jiyi Li, Yoshimi Suzuki |
ICONIP (4) | 2 |
| 2020 | A Neural Local Coherence Analysis Model for Clarity Text ScoringabstractLocal coherence relation between two phrases/sentences, such as cause-effect, and contrast gives a strong influence on whether a text is well-structured or not.This paper follows the assumption and presents a method for scoring text clarity by utilizing local coherence between adjacent sentences.We hypothesize that the coherence knowledge learned from the different domain data is beneficial for capturing a well-structured text and thus helpful for scoring text clarity.We propose a text clarity scoring method that utilizes local coherence analysis with an out-domain setting, i.e., the training data for the source and target tasks are different from each other.The method based on the pre-trained language model BERT firstly trains the local coherence model as an auxiliary manner and then re-trains it together with the text clarity scoring model.The experimental results by using the PeerRead benchmark dataset show the improvement compared with a single model, scoring text clarity model 1 . Panitan Muangkammuen, Fumiyo Fukumoto, Kanda Runapongsa Saikaew, Jiyi Li |
COLING | 3 |
| 2020 | Sentiment analysis using semi-supervised learning with few labeled dataabstractSentiment analysis has been widely explored in many text domains, including tweets, movie reviews, shop/restaurant reviews, product reviews, and peer reviews for scholarly papers. However, it is very costly to manually label the training data for sentiment analysis. We focus on the problem and presents an approach for leveraging contextual features from unlabeled movie and restaurant reviews with a neural-network-based learning model, Ladder network. The experimental results by using two benchmark datasets, IMDb and YelpNYC, show that our model outperforms the baseline models including LSTM and SVM. Especially we verified that our model is better performance gaining on limited training datasets with 1% data labeled. Our source codes are available online.11Our source code can be obtained from https://github.com/jepyh/sentiment_analysis_few_labeled. Yuhao Pan, Zhiqun Chen, Yoshimi Suzuki, Fumiyo Fukumoto, Hiromitsu Nishizaki |
CW | 4 |
| 2020 | HSCNN: A Hybrid-Siamese Convolutional Neural Network for Extremely Imbalanced Multi-label Text ClassificationabstractThe data imbalance problem is a crucial issue for the multi-label text classification.Some existing works tackle it by proposing imbalanced loss objectives instead of the vanilla cross-entropy loss, but their performances remain limited in the cases of extremely imbalanced data.We propose a hybrid solution which adapts general networks for the head categories, and few-shot techniques for the tail categories.We propose a Hybrid-Siamese Convolutional Neural Network (HSCNN) with additional technical attributes, i.e., a multi-task architecture based on Single and Siamese networks; a category-specific similarity in the Siamese structure; a specific sampling method for training HSCNN.The results using two benchmark datasets and three loss objectives show that our method can improve the performance of Single networks with diverse loss objectives on the tail or entire categories. Wenshuo Yang, Jiyi Li, Fumiyo Fukumoto, Yanming Ye |
EMNLP (1) | 3 |
| 2020 | Semi-Automatic Construction and Refinement of an Annotated Corpus for a Deep Learning Framework for Emotion ClassificationabstractIn the case of using a deep learning (machine learning) framework for emotion classification, one significant difficulty faced is the requirement of building a large, emotion corpus in which each sentence is assigned emotion labels. As a result, there is a high cost in terms of time and money associated with the construction of such a corpus. Therefore, this paper proposes a method of creating a semi-automatically constructed emotion corpus. For the purpose of this study sentences were mined from Twitter using some emotional seed words that were selected from a dictionary in which the emotion words were well-defined. Tweets were retrieved by one emotional seed word, and the retrieved sentences were assigned emotion labels based on the emotion category of the seed word. It was evident from the findings that the deep learning-based emotion classification model could not achieve high levels of accuracy in emotion classification because the semi-automatically constructed corpus had many errors when assigning emotion labels. In this paper, therefore, an approach for improving the quality of the emotion labels by automatically correcting the errors of emotion labels is proposed and tested. The experimental results showed that the proposed method worked well, and the classification accuracy rate was improved to 55.1% from 44.9% on the Twitter emotion classification task. Kyosuke Masuda, Hiromitsu Nishizaki, Fumiyo Fukumoto, Yoshimi Suzuki |
LREC | 4 |
| 2019 | Text Categorization by Learning Predominant Sense of Words as Auxiliary TaskabstractDistributions of the senses of words are often highly skewed and give a strong influence of the domain in a document.This paper follows the assumption and presents a method for text categorization by leveraging the predominant sense of words depending on the domain, i.e., domain-specific senses.The key idea is that the features learned from predominant senses are possible to discriminate the domain of the document and thus improve the overall performance of text categorization.We propose a multi-task learning framework based on the neural network model, transformer, which trains a model to simultaneously categorize documents and predicts a predominant sense for each word.The experimental results using four benchmark datasets including RCV1 show that our method is comparable to the state-of-the-art categorization approach, especially our model works well for categorization of multi-label documents. Kazuya Shimura, Jiyi Li, Fumiyo Fukumoto |
ACL (1) | 3 |
| 2019 | Acquisition of Domain-Specific Senses and Its Extrinsic Evaluation Through Text Categorization
Attaporn Wangpoonsarp, Kazuya Shimura, Fumiyo Fukumoto |
CICLing (2) | 3 |
| 2019 | Integrating Internet Directories by Estimating Category CorrespondencesabstractThis paper focuses on two existing category hierarchies and proposes a method for integrating these hierarchies into one. Integration of hierarchies is proceeded based on semantically related categories which are extracted by using text categorization. We extract semantically related category pairs by estimating category correspondences. Some categories within hierarchies are merged based on the extracted category pairs. We assign the remaining categories to a newly constructed hierarchy. To evaluate the method, we applied the results of new hierarchy to text categorization task. The results showed that the method was effective for categorization. Yoshimi Suzuki, Fumiyo Fukumoto |
KEOD | 2 |
| 2018 | HFT-CNN: Learning Hierarchical Category Structure for Multi-label Short Text CategorizationabstractWe focus on the multi-label categorization task for short texts and explore the use of a hierarchical structure (HS) of categories.In contrast to the existing work using non-hierarchical flat model, the method leverages the hierarchical relations between the categories to tackle the data sparsity problem.The lower the HS level, the worse the categorization performance.Because lower categories are fine-grained and the amount of training data per category is much smaller than that in an upper level.We propose an approach which can effectively utilize the data in the upper levels to contribute categorization in the lower levels by applying a Convolutional Neural Network (CNN) with a finetuning technique.The results using two benchmark datasets show that the proposed method, Hierarchical Fine-Tuning based CNN (HFT-CNN) is competitive with the state-of-the-art CNN based methods. Kazuya Shimura, Jiyi Li, Fumiyo Fukumoto |
EMNLP | 3 |
| 2017 | Is (President, 大統領) a Correct Sense Pair? - Linking and Creating Bilingual Sense Correspondences
Fumiyo Fukumoto, Yoshimi Suzuki, Attaporn Wangpoonsarp |
KEOD | 1 |
| 2017 | Integrating Local and Global Data View for Bilingual Sense Correspondences
Fumiyo Fukumoto, Yoshimi Suzuki, Attaporn Wangpoonsarp, Meng Ji |
IC3K | 1 |
| 2015 | Learning Timeline Difference for Text CategorizationabstractThis paper addresses text categorization problem that training data may derive from a different time period from the test data.We present a learning framework which extends a boosting technique to learn accurate model for timeline adaptation.The results showed that the method was comparable to the current state-of-theart biased-SVM method, especially the method is effective when the creation time period of the test data differs greatly from the training data. Fumiyo Fukumoto, Yoshimi Suzuki |
EMNLP | 1 |
| 2015 | Exploiting Guest Preferences with Aspect-Based Sentiment Analysis for Hotel Recommendation
Fumiyo Fukumoto, Hiroki Sugiyama, Yoshimi Suzuki, Suguru Matsuyoshi |
IC3K | 1 |
| 2015 | Opinion Extraction from Editorial Articles based on Context InformationabstractOpinion extraction supports various tasks such as sentiment analysis in user reviews for recommendations and editorial summarization. In this paper, we address the problem of opinion extraction from newspaper editorials. To extract author’s opinion, we used context information addition to the features within a single sentence only. Context information are a location of the target sentence, and its preceding, and succeeding sentences. We defined the opinion extraction task as a sequence labeling problem, using conditional random fields (CRF). We used Japanese newspaper editorials in the experiments, and used multiple combination of features of CRF to reveal which features are effective for opinion extraction. The experimental results show the effectiveness of the method, especially, predicate expression, location and previous sentence are effective for opinion extraction. Yoshimi Suzuki, Fumiyo Fukumoto |
KEOD | 2 |
| 2014 | Annotating the Focus of Negation in Japanese Text
Suguru Matsuyoshi, Ryo Otsuki, Fumiyo Fukumoto |
LREC | 3 |
| 2013 | Timeline adaptation for text classificationabstractIn this paper, we address the text classification problem that a period of time created test data is different from the training data, and present a method for text classification based on temporal adaptation. We first applied lexical chains for the training data to collect terms with semantic relatedness, and created sets (we call these Sem sets). Semantically related terms in the documents are replaced to their representative term. For the results, we identified short terms that are salient for a specific period of time. Finally, we trained SVM classifiers by applying a temporal weighting function to each selected short terms within the training data, and classified test data. Temporal weighting function is weighted each short term in the training data according to the temporal distance between training and test data. The results using MedLine data showed that the method was comparable to the current state-of-the-art biased-SVM method, especially the method is effective when testing on data far from the training data. Fumiyo Fukumoto, Yoshimi Suzuki, Atsuhiro Takasu |
CIKM | 1 |
| 2012 | Text classification with relatively small positive documents and unlabeled dataabstractThis paper addresses the problem of dealing with a collection of negative training documents which is suitable for relatively small number of positive documents, and presents a method for eliminating the need for manually collecting negative training documents based on supervised machine learning techniques. We applied an error correction technique to the results of negative training data obtained by the Positive Example Based Learning (PEBL). Moreover, we used a boosting technique to learn a set of negative data to train classifiers. The results using Japanese newspaper documents showed that the method contributes for reducing the cost of manual collection of negative training documents. Fumiyo Fukumoto, Takeshi Yamamoto, Suguru Matsuyoshi, Yoshimi Suzuki |
CIKM | 1 |
| 2012 | Topic and Subject Detection in News Streams for Multi-document Summarization
Fumiyo Fukumoto, Yoshimi Suzuki, Atsuhiro Takasu |
KEOD | 1 |
| 2012 | Segmentation of Review Texts by using Thesaurus and Corpus-based Word Similarity
Yoshimi Suzuki, Fumiyo Fukumoto |
KEOD | 2 |
| 2011 | Multi-document Summarization Using Link Analysis Based on Rhetorical Relations between Sentences
Nik Adilah Hanin Binti Zahri, Fumiyo Fukumoto |
CICLing (2) | 2 |
| 2011 | Semantic Classification of Unknown Words based on Graph-based Semi-supervised Clustering
Fumiyo Fukumoto, Yoshimi Suzuki |
KEOD | 1 |
| 2011 | Graph-Based Semi-supervised Clustering for Semantic Classification of Unknown Words
Fumiyo Fukumoto, Yoshimi Suzuki |
IC3K | 1 |
| 2011 | Multi-labeled Patent Document Classification using Technical Term Thesaurus
Yoshimi Suzuki, Fumiyo Fukumoto |
KEOD | 2 |
| 2011 | Cluster Labelling based on Concepts in a Machine-Readable Dictionary
Fumiyo Fukumoto, Yoshimi Suzuki |
IJCNLP | 1 |
| 2010 | Identifying Domain-specific Senses and Its Application to Text Classification
Fumiyo Fukumoto, Yoshimi Suzuki |
KEOD | 1 |
| 2008 | Retrieving Bilingual Verb-Noun Collocations by Integrating Cross-Language Category Hierarchies
Fumiyo Fukumoto, Yoshimi Suzuki, Kazuyuki Yamashita |
COLING | 1 |
| 2008 | Integrating Cross-Language Hierarchies and Its Application to Retrieving Relevant DocumentsabstractInternet directories such as Yahoo! are an approach to improve the efficacy and efficiency of Information Retrieval (IR) on the Web, as pages (documents) are organized into hierarchical categories, and similar pages are grouped together. Most of the search engines on the Web service find documents that are assigned to a single classification hierarchy. Categories in the hierarchy are carefully defined by human experts and documents are well organized. However, a single hierarchy in one language is often insufficient to find all relevant material, as each hierarchy tends to have some bias in both defining hierarchical structure and classifying documents. Moreover, documents written in a language other than the users native language often include large amounts of information related to the users request. In this article, we propose a method of integrating cross-language (CL) category hierarchies, that is, Reuters 96 hierarchy and UDC code hierarchy of Japanese by estimating category similarities. The method does not simply merge two different hierarchies into one large hierarchy but instead extracts sets of similar categories, where each element of the sets is relevant with each other. It consists of three steps. First, we classify documents from one hierarchy into categories with another hierarchy using a cross-language text classification (CLTC) technique, and extract category pairs of two hierarchies. Next, we apply Ç 2 statistics to these pairs to obtain similar category pairs, and finally we apply the generating function of the Apriori algorithm (Apriori-Gen) to the category pairs, and find sets of similar categories. Moreover, we examined whether integrating hierarchies helps to support retrieval of documents with similar contents. The retrieval results showed a 42.7% improvement over the baseline nonhierarchy model, and a 21.6% improvement over a single hierarchy. Fumiyo Fukumoto, Yoshimi Suzuki |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2007 | Topic tracking based on bilingual comparable corpora and semisupervised clusteringabstractIn this paper, we address the problem of skewed data in topic tracking: the small number of stories labeled positive as compared to negative stories and propose a method for estimating effective training stories for the topic-tracking task. For a small number of labeled positive stories, we use bilingual comparable, i.e., English, and Japanese corpora, together with the EDR bilingual dictionary, and extract story pairs consisting of positive and associated stories. To overcome the problem of a large number of labeled negative stories, we classified them into clusters. This is done using a semisupervised clustering algorithm, combining k means with EM. The method was tested on the TDT English corpus and the results showed that the system works well when the topic under tracking is talking about an event originating in the source language country, even for a small number of initial positive training stories. Fumiyo Fukumoto, Yoshimi Suzuki |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2006 | Using Bilingual Comparable Corpora and Semi-supervised Clustering for Topic Tracking
Fumiyo Fukumoto, Yoshimi Suzuki |
ACL | 1 |
| 2006 | Thesaurus expansion using similar word pairs from patent documents
Yoshimi Suzuki, Fumiyo Fukumoto |
INTERSPEECH | 2 |
| 2005 | Topic Tracking Based on Linguistic Features
Fumiyo Fukumoto, Yusuke Yamaji |
IJCNLP | 1 |
| 2004 | Correcting Category Errors in Text Classification
Fumiyo Fukumoto, Yoshimi Suzuki |
COLING | 1 |
| 2004 | A Comparison of Manual and Automatic Constructions of Category Hierarchy for Classifying Large Corpora
Fumiyo Fukumoto, Yoshimi Suzuki |
CoNLL | 1 |
| 2004 | Learning subject drift for topic trackingabstractFor topic tracking where data is collected over an extended period of time, the discussion of a topic, i.e. the in a story changes over time. This paper focuses on subject drift and presents a method for topic tracking on broadcast news stories to handle subject drift. The basic idea is to automatically extract the optimal positive training data of the target topic so as to include only the data which are sufficiently related to the current subject. The method was tested on the TDT1 and TDT2, and the results show the effectiveness of the method. Fumiyo Fukumoto, Yoshimi Suzuki |
INTERSPEECH | 1 |
| 2004 | Clustering similar nouns for selecting related news articlesabstractIn both written language and spoken language, we sometimes use different words in order to express the same meaning. For instance, we use “candidacy” and “running in an election” as the same meaning. This makes text classification and event tracking difficult. To do this, we have to identify the words which are semantically similar to each other accurately. In this paper, we propose a method to extract the words which are semantically similar to other words. Using the method, we extracted similar word pairs on newspaper articles. Further, we extracted news articles which are related to a news article. By using a similar word pair list, we obtained better result than that without it. The results suggest that the method is useful for extracting related document, text tracking and so on. Yoshimi Suzuki, Fumiyo Fukumoto, Yoshihiro Sekiguchi |
INTERSPEECH | 2 |
| 2002 | Detecting Shifts in News Stories for Paragraph Extraction
Fumiyo Fukumoto, Yoshimi Suzuki |
COLING | 1 |
| 2002 | Topic Tracking using Subject Templates and Clustering Positive Training Instances
Yoshimi Suzuki, Fumiyo Fukumoto, Yoshihiro Sekiguchi |
COLING | 2 |
| 2002 | Manipulating Large Corpora for Text ClassificationabstractIn this paper, we address the problem of dealing with a large collection of data and propose a method for text classification which manipulates data using two well-known machine learning techniques, Naive Bayes(NB) and Support Vector Machines(SVMs).NB is based on the assumption of word independence in a text, which makes the computation of it far more efficient.SVMs, on the other hand, have the potential to handle large feature spaces, which makes it possible to produce better performance.The training data for SVMs are extracted using NB classifiers according to the category hierarchies, which makes it possible to reduce the amount of computation necessary for classification without sacrificing accuracy. Fumiyo Fukumoto, Yoshimi Suzuki |
EMNLP | 1 |
| 2002 | Topic tracking using subject templates
Yoshimi Suzuki, Fumiyo Fukumoto, Yoshihiro Sekiguchi |
INTERSPEECH | 2 |
| 2000 | Selecting TV news stories and newswire articles related to a target article of newswire using SVM
Yoshimi Suzuki, Fumiyo Fukumoto, Yoshihiro Sekiguchi |
INTERSPEECH | 2 |
| 2000 | Event tracking based on domain dependencyabstractThis paper proposes a method for event tracking on broadcast news stories based on distinction between a topic and an event. A topic and an event are identified using a simple criterion called domain dependency of words: how greatly a word features a given set of data. The method was tested on the TDT corpus which has been developed by the TDT Pilot Study and the result can be regarded as promising the usefulness of the method. Fumiyo Fukumoto, Yoshimi Suzuki |
SIGIR | 1 |
| 1999 | Word Sense Disambiguation in Untagged Text based on Term Weight Learning
Fumiyo Fukumoto, Yoshimi Suzuki |
EACL | 1 |
| 1998 | An Empirical Approach to Text Categorization Based on Term Weight Learning
Fumiyo Fukumoto, Yoshimi Suzuki |
EMNLP | 1 |
| 1998 | Keyword extraction of radio news using domain identification based on categories of an encyclopedia
Yoshimi Suzuki, Fumiyo Fukumoto, Yoshihiro Sekiguchi |
ICSLP | 2 |
| 1998 | Keyword Extraction of Radio News Using Term Weighting with an Encyclopedia and Newspaper Articles
Yoshimi Suzuki, Fumiyo Fukumoto, Yoshihiro Sekiguchi |
SIGIR | 2 |
| 1996 | An Automatic Clustering of Articles Using Dictionary Definitions
Fumiyo Fukumoto, Yoshimi Suzuki |
COLING | 1 |
| 1994 | Automatic Recognition of Verbal Polysemy
Fumiyo Fukumoto, Jun'ichi Tsujii |
COLING | 1 |