EDBT 2026 Demo / reviewers in the wild / expert
Chikashi Nobata
dblp:36/883
· DBLP profile ↗
16ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 69% Data mining · 11% Query processing and optimization · 10% | |
| Artificial intelligence
3 papers |
Information extraction and text analysis · 94% Question answering and dialogue systems · 4% Speech recognition and synthesis · 2% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 56% Geometric modeling and processing · 44% |
Topics — the 20 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
abusive language detection |
0.2 | 1 | 2016 | Abusive Language Detection in Online User Content · WWW 2016 |
Natural language and speech › Information extraction and text analysis › abusive language detection
hate speech detection |
0.2 | 1 | 2016 | Abusive Language Detection in Online User Content · WWW 2016 |
Data mining
anomaly detection |
0.2 | 1 | 2016 | Abusive Language Detection in Online User Content · WWW 2016 |
Web and social media mining
content moderation |
0.2 | 1 | 2016 | Abusive Language Detection in Online User Content · WWW 2016 |
Information retrieval › ranking › context-aware ranking
location-based ranking |
0.2 | 1 | 2016 | Ranking Relevance in Yahoo Search · KDD 2016 |
Query processing and optimization
query rewriting |
0.2 | 1 | 2016 | Ranking Relevance in Yahoo Search · KDD 2016 |
Information retrieval
ranking |
0.2 | 1 | 2016 | Ranking Relevance in Yahoo Search · KDD 2016 |
Information retrieval › ranking › context-aware ranking
recency ranking |
0.2 | 1 | 2016 | Ranking Relevance in Yahoo Search · KDD 2016 |
Information retrieval › ranking › search ranking
relevance ranking |
0.2 | 1 | 2016 | Ranking Relevance in Yahoo Search · KDD 2016 |
Information retrieval
retrieval models |
0.2 | 1 | 2016 | Ranking Relevance in Yahoo Search · KDD 2016 |
Information retrieval
search engines |
0.2 | 1 | 2016 | Ranking Relevance in Yahoo Search · KDD 2016 |
Information retrieval › document retrieval › domain-specific retrieval
biomedical information retrieval |
0.1 | 1 | 2008 | Kleio: a knowledge-enriched information retrieval system for biology · SIGIR 2008 |
Information retrieval › document retrieval
domain-specific retrieval |
0.1 | 1 | 2008 | Kleio: a knowledge-enriched information retrieval system for biology · SIGIR 2008 |
Geometric modeling and processing › shape analysis
morphological analysis |
0.0 | 1 | 2004 | Morphological analysis of the corpus of spontaneous Japanese · IEEE Trans. Speech Audio Process. 2004 |
Audio and music processing
speech corpus |
0.0 | 1 | 2004 | Morphological analysis of the corpus of spontaneous Japanese · IEEE Trans. Speech Audio Process. 2004 |
Natural language and speech › Information extraction and text analysis
morphological analysis |
0.0 | 1 | 2003 | Morphological Analysis of a Large Spontaneous Speech Corpus in Japanese · ACL 2003 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.0 | 1 | 2000 | Difficulty Indices for the Named Entity Task in Japanese · ACL 2000 |
Natural language and speech › Question answering and dialogue systems
question difficulty estimation |
0.0 | 1 | 2000 | Difficulty Indices for the Named Entity Task in Japanese · ACL 2000 |
Information retrieval › document retrieval
metadata-based retrieval |
0.0 | 1 | 2008 | Kleio: a knowledge-enriched information retrieval system for biology · SIGIR 2008 |
Data mining
text mining |
0.0 | 1 | 2008 | Kleio: a knowledge-enriched information retrieval system for biology · SIGIR 2008 |
Methods — techniques the papers use, named apart from their topics
supervised machine learning · 0.5corpus annotation · 0.5semantic matching · 0.2query rewriting · 0.2semi-automatic analysis · 0.1terminology management · 0.1corpus-based index · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Ranking Relevance in Yahoo SearchabstractSearch engines play a crucial role in our daily lives. Relevance is the core problem of a commercial search engine. It has attracted thousands of researchers from both academia and industry and has been studied for decades. Relevance in a modern search engine has gone far beyond text matching, and now involves tremendous challenges. The semantic gap between queries and URLs is the main barrier for improving base relevance. Clicks help provide hints to improve relevance, but unfortunately for most tail queries, the click information is too sparse, noisy, or missing entirely. For comprehensive relevance, the recency and location sensitivity of results is also critical. In this paper, we give an overview of the solutions for relevance in the Yahoo search engine. We introduce three key techniques for base relevance -- ranking functions, semantic matching features and query rewriting. We also describe solutions for recency sensitive relevance and location sensitive relevance. This work builds upon 20 years of existing efforts on Yahoo search, summarizes the most recent advances and provides a series of practical relevance solutions. The performance reported is based on Yahoo's commercial search engine, where tens of billions of urls are indexed and served by the ranking system. Dawei Yin 0001, Yuening Hu, Jiliang Tang, Tim Daly Jr., Mianwei Zhou, Hua Ouyang, Changsung Kang, Hongbo Deng, Chikashi Nobata, Jean-Marc Langlois, Yi Chang 0001 |
KDD | 10 |
| 2016 | Abusive Language Detection in Online User ContentabstractDetection of abusive language in user generated online content has become an issue of increasing importance in recent years. Most current commercial methods make use of blacklists and regular expressions, however these measures fall short when contending with more subtle, less ham-fisted examples of hate speech. In this work, we develop a machine learning based method to detect hate speech on online user comments from two domains which outperforms a state-of-the-art deep learning approach. We also develop a corpus of user comments annotated for abusive language, the first of its kind. Finally, we use our detection tool to analyze abusive language over time and in different settings to further enhance our knowledge of this behavior. Chikashi Nobata, Joel R. Tetreault, Achint Oommen Thomas, Yashar Mehdad, Yi Chang 0001 |
WWW | 1 |
| 2015 | Propagation-based Sentiment Analysis for Microblogging DataabstractThe explosive popularity of microblogging services encourages more and more online users to share their opinions, and sentiment analysis on such opinion-rich resources has been proven to be an effective way to understand public opinions. On the one hand, the brevity and informality of microblogging data plus its wide variety and rapid evolution of language in microblogging pose new challenges to the vast majority of existing methods. On the other hand, microblogging texts contain various types of emotional signals strongly associated with their sentiment polarity, which brings about new opportunities for sentiment analysis. In this paper, we investigate propagation-based sentiment analysis for microblogging data. In particular, we provide a propagating process to incorporate various types of emotional signals in microblogging data into a coherent model, and propose a novel sentiment analysis framework PSA which learns from both labeled and unlabeled data by iteratively alternating a propagating process and a fitting process. We conduct experiments on real-world microblogging datasets, and the results demonstrate the effectiveness of the proposed framework. Further experiments are conducted to probe the working of the key components of the proposed framework. Jiliang Tang, Chikashi Nobata, Anlei Dong, Yi Chang 0001, Huan Liu 0001 |
SDM | 2 |
| 2011 | Detecting experimental techniques and selecting relevant documents for protein-protein interactions from biomedical literatureabstractBACKGROUND: The selection of relevant articles for curation, and linking those articles to experimental techniques confirming the findings became one of the primary subjects of the recent BioCreative III contest. The contest's Protein-Protein Interaction (PPI) task consisted of two sub-tasks: Article Classification Task (ACT) and Interaction Method Task (IMT). ACT aimed to automatically select relevant documents for PPI curation, whereas the goal of IMT was to recognise the methods used in experiments for identifying the interactions in full-text articles. RESULTS: We proposed and compared several classification-based methods for both tasks, employing rich contextual features as well as features extracted from external knowledge sources. For IMT, a new method that classifies pair-wise relations between every text phrase and candidate interaction method obtained promising results with an F1 score of 64.49%, as tested on the task's development dataset. We also explored ways to combine this new approach and more conventional, multi-label document classification methods. For ACT, our classifiers exploited automatically detected named entities and other linguistic information. The evaluation results on the BioCreative III PPI test datasets showed that our systems were very competitive: one of our IMT methods yielded the best performance among all participants, as measured by F1 score, Matthew's Correlation Coefficient and AUC iP/R; whereas for ACT, our best classifier was ranked second as measured by AUC iP/R, and also competitive according to other metrics. CONCLUSIONS: Our novel approach that converts the multi-class, multi-label classification problem to a binary classification problem showed much promise in IMT. Nevertheless, on the test dataset the best performance was achieved by taking the union of the output of this method and that of a multi-class, multi-label document classifier, which indicates that the two types of systems complement each other in terms of recall. For ACT, our system exploited a rich set of features and also obtained encouraging results. We examined the features with respect to their contributions to the classification results, and concluded that contextual words surrounding named entities, as well as the MeSH headings associated with the documents were among the main contributors to the performance. Xinglong Wang, Rafal Rak, Angelo C. Restificar, Chikashi Nobata, C. J. Rupp, Riza Theresa Batista-Navarro, Raheel Nawaz, Sophia Ananiadou |
BMC Bioinform. | 4 |
| 2008 | Kleio: a knowledge-enriched information retrieval system for biologyabstractKleio is an advanced information retrieval (IR) system developed at the UK National Centre for Text Mining (NaCTeM)1. The system offers textual and metadata searches across MEDLINE and provides enhanced searching functionality by leveraging terminology management technologies. Chikashi Nobata, Philip Cotter, Naoaki Okazaki, Brian Rea, Yutaka Sasaki, Yoshimasa Tsuruoka, Jun'ichi Tsujii, Sophia Ananiadou |
SIGIR | 1 |
| 2004 | Corpus and Evaluation Measures for Multiple Document Summarization with Multiple Sources
Tsutomu Hirao, Takahiro Fukusima, Manabu Okumura, Chikashi Nobata, Hidetsugu Nanba |
COLING | 4 |
| 2004 | Definition, Dictionaries and Tagger for Extended Named Entity Hierarchy
Satoshi Sekine, Chikashi Nobata |
LREC | 2 |
| 2004 | Morphological analysis of the corpus of spontaneous JapaneseabstractThis paper describes two methods for detecting word segments and their morphological information in a Japanese spontaneous speech corpus, and describes how to tag a large spontaneous speech corpus accurately by using the two methods. The first method is used to detect any type of word segments. The second method is used when there are several definitions for word segments and their POS categories, and when one type of word segments includes another type of word segments. In this paper, we show that by using semi-automatic analysis, we achieve a precision of better than 99% for detecting and tagging short-unit words and 97% for long-unit words; the two types of words that comprise the corpus. We also show that better accuracy is achieved by using both methods than by using only the first. Kiyotaka Uchimoto, Kazuma Takaoka, Chikashi Nobata, Atsushi Yamada, Satoshi Sekine, Hitoshi Isahara |
IEEE Trans. Speech Audio Process. | 3 |
| 2003 | Morphological Analysis of a Large Spontaneous Speech Corpus in JapaneseabstractThis paper describes two methods for detecting word segments and their morphological information in a Japanese spontaneous speech corpus, and describes how to tag a large spontaneous speech corpus accurately by using the two methods. The first method is used to detect any type of word segments. The second method is used when there are several definitions for word segments and their POS categories, and when one type of word segments includes another type of word segments. In this paper, we show that by using semi-automatic analysis we achieve a precision of better than 99% for detecting and tagging short words and 97% for long words; the two types of words that comprise the corpus. We also show that better accuracy is achieved by using both methods than by using only the first. Kiyotaka Uchimoto, Chikashi Nobata, Atsushi Yamada, Satoshi Sekine, Hitoshi Isahara |
ACL | 2 |
| 2002 | Morphological Analysis of the Spontaneous Speech Corpus
Kiyotaka Uchimoto, Chikashi Nobata, Atsushi Yamada, Satoshi Sekine, Hitoshi Isahara |
COLING | 2 |
| 2002 | Progress on Multi-lingual Named Entity Annotation Guidelines using RDF (S)
Nigel Collier, Koichi Takeuchi, Chikashi Nobata, Jun-ichi Fukumoto, Norihiro Ogata |
LREC | 3 |
| 2002 | Summarization System Integrated with Named Entity Tagging and IE pattern Discovery
Chikashi Nobata, Satoshi Sekine, Hitoshi Isahara, Ralph Grishman |
LREC | 1 |
| 2002 | Extended Named Entity Hierarchy
Satoshi Sekine, Kiyoshi Sudo, Chikashi Nobata |
LREC | 3 |
| 2000 | Difficulty Indices for the Named Entity Task in JapaneseabstractWe propose indices to measure the difficulty of the named entity (NE) task by looking at test corpora, based on expressions inside and outside the NEs. These indices are intended to estimate the difficulty of each task without actually using an NE system and to be unbiased towards a specific system. The values of the indices are compared with the systems' performance in Japanese documents. We also discuss the difference between NE classes with the indices and show useful clues which will make it easier to recognize NEs. Chikashi Nobata, Satoshi Sekine, Jun'ichi Tsujii |
ACL | 1 |
| 2000 | Extracting the Names of Genes and Gene Products with a Hidden Markov Model
Nigel Collier, Chikashi Nobata, Jun'ichi Tsujii |
COLING | 2 |
| 1999 | The GENIA project: corpus-based knowledge acquisition and information extraction from genome research papers
Nigel Collier, Hyun Seok Park, Norihiro Ogata, Yuka Tateisi, Chikashi Nobata, Tomoko Ohta, Tateshi Sekimizu, Hisao Imai, Katsutoshi Ibushi, Jun'ichi Tsujii |
EACL | 5 |