Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Daichi Mochihashi

dblp:86/2846 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-0344-5382ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Information extraction and text analysis · 65% Probabilistic and Bayesian machine learning · 11% Language models and text generation · 8%
Databases, data mining, and information retrieval
2 papers
Data mining · 93% Machine learning and data management · 7%

Topics — the 19 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › syntactic parsing › grammar-based parsing
combinatory categorial grammar parsing
0.712023
Holographic CCG Parsing · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.712023
Holographic CCG Parsing · ACL (1) 2023
Natural language and speech › Information extraction and text analysis › topic model
dynamic topic models
0.612022
Infinite SCAN: An Infinite Model of Diachronic Semantic Change · EMNLP 2022
Natural language and speech › Information extraction and text analysis › lexical semantics
semantic change detection
0.612022
Infinite SCAN: An Infinite Model of Diachronic Semantic Change · EMNLP 2022
Natural language and speech › Information extraction and text analysis
word segmentation
0.322015
Inducing Word and Part-of-Speech with Pitman-Yor Hidden Semi-Markov Models · ACL (1) 2015
Bayesian Unsupervised Word Segmentation with Nested Pitman-Yor Language Modeling · ACL/IJCNLP 2009
Data mining
dependence maximization
0.312017
Learning Co-Substructures by Kernel Dependence Maximization · IJCAI 2017
Data mining
pattern mining
0.312017
Learning Co-Substructures by Kernel Dependence Maximization · IJCAI 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.332013
Infinite Positive Semidefinite Tensor Factorization for Source Separation of Mixture Signals · ICML (3) 2013
The Infinite Markov Model · NIPS 2007
Bayesian Unsupervised Word Segmentation with Nested Pitman-Yor Language Modeling · ACL/IJCNLP 2009
Computer vision › Vision and language
grounded language learning
0.212015
Learning Word Meanings and Grammar for Describing Everyday Activities in Smart Environments · EMNLP 2015
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging
0.212015
Inducing Word and Part-of-Speech with Pitman-Yor Hidden Semi-Markov Models · ACL (1) 2015
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
gamma process
0.212013
Infinite Positive Semidefinite Tensor Factorization for Source Separation of Mixture Signals · ICML (3) 2013
Natural language and speech › Language models and text generation › language modeling
statistical language modeling
0.212013
Improvements to the Bayesian Topic N-Gram Models · EMNLP 2013
Audio and music processing
source separation
0.212013
Infinite Positive Semidefinite Tensor Factorization for Source Separation of Mixture Signals · ICML (3) 2013
Machine learning › Deep learning architectures and training
sequence modeling
0.112007
The Infinite Markov Model · NIPS 2007
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › markov processes
variable-order markov model
0.112007
The Infinite Markov Model · NIPS 2007
Machine learning › Generative modeling
generative model
0.112005
Context as Filtering · NIPS 2005
Natural language and speech › Language models and text generation
language modeling
0.112005
Context as Filtering · NIPS 2005
Data mining
clustering
0.012004
Learning Nonstructural Distance Metric by Minimum Cluster Distortion · EMNLP 2004
Machine learning and data management
metric learning
0.012004
Learning Nonstructural Distance Metric by Minimum Cluster Distortion · EMNLP 2004

Methods — techniques the papers use, named apart from their topics

bayesian modeling · 0.7span-based parsing · 0.7holographic embeddings · 0.7metropolis-hastings algorithm · 0.6logistic stick-breaking process · 0.6kernel dependence maximization · 0.6hilbert-schmidt independence criterion · 0.6gaussian markov random field · 0.6dirichlet process · 0.6pitman-yor hidden semi-markov model · 0.2variational bayesian inference · 0.2positive semidefinite tensor factorization · 0.2multiplicative update · 0.2minimum cluster distortion · 0.0
YearPublicationVenuePosition
2025 Analyzing Continuous Semantic Shifts with Diachronic Word Similarity Matrices
abstract
The meanings and relationships of words shift over time. This phenomenon is referred to as semantic shift. Research focused on understanding how semantic shifts occur over multiple time periods is essential for gaining a detailed understanding of semantic shifts. However, detecting change points only between adjacent time periods is insufficient for analyzing detailed semantic shifts, and using BERT-based methods to examine word sense proportions incurs a high computational cost. To address those issues, we propose a simple yet intuitive framework for how semantic shifts occur over multiple time periods by utilizing similarity matrices based on word embeddings. We calculate diachronic word similarity matrices using fast and lightweight word embeddings across arbitrary time periods, making it deeper to analyze continuous semantic shifts. Additionally, by clustering the resulting similarity matrices, we can categorize words that exhibit similar behavior of semantic shift in an unsupervised manner.
Hajime Kiyama, Taichi Aida, Mamoru Komachi, Toshinobu Ogiso, Hiroya Takamura, Daichi Mochihashi
COLING6
2025 Scalable Unsupervised Segmentation via Random Fourier Feature-based Gaussian Process
abstract
In this paper, we propose RFF-GP-HSMM, a fast unsupervised time-series segmentation method that incorporates random Fourier features (RFF) to address the high computational cost of the Gaussian process hidden semi-Markov model (GP-HSMM). GP-HSMM models time-series data using Gaussian processes, requiring inversion of an N × N kernel matrix during training, where N is the number of data points. As the scale of the data increases, matrix inversion incurs a significant computational cost. To address this, the proposed method approximates the Gaussian process with linear regression using RFF, preserving expressive power while eliminating the need for inversion of the kernel matrix. Experiments on the Carnegie Mellon University (CMU) motion-capture dataset demonstrate that the proposed method achieves segmentation performance comparable to that of conventional methods, with approximately 278 times faster segmentation on time-series data comprising 39,200 frames.
Issei Saito, Masatoshi Nagano, Tomoaki Nakamura, Daichi Mochihashi, Koki Mimura
IECON4
2023 Holographic CCG Parsing
abstract
We propose a method for formulating CCG as a recursive composition in a continuous vector space.Recent CCG supertagging and parsing models generally demonstrate high performance, yet rely on black-box neural architectures to implicitly model phrase structure dependencies.Instead, we leverage the method of holographic embeddings (Nickel et al., 2016) as a compositional operator to explicitly model the dependencies between words and phrase structures in the embedding space.Experimental results revealed that holographic composition effectively improves the supertagging accuracy to achieve state-of-the-art parsing performance when using a C&C parser.The proposed span-based parsing algorithm using holographic composition achieves performance comparable to state-of-the-art neural parsing with Transformers.Furthermore, our model can semantically and syntactically infill text at the phrase level due to the decomposability of holographic composition.
Ryosuke Yamaki, Tadahiro Taniguchi, Daichi Mochihashi
ACL (1)3
2023 Investigation of Information Processing Mechanisms in the Human Brain During Reading Tanka Poetry
Anna Sato, Junichi Chikazoe, Shotaro Funai, Daichi Mochihashi, Yutaka Shikano, Masayuki Asahara, Satoshi Iso, Ichiro Kobayashi 0001
ICANN (8)4
2022 Infinite SCAN: An Infinite Model of Diachronic Semantic Change
abstract
In this study, we propose a Bayesian model that can jointly estimate the number of senses of words and their changes through time.The model combines a dynamic topic model on Gaussian Markov random fields (Frermann and Lapata, 2016) with a logistic stick-breaking process that realizes the Dirichlet process.In the experiments, we evaluated the proposed model in terms of interpretability, accuracy in estimating the number of senses, and tracking their changes using both artificial data and real data.We quantitatively verified that the model behaves as expected through evaluation using artificial data.Using the CCOHA corpus, we showed that our model outperforms the baseline model and investigated the semantic changes of several well-known target words.
Seiichi Inoue, Mamoru Komachi, Toshinobu Ogiso, Hiroya Takamura, Daichi Mochihashi
EMNLP5
2022 Nonparametric Bayesian Deep Visualization
Haruya Ishizuka, Daichi Mochihashi
ECML/PKDD (1)2
2021 A Comprehensive Analysis of PMI-based Models for Measuring Semantic Differences
Taichi Aida, Mamoru Komachi, Toshinobu Ogiso, Hiroya Takamura, Daichi Mochihashi
PACLIC5
2020 How LSTM Encodes Syntax: Exploring Context Vectors and Semi-Quantization on Natural Text
abstract
Long Short-Term Memory recurrent neural network (LSTM) is widely used and known to capture informative long-term syntactic dependencies.However, how such information are reflected in its internal vectors for natural text has not yet been sufficiently investigated.We analyze them by learning a language model where syntactic structures are implicitly given.We empirically show that the context update vectors, i.e. outputs of internal gates, are approximately quantized to binary or ternary values to help the language model to count the depth of nesting accurately, as Suzgun et al. (2019) recently showed for synthetic Dyck languages.For some dimensions in the context vector, we show that their activations are highly correlated with the depth of phrase structures, such as VP and NP.Moreover, with an L 1 regularization, we also found that it can be accurately predicted whether a word is inside a phrase structure or not from a small number of components of the context vector.Even for the case of learning from raw text, context vectors are still shown to correlate well with the phrase structures.Finally, we show that natural clusters of the functional words and the parts of speech that trigger phrases are represented in a small but principal subspace of the context-update vector of LSTM.
Chihiro Shibata, Kei Uchiumi, Daichi Mochihashi
COLING3
2019 High-dimensional Motion Segmentation by Variational Autoencoder and Gaussian Processes
abstract
Humans perceive continuous high-dimensional information by dividing it into significant segments such as words and units of motion. We believe that such unsupervised segmentation is also important for robots to learn topics such as language and motion. To this end, we previously proposed a hierarchical Dirichlet process-Gaussian process-hidden semi-Markov model (HDP-GP-HSMM). However, an important drawback to this model is that it cannot divide high-dimensional time-series data. Further, low-dimensional features must be extracted in advance. Segmentation largely depends on the design of features, and it is difficult to design effective features, especially in the case of high-dimensional data. To overcome this problem, this paper proposes a hierarchical Dirichlet process-variational autoencoder-Gaussian process-hidden semi-Markov model (HVGH). The parameters of the proposed HVGH are estimated through a mutual learning loop of the variational autoencoder and our previously proposed HDP-GP-HSMM. Hence, HVGH can extract features from high-dimensional time-series data, while simultaneously dividing it into segments in an unsupervised manner. In an experiment, we used various motion-capture data to show that our proposed model estimates the correct number of classes and more accurate segments than baseline methods. Moreover, we show that the proposed method can learn latent space suitable for segmentation.
Masatoshi Nagano, Tomoaki Nakamura, Takayuki Nagai, Daichi Mochihashi, Ichiro Kobayashi 0001, Wataru Takano
IROS4
2018 Sequence Pattern Extraction by Segmenting Time Series Data Using GP-HSMM with Hierarchical Dirichlet Process
abstract
Humans recognize perceived continuous information by dividing it into significant segments such as words and unit motions. We believe that such unsupervised segmentation is also an important ability that robots need to learn topics such as language and motions. Hence, in this paper, we propose a method for dividing continuous time-series data into segments in an unsupervised manner. To this end, we proposed a method based on a hidden semi-Markov model with Gaussian process (GP-HSMM). If Gaussian processes, which are nonparametric models, are used, unit motion patterns can be extracted from complicated continuous motion. However, this approach requires the number of classes of segments in the time-series data in advance. To overcome this problem, in this paper, we extend GP-HSMM to a nonparametric Bayesian model by introducing a hierarchical Dirichlet process (HDP) and propose the hierarchical Dirichlet processes-Gaussian process-hidden semi-Markov model (HDP-GP-HSMM). In the nonparametric Bayesian model, an infinite number of classes is assumed and it becomes difficult to estimate the parameters naively. Instead, the parameters of the proposed HDP-GP-HSMM are estimated by applying slice sampling. In the experiments, we use various synthetic and motion-capture data to show that our proposed model can estimate a more correct number of classes and achieve more accurate segmentation than baseline methods.
Masatoshi Nagano, Tomoaki Nakamura, Takayuki Nagai, Daichi Mochihashi, Ichiro Kobayashi 0001, Masahide Kaneko
IROS4
2017 Learning Co-Substructures by Kernel Dependence Maximization
abstract
Modeling associations between items in a dataset is a problem that is frequently encountered in data and knowledge mining research. Most previous studies have simply applied a predefined fixed pattern for extracting the substructure of each item pair and then analyzed the associations between these substructures. Using such fixed patterns may not, however, capture the significant association. We, therefore, propose the novel machine learning task of extracting a strongly associated substructure pair (co-substructure) from each input item pair. We call this task dependent co-substructure extraction (DCSE), and formalize it as a dependence maximization problem. Then, we discuss critical issues with this task: the data sparsity problem and a huge search space. To address the data sparsity problem, we adopt the Hilbert--Schmidt independence criterion as an objective function. To improve search efficiency, we adopt the Metropolis--Hastings algorithm. We report the results of empirical evaluations, in which the proposed method is applied for acquiring and predicting narrative event pairs, an active task in the field of natural language processing.
Sho Yokoi, Daichi Mochihashi, Naoaki Okazaki, Kentaro Inui
IJCAI2
2017 MIPA: Mutual Information Based Paraphrase Acquisition via Bilingual Pivoting
abstract
We present a pointwise mutual information (PMI)-based approach to formalize paraphrasability and propose a variant of PMI, called MIPA, for the paraphrase acquisition. Our paraphrase acquisition method first acquires lexical paraphrase pairs by bilingual pivoting and then reranks them by PMI and distributional similarity. The complementary nature of information from bilingual corpora and from monolingual corpora makes the proposed method robust. Experimental results show that the proposed method substantially outperforms bilingual pivoting and distributional similarity themselves in terms of metrics such as MRR, MAP, coverage, and Spearman’s correlation.
Tomoyuki Kajiwara, Mamoru Komachi, Daichi Mochihashi
IJCNLP(1)3
2017 Semi-Supervised Learning of a Pronunciation Dictionary from Disjoint Phonemic Transcripts and Text
Takahiro Shinozaki, Shinji Watanabe 0001, Daichi Mochihashi, Graham Neubig
INTERSPEECH3
2017 Nonparametric Bayesian Semi-supervised Word Segmentation
abstract
This paper presents a novel hybrid generative/discriminative model of word segmentation based on nonparametric Bayesian methods. Unlike ordinary discriminative word segmentation which relies only on labeled data, our semi-supervised model also leverages a huge amounts of unlabeled text to automatically learn new “words”, and further constrains them by using a labeled data to segment non-standard texts such as those found in social networking services. Specifically, our hybrid model combines a discriminative classifier (CRF; Lafferty et al. (2001) and unsupervised word segmentation (NPYLM; Mochihashi et al. (2009)), with a transparent exchange of information between these two model structures within the semi-supervised framework (JESS-CM; Suzuki and Isozaki (2008)). We confirmed that it can appropriately segment non-standard texts like those in Twitter and Weibo and has nearly state-of-the-art accuracy on standard datasets in Japanese, Chinese, and Thai.
Ryo Fujii, Ryo Domoto, Daichi Mochihashi
Trans. Assoc. Comput. Linguistics3
2015 Inducing Word and Part-of-Speech with Pitman-Yor Hidden Semi-Markov Models
abstract
Kei Uchiumi, Hiroshi Tsukahara, Daichi Mochihashi. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Kei Uchiumi, Hiroshi Tsukahara, Daichi Mochihashi
ACL (1)3
2015 Learning Word Meanings and Grammar for Describing Everyday Activities in Smart Environments
abstract
Muhammad Attamimi, Yuji Ando, Tomoaki Nakamura, Takayuki Nagai, Daichi Mochihashi, Ichiro Kobayashi, Hideki Asoh. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015.
Muhammad Attamimi, Yuji Ando, Tomoaki Nakamura, Takayuki Nagai, Daichi Mochihashi, Ichiro Kobayashi 0001, Hideki Asoh
EMNLP5
2014 Mixture of Gaussian process experts for predicting sung melodic contour with expressive dynamic fluctuations
abstract
We present a generative model for predicting the sung melodic contour, i.e., F0contour, with expressive dynamic fluctuations, such as vibrato and portamento, for a given musical score. Although several studies have attempted to characterize such fluctuations, no systematic method has been developed for generating the F0contour with them in connection with musical notes. In our model, the relationship between a musical note sequence and F0contour is directly learned by a mixture of Gaussian process experts. This approach allows us to automatically characterize the fluctuations by utilizing the kernel function for each Gaussian process expert and predict the F0contour for an arbitrary musical note sequence. Experimental results show that our model can better predict the F0contour than a baseline method can. Additionally, we discuss the effective musical contexts and the amount of training data for the prediction.
Yasunori Ohishi, Daichi Mochihashi, Hirokazu Kameoka, Kunio Kashino
ICASSP2
2013 Improvements to the Bayesian Topic N-Gram Models
abstract
One of the language phenomena that n-gram language model fails to capture is the topic information of a given situation.We advance the previous study of the Bayesian topic language model by Wallach (2006) in two directions: one, investigating new priors to alleviate the sparseness problem caused by dividing all ngrams into exclusive topics, and two, developing a novel Gibbs sampler that enables moving multiple n-grams across different documents to another topic.Our blocked sampler can efficiently search for higher probability space even with higher order n-grams.In terms of modeling assumption, we found it is effective to assign a topic to only some parts of a document.
Hiroshi Noji, Daichi Mochihashi, Yusuke Miyao
EMNLP2
2013 Bayesian semi-supervised audio event transcription based on Markov indian buffet process
abstract
We present a novel generative model for audio event transcription that recognizes “events” on audio signals including multiple kinds of overlapping sounds. In the proposed model, firstly, the overlapping audio events are modeled based on nonnegative matrix factorization into which Bayesian nonparametric approaches: the Markov Indian buffet process and the Chinese restaurant process, are incorporated. This approach allows us to automatically transcribe the events while avoiding the model selection problem by assuming a countably infinite number of possible audio events in the input signal. Then, Bayesian logistic regression annotates the audio frames with the multiple event labels in a semi-supervised learning setup. Experimental results show that our model can better annotate an audio signal in comparison with a baseline method. Additionally, we verify that our infinite generative model is also able to detect unknown audio events that are not included in the training data.
Yasunori Ohishi, Daichi Mochihashi, Tomoko Matsui, Masahiro Nakano, Hirokazu Kameoka, Tomonori Izumitani, Kunio Kashino
ICASSP2
2013 Infinite Positive Semidefinite Tensor Factorization for Source Separation of Mixture Signals
abstract
This paper presents a new class of tensor factorization called positive semidefinite tensor factorization (PSDTF) that decomposes a set of positive semidefinite (PSD) matrices into the convex combinations of fewer PSD basis matrices. PSDTF can be viewed as a natural extension of nonnegative matrix factorization. One of the main problems of PSDTF is that an appropriate number of bases should be given in advance. To solve this problem, we propose a nonparametric Bayesian model based on a gamma process that can instantiate only a limited number of necessary bases from the infinitely many bases assumed to exist. We derive a variational Bayesian algorithm for closed-form posterior inference and a multiplicative update rule for maximum-likelihood estimation. We evaluated PSDTF on both synthetic data and real music recordings to show its superiority.
Kazuyoshi Yoshii, Ryota Tomioka, Daichi Mochihashi, Masataka Goto
ICML (3)3
2012 A Stochastic Model of Singing Voice F0 Contours for Characterizing Expressive Dynamic Components
abstract
We present a novel stochastic model of singing voice fundamental frequency (F0) contours for characterizing expressive dynamic components, such as vibrato and portamento. Although dynamic components can be important features for any singing voice applications, modeling and extracting these components from a raw F0 contour have yet to be accomplished. Therefore, we describe a process for generating dynamic components explicitly and represent the process as a stochastic model. Then we develop an algorithm for estimating the model parameters based on statistical techniques. Experimental results show that our method successfully extracts the expressive components from raw F0 contours.
Yasunori Ohishi, Hirokazu Kameoka, Daichi Mochihashi, Kunio Kashino
INTERSPEECH3
2011 Gibbs sampling based Multi-scale Mixture Model for speaker clustering
abstract
The aim of this work is to apply a sampling approach to speech modeling, and propose a Gibbs sampling based Multi-scale Mixture Model (M3). The proposed approach focuses on the multi-scale property of speech dynamics, i.e., dynamics in speech can be observed on, for instance, short-time acoustical, linguistic-segmental, and utterance-wise temporal scales. M3is an extension of the Gaussian mixture model and is considered a hierarchical mixture model, where mixture components in each time scale will change at intervals of the corresponding time unit. We derive a fully Bayesian treatment of the multi-scale mixture model based on Gibbs sampling. The advantage of the proposed model is that each speaker cluster can be precisely modeled based on the Gaussian mixture model unlike conventional single-Gaussian based speaker clustering (e.g., using the Bayesian Information Criterion (BIC)). In addition, Gibbs sampling offers the potential to avoid a serious local optimum problem. Speaker clustering experiments confirmed these advantages and obtained a significant improvement over the conventional BIC based approaches.
Shinji Watanabe 0001, Daichi Mochihashi, Takaaki Hori, Atsushi Nakamura
ICASSP2
2010 Statistical modeling of F0 dynamics in singing voices based on Gaussian processes with multiple oscillation bases
abstract
We present a novel statistical model for dynamics of various singing behaviors, such as vibrato and overshoot, in a fundamental frequency (F0) contour. These dynamics are the important cues for perceiving individuality of a singer, and can be a useful measure for various applications, such as singing skill evaluation and singing voice synthesis. While most previous studies have modeled the dynamics using a second-order linear system, the automatic and accurate estimation of model parameters has yet to be accomplished. In this paper, we first develop a complete stochastic representation of the second-order system with Gaussian processes from parametric discretization, and propose a complete, efficient scheme for parameter estimation using the Expectation-Maximization (EM) algorithm. Experimental results show that the proposed method can decompose an F0 contour into a musical component and a dynamics component. Finally, we discuss estimating singing styles from the model parameters for each singer.
Yasunori Ohishi, Hirokazu Kameoka, Daichi Mochihashi, Hidehisa Nagano, Kunio Kashino
INTERSPEECH3
2009 Bayesian Unsupervised Word Segmentation with Nested Pitman-Yor Language Modeling
Daichi Mochihashi, Takeshi Yamada, Naonori Ueda
ACL/IJCNLP1
2009 On the properties of von Neumann kernels for link analysis
Masashi Shimbo, Takahiko Ito, Daichi Mochihashi, Yuji Matsumoto 0001
Mach. Learn.3
2007 The Infinite Markov Model
abstract
We present a nonparametric Bayesian method of estimating variable order Markov processes up to a theoretically infinite order. By extending a stick-breaking prior, which is usually defined on a unit interval, “vertically” to the trees of infinite depth associated with a hierarchical Chinese restaurant process, our model directly infers the hidden orders of Markov dependencies from which each symbol originated. Experiments on character and word sequences in natural language showed that the model has a comparative performance with an exponentially large full-order model, while computationally much efficient in both time and space. We expect that this basic model will also extend to the variable order hierarchical clustering of general data.
Daichi Mochihashi, Eiichiro Sumita
NIPS1
2006 Exploring Multiple Communities with Kernel-Based Link Analysis
Takahiko Ito, Masashi Shimbo, Daichi Mochihashi, Yuji Matsumoto 0001
PKDD3
2005 Context as Filtering
abstract
Long-distance language modeling is important not only in speech recognition and machine translation, but also in high-dimensional discrete sequence modeling in general. However, the problem of context length has almost been neglected so far and a nave bag-of-words history has been i employed in natural language processing. In contrast, in this paper we view topic shifts within a text as a latent stochastic process to give an explicit probabilistic generative model that has partial exchangeability. We propose an online inference algorithm using particle filters to recognize topic shifts to employ the most appropriate length of context automatically. Experiments on the BNC corpus showed consistent improvement over previous methods involving no chronological order.
Daichi Mochihashi, Yuji Matsumoto 0001
NIPS1
2004 Learning Nonstructural Distance Metric by Minimum Cluster Distortion
Daichi Mochihashi, Gen-ichiro Kikui, Kenji Kita
EMNLP1