Daniel Y. Takahashi

dblp:11/8919 · also Daniel Yasumasa Takahashi · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0003-4972-001XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Pathwise Guessing in Categorical Time Series With Unbounded Alphabets
abstract
The following learning problem arises naturally in various applications: Given a finite sample from a categorical or count time series, can we learn a function of the sample that (nearly) maximizes the probability of correctly guessing the values of a given portion of the data using the values from the remaining parts? Unlike classical approaches in statistical inference, our approach avoids explicitly estimating the conditional probabilities. We propose a non-parametric guessing function with a learning rate independent of the alphabet size. Our analysis focuses on a broad class of time series models that encompasses finite-order Markov chains, some hidden Markov chains, Poisson regression for count processes, and one-dimensional Gibbs measures. We provide a margin condition that controls the rate of convergence for the risk. Additionally, we establish a minimax lower bound for the convergence rate of the risk associated with our guessing problem. This lower bound matches the upper bound achieved by our estimator up to a logarithmic factor, demonstrating its near-optimality.
Jean-René Chazottes, Sandro Gallo, Daniel Y. Takahashi
IEEE Trans. Inf. Theory3
2023 Sparse Markov Models for High-dimensional Inference
abstract
Finite-order Markov models are well-studied models for dependent finite alphabet data. Despite their generality, application in empirical work is rare when the order $d$ is large relative to the sample size $n$ (e.g., $d = \mathcal{O}(n)$). Practitioners rarely use higher-order Markov models because (1) the number of parameters grows exponentially with the order, (2) the sample size $n$ required to estimate each parameter grows exponentially with the order, and (3) the interpretation is often difficult. Here, we consider a subclass of Markov models called Mixture of Transition Distribution (MTD) models, proving that when the set of relevant lags is sparse (i.e., $\mathcal{O}(\log(n))$), we can consistently and efficiently recover the lags and estimate the transition probabilities of high-dimensional ($d = \mathcal{O}(n)$) MTD models. Moreover, the estimated model allows straightforward interpretation. The key innovation is a recursive procedure for a priori selection of the relevant lags of the model. We prove a new structural result for the MTD and an improved martingale concentration inequality to prove our results. Using simulations, we show that our method performs well compared to other relevant methods. We also illustrate the usefulness of our method on weather data where the proposed method correctly recovers the long-range dependence.
Guilherme Ost, Daniel Y. Takahashi
J. Mach. Learn. Res.2
2022 A mechanism for punctuating equilibria during mammalian vocal development
abstract
Evolution and development are typically characterized as the outcomes of gradual changes, but sometimes (states of equilibrium can be punctuated by sudden change. Here, we studied the early vocal development of three different mammals: common marmoset monkeys, Egyptian fruit bats, and humans. Consistent with the notion of punctuated equilibria, we found that all three species undergo at least one sudden transition in the acoustics of their developing vocalizations. To understand the mechanism, we modeled different developmental landscapes. We found that the transition was best described as a shift in the balance of two vocalization landscapes. We show that the natural dynamics of these two landscapes are consistent with the dynamics of energy expenditure and information transmission. By using them as constraints for each species, we predicted the differences in transition timing from immature to mature vocalizations. Using marmoset monkeys, we were able to manipulate both infant energy expenditure (vocalizing in an environment with lighter air) and information transmission (closed-loop contingent parental vocal playback). These experiments support the importance of energy and information in leading to punctuated equilibrium states of vocal development.
Thiago T. Varella, Yisi S. Zhang, Daniel Y. Takahashi, Asif A. Ghazanfar
PLoS Comput. Biol.3
2014 A comparative study of statistical methods used to identify dependencies between gene expression signals
abstract
One major task in molecular biology is to understand the dependency among genes to model gene regulatory networks. Pearson's correlation is the most common method used to measure dependence between gene expression signals, but it works well only when data are linearly associated. For other types of association, such as non-linear or non-functional relationships, methods based on the concepts of rank correlation and information theory-based measures are more adequate than the Pearson's correlation, but are less used in applications, most probably because of a lack of clear guidelines for their use. This work seeks to summarize the main methods (Pearson's, Spearman's and Kendall's correlations; distance correlation; Hoeffding's D: measure; Heller-Heller-Gorfine measure; mutual information and maximal information coefficient) used to identify dependency between random variables, especially gene expression data, and also to evaluate the strengths and limitations of each method. Systematic Monte Carlo simulation analyses ranging from sample size, local dependence and linear/non-linear and also non-functional relationships are shown. Moreover, comparisons in actual gene expression data are carried out. Finally, we provide a suggestive list of methods that can be used for each type of data set.
Suzana de Siqueira Santos, Daniel Y. Takahashi, Asuka Nakata, André Fujita
Briefings Bioinform.2