Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tatsuya Asai

dblp:54/6721 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
3since 2021 · last 2025
0009-0003-9708-2929ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data integration and cleaning · 91% Data mining · 9%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
data preprocessing
0.812024
AutoDW: Automatic Data Wrangling Leveraging Large Language Models · ASE 2024
Data integration and cleaning
data wrangling
0.812024
AutoDW: Automatic Data Wrangling Leveraging Large Language Models · ASE 2024
Data mining
data stream mining
0.012002
Online Algorithms for Mining Semi-structured Data Stream · ICDM 2002
Data mining › pattern mining
frequent pattern mining
0.012002
Online Algorithms for Mining Semi-structured Data Stream · ICDM 2002
Data mining
semi-structured data mining
0.012002
Online Algorithms for Mining Semi-structured Data Stream · ICDM 2002
Data mining › pattern mining › tree mining
tree pattern mining
0.012002
Online Algorithms for Mining Semi-structured Data Stream · ICDM 2002

Methods — techniques the papers use, named apart from their topics

large language model · 1.5code generation · 1.5tree sweeping · 0.0incremental maintenance · 0.0
YearPublicationVenuePosition
2025 AutoDW-TS: Automated Data Wrangling for Time-Series Data
Lei Liu 0061, So Hasegawa, Shailaja Sampat, Mehdi Bahrami, Wei-Peng Chen, Kodai Toyota, Takashi Kato, Takumi Akazaki, Akira Ura, Tatsuya Asai
CIKM10
2024 AutoDW: Automatic Data Wrangling Leveraging Large Language Models
abstract
Data wrangling is a critical yet often labor-intensive process, essential for transforming raw data into formats suitable for downstream tasks such as machine learning or data analysis. Traditional data wrangling methods can be time-consuming, resource-intensive, and prone to errors, limiting the efficiency and effectiveness of subsequent downstream tasks. In this paper, we introduce AutoDW: an end-to-end solution for automatic data wrangling that leverages the power of Large Language Models (LLMs) to enhance automation and intelligence in data preparation. AutoDW distinguishes itself through several innovative features, including comprehensive automation that minimizes human intervention, the integration of LLMs to enable advanced data processing capabilities, and the generation of source code for the entire wrangling process, ensuring transparency and reproducibility. These advancements position AuoDW as a superior alternative to existing data wrangling tools, offering significant improvements in efficiency, accuracy, and flexibility. Through detailed performance evaluations, we demonstrate the effectiveness of AutoDW for data wrangling. We also discuss our experience and lessons learned from the industrial deployment of AutoDW, showcasing its potential to transform the landscape of automated data preparation.
Lei Liu 0061, So Hasegawa, Shailaja Sampat, Maria Xenochristou, Wei-Peng Chen, Takashi Kato, Taisei Kakibuchi, Tatsuya Asai
ASE8
2022 Incorporating AI Methods in Micro-dynamic Analysis to Support Group-Specific Policy-Making
Shuang Chang, Tatsuya Asai, Yusuke Koyanagi, Kento Uemura, Koji Maruhashi, Kotaro Ohori
PRIMA2
2014 Discovery of Areas with Locally Maximal Confidence from Location Data
Hiroya Inakoshi, Hiroaki Morikawa, Tatsuya Asai, Nobuhiro Yugami, Seishi Okamoto
DASFAA (1)3
2012 EVIS: A Fast and Scalable Episode Matching Engine for Massively Parallel Data Streams
Shin-ichiro Tago, Tatsuya Asai, Takashi Katoh, Hiroaki Morikawa, Hiroya Inakoshi
DASFAA (2)2
2010 Chimera: Stream-Oriented XML Filtering/Querying Engine
Tatsuya Asai, Shin-ichiro Tago, Hiroya Inakoshi, Seishi Okamoto, Masayuki Takeda
DASFAA (2)1
2010 Training Parse Trees for Efficient VF Coding
Takashi Uemura, Satoshi Yoshida, Takuya Kida, Tatsuya Asai, Seishi Okamoto
SPIRE4
2004 An Efficient Algorithm for Enumerating Closed Patterns in Transaction Databases
Takeaki Uno, Tatsuya Asai, Yuzo Uchida, Hiroki Arimura
Discovery Science2
2003 Discovering Frequent Substructures in Large Unordered Trees
Tatsuya Asai, Hiroki Arimura, Takeaki Uno, Shin-Ichi Nakano
Discovery Science1
2002 Online Algorithms for Mining Semi-structured Data Stream
abstract
In this paper, we study an online data mining problem from streams of semi-structured data such as XML data. Modeling semi-structured data and patterns as labeled ordered trees, we present an online algorithm StreamT that receives fragments of an unseen possibly infinite semi-structured data in the document order through a data stream, and can return the current set of frequent patterns immediately on request at any time. A crucial part of our algorithm is the incremental maintenance of the occurrences of possibly frequent patterns using a tree sweeping technique. We give modifications of the algorithm to other online mining model. We present theoretical and empirical analyses to evaluate the performance of the algorithm.
Tatsuya Asai, Hiroki Arimura, Kenji Abe, Shinji Kawasoe, Setsuo Arikawa
ICDM1
2002 Optimized Substructure Discovery for Semi-structured Data
Kenji Abe, Shinji Kawasoe, Tatsuya Asai, Hiroki Arimura, Setsuo Arikawa
PKDD3
2002 Efficient Substructure Discovery from Large Semi-structured Data
abstract
1 Introduction By rapid progress of network and storage technologies, a huge amount of electronic data such as Web pages and XML data [23] has been available on intra and internet. These electronic data are heterogeneous collection of ill-structured data that have no rigid structures, and often called semi-structured data [1]. Hence, there have been increasing demands for automatic methods for extracting useful information, particularly, for discovering rules or patterns from large collections of semi-structured data, namely, semi-structured data mining [6, 11, 18, 19, 21, 25].
Tatsuya Asai, Kenji Abe, Shinji Kawasoe, Hiroki Arimura, Hiroshi Sakamoto, Setsuo Arikawa
SDM1