Yukio Ohsawa

dblp:24/4750 · DBLP profile ↗
← Back
25ranked-venue papers in the field
6as first author
14since 2021 · last 2025
0000-0003-2943-2547ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 17 (4 first)Knowledge Engineering, Semantic Web & Information Systems · 3 (1 first)Other / Interdisciplinary · 3Data Mining & Knowledge Discovery · 1Information Retrieval & Web Search · 1 (1 first)
YearPublicationVenuePosition
2025 Small Data for Walkability: Street-Furniture Experiments and Mobility Sensors in a Japanese Compact City
Sae Kondo, Kazuma Suzuki, Yukio Ohsawa
IEEE Big Data3
2025 Data Users' Tsugoes Fit Data Free Flow with Trust - Lessons by Large Language Model -
Yukio Ohsawa
IEEE Big Data1
2025 Latent Interestingness of Regions Externalized on Multiscale Trend-Reach Entropy
Yukio Ohsawa, Sae Kondo
IEEE Big Data1
2025 MMPEP: A Multi-Modal Framework for Post-Earnings Stock Movement Prediction
Dingming Xue, Kaira Sekiguchi, Yukio Ohsawa, Andi Li, Cecilia Melin
IEEE Big Data4
2024 LLM as a Tool for Trust-Creating Communication in the Data Marketplace
abstract
Methods for exchanging datasets and values as payment have been studied so far, where the communication of user’s intension has been positioned as a key for creating trust. However, user’s intension does not stand alone without logical description of the user’s current and desired situations. Here, with referring to the authors’ previous study about tsugo network, it is proposed to provide a prompt including the statement about the intention of data user, composed of his requirement and situation. This design of prompt is shown to play a significant role for obtaining a satisfactory use plan of datasets.
Son Yeon Hyuk, Kaira Sekiguchi, Yukio Ohsawa
IEEE Big Data3
2024 Does Station-Town Space Design Change Pedestrians' Mobility in Urban Area?
abstract
In Japan, a new concept of 'Station and Town Special Design’, an urban development that considers the station and the surrounding town block as a whole and revitalises the vitality of the town as a whole, is attracting attention. We hypothesised that people's movement behaviour in the area would become more active via this Station-Town Space Design project. We analysed the effects using movement direction entropy (MDE), which measures the diversity of people's movement directions. The results showed that MDE increased in the project area. Furthermore, MDE tended to be higher in the surroundings of stations. This suggests that the Station-Town Space Design stimulates people's movement behaviour in the station, station square, and surrounding area. These findings have significant implications for urban planners and developers, as they indicate that MDE can be a useful indicator for measuring the liveliness of a town, and that the Station-Town Space Design design can enhance the vitality of urban areas.
Sae Kondo, Shinpei Nomura, Kaira Sekiguchi, Yukio Ohsawa
IEEE Big Data5
2024 Structure Estimation of Financial Networks and Its Correlation with Mobility Data
abstract
During the COVID-19 pandemic, the human mobility was restricted and significantly impacted, which also had a serious effect on financial markets. Specifically, industries such as energy, utilities, travel, retail, and hospitality were negatively affected by the decrease in population movement, leading to reduced consumer spending and lower demand for services in these sectors, while other sectors, such as healthcare and technology, saw significant growth due to increased demand during the pandemic. Additionally, the challenging market environment led to multiple anomalies in financial markets. When financial markets are abnormal, signs of distress in the market structure often appear before a crisis fully erupts. This study proposes a model based on the Stochastic Block Model (SBM) to detect anomalies by analyzing market changes and underlying spatial structure. We then extended the model to create a multi-layer model, treating the blocks in the lower layer as nodes in the upper layer, using a two-layer approach for cause analysis. We analyzed the structure of the S&P 100 and TOPIX 100 indices during the pandemic and compared it with mobility data from the same period. The results indicate a correlation between population mobility and changes in market structure during the pandemic, offering a new perspective for analyzing financial markets.
Kaira Sekiguchi, Yoshiyuki Nakata, Toshiaki Sugie, Takaaki Yoshino, Yukio Ohsawa
IEEE Big Data6
2024 Word Archipelagoes for Explaining Contextual Shifts in Sequential Actions and Events
abstract
Text, as data which can represent actors’ behaviors with the reflection of emotional dimension, is expected to be used combined with physical mobility data for explaining contextual continuity and transitions. Even in the era of LLM, we still desire methods for the preparation of essential parts of texts for finely obtaining the contextual shift which may go from and to interleaving meaningful key-terms rather than sheer change of topics corresponding to word clusters. In this study, the explanatory relay of words in text is represented in the form of archipelagoes. Here, each archipelago is a sequence of islands composed of the occurrences of a certain word. An island here is interpreted as the local sequence where the word is emphasized, and an archipelago of a length comparable to the target text is extracted by using the variation of entropy type-A, the window-based entropy based on the distribution of the word’s occurrences with the width of each time window. The results show that the parts of the target text including the words forming archipelagoes thus extracted, without pre-learned knowledge, form an explanatory part of the text that is of smaller entropy B than the parts extracted by the baseline methods. Here, entropy type-B is the graph-based entropy representing the contextual dispersion of the text. This result means archipelagoes work, for extracting contextual flow with discarding noisy words in the text without a prepared knowledge about stop words in text or noises in sequential data.
Yukio Ohsawa
IEEE Big Data1
2024 Exploring Tourism Value Using Boundary Objects and Edges
abstract
Tourism information is crucial for enhancing tourist experiences, yet conventional visualization methods often overlook intersections and peripheral data that could reveal new insights. Boundary Objects and Edges are identified as key elements in discovering these overlooked aspects. Boundary Objects act as bridges between different information sources, facilitating new perspectives, while Edges offer new dimensions of value beyond the usual range of information. Here, we show that by applying our algorithm to identify these elements in tourism data, tourists can discover novel insights and enhanced experiences. This approach contrasts with previous methods that focus primarily on central data, often missing the potential insights at the boundaries. Our findings suggest that integrating Boundary Objects and Edges into tourism information systems can significantly enrich the tourist experience, offering new opportunities for value creation. These results pave the way for more innovative and comprehensive tool in the tourism industry, with broader implications for how information is processed and utilized across various fields.
Riko Tsubaki, Kaira Sekiguchi, Yukio Ohsawa
IEEE Big Data3
2024 Interacting Evolution of Modeling and Data Exchange : A Case for Predicting Chemical-Protein Interactive Visualization
abstract
In the process of data exchange, we assert that the needs of data users should be prioritized and clearly articulated for a trustworthy data marketplace. Here, we present a case study on predicting and discovering chemical-protein interactions, which are crucial for drug discovery, by collecting data tailored to these requirements. We developed a linear sequence model based on a list of amino acid residues to meet this need, and the data were selected for analysis with this model. A convolutional neural network is employed to capture the spatial characteristics of proteins, while also integrating the linear molecular features of drugs and the network topology. The framework’s 3D visualization provides intuitive representations of chemical-protein interactions, enhancing biological interpretation. This approach effectively balances data processing efficiency with rich structural insights, enabling a deeper analysis of critical protein-related features.
Jianshi Wang, Yukio Ohsawa
IEEE Big Data2
2024 BERT-MLTKE: A Multi-task Deep Learning Framework for Keyphrase Extraction in Social Media
abstract
Keyphrase extraction is a critical task in natural language processing (NLP) aimed at automatically identifying and extracting essential phrases from a given text. Traditionally, keyphrase extraction has been extensively studied in the context of structured text sources such as articles, documents, and web pages. With the growth of social media platforms like Twitter (X), the demand for effective keyphrase extraction has become increasingly important in not only public opinion and sentiment understanding but also facilitating real-time trend identification and social governance. Keyphrase extraction from social media presents unique challenges due to the informal, unstructured, and highly diverse nature of the data. Existing methods mainly rely on either unsupervised (statistical based) approaches or supervised (deep learning based) techniques, each of which has its own advantages and disadvantages. In order to leverage the strengths of both supervised and unsupervised learning, we propose BERT-MLTKE, a BERT-based multi-task framework designed to significantly enhance keyphrase extraction performance on social media data. Our framework fine-tunes the pre-trained BERT model to effectively capture both contextual information and syntactic structures within sentences. Instead of using traditional unsupervised approaches of extracting candidate phrases and ranking them based on statistical metrics, we introduce a multi-task module with supervised deep learning models that addresses both keyword identification and keyphrase extraction as two progressive tasks, enabling the framework to share essential information while tackling multiple subtasks concurrently. During the implementation of deep learning models, we used several techniques that contribute to model stability. For further enhancement in framework’s performance, we conducted several comparative experiments, optimizing both the architecture of BERT-MLTKE and deep model candidates selection in embedding and feature transformation layers. The results of these experiments were thoroughly analyzed, and we offer possible explanations for the observed variations for different deep models in task performance. We also propose strict evaluation metrics to ensure a more precise assessment of the experimental results in an objective manner.
Dingming Xue, Kaira Sekiguchi, Yukio Ohsawa
IEEE Big Data4
2024 Identification of factors that induce recreational cycling using a route generation model
abstract
The study of route selection for cyclists is a field aimed at promoting bicycle use and fostering the development of eco-friendly cities. Such research often focuses on investigating bicycle routes used for commuting. However, there is a lack of research on routes used for recreational cycling, and it is a missed opportunity not to focus on the strolling behaviors of people who enjoy such freedom. This study presents the results of utilizing a route generation model for recreational cycling, with a focus on identifying the elements of roads and surrounding environments that attract individuals to participate in this activity. We discovered how much people value waterfronts and recreational spaces, as well as the impact of railway density on recreational cycling across different areas. These findings could be used to inform policies for creating cities where recreational cycling is easier to enjoy and bicycles, as an eco-friendly mode of transportation, are more readily utilized.
Hiro Yoshida, Kaira Sekiguchi, Yukio Ohsawa
IEEE Big Data3
2022 Data Leaves as Scenario-oriented Metadata for Data Federative Innovation on Trust
abstract
Communication regarding data use enhances and sustains trust in the data market. Showing this principle based on the literature on trust, this study proposes a novel method for representing the digest information of datasets to foster the thoughts of potential data users, who attempt to create valuable products, services, or business models using datasets and aid their communication with data providers. Compared with existing metadata, where variable labels in a dataset are listed, the presented metadata, called data leaf (DL), includes events, situations, and/or actions that, as a set, compose scenarios supposed to be active in the target real world of a dataset. This method considers the fitness of metadata corresponding to various datasets to a feature concept (FC), which provides an abstract illustration of the knowledge acquired or expected to be acquired from the data, and plays an essential role in data utilization. In experiments using metadata as elements to be combined in human thought for data federative innovation, DLs significantly outperformed cases using data jackets (DJs) where variable labels were used as attributes of the datasets.
Yukio Ohsawa, Kaira Sekiguchi, Tomohide Maekawa, Hiroki Yamaguchi, Son Yeon Hyuk, Sae Kondo
IEEE Big Data1
2022 The impact of product descriptions on customers' trust in sellers and purchase intention
abstract
A key issue is how to help buyers and sellers build a trusting relationship with each other on e-commerce platforms. In this regard, researchers have been interested in addressing the information asymmetry between buyers and sellers. This study focuses on figure products often sold at higher than original prices and whether the signals hidden in sellers’ presentations of such products can alleviate this information asymmetry. To this end, we use text analysis and logic models to discover new signals that can reduce information asymmetry, and the effect of the signals varies with price. Such findings may have managerial implications for platform owners, helping them to maintain better buyer-seller relationships and increase trust in the platform.
Yukio Ohsawa
IEEE Big Data2
2020 Data Requests and Scenarios for Data Design of Unobserved Events in Corona-related Confusion Using TEEDA
abstract
Due to the global violence of the novel coronavirus, various industries have been affected and the breakdown between systems has been apparent. To understand and overcome the phenomenon related to this unprecedented crisis caused by the coronavirus infectious disease (COVID-19), the importance of data exchange and sharing across fields has gained social attention. In this study, we use the interactive platform called treasuring every encounter of data affairs (TEEDA) to externalize data requests from data users, which is a tool to exchange not only the information on data that can be provided but also the call for data-what data users want and for what purpose. Further, we analyze the characteristics of missing data in the corona-related confusion stemming from both the data requests and the providable data obtained in the workshop. We also create three scenarios for the data design of unobserved events focusing on variables.
Teruaki Hayashi, Nao Uehara, Daisuke Hase, Yukio Ohsawa
IEEE BigData4
2020 Verification of Data Similarity using Metadata on a Data Exchange Platform
abstract
With the development of computers and the rise of data exchange, the expectations for innovation by combining data from different industries, and data exchange platforms that handle different types of data are sprouting up. However, the data handled on such platforms have been obtained and are stored independently by data providers with different purposes, and the maintenance of data catalogs and the spread of schemata are not currently insufficient, making it difficult to understand the relationships between the data on such platforms. In this study, in order to derive the relationships among datasets and to discuss the similarity and combinability of the data on the data exchange platform, we analyze and discuss the relationships between datasets on the platform. We focussed on the data outlines and variables using the metadata of the data exchange platform service, D-Ocean, and found that the similarity of data cannot be measured by a single indicator, and that the items to be referred to calculate the similarity depending on the indicator.
Hiroki Sakaji, Teruaki Hayashi, Kiyoshi Izumi, Yukio Ohsawa
IEEE BigData4
2020 Hierarchical Graph Convolutional Network for Data Evaluation of Dynamic Graphs
abstract
As data are being generated at an incredible speed and scale, the market of data that aims to fully discover the value of data is becoming increasingly important. Data evaluation is a vital part of the data market. It can discover the data's flaws and the value hidden in the data. Nevertheless, there have been few studies investigating how to evaluate data in the data market. Existing methods seldom utilize the multi-level structure in data. Our work proposes a novel hierarchical graph convolutional network for the data evaluation of dynamic graphs, following the anomaly detection paradigm. Our model performs significantly better than existing models on several benchmark datasets. This study provides a more powerful tool for processing dynamic graphs. It could also guide the direction of data evaluation for a broader range of data categories in the data market.
Teruaki Hayashi, Yukio Ohsawa
IEEE BigData3
2017 Matrix-Based Method for Inferring Variable Labels Using Outlines of Data in Data Jackets
Teruaki Hayashi, Yukio Ohsawa
PAKDD (2)2
2009 A shift of mind - Introducing a concept creation model
Jun Nakamura 0001, Yukio Ohsawa
Inf. Sci.2
2009 Special issue on chance discovery - Discovery of significant events for decision making
Yukio Ohsawa, Katsutoshi Yada
Inf. Sci.1
2003 Average-Clicks: A New Measure of Distance on the World Wide Web
Yutaka Matsuo, Yukio Ohsawa, Mitsuru Ishizuka
J. Intell. Inf. Syst.2
2002 Featuring web communities based on word co-occurrence structure of communications: 736
abstract
Textual communication in message boards is analyzed for classifying Web communities. We present a communication-content based generalization of an existing business-oriented classification of Web communities, using KeyGraph, a method for visualizing the co-occurrence relations between words and word clusters in text. Here, the text in a message board is analyzed with KeyGraph, and the structure obtained is shown to reflect the essence of the content-flow. The relation of this content-flow with participants' interests is then formalized. Three structure-features of relations between participants and words, determining the type of the community, are shown to be computed and visualized: (1) centralization (2) context coherence and (3) creative decisions. This helps in surveying the essence of a community, e.g. whether the community creates useful knowledge, how easy it is to join the community, and whether/why the community is good for making commercial advertisement.
Yukio Ohsawa, Hirotaka Soma, Yutaka Matsuo, Naohiro Matsumura, Masaki Usui
WWW1
2001 Discovering Seeds of New Interest Spread from Premature Pages Cited by Multiple Communities
Naohiro Matsumura, Yukio Ohsawa, Mitsuru Ishizuka
Web Intelligence2
2001 Discovery of Emerging Topics between Communities on WWW
Naohiro Matsumura, Yukio Ohsawa, Mitsuru Ishizuka
Web Intelligence2
2001 Average-Clicks: A New Measure of Distance on the World Wide Web
Yutaka Matsuo, Yukio Ohsawa, Mitsuru Ishizuka
Web Intelligence2