Katsumi Tanaka

dblp:44/3648 · DBLP profile ↗
← Back
215ranked-venue papers
15as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 165 · 11 first-author · 2 since 2021Artificial intelligence and machine learning · 67 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 6 · 1 first-authorSystems, architecture and hardware · 2Theory of computation · 2 · 1 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
31 papers
Information retrieval · 77% Data mining · 16% Web and social media mining · 3%
Human-computer interaction and pervasive computing
3 papers
Health and well-being technologies · 51% Immersive interaction · 31% Collaborative and social computing · 15%
Artificial intelligence
5 papers
Representation and self-supervised learning · 50% Information extraction and text analysis · 49% Efficient and distributed learning · 2%
Computer graphics and multimedia
8 papers
Visualization and visual analytics · 42% Multimedia analysis and retrieval · 32% Multimedia systems and quality of experience · 16%

Topics — the 30 heaviest of 77, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Health and well-being technologies › physical activity promotion
running support
0.712023
Development of an Online Marathon System using Acoustic AR · ACM Multimedia 2023
Information retrieval
query suggestion
0.422016
ScentBar: A Query Suggestion Interface Visualizing the Amount of Missed Relevant Information for Intrinsically Diverse Search · SIGIR 2016
Structured query suggestion for specialization and parallel movement: effect on search behaviors · WWW 2012
Information retrieval
multimedia analysis and retrieval
0.312017
Search by Screenshots for Universal Article Clipping in Mobile Apps · ACM Trans. Inf. Syst. 2017
Data mining › text mining › information extraction
temporal information extraction
0.312017
Timestamping Entities using Contextual Information · SIGIR 2017
Information retrieval
retrieval models
0.332013
Estimating content concreteness for finding comprehensible documents · WSDM 2013
Proposal of integrated search engine of web and TV contents · WWW 2006
A comparative web browser (CWB) for browsing and comparing web pages · WWW 2003
Information retrieval › search interfaces
search result presentation
0.322016
To Suggest, or Not to Suggest for Queries with Diverse Intents: Optimizing Search Result Presentation · WSDM 2016
Toward tighter integration of web search with a geographic information system · WWW 2006
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.212016
The Past is Not a Foreign Country: Detecting Semantically Similar Terms across Time · IEEE Trans. Knowl. Data Eng. 2016
Machine learning › Representation and self-supervised learning
word representation
0.212016
The Past is Not a Foreign Country: Detecting Semantically Similar Terms across Time · IEEE Trans. Knowl. Data Eng. 2016
Data mining › text mining › sentiment analysis
review mining
0.212016
Detecting Evolution of Concepts based on Cause-Effect Relationships in Online Reviews · WWW 2016
Information retrieval
search result diversification
0.212016
ScentBar: A Query Suggestion Interface Visualizing the Amount of Missed Relevant Information for Intrinsically Diverse Search · SIGIR 2016
Information retrieval › query understanding
subtopic mining
0.212016
ScentBar: A Query Suggestion Interface Visualizing the Amount of Missed Relevant Information for Intrinsically Diverse Search · SIGIR 2016
Information retrieval
web search
0.232014
Enhancing credibility judgment of web search results · CHI 2011
Toward tighter integration of web search with a geographic information system · WWW 2006
Investigating users' query formulations for cognitive search intents · SIGIR 2014
Natural language and speech › Information extraction and text analysis › lexical semantics
semantic change detection
0.212015
Omnia Mutantur, Nihil Interit: Connecting Past with Present by Finding Corresponding Terms across Time · ACL (1) 2015
Data mining › text mining
temporal text analysis
0.212015
Omnia Mutantur, Nihil Interit: Connecting Past with Present by Finding Corresponding Terms across Time · ACL (1) 2015
Immersive interaction › augmented reality
audio augmented reality
0.212023
Development of an Online Marathon System using Acoustic AR · ACM Multimedia 2023
Immersive interaction
augmented reality
0.212023
Development of an Online Marathon System using Acoustic AR · ACM Multimedia 2023
Information retrieval › query reformulation
query expansion
0.212014
Investigating users' query formulations for cognitive search intents · SIGIR 2014
Information retrieval
query formulation
0.212014
Investigating users' query formulations for cognitive search intents · SIGIR 2014
Information retrieval › evaluation
document quality
0.212013
Estimating content concreteness for finding comprehensible documents · WSDM 2013
Information retrieval
ranking
0.212013
Estimating content concreteness for finding comprehensible documents · WSDM 2013
Information retrieval › digital libraries
web archiving
0.122008
Visualizing historical content of web pages · WWW 2008
A browser for browsing the past web · WWW 2006
Visualization and visual analytics
temporal data visualization
0.122008
Visualizing historical content of web pages · WWW 2008
A browser for browsing the past web · WWW 2006
Information retrieval
content-based retrieval
0.112012
Content-based retrieval for heterogeneous domains: domain adaptation by relative aggregation points · SIGIR 2012
Information retrieval
domain adaptation
0.112012
Content-based retrieval for heterogeneous domains: domain adaptation by relative aggregation points · SIGIR 2012
Information retrieval
query reformulation
0.112012
Structured query suggestion for specialization and parallel movement: effect on search behaviors · WWW 2012
Information retrieval › web search › web information retrieval › web content quality
credibility assessment
0.112011
Enhancing credibility judgment of web search results · CHI 2011
Information retrieval
image retrieval
0.122005
3D viewpoint-based photo search and information browsing · SIGIR 2005
RelaxImage: A Cross-Media Meta-Search Engine for Searching Images from Web Based on Query Relaxation · ICDE 2005
Information retrieval › reranking
search result re-ranking
0.112010
RerankEverything: a reranking interface for browsing search results · WWW 2010
Information retrieval › user behavior
user search behavior
0.122014
Investigating users' query formulations for cognitive search intents · SIGIR 2014
Structured query suggestion for specialization and parallel movement: effect on search behaviors · WWW 2012
Recommender systems
user modeling
0.122013
Estimating content concreteness for finding comprehensible documents · WSDM 2013
Enhancing credibility judgment of web search results · CHI 2011

Methods — techniques the papers use, named apart from their topics

virtual runner · 0.7acoustic AR · 0.7term classification · 0.5popularity tracking · 0.5screenshot segmentation · 0.3query formulation · 0.3link structure propagation · 0.3learning to rank · 0.3word embeddings · 0.2vector space transformation · 0.2probabilistic user model · 0.2intent-aware metrics · 0.2gain estimation · 0.2frequency-based change detection · 0.2context-based change detection · 0.2aspect importance estimation · 0.2temporal term cloud · 0.2interactive visualization · 0.2
YearPublicationVenuePosition
2026 Asymmetric Pipeline for Dataset Construction and Situation-aware Generative Outfit Retrieval Leveraging Differences in Task Difficulty
abstract
This paper proposes a method to verbalize and generate corresponding outfit images based on natural language inputs describing everyday situations. To achieve this, an asymmetric pipeline is constructed for both dataset construction and situation-aware generative outfit retrieval. The design of this pipeline is grounded in the difference in task difficulty between generating images from situations and captioning images. A Vision-Language Model (VLM) is employed to caption outfit images, producing both the constituent elements of the outfit and the situations in which the outfit might be worn. These outputs are paired to construct a situation–outfit description dataset named OOTM (Outfit Of The Moment). An LLM trained on this dataset learns to transform situation descriptions into appropriate outfit descriptions. The generated outfit descriptions are then provided as input to a text-to-image model, which synthesizes outfit images. Through this process, the framework enables situation-aware outfit retrieval by generating outfit images that align with the given situational query. The appropriateness of the generated text and images for their respective situations was evaluated through both automatic evaluation using the constructed dataset and human subject experiments.
Yuma Oe, Katsumi Tanaka, Yoshiyuki Shoji
ICMR2
2023 User Latent Interest Estimation in Real Space: A Comparative Analysis of Time-Series and Non-Time-Series Processing Algorithms
abstract
Web advertising services have exhibited consistent growth over the years. However, the conventional methods of web advertising recommendations, relying on keyword matching with search queries and browsing histories, encounter challenges when it comes to effectively targeting users with hidden or latent interests. In contrast, the use of mobile device location data in advertising recommendations often centers around physical store proximity. To address these limitations, our research aims to enhance web advertising recommendations by analyzing latent user interests through real-world behavioral data. This study specifically investigates the influence of area size on user behavioral analysis and its subsequent impact on the accuracy of predicting visit probabilities. We achieve this by extracting the user’s activity range from user behavior (movement) log data and geotagged tweets. Subsequently, we tally the places visited by the user, considering spot attributes, and convert this data into feature vectors. Utilizing these feature vectors in conjunction with various classification methods, we build learning models. In this paper, we present and evaluate these learning models employing different area sizes, verifying their accuracy in predicting user visits to specific stores.
Takanobu Omura, Da Li 0008, Panote Siriaraya, Katsumi Tanaka, Yukiko Kawai, Shinsuke Nakajima
IEEE Big Data4
2023 Development of an Online Marathon System using Acoustic AR
abstract
In recent years, the number of people who run for the purpose of improving their health and physical fitness has been increasing, but it is not easy to continue running. Therefore, we believe that it is very important to develop a running support system. In our previous study, we developed a running support system and an application that enables users to run with a virtual runner created by using their past running data in an acoustic augmented reality space. However, this system has the problem that it can only race against the past running records of oneself or one's acquaintance, and can only race against a single running record. Therefore, we considered it necessary to develop a system that allows users to race against an unspecified number of users and that allows users to arbitrarily select the distance and number of times they wish to race. In this paper, we propose an online marathon system that enables large-scale online races and examine the effect of the number of competitors on user motivation.
Yuki Konishi, Panote Siriaraya, Da Li 0008, Katsumi Tanaka, Yukiko Kawai, Shinsuke Nakajima
ACM Multimedia4
2020 Context-Guided Learning to Rank Entities
Makoto P. Kato, Wiradee Imrattanatrai, Takehiro Yamamoto, Hiroaki Ohshima, Katsumi Tanaka
ECIR (1)5
2019 Discovering Latent Threads in Entity Histories
abstract
Abstract Knowledge of entity histories is often necessary for comprehensive understanding and characterization of entities. Yet, the analysis of an entity’s history is often most meaningful when carried out in comparison with the histories of other entities. In this paper, we describe a novel task ofhistory-based entity categorizationandcomparison. Based on a set of entity-related documents which are assumed as an input, we determine latent entity categories whose members share similar histories; hence, we are effectively grouping entities based on the correspondences in their historical developments. Next, we generate comparative timelines for each determined group allowing users to elucidate similarities and differences in the histories of entities. We evaluate our approach on several datasets of different entity types demonstrating its effectiveness against competitive baselines.
Yijun Duan, Adam Jatowt, Katsumi Tanaka
Data Sci. Eng.3
2018 Guest editorial - Special issue on conceptual modeling - 35th International Conference on Conceptual Modeling (ER2016)
Isabelle Comyn-Wattiau, Il-Yeol Song, Katsumi Tanaka, Motoshi Saeki, Shuichiro Yamamoto
Data Knowl. Eng.3
2017 Temporal Analog Retrieval using Transformation over Dual Hierarchical Structures
abstract
In recent years, we have witnessed a rapid increase of text con- tent stored in digital archives such as newspaper archives or web archives. Many old documents have been converted to digital form and made accessible online. Due to the passage of time, it is however difficult to effectively perform search within such collections. Users, especially younger ones, may have problems in finding appropriate keywords to perform effective search due to the terminology gap arising between their knowledge and the unfamiliar domain of archival collections. In this paper, we provide a general framework to bridge different domains across-time and, by this, to facilitate search and comparison as if carried in user's familiar domain (i.e., the present). In particular, we propose to find analogical terms across temporal text collections by applying a series of transformation procedures. We develop a cluster-biased transformation technique which makes use of hierarchical cluster structures built on the temporally distributed document collections. Our methods do not need any specially prepared training data and can be applied to diverse collections and time periods. We test the performance of the proposed approaches on the collections separated by both short (e.g., 20 years) and long time gaps (70 years), and we report improvements in range of 18%-27% over short and 56%-92% over long periods when compared to state-of-the-art baselines.
Adam Jatowt, Katsumi Tanaka
CIKM3
2017 Timestamping Entities using Contextual Information
abstract
Wikipedia is the result of collaborative effort aiming to represent human knowledge and to make it accessible to the public. Many Wikipedia articles however lack key metadata information. For example, relatively large number of people described in Wikipedia have no information on their birth and death dates. We propose in this paper to estimate entity's lifetimes using link structure in Wikipedia focusing on person entities. Our approach is based on propagating temporal information over links between Wikipedia articles.
Adam Jatowt, Daisuke Kawai, Katsumi Tanaka
SIGIR3
2017 Entity search by leveraging attributive terms in sentential queries over RDF data
abstract
This paper proposes methods of finding a ranked list of entities using RDF data for a given sentential query (e.g. "Cars 3", "Toy Story 4", or "The Incredibles 2" for the query "upcoming animated films pixar") by leveraging different types of modifiers in the query through identifying corresponding properties (e.g. released and movie type for the modifiers "upcoming" and "animated", respectively). While major search engines provide the entity search functionality that returns a list of entities based on users' queries, entities are neither presented for a wide variety of search queries, nor in the order that users expect. To enhance the efficiency of entity search, we propose two entity ranking methods. Our first proposed method is a Web-based entity ranking that directly finds highly relevant entities from Web search results returned in response to the query as a whole, and propagates the estimated relevance to the other entities. The second proposed method is a property-based entity ranking that ranks entities based on properties corresponding to modifier terms in the query. To this end, we propose a novel method that identifies a set of relevant properties based on the combination of the frequency of property values containing the modifier, co-occurrence of the modifier and property names, and difference in property value distributions of entities in the search results for a query. The experimental results showed that our proposed property identification method could predict more relevant properties than using each criterion separately. Moreover, we achieved the best performance for returning a ranked list of relevant entities when using both of the Web-based and property-based entity ranking methods.
Wiradee Imrattanatrai, Makoto P. Kato, Katsumi Tanaka
WI3
2017 Context-aware relevance feedback over SNS graph data
abstract
This study proposes a method for retrieving and ranking posts from social network services(SNSs) by specifying and providing feedback on the context of posts. Current search systems for SNS posts cannot handle user intent with regard to the context of posts to be retrieved, mainly owing to the incompleteness of SNS posts, i.e., they do not contain the users' contexts (e.g., situations or preferences) of users posting messages. Hence, we propose a search method that accepts two kinds of queries, namely, content queries and context queries, and that updates these queries based on the user feedback with special attention to the contexts of posts. Our search method considers the whole SNS dataset as a graph and the nodes surrounding each post as its context; to find relevant posts in terms of content and context, our method propagates user feedback via this graph. Our experimental results based on a Twitter test collection revealed that our proposed method showed improved retrieval performance as compared with conventional SNS retrieval and relevance feedback. In addition, we could detect the optimal parameters for feedback propagating.
Daisuke Kataoka, Makoto P. Kato, Takehiro Yamamoto, Hiroaki Ohshima, Katsumi Tanaka
WI5
2017 Mining alternative actions from community Q&A corpus for task-oriented web search
abstract
Web searchers often use a Web search engine to find a way or means to achieve his/her goal. For example, a user intending to solve his/her sleeping problem, the query "sleeping pills" may be used. However, there may be another solution to achieve the same goal, such as "have a cup of hot milk" or "stroll before bedtime." The problem is that the user may not be aware that these solutions exist. Thus, he/she will probably choose to take a sleeping pill without considering these solutions. In this study, we define and tackle the alternative action mining problem. In particular, we attempt to develop a method for mining alternative actions for a given query. We define alternative actions as actions which share the same goal and define the alternative action mining problem as similar in the search result diversification. To tackle the problem, we propose leveraging a community Q&A (cQA) corpus for mining alternative actions. We propose a method to compute how well two actions can be alternative actions by using a question-answer structure in a cQA corpus. Our method builds a question-action bipartite graph and recursively computes how well two actions can be alternative actions. We conducted experiments to investigate the effectiveness of our method using two newly built test collections, each containing 50 queries. The experimental results indicated that our proposed method outperformed the query suggestion methods provided by the commercial search engines in terms of D#-nDCG.
Suppanut Pothirattanachaikul, Takehiro Yamamoto, Sumio Fujita, Akira Tajima, Katsumi Tanaka
WI5
2017 Search by Screenshots for Universal Article Clipping in Mobile Apps
abstract
To address the difficulty in clipping articles from various mobile applications (apps), we propose a novel framework called UniClip, which allows a user to snap a screen of an article to save the whole article in one place. The key task of the framework is search by screenshots , which has three challenges: (1) how to represent a screenshot; (2) how to formulate queries for effective article retrieval; and (3) how to identify the article from search results. We solve these by (1) segmenting a screenshot into structural units called blocks, (2) formulating effective search queries by considering the role of each block, and (3) aggregating the search result lists of multiple queries. To improve efficiency, we also extend our approach with learning-to-rank techniques so that we can find the desired article with only one query. Experimental results show that our approach achieves high retrieval performance ( F 1 = 0.868), which outperforms baselines based on keyword extraction and chunking methods. Learning-to-rank models improve our approach without learning by about 6%. A user study conducted to investigate the usability of UniClip reveals that ours is preferred by 21 out of 22 participants for its simplicity and effectiveness.
Kazutoshi Umemoto, Ruihua Song, Jian-Yun Nie, Xing Xie 0001, Katsumi Tanaka, Yong Rui
ACM Trans. Inf. Syst.5
2016 Towards understanding word embeddings: Automatically explaining similarity of terms
abstract
Word embedding techniques (e.g., Word2Vec, GloVe) have been recently used for variety of applications with quite good rate of success. They allow to capture word semantics and syntactics with decreased dimensionality based on the concept of distributional vector representations. Vector representations can be then used for similarity comparison. However, if we treat the word embeddings as a kind of encryption process, it is difficult to decrypt their meaning. This makes it problematic to justify why particular terms should be considered similar as well as to prove that the overall quality of the trained vector space is high. Evaluating the accuracy of the similarity computation between any two given terms is difficult due to the lack of concrete evidences to explain and support the similarity. In this paper, we propose a novel way to automatically extract evidences represented as term pairs to explain the similarity of arbitrary terms. Our approach is unsupervised and can be applied to either homogeneous or heterogeneous vector spaces.
Adam Jatowt, Katsumi Tanaka
IEEE BigData3
2016 Predicting Importance of Historical Persons using Wikipedia
abstract
Wikipedia contains a lot of contemporary as well as history-related information, and given its vast coverage and richness, it can be used to rank entities in a variety of different ways. In this work, we are interested in utilizing Wikipedia for judging historical person's importance. Based on the two well-known lists of the most important people in the last millennium, we look closely into factors that determine significance of historical persons. We predict person's importance using six classifiers equipped with features derived from link structure, visit logs and article content.
Adam Jatowt, Daisuke Kawai, Katsumi Tanaka
CIKM3
2016 ScentBar: A Query Suggestion Interface Visualizing the Amount of Missed Relevant Information for Intrinsically Diverse Search
abstract
For intrinsically diverse tasks, in which collecting extensive information from different aspects of a topic is required, searchers often have difficulty formulating queries to explore diverse aspects and deciding when to stop searching. With the goal of helping searchers discover unexplored aspects and find the appropriate timing for search stopping in intrinsically diverse tasks, we propose ScentBar, a query suggestion interface visualizing the amount of important information that a user potentially misses collecting from the search results of individual queries. We define the amount of missed information for a query as the additional gain that can be obtained from unclicked search results of the query, where gain is formalized as a set-wise metric based on aspect importance, aspect novelty, and per-aspect document relevance and is estimated by using a state-of-the-art algorithm for subtopic mining and search result diversification. Results of a user study involving 24 participants showed that the proposed interface had the following advantages when the gain estimation algorithm worked reasonably: (1) ScentBar users stopped examining search results after collecting a greater amount of relevant information; (2) they issued queries whose search results contained more missed information; (3) they obtained higher gain, particularly at the late stage of their sessions; and (4) they obtained higher gain per unit time. These results suggest that the simple query visualization helps make the search process of intrinsically diverse tasks more efficient, unless inaccurate estimates of missed information are visualized.
Kazutoshi Umemoto, Takehiro Yamamoto, Katsumi Tanaka
SIGIR3
2016 Query Suggestion for Struggling Search by Struggling Flow Graph
abstract
We propose a method to generate effective query suggestions aiming to help struggling search, where users experience difficulty in locating information that is relevant to their information need in the search session. The core is identifying struggling component of an on-going struggling session and mining the effective representations of it. The struggling component is the semantic component of information need for which the user struggled to find an effective representation during the struggling session. The proposed method identifies the struggling component of given on-going struggling session and mines the sessions containing the identified struggling component from a query log to build a struggling flow graph. The struggling flow graph records users' reformulation behaviors for the terms of the struggling component, through struggling flow graph we can mine effective representations of the struggling component. The experimental results demonstrate that the proposed method outperforms the baseline methods when it can use two or more queries in a struggling session.
Zebang Chen, Takehiro Yamamoto, Katsumi Tanaka
WI3
2016 Supporting News Article Understanding by Detecting Subject-Background Event Relations
abstract
Typically, news articles mention not just one but multiple events. These events can be classified into subject or background events. The former are events that the article is written about, while the latter are additional events referred to in order to explain the background of the subject events (e.g., causal relations, circumstances or the consequences of the main event). Background events are considered to play an important role in helping to understand articles. In this paper, we first propose to classify content of news articles into subject or background event descriptions. In the second part of the paper, we demonstrate a novel solution for improving the news article search. Based on the subject and background relationship structure between events and articles, our method outputs news articles that help with understanding of a given target article.
Shotaro Tanaka, Adam Jatowt, Katsumi Tanaka
WI3
2016 To Suggest, or Not to Suggest for Queries with Diverse Intents: Optimizing Search Result Presentation
abstract
We propose a method of optimizing search result presentation for queries with diverse intents, by selectively presenting query suggestions for leading users to more relevant search results. The optimization is based on a probabilistic model of users who click on query suggestions in accordance with their intents, and modified versions of intent-aware evaluation metrics that take into account the co-occurrence between intents. Showing many query suggestions simply increases a chance to satisfy users with diverse intents in this model, while it in fact requires users to spend additional time for scanning and selecting suggestions, and may result in low satisfaction for some users. Therefore, we measured the loss of time caused by query suggestion presentation by conducting a user study in different settings, and included its negative effects in our optimization problem. Our experiments revealed that the optimization of search result presentation significantly improved that of a single ranked list, and was beneficial especially for patient users. Moreover, experimental results showed that our optimization was effective particularly when intents of a query often co-occur with a small subset of intents.
Makoto P. Kato, Katsumi Tanaka
WSDM2
2016 Detecting Evolution of Concepts based on Cause-Effect Relationships in Online Reviews
abstract
Analyzing how technology evolves is important for understanding technological progress and its impact on society. Although the concept of evolution has been explored in many domains (e.g., evolution of topics, events or terminology, evolution of species), little research has been done on automatically analyzing the evolution of products and technology in general. In this paper, we propose a novel approach for investigating the technology evolution based on collections of product reviews. We are particularly interested in understanding social impact of technology and in discovering how changes of product features influence changes in our social lives. We address this challenge by first distinguishing two kinds of product-related terms: physical product features and terms describing situations when products are used. We then detect changes in both types of terms over time by tracking fluctuations in their popularity and usage. Finally, we discover cases when changes of physical product features trigger the changes in product's use. We experimentally demonstrate the effectiveness of our approach on the Amazon Product Review Dataset that spans over 18 years.
Adam Jatowt, Katsumi Tanaka
WWW3
2016 UniClip: Leveraging Web Search for Universal Clipping of Articles on Mobile
abstract
In this paper we address the difficulty of clipping articles from mobile apps. We propose a service called UniClip that allows a user to save the full content of an article by snapping a screenshot part of it. UniClip leverages a huge amount of indexed web data to mine the article by starting with a snapped screenshot. We propose approaches to solve three challenges: (1) how to represent a screenshot; (2) how to formulate effective queries for retrieving a full article; and (3) how to rank the best URL at the top from multiple search result lists. Experimental results indicate that our approach is effective in achieving as high an $$F_1$$ F 1 measure as 0.905, which outperforms the best of three baseline methods by 18 points.
Ruihua Song, Kazutoshi Umemoto, Jian-Yun Nie, Xing Xie 0001, Katsumi Tanaka, Yong Rui
Data Sci. Eng.5
2016 The Past is Not a Foreign Country: Detecting Semantically Similar Terms across Time
abstract
Numerous archives and collections of past documents have become available recently thanks to mass scale digitization and preservation efforts. Libraries, national archives, and other memory institutions have started opening up their collections to interested users. Yet, searching within such collections usually requires knowledge of appropriate keywords due to different context and language of the past. Thus, non-professional users may have difficulties with conceptualizing suitable queries, as, typically, their knowledge of the past is limited. In this paper, we propose a novel approach for the temporal correspondence detection task that requires finding terms in the past which are semantically closest to a given input present term. The approach we propose is based on vector space transformation that maps the distributed word representation in the present to the one in the past. The key problem in this approach is obtaining correct training set that could be used for a variety of diverse document collections and arbitrary time periods. To solve this problem, we propose an effective technique for automatically constructing seed pairs of terms to be used for finding the transformation. We test the performance of proposed approaches over short as well as long time frames such as 100 years. Our experiments demonstrate that the proposed methods outperform the best-performing baseline by 113 percent for the New York Times Annotated Corpus and by 28 percent for the Times Archive in MRR on average, when the query has a different literal form from its temporal counterpart.
Adam Jatowt, Sourav S. Bhowmick, Katsumi Tanaka
IEEE Trans. Knowl. Data Eng.4
2016 Causal Relationship Detection in Archival Collections of Product Reviews for Understanding Technology Evolution
abstract
Technology progress is one of the key reasons behind today's rapid changes in lifestyles. Knowing how products and objects evolve can not only help with understanding the evolutionary patterns in our society but can also provide clues on effective product design and can offer support for predicting the future. We propose a general framework for analyzing technology's impact on our lives through detecting cause--effect relationships, where causes represent changes in technology while effects are changes in social life, such as new activities or new ways of using products. We address the challenge of viewing technology evolution through the “social impact lens” by mining causal relationships from the long-term collections of product reviews. In particular, we first propose dividing vocabulary into two groups: terms describing product features (called physical terms ) and terms representing product usage (called conceptual terms ). We then search for two kinds of changes related to the appearance of terms: frequency-based and context-based changes. The former indicate periods when a word was significantly more frequently used, whereas the latter indicate periods of high change in the word's context. Based on the detected changes, we then search for causal term pairs such that the change in the physical term triggers the change in the conceptual term. We next extend our approach to finding causal relationships between word groups such as a group of words representing the same technology and causing a given conceptual change or group of words representing two different technologies that simultaneously “co-cause” a conceptual change. We conduct experiments on different product types using the Amazon Product Review Dataset, which spans 1995 to 2013, and we demonstrate that our approaches outperform state-of-the-art baselines.
Adam Jatowt, Katsumi Tanaka
ACM Trans. Inf. Syst.3
2015 Omnia Mutantur, Nihil Interit: Connecting Past with Present by Finding Corresponding Terms across Time
abstract
Yating Zhang, Adam Jatowt, Sourav Bhowmick, Katsumi Tanaka. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Adam Jatowt, Sourav S. Bhowmick, Katsumi Tanaka
ACL (1)4
2015 NWSearch 2015: International Workshop on Novel Web Search Interfaces and Systems
abstract
Held for the first time in conjunction with the ACM International Conference on Information and Knowledge Management (CIKM), NWSearch 2015 aims to bring together researchers, developers and practitioners who are interested in pushing the search boundary on the Web and exploring more novel forms of searches, interfaces, task formulations, and result organizations and presentations. In particular, the workshop seeks to identify some of the problems and challenges facing the development of such tools and interfaces and to flourish new ideas and findings that can shape or influence future research directions and developments. The workshop organizers solicited contributions that would fall within the large spectrum of human-computer interaction in one extreme and system production and development in the other extreme.
Davood Rafiei, Katsumi Tanaka
CIKM2
2015 Web page revisiting by coordinate page discovery
abstract
A recent study on information refinding reported that 44% of Web page visits and 33% of Web queries involved revisiting previously browsed pages. We propose methods for finding previously browsed pages regarded as coordinate pages of currently browsed pages. Intuitively, the notion of coordinate pages means that both of them belong to an identical class. To find the coordinate pages for given pages, we use a user's browsing and search behavior, such as her query log and tab usage, as well as link navigation. Our page revisiting methods were implemented within a Web browser, so that users can find those previously browsed pages while browsing and searching. We conducted experiments in which our methods outperformed conventional baseline methods in terms of page revisiting.
Yusuke Takeda, Hiroaki Ohshima, Katsumi Tanaka
iiWAS3
2015 Sentential query rewriting via mutual reinforcement of paraphrase-coordinate relationships
abstract
The effectiveness of retrieval decreases with the increase in query length. We target at sentential queries and propose a method for improving their retrieval performance, called query rewriting. Briefly, given a sentential query, our method acquires paraphrases from the noisy Web and uses them to avoid returning no answers. In particular, since a relation can be represented either intensionally (referred to as paraphrase templates) or extensionally (referred to as coordinate tuples), the mutual reinforcement between them are taken into account. The experimental results show that for declarative sentences, the average precision of our method is 68.1%, compared to 44.2% of the baseline. Besides, the relative recall of our method is 95.9%, nearly 3 times compared to that of the baseline. While for questions, the average precision of our method is 46.9%, compared to 39.9% of the baseline. We also show the effectiveness of query rewriting in two applications.
Hiroaki Ohshima, Katsumi Tanaka
iiWAS3
2015 Generic method for detecting focus time of documents
Adam Jatowt, Ching-man Au Yeung, Katsumi Tanaka
Inf. Process. Manag.3
2014 Re-call and Re-cognition in Episode Re-retrieval: A User Study on News Re-finding a Fortnight Later
abstract
This study investigates recall and recognition in a news refinding task where participants were asked to read news articles and then to search for the same articles a fortnight later. Recall, which is a task to express what a person remembers, corresponds to query formulations, while recognition, which is a task to judge whether a presented item has been shown before, corresponds to a user's relevance judgment on search results in a refinding task. Our four main contributions can be summarized as follows: (i) we developed a method to investigate the effects of memory loss on episode refinding tasks on a large scale; (ii) our user study revealed a big drop on search performances in the refinding task after a fortnight and several differences between search queries input immediately after news browsing and ones at a later time; (iii) we found that asking questions and expanding input queries on the basis of the answers significantly improved the search performance in the news refinding task; and (iv) the users' recognition abilities were different than their recall abilities, e.g. object names in a news story could be correctly recognized even though they were rarely recalled. Our findings support several findings in cognitive psychology from the viewpoint of information refinding and also have several implications for search algorithms for assisting user refinding.
Shuya Ochiai, Makoto P. Kato, Katsumi Tanaka
CIKM3
2014 Learning to Generate Coherent Summary with Discriminative Hidden Semi-Markov Model
Hitoshi Nishikawa, Kazuho Arita, Katsumi Tanaka, Tsutomu Hirao, Toshiro Makino, Yoshihiro Matsuo
COLING3
2014 Investigating users' query formulations for cognitive search intents
abstract
This study investigated query formulations by users with {\it Cognitive Search Intents} (CSIs), which are users' needs for the cognitive characteristics of documents to be retrieved, {\em e.g. comprehensibility, subjectivity, and concreteness. Our four main contributions are summarized as follows (i) we proposed an example-based method of specifying search intents to observe query formulations by users without biasing them by presenting a verbalized task description;(ii) we conducted a questionnaire-based user study and found that about half our subjects did not input any keywords representing CSIs, even though they were conscious of CSIs;(iii) our user study also revealed that over 50\% of subjects occasionally had experiences with searches with CSIs while our evaluations demonstrated that the performance of a current Web search engine was much lower when we not only considered users' topical search intents but also CSIs; and (iv) we demonstrated that a machine-learning-based query expansion could improve the performances for some types of CSIs.Our findings suggest users over-adapt to current Web search engines,and create opportunities to estimate CSIs with non-verbal user input.
Makoto P. Kato, Takehiro Yamamoto, Hiroaki Ohshima, Katsumi Tanaka
SIGIR4
2014 Finding Photo Sets of Events by Minimizing Misrecognition from Neighbor Events
Bei Liu 0001, Makoto P. Kato, Katsumi Tanaka
WAIM3
2014 A Query Suggestion Interface with Features of Queries and Search Results
Shuhei Shogen, Takehiro Yamamoto, Katsumi Tanaka
WAIM3
2013 Estimating document focus time
abstract
Temporality is an important characteristic of text documents. While some documents are clearly atemporal, many have temporal character and can be mapped to certain time periods. In this paper, we introduce the problem of estimating focus time of documents. Document focus time is defined as the time to which the content of a document refers to and is considered as a complementary dimension to its creation time or timestamp. We propose several estimators of focus time by utilizing external knowledge bases such as news article collections which contain explicit temporal references. We then evaluate the effectiveness of our methods on diverse datasets of documents about historical events in five countries.
Adam Jatowt, Ching-man Au Yeung, Katsumi Tanaka
CIKM3
2013 Can We Predict User Intents from Queries?
Katsumi Tanaka
DASFAA (1)1
2013 Kcanvas: An Application for Modern PIM - Organizing Daily Fragments of Information and Telling Story of Personal Interest
Akiko Takahashi, Katsumi Tanaka
WEBIST2
2013 Estimating content concreteness for finding comprehensible documents
abstract
Document comprehensibility is one of key factors determining document quality and, in result, user's satisfaction. Relevant web pages are of little utility if they are incomprehensible or impose too much cognitive burden on readers. Traditional measures of text difficulty focus often on syntactic factors of text such as sentence length, word length, syllable count, or they utilize fixed list of common terms. However, document comprehensibility depends on many factors, of which concreteness and the ease of concept visualization are crucial ones. In this paper, we first propose a method for predicting the concreteness of terms using SVM regression. We then extend it to calculating document concreteness level. The experimental results indicate satisfactory accuracy in estimating both term and document concreteness as well as demonstrate positive correlation between the document concreteness and comprehensibility. Our ultimate goal is to enable comprehension-driven search, which will return both relevant and comprehensible results.
Shinya Tanaka, Adam Jatowt, Makoto P. Kato, Katsumi Tanaka
WSDM4
2013 Extending information unit across media streams for improving retrieval effectiveness
Nimit Pattanasri, Sujeet Pradhan, Katsumi Tanaka
Data Knowl. Eng.3
2013 When do people use query suggestion? A query suggestion log analysis
Makoto P. Kato, Tetsuya Sakai, Katsumi Tanaka
Inf. Retr.3
2012 Search intent estimation from user's eye movements for supporting information seeking
abstract
In this paper, we propose a two-stage system using user's eye movements to accommodate the increasing demands to obtain information from the Web in an efficient way. In the first stage the system estimates a user's search intent as a set of weighted terms extracted based on the user's eye movements while browsing Web pages. Then in the second stage, the system shows relevant information to the user by using the estimated intent for re-ranking search results, suggesting intent-based queries, and emphasizing relevant parts of Web pages. The system aims to help users to efficiently obtain what they need by repeating these steps throughout the information seeking process. We proposed four types of search intent estimation methods (MLT, nMLT, DLT and nDLT) considering the relationship among intents, term frequencies and eye movements. As a result of an experiment designed for evaluating the accuracy of each method with a prototype system, we confirmed that the nMLT method works best. In addition, by analyzing the extracted intent terms for eight subjects in the experiment, we found that the system could estimate the unique search intent of each user even if they performed the same search tasks.
Kazutoshi Umemoto, Takehiro Yamamoto, Satoshi Nakamura 0002, Katsumi Tanaka
AVI4
2012 Large scale analysis of changes in english vocabulary over recent time
abstract
Recently many historical texts have become digitized and made accessible for search and browsing. As human language is subject to constant evolution, these texts pose varying challenges to current users. In this paper we report the results of large-scale studies on the usage of words and the evolution of English language vocabulary over the last two centuries to help with understanding its impact on readability and retrieval of historical documents. We perform analysis of several lexical factors which may influence accessibility and readability of historical texts based on two large scale lexical corpora: the Corpus of Historical American English and Google Books 1-gram.
Adam Jatowt, Katsumi Tanaka
CIKM2
2012 Is wikipedia too difficult?: comparative analysis of readability of wikipedia, simple wikipedia and britannica
abstract
Readability is one of key factors determining document quality and reader's satisfaction. In this paper we analyze readability of Wikipedia, which is a popular source of information for searchers about unknown topics. Although Wikipedia articles are frequently listed by search engines on top ranks, they are often too difficult for average readers searching information about difficult queries. We examine the average readability of content in Wikipedia and compare it to the one in Simple Wikipedia and Britannica. Next, we investigate readability of selected categories in Wikipedia. Apart from standard readability measures we use some new metrics based on words' popularity and their distributions across different document genres and topics.
Adam Jatowt, Katsumi Tanaka
CIKM2
2012 The wisdom of advertisers: mining subgoals via query clustering
abstract
This paper tackles the problem of mining subgoals of a given search goal from data. For example, when a searcher wants to travel to London, she may need to accomplish several subtasks such as "book flights," "book a hotel," "find good restaurants" and "decide which sightseeing spots to visit." As another example, if a searcher wants to lose weight, there may exist several alternative solutions such as "do physical exercise," "take diet pills," and "control calorie intake." In this paper, we refer to such subtasks or solutions as subgoals, and propose to utilize sponsored search data for finding subgoals of a given query by means of query clustering. Advertisements (ads) reflect advertisers' tremendous efforts in trying to match a given query with implicit user needs. Moreover, ads are usually associated with a particular action or transaction. We therefore hypothesized that they are useful for subgoal mining. To our knowledge, our work is the first to use sponsored search data for this purpose. Our experimental results show that sponsored search data is a good resource for obtaining related queries and for identifying subgoals via query clustering. In particular, our method that combines ad impressions from sponsored search data and query co-occurrences from session data outperforms a state-of-the-art query clustering method that relies on document clicks rather than ad impressions in terms of purity, NMI, Rand Index, F1-measure and subgoal recall.
Takehiro Yamamoto, Tetsuya Sakai, Mayu Iwata, Ji-Rong Wen, Katsumi Tanaka
CIKM6
2012 On-the-Fly Generation of Facets as Navigation Signs for Web Objects
Yu Kawano, Hiroaki Ohshima, Katsumi Tanaka
DASFAA (1)3
2012 Data Management Challenges and Opportunities in Cloud Computing
Kyuseok Shim, Sang Kyun Cha, Lei Chen 0002, Wook-Shin Han, Divesh Srivastava, Katsumi Tanaka, Hwanjo Yu, Xiaofang Zhou 0001
DASFAA (2)6
2012 Relative Relevance Feedback in Image Retrieval
abstract
We propose a relative relevance feedback method for image retrieval systems. Relevance feedback is an effective method to modify a user's query by selecting relevant and irrelevant items in the search result. However, users cannot always find exactly relevant items in the first few search result pages, especially when the initial query is not specified due to the lack of user's knowledge. Thus, we propose relative relevance feedback in the present paper, which allows users to select relatively relevant and irrelevant items, and modifies a query by taking into account the relativity of user's feedback. Our experimental result shows that the relative relevance feedback outperforms a conventional relevance feedback for image retrieval tasks.
Yuki Sugiyama, Makoto P. Kato, Hiroaki Ohshima, Katsumi Tanaka
ICME4
2012 Content-based retrieval for heterogeneous domains: domain adaptation by relative aggregation points
abstract
We introduce the problem of domain adaptation for content-based retrieval and propose a domain adaptation method based on relative aggregation points (RAPs). Content-based retrieval including image retrieval and spoken document retrieval enables a user to input examples as a query, and retrieves relevant data based on the similarity to the examples. However, input examples and relevant data can be dissimilar, especially when domains from which the user selects examples and from which the system retrieves data are different. In content-based geographic object retrieval, for example, suppose that a user who lives in Beijing visits Kyoto, Japan, and wants to search for relatively inexpensive restaurants serving popular local dishes by means of a content-based retrieval system. Since such restaurants in Beijing and Kyoto are dissimilar due to the difference in the average cost and areas' popular dishes, it is difficult to find relevant restaurants in Kyoto based on examples selected in Beijing. We propose a solution for this problem by assuming that RAPs in different domains correspond, which may be dissimilar but play the same role. A RAP is defined as the expectation of instances in a domain that are classified into a certain class, e.g. the most expensive restaurant, average restaurant, and restaurant serving the most popular dishes. Our proposed method constructs a new feature space based on RAPs estimated in each domain and bridges the domain difference for improving content-based retrieval in heterogeneous domains. To verify the effectiveness of our proposed method, we evaluated various methods with a test collection developed for content-based geographic object retrieval. Experimental results show that our proposed method achieved significant improvements over baseline methods. Moreover, we observed that the search performance of content-based retrieval in heterogeneous domains was significantly lower than that in homogeneous domains. This finding suggests that relevant data for the same search intent depend on the search context, that is, the location where the user searches and the domain from which the system retrieves data.
Makoto P. Kato, Hiroaki Ohshima, Katsumi Tanaka
SIGIR3
2012 Search Intent Discovery by Structurization of Community QA Contents
Soungwoong Yoon, Adam Jatowt, Katsumi Tanaka
WISE3
2012 Panoramic Image Search by Similarity and Adjacency for Similar Landscape Discovery
Hiroaki Ohshima, Katsumi Tanaka
WISE3
2012 Structured query suggestion for specialization and parallel movement: effect on search behaviors
abstract
Query suggestion, which enables the user to revise a query with a single click, has become one of the most fundamental features of Web search engines. However, it is often difficult for the user to choose from a list of query suggestions, and to understand the relation between an input query and suggested ones. In this paper, we propose a new method to present query suggestions to the user, which has been designed to help two popular query reformulation actions, namely, specialization (e.g. from "nikon" to "nikon camera" ) and parallel movement (e.g. from "nikon camera" to "canon camera"). Using a query log collected from a popular commercial Web search engine, our prototype called SParQS classifies query suggestions into automatically generated categories and generates a label for each category. Moreover, SParQS presents some new entities as alternatives to the original query (e.g. "canon" in response to the query "nikon"), together with their query suggestions classified in the same way as the original query's suggestions. We conducted a task-based user study to compare SParQS with a traditional "flat list" query suggestion interface. Our results show that the SParQS interface enables subjects to search more successfully than the flat list case, even though query suggestions presented were exactly the same in the two interfaces. In addition, the subjects found the query suggestions more helpful when they were presented in the SParQS interface rather than in a flat list.
Makoto P. Kato, Tetsuya Sakai, Katsumi Tanaka
WWW3
2011 Enhancing credibility judgment of web search results
abstract
In this paper, we propose a system for helping users to judge the credibility of Web search results and to search for credible Web pages. Conventional Web search engines present only titles, snippets, and URLs for users, which give few clues to judge the credibility of Web search results. Moreover, ranking algorithms of the conventional Web search engines are often based on relevance and popularity of Web pages. Towards credibility-oriented Web search, our proposed system provides users with the following three functions: (1) calculation and visualization of several scores of Web search results on the main credibility aspects, (2) prediction of user's credibility judgment model through user's credibility feedback for Web search results, and (3) re-ranking of Web search results based on user's predicted credibility model. Experimental results suggest that our system enables users - in particular, users with knowledge about search topics - to find credible Web pages from a list of Web search results more efficiently than conventional Web search interfaces.
Yusuke Yamamoto, Katsumi Tanaka
CHI2
2011 RerankEverything: a reranking interface for exploring search results
abstract
This paper proposes a system called "RerankEverything", which enables users to rerank search results in any search service, such as a Web search engine, an e-commerce site, a hotel reservation site, and so on. This system helps users explore diverse search results. In conventional search services, interactions between users and systems are quite limited and complicated. By using RerankEverything, users can interactively explore search results in accordance with their interests by reranking search results from various viewpoints. Experimental results show that our system potentially help users search more proactively. When using our system, users were more likely to click search results that were initially low ranked. Users also browsed through more diverse search results by reranking search results after giving various types of feedback with our system.
Takehiro Yamamoto, Satoshi Nakamura 0002, Katsumi Tanaka
CIKM3
2011 Extracting adjective facets from community Q&A corpus
abstract
In this paper, we propose a method for helping users explore information via Web searches by using a question and answer (Q&A) corpus archived in a community Q&A site. When users do not have clear information needs and have little knowledge about the task domain, it is difficult for them to create queries that adequately reflect their information needs. We focused on terms like "famous temples," "historical townscapes," and "delicious sweets," which we call "adjective facets", and developed a method of extracting these facets from question and answer archives at a community Q&A site. We evaluated the effectiveness of our adjective facets by comparing them with several baselines.
Takehiro Yamamoto, Satoshi Nakamura 0002, Katsumi Tanaka
CIKM3
2011 Towards Web Search by Sentence Queries: Asking the Web for Query Substitutions
Yusuke Yamamoto, Katsumi Tanaka
DASFAA (2)2
2011 Structural Approach to Service Composition Based on Relational Model
abstract
We propose a structural approach to service composition using a relational model. We focus on structural aspects of service composition and apply the relational model. Services in our model are defined and organized by relations, and service manipulations such as service invoke and service composition, are achieved by relational operations to the corresponding relations. Since conventional relational algebra is insufficient for service management, we extend relational algebra and introduce a new θ-operator to deal with services.
Yuhei Akahoshi, Koji Zettsu, Yutaka Kidawara, Katsumi Tanaka
EJC4
2011 Node-First Causal Network Extraction for Trend Analysis Based on Web Mining
Hideki Kawai, Katsumi Tanaka, Kazuo Kunieda, Keiji Yamada
KES (2)2
2011 Extraction and Geographical Navigation of Important Historical Events in the Web
Mitsuo Yamamoto, Yuku Takahashi, Hirotoshi Iwasaki, Satoshi Oyama, Hiroaki Ohshima, Katsumi Tanaka
W2GIS6
2011 Measuring Comprehensibility of Web Pages Based on Link Analysis
abstract
We put forward a hypothesis that if there is a link from one page to another, it is likely that comprehensibility of the two pages is similar. To investigate whether this hypothesis is true or not, we conduct experiments using existing readability measures. We investigate the relationship between links and readability of text extracted from web pages for two datasets, set of English and Japanese pages. We could find that links and readability of text extracted from web pages are correlated. Based on the hypothesis, we propose a link analysis algorithm to measure comprehensibility of web pages. Our method is based on the Trust Rank algorithm which is originally used for combating web spam. We use link structure to propagate readability scores from source pages selected based on their comprehensibility. The results of experimental evaluation demonstrate that our method could improve estimation of comprehensibility of pages.
Kouichi Akamatsu, Nimit Pattanasri, Adam Jatowt, Katsumi Tanaka
Web Intelligence4
2011 Improving Retrieval of Future-Related Information in Text Collections
abstract
People often want to know expected future events related to given real world entities. For supporting users in the process of future scenario analysis, we propose several methods that enable to retrieve and analyze future-related opinions from large text collections. In particular, we focus on time-unreferenced predictions, which do not contain any explicit future time reference and hence are more difficult to be retrieved. As a second contribution, we propose estimating validity of predictions by automatically searching for real world events corresponding to the predictions. This kind of analysis aims to help detect predictions that are no longer valid as well as help estimating prediction accuracy of information sources.
Kensuke Kanazawa, Adam Jatowt, Katsumi Tanaka
Web Intelligence3
2011 Detecting Intent of Web Queries Using Questions and Answers in CQA Corpus
abstract
Detecting intent in Web search activity is important task for finding relevant Web information. However extracting intents from users' queries is difficult as users express their intent by issuing short and often ambiguous queries, yet at the same time it is crucial factor for enhancing user satisfaction. Showing the variety of candidate intents behind a query could help users choose correct intent expressions and improve the Web search. In this paper, we propose the methodology for detecting intent of Web queries using Community Question-Answer (CQA)information. Our assumption is that questions and its answers in CQA corpus reflect intents of questioners. To detect these intents, we use the semantic connections between questions and its answers. We categorize questions to find the connections of features within a question and its answers, detect intent words in answers by calculating supports of concerned CQA contents, and cluster questions and their answers by these intent words. Experimental results show that the variety of Web query intents can be found with satisfactory performance.
Soungwoong Yoon, Adam Jatowt, Katsumi Tanaka
Web Intelligence3
2011 Supporting Sharing of Browsing Information and Search Results in Mobile Collaborative Searches
Daisuke Kotani, Satoshi Nakamura 0002, Katsumi Tanaka
WISE3
2010 Search as if you were in your home town: geographic search by regional context and dynamic feature-space selection
abstract
We propose a query-by-example geographic object search method for users that do not know well about the place they are in. Geographic objects, such as restaurants, are often retrieved using an attribute-based or keyword query. These queries, however, are difficult to use for users that have little knowledge on the place where they want to search. The proposed query-by-example method allows users to query by selecting examples in familiar places for retrieving objects in unfamiliar places. One of the challenges is to predict an effective distance metric, which varies for individuals. Another challenge is to calculate the distance between objects in heterogeneous domains considering the feature gap between them, for example, restaurants in Japan and China. Our proposed method is used to robustly estimate the distance metric by amplifying the difference between selected and non-selected examples. By using the distance metric, each object in a familiar domain is evenly assigned to one in an unfamiliar domain to eliminate the difference between those domains. We developed a restaurant search using data obtained from a Japanese restaurant Web guide to evaluate our method.
Makoto P. Kato, Hiroaki Ohshima, Satoshi Oyama, Katsumi Tanaka
CIKM4
2010 Cloud as Virtual Databases: Bridging Private Databases and Web Services
Hiroaki Ohshima, Satoshi Oyama, Katsumi Tanaka
DASFAA (1)3
2010 Plus One or Minus One: A Method to Browse from an Object to Another Object by Adding or Deleting an Element
Kosetsu Tsukuda, Takehiro Yamamoto, Satoshi Nakamura 0002, Katsumi Tanaka
DEXA (2)4
2010 Evaluating Truthfulness of Modifiers Attached to Web Entity Names
Ryohei Takahashi, Satoshi Oyama, Hiroaki Ohshima, Katsumi Tanaka
WAIM4
2010 Searching the Web for Alternative Answers to Questions on WebQA Sites
Natsuki Takata, Hiroaki Ohshima, Satoshi Oyama, Katsumi Tanaka
WAIM4
2010 Web Information Credibility
Katsumi Tanaka
WAIM1
2010 Estimating News Coverage of Web Search Results
abstract
The abundance of content on the web and the lack of quality control require more refined approaches in analyzing online information. In this paper, we propose evaluating the extent to which web search results cover important and recent news related to real-world objects. Our method allows for identifying search results that provide comprehensive overviews of major events related to user queries or that contain most recent information.
Adam Jatowt, Yukiko Kawai, Katsumi Tanaka
Web Intelligence3
2010 Analyzing collective view of future, time-referenced events on the web
abstract
Humans have always desired to guess the future in order to adapt their behavior and maximize chances of success. In this paper, we conduct exploratory analysis of future-related information on the web. We focus on the future-related information which is grounded in time, that is, the information on forthcoming events whose expected occurrence dates are already known. We collect data by crawling search engine index and analyze collective view of future time-referenced events discussed on the web.
Adam Jatowt, Hideki Kawai, Kensuke Kanazawa, Katsumi Tanaka, Kazuo Kunieda, Keiji Yamada
WWW4
2010 RerankEverything: a reranking interface for browsing search results
abstract
This paper proposes a system called RerankEverything, which enables users to rerank search results in any search service, such as a Web search engine, an e-commerce site, a hotel reservation site and so on. In conventional search services, interactions between users and services are quite limited and complicated. In addition, search functions and interactions to refine search results differ depending on the services. By using RerankEverything, users can interactively explore search results in accordance with their interests by reranking search results from various viewpoints.
Takehiro Yamamoto, Satoshi Nakamura 0002, Katsumi Tanaka
WWW3
2009 Query by analogical example: relational search using web search engine indices
abstract
We describe methods to search with a query by example in a known domain for information in an unknown domain by exploiting Web search engines. Relational search is an effective way to obtain information in an unknown field for users. For example, if an Apple user searches for Microsoft products, similar Apple products are important clues for the search. Even if the user does not know keywords to search for specific Microsoft products, the relational search returns a product name by querying simply an example of Apple products. More specifically, given a tuple containing three terms, such as (Apple, iPod, Microsoft), the term Zune can be extracted from the Web search results, where Apple is to iPod what Microsoft is to Zune. As a previously proposed relational search requires a huge text corpus to be downloaded from the Web, the results are not up-to-date and the corpus has a high construction cost. We introduce methods for relational search by using Web search indices. We consider methods based on term co-occurrence, on lexico-syntactic patterns, and on combinations of the two approaches. Our experimental results showed that the combination methods got the highest precision, and clarified the characteristics of the methods.
Makoto P. Kato, Hiroaki Ohshima, Satoshi Oyama, Katsumi Tanaka
CIKM4
2009 Easiest-first search: towards comprehension-based web search
abstract
Although Web search engines have become information gateways to the Internet, for queries containing technical terms, search results often contain pages that are difficult to be understood by non-expert users. Therefore, re-ranking search results in a descending order of their comprehensibility should be effective for non-expert users. In our approach, the comprehensibility of Web pages is estimated considering both the document readability and the difficulty of technical terms in the domain of search queries. To extract technical terms, we exploit the domain knowledge extracted from Wikipedia. Our proposed method can be applied to general Web search engines as Wikipedia includes nearly every field of human knowledge. We demonstrate the usefulness of our approach by user experiments.
Makoto Nakatani, Adam Jatowt, Katsumi Tanaka
CIKM3
2009 Quality Evaluation of Search Results by Typicality and Speciality of Terms Extracted from Wikipedia
Makoto Nakatani, Adam Jatowt, Hiroaki Ohshima, Katsumi Tanaka
DASFAA4
2009 Reranking and Classifying Search Results Exhaustively Based on Edit-and-Propagate Operations
Takehiro Yamamoto, Satoshi Nakamura 0002, Katsumi Tanaka
DEXA3
2009 Video Search by Impression Extracted from Social Annotation
Satoshi Nakamura 0002, Katsumi Tanaka
WISE2
2009 Towards Improving Web Search: A Large-Scale Exploratory Study of Selected Aspects of User Search Behavior
Hiroaki Ohshima, Adam Jatowt, Satoshi Oyama, Satoshi Nakamura 0002, Katsumi Tanaka
WISE5
2009 Seeing Past Rivals: Visualizing Evolution of Coordinate Terms over Time
Hiroaki Ohshima, Adam Jatowt, Satoshi Oyama, Katsumi Tanaka
WISE4
2009 TermCloud for Enhancing Web Search
Takehiro Yamamoto, Satoshi Nakamura 0002, Katsumi Tanaka
WISE3
2009 Finding Comparative Facts and Aspects for Judging the Credibility of Uncertain Facts
Yusuke Yamamoto, Katsumi Tanaka
WISE2
2009 Intent-Based Categorization of Search Results Using Questions from Web Q&A Corpus
Soungwoong Yoon, Adam Jatowt, Katsumi Tanaka
WISE3
2009 A Novel Visualization Method for Distinction of Web News Sentiment
Jianwei Zhang 0002, Yukiko Kawai, Tadahiko Kumamoto, Katsumi Tanaka
WISE4
2008 Mining the Web for Hyponymy Relations Based on Property Inheritance
Shun Hattori, Hiroaki Ohshima, Satoshi Oyama, Katsumi Tanaka
APWeb4
2008 How Many Objects?: Determining the Number of Clusters with a Skewed Distribution
abstract
We propose a supervised approach to enable accurate determination of the number of clusters in object identification. We use the aggregated attribute values of the data set to be clustered as explanatory variables in the prediction model. Attribute aggregation can be done in linear time with respect to the number of data items, so our method can be used to predict the number of clusters with a low computational burden. To deal with skewed target values, we introduce a two-stage method as well as a method using a higher-order combination of explanatory variables. Experiments demonstrate our methods enable more accurate prediction than existing methods.
Satoshi Oyama, Katsumi Tanaka
ECAI2
2008 Automatic Generation of Computer Animation Conveying Impressions of News Articles
Tadahiko Kumamoto, Akiyo Nadamoto, Katsumi Tanaka
KES (1)3
2008 Estimation of Geographic Relevance for Web Objects Using Probabilistic Models
Taro Tezuka, Hiroyuki Kondo, Katsumi Tanaka
W2GIS3
2008 Extracting Concept Hierarchy Knowledge from the Web Based on Property Inheritance and Aggregation
abstract
Concept hierarchy knowledge, such as hyponymy and meronymy, is very important for various natural language processing systems. While WordNet and Wikipedia are being manually constructed and maintained as lexical ontologies, many researchers have tackled how to extract concept hierarchies from very large corpora of text documents such as the Web not manually but automatically. However, their methods are mostly based on lexico-syntactic patterns as not necessary but sufficient conditions of hyponymy and meronymy, so they can achieve high precision but low recall when using stricter patterns or they can achieve high recall but low precision when using looser patterns. Therefore, we need necessary conditions of hyponymy and meronymy to achieve high recall and not low precision. In this paper, not only "Property Inheritance'' from a target concept to its hyponyms but also "Property Aggregation'' from its hyponyms to the target concept is assumed to be necessary and sufficient conditions of hyponymy, and we propose a method to extract concept hierarchy knowledge from the Web based on property inheritance and property aggregation.
Shun Hattori, Katsumi Tanaka
Web Intelligence2
2008 Unsupervised Discovery of Coordinate Terms for Multiple Aspects from Search Engine Query Logs
abstract
A method is described for discovering coordinate terms, such as "Honda'' and "Nissan,'' for a given term, such as "Toyota,'' as well as their common topic terms, from the query logs of a Web search engine. Coordinate terms are good candidates for use in making comparisons. A HITS-based algorithm is applied to a bipartite graph between coordinate term candidates and co-occurrence patterns to identify coordinate and topic terms. Spectral analysis is used to distinguish coordinate terms corresponding to different aspects of the search term. As a result, we can discover terms related to the terms in a search engine query that reflect the needs and interests of the user.
Masashi Yamaguchi, Hiroaki Ohshima, Satoshi Oyama, Katsumi Tanaka
Web Intelligence4
2008 Can Social Tagging Improve Web Image Search?
Makoto P. Kato, Hiroaki Ohshima, Satoshi Oyama, Katsumi Tanaka
WISE4
2008 SyncRerank: Reranking Multi Search Results Based on Vertical and Horizontal Propagation of User Intention
Satoshi Nakamura 0002, Takehiro Yamamoto, Katsumi Tanaka
WISE3
2008 Supporting Judgment of Fact Trustworthiness Considering Temporal and Sentimental Aspects
Yusuke Yamamoto, Taro Tezuka, Adam Jatowt, Katsumi Tanaka
WISE4
2008 Visualizing historical content of web pages
abstract
Recently, along with the rapid growth of the Web, the preservation efforts have also increased. As a consequence, large amounts of past Web data are stored in Web archives. This historical data can be used for better understanding of long-term page topics and characteristics. In this paper, we propose an interactive visualization system called Page History Explorer for exploring page histories. It allows for roughly portraying evolution of pages and summarizing their content over time. We use a temporal term cloud as a structure for visualizing prevailing and active terms appearing on pages in the past.
Adam Jatowt, Yukiko Kawai, Katsumi Tanaka
WWW3
2007 Mining the Web for Appearance Description
Shun Hattori, Taro Tezuka, Katsumi Tanaka
DEXA3
2007 Rerank-by-Example: Efficient Browsing of Web Search Results
Takehiro Yamamoto, Satoshi Nakamura 0002, Katsumi Tanaka
DEXA3
2007 Secure Spaces: Protecting Freedom of Information Access in Public Places
Shun Hattori, Katsumi Tanaka
ICOST2
2007 Creating Personal Histories from the Web Using Namesake Disambiguation and Event Extraction
Rui Kimura, Satoshi Oyama, Hiroyuki Toda, Katsumi Tanaka
ICWE4
2007 Towards Improving Web Search by Utilizing Social Bookmarks
Yusuke Yanbe, Adam Jatowt, Satoshi Nakamura 0002, Katsumi Tanaka
ICWE4
2007 Towards New Content Services by Fusion of Web and Broadcasting Contents
Katsumi Tanaka
IEA/AIE1
2007 Temporal filtering system to reduce the risk of spoiling a user's enjoyment
abstract
This paper proposes a temporal filtering system called the Anti-Spoiler system. The system changes filters dynamically based on user-specified preferences and the user's timetable. The system then blocks contents that would spoil the user's enjoyment of a previously unwatched content. The system analyzes a user-requested Web content, and then uses filters to prevent portions of the content being displayed that might spoil user's enjoyment. For example, the system hides the final score of football from the Web content before watching it on TV.
Satoshi Nakamura 0002, Katsumi Tanaka
IUI2
2007 Modeling Omni-Directional Video
Shumian He, Katsumi Tanaka
MMM (1)2
2007 Automatic Generation of Multimedia Tour Guide from Local Blogs
Hiroshi Kori, Shun Hattori, Taro Tezuka, Katsumi Tanaka
MMM (1)4
2007 Enhancing Comprehension of Events in Video Through Explanation-on-Demand Hypervideo
Nimit Pattanasri, Adam Jatowt, Katsumi Tanaka
MMM (1)3
2007 Presentation of Dynamic Maps by Estimating User Intentions from Operation History
Taro Tezuka, Katsumi Tanaka
MMM (1)2
2007 WeBrowSearch: Toward Web Browser with Autonomous Search
Taiga Yoshida, Satoshi Nakamura 0002, Katsumi Tanaka
WISE3
2006 Using Web Archive for Improving Search Engine Results
Adam Jatowt, Yukiko Kawai, Katsumi Tanaka
APWeb3
2006 Context Matcher: Improved Web Search Using Query Term Context in Source Document and in Search Results
Takahiro Kawashige, Satoshi Oyama, Hiroaki Ohshima, Katsumi Tanaka
APWeb4
2006 Identifying Agitators as Important Blogger Based on Analyzing Blog Threads
Shinsuke Nakajima, Jun'ichi Tatemura, Yoshinori Hara, Katsumi Tanaka, Shunsuke Uemura
APWeb4
2006 Extracting Semantic Relationships Between Terms from PC Documents and Its Applications to Web Search Personalization
Hiroaki Ohshima, Satoshi Oyama, Katsumi Tanaka
APWeb3
2006 Visual Description Conversion for Enhancing Search Engines and Navigational Systems
Taro Tezuka, Katsumi Tanaka
APWeb2
2006 Automated Content Transformation with Adjustment for Visual Presentation Related to Terminal Types
Hiromi Uwada, Akiyo Nadamoto, Tadahiko Kumamoto, Toru Hamabe, Makoto Yokozawa, Katsumi Tanaka
APWeb6
2006 Personalized Detection of Fresh Content and Temporal Annotation for Improved Page Revisiting
Adam Jatowt, Yukiko Kawai, Katsumi Tanaka
DEXA3
2006 User Preference Modeling Based on Interest and Impressions for News Portal Site Systems
Yukiko Kawai, Tadahiko Kumamoto, Katsumi Tanaka
DEXA3
2006 Mining and Visualizing Local Experiences from Blog Entries
Takeshi Kurashima, Taro Tezuka, Katsumi Tanaka
DEXA3
2006 WebDriving: Web Browsing Based on a Driving Metaphor for Improved Children's e-Learning
Mika Nakaoka, Taro Tezuka, Katsumi Tanaka
DEXA3
2006 Improving Web Retrieval Precision Based on Semantic Relationships and Proximity of Query Keywords
Chi Tian, Taro Tezuka, Satoshi Oyama, Keishi Tajima, Katsumi Tanaka
DEXA5
2006 Mining Text and Visual Links to Browse TV Programs in a Web-Like Way
abstract
As the amount of receded TV content is increasing rapidly, people need active and interactive browsing methods. In this paper, we use both text information from closed captions and visual information from video frames to generate links to enable one to explore not only the original video content but also augmented information from the Web. This solution especially shows its superiority when the video content cannot be well represented only by closed captions. A prototype system was implemented and some experiments were carried out to prove the effectiveness and efficiency
Xin Fan 0001, Hisashi Miyamori, Katsumi Tanaka, Mingjing Li
ICME3
2006 Query Modification Based on Real-World Contexts for Mobile and Ubiquitous Computing Environments
abstract
With the growing amount of information on the WWW and the improvement of mobile computing environments, mobile Web search engines will increase significance or more in the future. Because mobile devices have the restriction of output performance and we have little time for browsing information slowly while moving or doing some activities in the real world, it is necessary to refine the retrieval results in mobile computing environments better than in fixed ones. However, since a mobile user’s query is often shorter and more ambiguous than a fixed user’s query which is not enough to guess his/her information demand accurately, too many results might be retrieved by commonly used Web search engines. This paper proposes two novel methods for query modification based on real-world contexts of a mobile user, such as his/her geographic location and the objects surrounding him/her, aiming to enhance location-awareness, and moreover, context-awareness, to the existing location-free information retrieval systems.
Shun Hattori, Taro Tezuka, Katsumi Tanaka
MDM3
2006 Content-Based Entry Control for Secure Spaces
abstract
We define "Secure Space" as a physical space in which any resource is always protected from its unauthorized users in terms of enforcing its authorization policies assuredly. Aiming to build such secure spaces, this paper proposes an architecture and a model for space entry control based on its dynamically changing contents, such as users, physical resources and virtual resources outputted by some embedded devices. We first describe the architecture and then formalize the model and mechanism for secure spaces.
Shun Hattori, Taro Tezuka, Katsumi Tanaka
MDM3
2006 u-Cam: A User-Driven Control Mechanism for Ubiquitous Cameras and Its Content Management
abstract
We developed a mechanism for photographing people and annotating their behavior along with nearby elements using multiple embedded cameras. In addition, we developed a method of dynamically integrating and presenting the recorded content. As conventional cameras are used to photograph objects selected by users, taking pictures that include users while holding the camera is difficult. Security camera systems take pictures that include people and nearby elements, but such systems cannot show the intentions of the people being photographed. After detecting the intentions and behaviors of subjects with radio frequency identification (RFID) tags, our system selects the best camera from cameras located in an area, and then the camera photographs the subject and the surrounding area. In addition, these photos can be annotated with information about the context and movement history of the subject. We created a prototype of our system and determined its effectiveness experimentally.
Shumian He, Yukiko Kawai, Yutaka Kidawara, Koji Zettsu, Katsumi Tanaka
MDM5
2006 u-Cam: Ubiquitous Camera in Real World with User-Driven Control
abstract
We developed a mechanism for photographing people and annotating their behavior along with nearby elements using multiple embedded cameras. In addition, we developed a method of dynamically integrating and presenting the recorded content. As a conventional camera is used to photograph objects selected by a photographer, taking pictures that include him while holding the camera is difficult. Security camera systems take pictures that include people and nearby elements, but such systems cannot take the user aspect of intentions and interest of the people. After detecting the intentions and behaviors of subjects using radio frequency identification (RFID) tags, our system selects the best camera from cameras located in an area, and then the camera photographs the subject and the surrounding area. In addition, these photos can be annotated with meta datainformation about the context and interest of the subject.
Shumian He, Yukiko Kawai, Koji Zettsu, Katsumi Tanaka
MDM4
2006 Cooperative Device Browsing through Portable Private Area Network
abstract
In next-generation networking environments, ubiquitous networks will be available both indoors and outdoors. Various devices will be ubiquitously embedded in our homes and cityscape. Digital content will be stored not only by servers on the Internet, but also in embedded devices belonging to ubiquitous networks. In this paper, we propose a content-processing mechanism for use in environment-enabling collaborative acquisition of embedded digital content in real-world situations. We have developed a network management device that makes it possible to acquire embedded content using coordinated ubiquitous devices. This management device actively configures networks that include content-providing devices and browsing devices to enable sharing of a range of digital content. To demonstrate our system, we built a practical prototype called the "Virtual Insect Catching System", which is simple enough for children to use. In a test that 48 children took part in we demonstrated that the system can be used to find embedded devices, build peer-topeer networks, acquire embedded digital content, retrieve related content from the Internet, and then create new web content.
Yutaka Kidawara, Katsumi Tanaka
MDM2
2006 u-PaV: Automatic Transformation of Web Content into TV-like Video Content for Ubiquitous Environment
abstract
We propose a system that automatically transforms web content into TV-like video content for ubiquitous environments. We call this system the ubiquitous/universal passive viewer (u-PaV). The u-PaV consists mainly of audio and visual components. The audio component uses synthesized speech to read out titles and lines extracted from the target web content. Simultaneously, the visual component of the u- PaV presents the titles and lines to a user display through a ticker. Keywords and images extracted from the web content are animated on the display. A suitable background color is determined based on the overall impression of the content. The u-PaV synchronizes the ticker, animation, and speech. We introduce the u-PaV and explain how the keywords are extracted and how the impression value of the web content is determined. A test with 50 users showed that the u-PaV is easier to use and understand than browsing web content alone.
Akiyo Nadamoto, Tadahiko Kumamoto, Hiromi Uwada, Toru Hamabe, Makoto Yokozawa, Katsumi Tanaka
MDM6
2006 Learning a Distance Metric for Object Identification Without Human Supervision
Satoshi Oyama, Katsumi Tanaka
PKDD2
2006 Buzz Network on the Weblog Community
Shinya Takami, Katsumi Tanaka
WEBIST (2)2
2006 Web Driving: An Image-Based Opportunistic Web Browser That Visualizes a Peripheral Information Space
Mika Nakaoka, Taro Tezuka, Katsumi Tanaka
WISE3
2006 Searching Coordinate Terms with Their Context from the Web
Hiroaki Ohshima, Satoshi Oyama, Katsumi Tanaka
WISE3
2006 Towards Next-Generation Search Engines and Browsers - Search Beyond Media Types and Places
Katsumi Tanaka
WISE1
2006 A browser for browsing the past web
abstract
We describe a browser for the past web. It can retrieve data from multiple past web resources and features a passive browsing style based on change detection and presentation. The browser shows past pages one by one along a time line. The parts that were changed between consecutive page versions are animated to reflect their deletion or insertion, thereby drawing the user's attention to them. The browser enables automatic skipping of changeless periods and filtered browsing based on user specified query.
Adam Jatowt, Yukiko Kawai, Satoshi Nakamura 0002, Yutaka Kidawara, Katsumi Tanaka
WWW5
2006 Proposal of integrated search engine of web and TV contents
abstract
A search engine that can handle TV programs and Web content in an integrated way is proposed. Conventional search engines have been able to handle Web content and/or data stored in a PC desktop as target information. In the future, however, the target information is expected to be stored in various places such as in hard-disk (HD)/DVD recorders, digital cameras, mobile devices, and even in real space as ubiquitous content, and a search engine that can search across such heterogeneous resources will become essential. Therefore, as a first step towards developing such next-generation search engine, a prototype search system for Web and TV programs is developed that performs integrated search of those content, and that allows chain search where related content can be accessed from each search result. The integrated search is achieved by generating integrated indices for Web and TV content based on vector space model and by computing similarity between the query and all the content described by the indices. The chain search of related content is done by computing similarity between the selected result and all other content based on the integrated indices. Also, the zoom-based display of the search results enables to control media transition and level of details of the contents to acquire information efficiently. In this paper, testing of a prototype of the integrated search engine validated the approach taken by the proposed method.
Hisashi Miyamori, Mitsuru Minakuchi, Zoran Stejic, Qiang Ma 0001, Tadashi Araki, Katsumi Tanaka
WWW6
2006 Toward tighter integration of web search with a geographic information system
abstract
Integration of Web search with geographic information has recently attracted much attention. There are a number of local Web search systems enabling users to find location-specific Web content. In this paper, however, we point out that this integration is still at a superficial level. Most local Web search systems today only link local Web content to a map interface. They are extensions of a conventional stand-alone geographic information system (GIS), applied to a Web-based client-server architecture. In this paper, we discuss the directions available for tighter integration of Web search with a GIS, in terms of extraction, knowledge discovery, and presentation. We also describe implementations to support our argument that the integration must go beyond the simple map-and hyperlink architecture.
Taro Tezuka, Takeshi Kurashima, Katsumi Tanaka
WWW3
2006 Complementary information retrieval for cross-media news content
Qiang Ma 0001, Akiyo Nadamoto, Katsumi Tanaka
Inf. Syst.3
2005 Web opinion poll: extracting people's view by impression mining from the web
abstract
No abstract available.
Tadahiko Kumamoto, Katsumi Tanaka
CIKM2
2005 Landmark Extraction: A Web Mining Approach
Taro Tezuka, Katsumi Tanaka
COSIT2
2005 Zooming Cross-Media: A Zooming Description Language Coding LOD Control and Media Transition
Tadashi Araki, Hisashi Miyamori, Mitsuru Minakuchi, Ai Kato, Zoran Stejic, Yasushi Ogawa, Katsumi Tanaka
DEXA7
2005 My Portal Viewer: Integration System Based on User Preferences for News Web Sites
Yukiko Kawai, Daisuke Kanjo, Katsumi Tanaka
DEXA3
2005 Context-Sensitive Complementary Information Retrieval for Text Stream
Qiang Ma 0001, Katsumi Tanaka
DEXA2
2005 Webified Video: Media Conversion from TV Programs to Web Content for Cross-Media Information Integration
Hisashi Miyamori, Katsumi Tanaka
DEXA2
2005 RelaxImage: A Cross-Media Meta-Search Engine for Searching Images from Web Based on Query Relaxation
abstract
We introduce a cross-media meta-search engine RelaxImage for searching images from Web. Notable features of the RelaxImage are as follows: (1) each user's keyword query is "relaxed", that is, by gradually relaxing the search terms used for image search, we can solve the problem of conventional image search engine such as Google. (2) For searching images, our RelaxImage sends a different keyword-query to each search engine of different media-type. We show several examples of how the relaxation approach works as well as ways that it can be applied. That is, our RelaxImage shows a great improvement for increasing recall ratio without decreasing of precision ratio.
Akihiro Kuwabara, Katsumi Tanaka
ICDE2
2005 Automatic indexing of broadcast content using its live chat on the Web
abstract
A method of automatically indexing broadcast content using live chat on the Web is proposed. The live chat is a Web bulletin board where the viewers post messages in sync with a TV program. Statistical analysis and pattern recognition of these messages can effectively extract metadata related to viewer's viewpoints such as important scenes in the program or responses by a particular viewer. Preliminary experiments indicate that the proposed method can efficiently extract metadata such as the intensity of viewers' responses and degree of emotional delight or depression. They also indicate that a prototype TV viewing system using the extracted metadata enables a new way of viewing TV content from different perspectives reflecting viewers' viewpoints.
Hishasi Miyamori, Satoshi Nakamura 0002, Katsumi Tanaka
ICIP (3)3
2005 WA-TV: Webifying and Augmenting Broadcast Content for Next-Generation Storage TV
abstract
A method is proposed for viewing broadcast content that converts TV programs into Web content and integrates the results with complementary information retrieved using the Internet. Converting the programs into Web pages enables the programs to be skimmed over to get an overview and for particular scenes to be easily explored. Integrating complementary information enables the programs to be viewed efficiently with value-added content. An intuitive, user-friendly browsing interface enables the user to easily changing the level of detail displayed for the integrated information by zooming. Preliminary testing of a prototype system for next generation storage TV, “ WA-TV”, validated the approach taken by the proposed method.
Hisashi Miyamori, Qiang Ma 0001, Katsumi Tanaka
ICME3
2005 Automatic Conversion from E-Content into Animated Storytelling
Kaoru Sumi, Katsumi Tanaka
ICEC2
2005 Proposal of Impression Mining from News Articles
Tadahiko Kumamoto, Katsumi Tanaka
KES (1)2
2005 A collaborative environment for enhanced information access on small-form-factor devices
abstract
This paper proposes to establish a collaborative environment to achieve enhanced information access on the small-form-factor devices. We designed a distributed user interface that crosses devices to cooperatively present information adaptable to small displays. We applied our system to the document and webpage that are ubiquitously available media content on mobile devices.
Zhigang Hua, Yutaka Kidawara, Hanqing Lu, Katsumi Tanaka
Mobile HCI4
2005 Generation of views of TV content using TV viewers' perspectives expressed in live chats on the web
abstract
We propose a method of generating views of TV programs based on viewer's perspectives expressed in live chats on the Web. Important scenes in a program and responses by particular viewers can be extracted efficiently by statistically computing and/or recognizing live chat data obtained in sync with the broadcast content. We show that by using the computed results, views can be generated that indicate the momentum of reactions by viewers and scenes of interest to particular viewers whose preferences are similar to those of the viewer, etc. This is a new way of viewing TV content from various perspectives.
Hisashi Miyamori, Satoshi Nakamura 0002, Katsumi Tanaka
ACM Multimedia3
2005 Complementing your TV-viewing by web content automatically-transformed into TV-program-type content
abstract
Despite much talk about the fusion of broadcasting and the Internet, no technology has been established for fusing web and TV program content. In this paper, we propose ways to transform web content into TV-program-type content as a first step towards the fusion of these media. Our transformation method is based on two criteria - the transmitted information and the dialogue among character agents. The method deals with both an audio component and a visual component. By combining these techniques, we can transform web content into various forms of TV-program-type content depending on the user's aims. We present three different prototype systems, u-Pav which reads out the entire text of web content and presents image animation, Web2TV which reads out the entire text of web content and presents character agent animation, and Web2Talkshow which presents keyword-based dialogue and character agent animation. These prototype systems enable users to watch web content in the same way, they watch a TV program.
Akiyo Nadamoto, Katsumi Tanaka
ACM Multimedia2
2005 3D viewpoint-based photo search and information browsing
abstract
We propose a new photo search method that uses three-dimensional (3D) viewpoints as queries. 3D viewpoint-based image retrieval is especially useful for searching collections of archaeological photographs,which contain many different images of the same object. Our method is designed to enable users to retrieve images that contain the same object but show a different view, and to browse groups of images taken from a similar viewpoint. We also propose using 3D scenes to query by example, which means that users do not have the problem of trying to formulate appropriate queries. This combination gives users an easy way of accessing not only photographs but also archived information.
Rieko Kadobayashi, Katsumi Tanaka
SIGIR2
2005 Trajectory-Based Presentation of Heterogeneous Spatio-temporal Content
Taro Tezuka, Katsumi Tanaka
W2GIS2
2005 Referential Context Mining: Discovering Viewpoints from the Web
abstract
The Web is a vast playground for the propagation of information by individuals. The capability of Web users to leverage the value of Web content, in terms of factors such as usefulness, reputation, or reliability, has emerged as a requirement. The most significant characteristics of the Web are hyperlinks, which enables an author to refer to Web content published by other people. As a result, Web content can refer to other content on the Web in various contexts. The references provide important clues to understanding the roles or reputations of various Web pages according to the viewpoints of third parties. In this paper, we propose an approach to mining referential contexts on the Web.
Koji Zettsu, Katsumi Tanaka
Web Intelligence2
2005 Temporal Ranking of Search Engine Results
Adam Jatowt, Yukiko Kawai, Katsumi Tanaka
WISE3
2005 Blog Map of Experiences: Extracting and Geographically Mapping Visitor Experiences from Urban Blogs
Takeshi Kurashima, Taro Tezuka, Katsumi Tanaka
WISE3
2005 Approximate Intensional Representation of Web Search Results
Yasunori Matsuike, Satoshi Oyama, Katsumi Tanaka
WISE3
2005 Optimization Issues for Keyword Search over Tree-Structured Documents
Sujeet Pradhan, Katsumi Tanaka
WISE2
2005 Topic-structure-based complementary information retrieval and its application
abstract
A great deal of technology has been developed to help people access the information they require. With advances in the availability of information, information-seeking activities are becoming more sophisticated. This means that information technology must move to the next stage, i.e., enable users to acquire information from multiple perspectives to satisfy diverse needs. For instance, with the spread of digital broadcasting and broadband Internet connection services, infrastructure for the integration of TV programs and the Internet has been developed that enables users to acquire information from different media at the same time to improve information quality and the level of detail. In this paper, we propose a novel content-based join model for data streams (closed captions of videos or TV programs) and Web pages based on the concept of topic structures. We then propose a mechanism based on this model for retrieving complementary Web pages to augment the content of video or television programs. One of the most notable features of this complementary retrieval mechanism is that the retrieved information is not just similar to the video or TV program, but also provides additional information. In addition, we introduce an application system called WebTelop, which augments the content of TV programs in real time by using complementary Web pages. We also describe some experimental results.
Qiang Ma 0001, Katsumi Tanaka
ACM Trans. Asian Lang. Inf. Process.2
2005 B-CWB: Bilingual Comparative Web Browser Based on Content-Synchronization and Viewpoint Retrieval
Akiyo Nadamoto, Qiang Ma 0001, Katsumi Tanaka
World Wide Web3
2004 Topic-Structure Based Complementary Information Retrieval for Information Augmentation
Qiang Ma 0001, Katsumi Tanaka
APWeb2
2004 Query Modification by Discovering Topics from Web Page Structures
Satoshi Oyama, Katsumi Tanaka
APWeb2
2004 Aspect Discovery: Web Contents Characterization by Their Referential Contexts
Koji Zettsu, Yutaka Kidawara, Katsumi Tanaka
APWeb3
2004 Relative Queries and the Relative Cluster-Mapping Method
Shinsuke Nakajima, Katsumi Tanaka
DASFAA2
2004 Discovering Aspects of Web Pages from Their Referential Contexts in the Web
Koji Zettsu, Yutaka Kidawara, Katsumi Tanaka
DASFAA3
2004 Device Cooperative Web Browsing and Retrieving Mechanism on Ubiquitous Networks
Yutaka Kidawara, Koji Zettsu, Tomoyuki Uchiyama, Katsumi Tanaka
DEXA4
2004 Retrieving Relevant Portions from Structured Digital Documents
Sujeet Pradhan, Katsumi Tanaka
DEXA2
2004 Guiding Web Search by Third-party Viewpoints: Browsing Retrieval Results by Referential Contexts in Web
Koji Zettsu, Yutaka Kidawara, Katsumi Tanaka
DEXA3
2004 My portal viewer for content fusion based on user's preferences
abstract
A novel Web application called "my portal viewer (MPV)" has been developed to provide Web users with higher quality content, which is needed due to a rapidly growing amount of content on the Web. It provides fused news to the user, based on two viewpoints, through a user friendly interface and the user's preferences. MPV automatically selects and merges content from many news pages, based on the user's interest and knowledge, after gathering these pages from various news Web sites. Our unique approach is that the layout of the MPV page is applied to the users' favorite news portal page and a part of the original content is replaced by the fused content. Whenever a user accesses an MPV page after browsing other news pages, he/she can acquire the desired content efficiently because MPV presents a refreshed page, based on the user's behavior, which reflects his/her interests and knowledge. In addition to the MPV framework, methods that are based on user reference for replacing and selecting have been developed using an HTML table model.
Yukiko Kawai, Daisuke Kanjo, Katsumi Tanaka
ICME3
2004 Query relaxation and answer integration for cross-media meta-searches
abstract
The amount of information available on the Web has increased dramatically. This information comes in a wide variety of media types. Therefore, to retrieve information on a topic using conventional search engines, users must search many sites from several different aspects. Thus, it is important to integrate and organize this information from the different sites. We propose an integration and organization system based on a query relaxation approach for cross-media meta-search engines. Frequently, the parameters for information retrieval are too specific or exacting to generate any relevant sites. However, by gradually relaxing the search terms used for information retrieval, we can solve this problem while narrowing the search to sites that are most relevant to the subject being researched. We show several examples of how the relaxation approach works as well as ways in which it can be applied. We also demonstrate the advantages of our approach and future work for this research.
Akihiro Kuwabara, Kazutoshi Sumiya, Katsumi Tanaka
ICME3
2004 Discovering aspect-based correlation of Web contents for cross-media information retrieval
abstract
The main issue regarding cross-media information retrieval is the determination of correlations between different types of media objects. The conventional approach derives the correlations based on common properties extracted from media contents or synchronous presentation of multiple media pre-authored in a scheduled scenario. We propose a novel approach for determining the cross-media correlation derived from the referential contexts of media objects in the Web. A Web page links to the media objects distributed over the Web so that it aggregates them with respect to the page content. Our approach extracts the referential context by analyzing the logical structure of the Web and discovers the aspect of a media object, which means the latent semantics of the referential context. The aspect-based correlation reveals the relation between media objects regarding their reputations on the Web. In this paper, we propose an approach for discovering aspect-based correlations with an experimental implementation
Koji Zettsu, Yutaka Kidawara, Katsumi Tanaka
ICME3
2004 Temporal and Spatial Attribute Extraction from Web Documents and Time-Specific Regional Web Search System
Taro Tezuka, Katsumi Tanaka
W2GIS2
2004 Extraction of Cognitively-Significant Place Names and Regions from Web-Based Physical Proximity Co-occurrences
Taro Tezuka, Yusuke Yokota, Mizuho Iwaihara, Katsumi Tanaka
WISE4
2003 A Localness-Filter for Searched Web Pages
Qiang Ma 0001, Chiyako Matsumoto, Katsumi Tanaka
APWeb3
2003 Image Retrieval by Web Context: Filling the Gap between Image Keywords and Usage Keywords
Koji Zettsu, Yutaka Kidawara, Katsumi Tanaka
DEXA3
2003 WebTelop: dynamic TV-content augmentation by using Web pages
abstract
In this paper, we propose a method for dynamic integration of TV-program content and related Web content and demonstrate our prototype system called WebTelop. The primary source of information in WebTelop is the TV-program content. Web pages related to the TV-program content are retrieved automatically in real time and are used to augment the TV-program content in the form of captions. Our method enables: (1) real time integration of TV-program content with Web-page content, (2) information retrieval for TVprogram content augmentation based on the topic structures of the TV programs and Web pages, and (3) the use of a virtual character agent to help users navigate through the related Web pages of the TV programs.
Qiang Ma 0001, Katsumi Tanaka
ICME2
2003 Amplifying the differences between your positive samples and neighbors in image retrieval
abstract
A novel method for retrieving images based on relevance feedback and clustering has been developed. That is, by clustering sets of retrieved data, a user can select some good answers from them by considering the difference between the feature data of the selected images and the feature data of images placed in their neighborhood. This difference information improves previous queries since the user must have found some important difference between their-selected image and similar neighboring images. An image-retrieval system based on a relevance feedback by difference amplification is set up and shown to be more effective than conventional methods.
Shinsuke Nakajima, Shinichi Kinoshita, Katsumi Tanaka
ICME3
2003 Concurrent Browsing of Bilingual Web Sites by Content-Synchronization and Difference-Detection
abstract
We propose a new way of browsing bilingual Web sites through concurrent browsing with automatic similar-content synchronization and difference-detection facilities. Our prototype browser system is called the bilingual comparative Web browser (B-CWB) and it concurrently presents bilingual web pages in a way that enables their content of the Web pages to be automatically synchronized. The B-CWB allows users to browse two Web news sites concurrently and compare the similar news articles written in different languages (English and Japanese). The major characteristics of the B-CWB are its content synchronization and difference detection: content synchronization means that user operation (scrolling or clicking) on one Web page automatically invokes not necessarily same operations on the other Web page to preserve similarity of content between the two Web pages. For example, scrolling a Web page may involve passage-level similarity retrieval on the other Web page. Clicking a Web page (and obtaining a new Web page) invokes page-level similarity retrieval within the other Web site pages through the use of a English-Japanese dictionary. Difference detection means that the B-CWB analyzes two similar Web pages shown concurrently to discover the several "differences" between them. This facility is important in comparing two news articles that report the same affairs.
Akiyo Nadamoto, Qiang Ma 0001, Katsumi Tanaka
WISE3
2003 A Dynamic Content Integration Language for Video Data and Web Content
abstract
Dynamic content integration of multiple information sources is one way of providing richer content that will satisfy the diverse demands of users. In this paper, we propose an XML-based language to compose synchronized content from Web and video content. The notable features of this language are as follows: (1) dynamic unit identification of content that is composed into synchronized content; and (2) dynamic retrieval of content through pre-defined retrieval criteria. This dynamic identification and retrieval of composable units are based on the author's intentions. Content authors can specify the units of their content that are to be integrated into new content by describing the conditions concerning this content and the conditions concerning the surrounding content. Although the proposed language looks like SMIL (synchronized multimedia integration language), it differs in its dynamic identification and retrieval capabilities. Indeed, the proposed language works just like the meta-mechanism for conventional SMIL. That is, the script written by the proposed language can generate SMIL data as its output.
Takayuki Yumoto, Qiang Ma 0001, Kazutoshi Sumiya, Katsumi Tanaka
WISE4
2003 A comparative web browser (CWB) for browsing and comparing web pages
abstract
In this paper, we propose a new type of Web browser, called the Comparative Web Browser(CWB), which concurrently presents multiple Web pages in a way that enables the content of the Web pages to be automatically synchronized. The ability to view multiple Web pages at one time is useful when we wish to make a comparison on the Web, such as when we compare similar products or news articles from different newspapers. The CWB is characterized by (1) automatic content-based retrieval of passages from another Web page based on a passage of the Web page the user is reading, and (2) automatic transformation of a user's behavior (scrolling, clicking, or moving backward or forward) on a Web page into a series of behaviors on the other Web pages. The CWB tries to concurrently present "similar" passages from different Web pages, and for this purpose our CWB automatically navigates Web pages that contain passages similar to those of the initial Web page. Furthermore, we propose an enhancement to the CWB, which enables it to use linkage information to find related documents based on link structure.
Akiyo Nadamoto, Katsumi Tanaka
WWW2
2002 Web Information Retrieval Based on the Localness Degree
Chiyako Matsumoto, Qiang Ma 0001, Katsumi Tanaka
DEXA3
2002 MWM: Retrieval and Automatic Presentation of Manual Data for Mobile Terminals
Masaki Shikata, Akiyo Nadamoto, Kazutoshi Sumiya, Katsumi Tanaka
DEXA4
2002 Autonomous presentation of 3 dimensional CG contents on the Web
abstract
Web3D technologies have made it possible for users to browse and manipulate 3D CG model data on the Web. Usually, such data requires users to interact with it in order to see its attribute information and its attached behavior. To automatically present 3D CG data and its attribute information, it is necessary to create 3D CG animation in advance. However, much time and cost is needed to create 3D CG animation that introduces its shape and its attribute information. We aim at constructing a mechanism that automatically produces 3D CG animation from the original 3D CG model data and its attribute information. 3D CG data has become popular on mobile phones such as i-mode. It is, however, difficult for users to interact with such data on mobile phones, because of their limited interaction capability and small-size screen. To cope with the problem, we also need to have a mechanism that produces 3D CG animation in an autonomous manner. We propose a way to automatically generate 3D CG animation from the original 3D CG model data and its attribute information. With the proposed method, users can browse 3D CG data and its attribute information in a "less clicking and more watching" manner. The proposed system automatically generates CG animation synchronized with synthesized speech from 3D CG contents and their attribute data. We also propose a method to generate a presentation for multiple 3D CG objects, that emphasizes the major differences in their attribute information.
Akiyo Nadamoto, Takeshi Yabe, Masaki Shikata, Katsumi Tanaka
ICME (2)4
2002 Context-Dependent Web Bookmarks and Their Usage as Queries
abstract
Conventional Web bookmarks only contain URLs and titles of Web pages that users are interested in. This makes the process of remembering, sharing or ranking such pages difficult. The "context" of users' navigation can be described as collections of browsed pages. Conventional bookmarks do not contain such information. We believe that such context information conveys the users' intention and the importance of bookmarks. We introduce a notion of context-dependent Web bookmarks that reflects users' browsing histories. A context-dependent Web bookmark consists of (1) representative keywords of bookmarked pages and browsed pages, (2) the ranking value of bookmarked pages calculated by its context, as well as the URL and title of the page that the user bookmarked. Context-dependent bookmarks will make it possible for users to remember the situation of the bookmarking process, grasp the degree of significance of the bookmark, and share the bookmark among multiple users. Furthermore, it becomes possible to re-use context-dependent bookmarks as queries, which could be executed for unvisited Web pages. We also describe our Web browser prototype system based on the context-dependent bookmark function, and our experimental results.
Shinsuke Nakajima, Satoshi Oyama, Kazutoshi Sumiya, Katsumi Tanaka
WISE4
2001 Querying Multiple Perspective Video by Camera Metaphor
abstract
Suppose that, in a sports event, a user searches for and obtains video scenes that have "better" views of interesting objects from multiple-perspective video (a collection of mutually synchronized video data taken by a lot of cameras). Because members of the general public take videos and query them in such applications, one of the most important research issues is how to provide a mechanism by means of which a user can formulate appropriate queries easily while looking for and focusing on interesting objects. Another issue is how to search for video scenes that have "better" views of the objects. In this paper, we introduce a querying method based on the concept of "querying by camera", and we propose an algorithm that searches for video intervals that have spatio-temporally "better" views of the focused-on object.
Toshihiko Hata, Tatsuo Hirose, Yoshihiro Nakanishi, Katsumi Tanaka
DASFAA4
2001 WebCarrousel: Restructuring Web Search Results for Passive Viewing in Mobile Environments
abstract
In the present paper, we propose a new way of organizing Web search results and of viewing those results passively in the mobile environment which has limited display and limited interaction. Specifically, the system makes Carousel Components, that are composed of image and voice, from Web search result. Each time of user's interaction, the system automatically computes sets of similar, different, more-detailed, and more-abstracted pages, respectively and reorganizes them as carousels. We call this system WebCarousel.
Akiyo Nadamoto, Hiroyuki Kondo, Katsumi Tanaka
DASFAA3
2001 WebCarousel: Automatic Presentation and Semantic Restructuring of Web Search Result for Mobile Environments
Akiyo Nadamoto, Hiroyuki Kondo, Katsumi Tanaka
DEXA3
2001 WebSCAN: Discovering and Notifying Important Changes of Web Sites
Qiang Ma 0001, Shinya Miyazaki, Katsumi Tanaka
DEXA3
2001 Modeling and Structuring Multiple Perspective Video for Browsing
Yoshihiro Nakanishi, Tatsuo Hirose, Katsumi Tanaka
ER3
2001 A Query Model to Synthesize Answer Intervals from Indexed Video Units
abstract
While a query result in a traditional database is a subset of the database, in a video database, it is a set of subintervals extracted from the raw video sequence. It is very hard, if not impossible, to predetermine all the queries that will be issued in the future, and all the subintervals that will become necessary to answer them. As a result, conventional query frameworks are not applicable to video databases. We propose a new video query model that computes query results by dynamically synthesizing needed subintervals from fragmentary indexed intervals in the database. We introduce new interval operations required for that computation. We also propose methods to compute relative relevance of synthesized intervals to a given query. A query result is a list of synthesized intervals sorted in the order of their degree of relevance.
Sujeet Pradhan, Keishi Tajima, Katsumi Tanaka
IEEE Trans. Knowl. Data Eng.3
2001 Special Issue on The 2nd Web Information Systems Engineering Conference (WISE'01)
M. Tamer Özsu, Hans-Jörg Schek, Katsumi Tanaka, Yanchun Zhang
World Wide Web3
2000 InfoLOD and LandMark: Spatial Presentation of Attribute Information and Computing Representative Objects for Spatial Data
abstract
In this paper, we will propose a way of visualizing attribute information for spatial objects in the three-dimensional space and a calculation method for extracting a representative object from objects in a given region. In conventional three-dimensional visualizations such as architectural simulations, most of the attention has been paid to image data such as colors, shapes, and textures of spatial objects. In this research, we will focus on the attribute information of spatial objects including image data. We propose InfoLOD concept which introduces the notion of level of detail(LOD) to attribute information as well as image data such as photographs and computer graphics for controlling the visualization of attribute information in a three-dimensional space. The visualization is controlled based on distance and orientation, and we will also discuss the differentiation factor which visualizes the differences among the objects. In addition to visualization control, we will propose the LandMark algorithm for extracting a representative object from the objects in a given region based on their spatial occupancy ratio and the uniqueness of the attribute data. The region for browsing may be specified manually by the user or may be automatically specified by some algorithm. Here, we discuss the spatial glue operation which dynamically retrieves regions containing objects with user-specified attribute information unlike conventional method based on static mesh which are often used in GIS(Geographic Information System). We will also introduce some of our implementations in order to illustrate our ideas.
Kengo Koiso, Takehisa Mori, Hiroaki Kawagishi, Katsumi Tanaka, Takahiro Matsumoto
Int. J. Cooperative Inf. Syst.4
1999 An Interactive Classification of Web Documents by Self-Organizing Maps and Search Engines
abstract
We propose an effective classification view mechanism for hypertext data such as Web documents based on Kohonen's self-organizing map (SOM) and search engines. Web documents collected by search engines are automatically classified by SOM and the obtained SOMs are incrementally modified according to the interaction between users and SOMs. At present, various search engines are designed to retrieve Web documents. When we use search engines to retrieve Web documents we get a lot of answers and have to examine each Web document. Therefore, in order to make up for search engines, we need a function to classify Web documents corresponding to the user's point of view and their purposes. Furthermore, we cannot retrieve pertinent Web documents by conventional search engines when a specific topic is described by more than one Web document. To solve these problems, we exploited a content-based clustering system for Web documents. In this system, Web documents are automatically clustered by their feature vectors produced from Web documents or minimal subgraphs consisting of multiple Web documents, and their overview maps are dynamically generated by SOM. Furthermore, we propose a method by which an obtained SOM is modified by user's interaction such as feedback operations.
Kenji Hatano, Ryouichi Sano, Yiwei Duan, Katsumi Tanaka
DASFAA4
1999 Spatial Presentation and Aggregation of Georeferenced Data
abstract
In this paper, we introduce a method of spatial presentation of georeferenced data in a three-dimensional space. Photographs, Quicktime VRs, videos, and computer graphic renderings provide realistic presentations rich in visual information. We believe, however, there is an issue of visualizing georeferenced data such as attribute data for spatial objects as well as showing the objects themselves. We introduce an orientation-based visualization model for visualizing georeferenced data, and discuss abstraction of the georeferenced data through their spatial aggregation in the space specified by a user.
Kengo Koiso, Takahiro Matsumoto, Katsumi Tanaka
DASFAA3
1999 A Robust Selection System Using Real-Time Multi-Modal User-Agent Interactions
abstract
This paper presents a real-time object selection system which can deal with gaze and speech inputs that include uncertainty. Although much research has focused on integration of multi-modal information, most of it assumes that each input is accurately symbolized in advance. In addition, real-time interaction with the user is an important and desirable feature which most systems have overlooked. Unlike those systems, our system is intended to satisfy these two requirements. In our system, target objects are modeled by agents which react to user’s action in real-time. The agent’s reactions are based on integration of multi-modal inputs. We use gaze input which enables real-time detection of focus-of-attention but has low accuracy, whereas speech input has high accuracy but nonreal-time feature. ISghly accurate selection with robustness is achieved by complementary effect through probabilistic integration of these two modalities. Our first experiment shows that it is possible to select target object successfully in most cases, even if either of the modalities includes great uncertainty.
Katsumi Tanaka
IUI1
1998 Hypermedia Broadcasting with Temporal Links
Kazutoshi Sumiya, Reiko Noda, Katsumi Tanaka
DEXA3
1997 A SOM-Based Information Organizer for Text and Video Data
Kenji Hatano, Katsumi Tanaka
DASFAA3
1997 Encapsulating Multimedia Contents and a Copyright Protection Mechanism into Distributed Objects
Yutaka Kidawara, Katsumi Tanaka, Kuniaki Uehara
DEXA2
1997 A Time-Stamped Authoring Graph for Video Databases
Koji Zettsu, Kuniaki Uehara, Katsumi Tanaka, Nobuo Kimura
DEXA3
1997 A Dynamic Theory of Incentives in Multi-Agent Systems
Yoav Shoham, Katsumi Tanaka
IJCAI (1)2
1995 Incremental Data Organization for Ancient Document Databases
Shinichi Ueshima, Kazuhiro Ohtsuki, Jun-ya Morishita, Hiroaki Oiso, Katsumi Tanaka
DASFAA6
1995 Block permutation coding of images using cosine transform
abstract
We present the theory and practice of permutation coding as a new tool for very low-bit-rate image compression. Conventional source coding deals with the data information of signals, while the permutation coding achieves compression through efficiently representing the positional information (i.e., position permutation) caused by ordering the data information into order statistics. A set of four theorems is presented. The first one reveals the information-theoretic relationship between data and permutation information and the rest solves the efficient coding problem. For this, novel tools from finite group theory are applied to derive a compact form of representation for permutation, called permutation-cyclic-representation (PCR) vectors, with which various regularities and constraints in the structure of positional information are displayed, whereby the coding is made very easy using a runlength and Huffman method. A block DCT-based permutation coding algorithm (the BCPC) is developed attempting to combine the DCT's excellent features of energy packing and magnitude ordering that are found to be amenable to permutation coding. This mutually beneficial characteristic significantly reduces the coding bit-rate. Simulation results are provided for real images, showing an improvement by 3-4 dB in the peak-SNR index as compared to those representing the state-of-the-art.
Zhongshu Ji, Katsumi Tanaka, Shinzo Kitamura
IEEE Trans. Commun.2
1995 Interval Queries on Object Histories
Seymour Ginsburg, Katsumi Tanaka
Theor. Comput. Sci.2
1993 OVID: Design and Implementation of a Video-Object Database System
abstract
A video-object data model and the design and implementation of a prototype video-object database system named OVID based on the model are described. Notable features of the video-object data model are a mechanism to share common descriptional data among video-objects, called interval-inclusion based inheritance, and operations to composite video-objects. The OVID system offers a browsing/inspection tool called VideoChart, and adhoc query facility called VideoSQL, and a video-object definition tool.>
Eitetsu Oomoto, Katsumi Tanaka
IEEE Trans. Knowl. Data Eng.2
1991 Adding methods to relational database constructs
abstract
An approach is described to realize a hypertext-like navigation over relational database management systems (DBMSs) by encapsulating query methods to relational database constructs (attribute-values, tuples, columns, and relations) and by invoking those query methods directly over user-selected data. In order to handle structured query language (SQL) queries as generic/polymorphic methods, the authors extended the ordinary SQL and introduced a mechanism to attach query methods dynamically to several kinds of relational database constructs. They present their prototype system, called SQL-Navigator, which currently runs on the NeXT Sybase relational DBMS.>
Katsumi Tanaka, Norio Sanada
COMPSAC1
1991 Advanced Applications and Future Research Issues of OODBs
Katsumi Tanaka
DASFAA1
1991 Uncertainty Management in Object-Oriented Database Systems
Katsumi Tanaka, Susumu Kobayashi, Tomomi Sakanoue
DEXA1
1991 Query Pairs as Hypertext Links
abstract
A new idea is proposed for constructing object-oriented hypertext database systems: query pairs as hypertext links, where each query is defined over a collection of document objects that are classified using a class hierarchy. With this idea, users need not modify their hypertext links against the insertions, deletions and updates of document objects. Also, when a database schema (here, a class hierarchy) evolves, a systematic method can be considered to modify user-defined query-pair links. The TextLink-III system that was developed based on this idea and that is currently running is described. The notable features of TextLink-III are the following: document objects are classified by class hierarchy and the notion of class expression is used for formulating a query for a class hierarchy; multiple viewpoint support; and less link maintenance against data updates.>
Katsumi Tanaka, N. Nishikawa, S. Hirayama, K. Nanba
ICDE1
1989 Alternative Objects in Object-oriented Databases
Tae-Soo Chang, Katsumi Tanaka
DASFAA2
1989 HistoryChart: A Visual Language for Historical Databases
Katsumi Tanaka, Eitetsu Ohmoto
DASFAA1
1988 Schema Virtualization in Object-Oriented Databases
abstract
A description is given of the concept and implementation techniques of schema virtualization in object-oriented databases. The objective of schema virtualization is to provide users with multiple views of a database. First, the notions of virtual classes and virtual schemata, which are natural extension of views in relational databases, is introduced. Then, procedures to convert a schema into a virtual one, as well as rules schemata and their conversion should satisfy, are discussed. The key issue of the design of virtual classes and virtual schemata is 'regarding procedures as objects'. Closure properties of classes and schemata are also discussed. Finally, several implementation techniques for realizing schema virtualization in Smalltalk-80 are presented.>
Katsumi Tanaka, Masatoshi Yoshikawa, Kozo Ishihara
ICDE1
1988 Towards Abstracting Complex Database Objects: Generalization, Reduction and Unification of Set-type Objects (Extended Abstract)
Katsumi Tanaka, Masatoshi Yoshikawa
ICDT1
1986 Computation-Tuple Sequences and Object Histories
abstract
A record-based, algebraically-oriented model is introduced for describing data for “object histories” (with computation), such as checking accounts, credit card accounts, taxes, schedules, and so on. The model consists of sequences of computation tuples defined by a computation-tuple sequence scheme (CSS). The CSS has three major features (in addition to input data): computation (involving previous computation tuples), “uniform” constraints (whose satisfaction by a computation-tuple sequence u implies satisfaction by every interval of u ), and specific sequences with which to start the valid computation-tuple sequences. A special type of CSS, called “local,” is singled out for its relative simplicity in maintaining the validity of a computation-tuple sequence. A necessary and sufficient condition for a CSS to be equivalent to at least one local CSS is given. Finally, the notion of “local bisimulatability” is introduced for regarding two CSS as conveying the same information, and two results on local bisimulatability in connection with local CSS are established.
Seymour Ginsburg, Katsumi Tanaka
ACM Trans. Database Syst.2
1984 Interval Queries on Object Histories: Extended Abstract
Seymour Ginsburg, Katsumi Tanaka
VLDB2
1983 Synthesis of unnormalized relations incorporating more meaning
Yahiko Kambayashi, Katsumi Tanaka, Koichi Takeda 0002
Inf. Sci.2
1982 Performance Analysis for Parallel Processing Schemes of Relational Operations and a Relational Database Machine Architecture with Optimal Scheme Selection Mechanism
Yasushi Kiyoki, Michio Isoda, K. Kojima, Katsumi Tanaka, A. Minematsu, Hideo Aiso
ICDCS4
1981 Design and Evaluation of a Relational Data Base Machine Employing Advanced Data Structures and Algorithms
Yasushi Kiyoki, Katsumi Tanaka, Hideo Aiso, Noriyuki Kamibayashi
ISCA2
1981 Testing of Join Dependency Preserving by a Modified Chase Method
Katsumi Tanaka, Yahiko Kambayashi
MFCS1
1981 Logical Integration of Locally Independent Relational Databases into a Distributed Database
Katsumi Tanaka, Yahiko Kambayashi
VLDB1
1979 Semantic aspects of data dependencies and their application to relational database design
abstract
In this paper, we mainly discuss the semantic aspects of functional and multivalued dependencies in order to choose a better logical design of a relational database. We clarify the differences between a conceptual design and a logical design of a relational database. For example, a multivalued dependency does not always capture a conceptual dependency such that a data element semantically determines a set of other data elements. Moreover, some transitively specified multivalued dependencies are shown to often impose a semantically unnatural constraint. A sufficient condition for these multivalued dependencies to hold in a natural sense is provided. A mixed design approach of a conceptual and a logical designs is introduced, in which each design approach is compensated for its deficiency by the other one. Some generalization of a relational model is also suggested to handle some interface problems between a logical design and a conceptual design, such as a null value problem and an entity identification problem.
Yahiko Kambayashi, Katsumi Tanaka, Shuzo Yajima
COMPSAC2
1979 Use of abstracted characteristics of data in relational databases
abstract
In the database field, the requirement of high level facilities for retrieval/update operations is now increasing rapidly. Our approach from the relational database point of view for contribution to this problem is to provide 1) efficient processing of relational retrieval/ update operations, and 2) a high level user interface. In order to achieve this goal, a new concept concerning "abstracted characteristics" is presented. Abstracted characteristics are defined to be characteristics abstracted from sets of tuples in relations stored in the database. A classification of abstracted characteristics is presented. Functional dependencies, which play an important role in relational database design, and time-dependent functional dependencies are pointed out to be useful in processing retrieval/update operations. Some important applications of abstracted characteristics are discussed. Among them: 1) efficient processing of retrieval/update operations, 2) powerful view update checking facilities, 3) providing some rough meanings of null responses and 4) a high level user interface.
Chung Le Viet, Yahiko Kambayashi, Katsumi Tanaka, Shuzo Yajima
COMPSAC3
1979 Organization of quasi-consecutive retrieval files
Katsumi Tanaka, Yahiko Kambayashi, Shuzo Yajima
Inf. Syst.1
1977 A Relational Data Language with Simplified Binary Relation Handling Capability
Yahiko Kambayashi, Katsumi Tanaka, Shuzo Yajima
VLDB2