Yamei Tu

dblp:294/3186 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-0722-837XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 VADIS: A Visual Analytics Pipeline for Dynamic Document Representation and Information-Seeking
abstract
In the biomedical domain, visualizing the document embeddings of an extensive corpus has been widely used in information-seeking tasks. However, three key challenges with existing visualizations make it difficult for clinicians to find information efficiently. First, the document embeddings used in these visualizations are generated statically by pretrained language models, which cannot adapt to the user's evolving interest. Second, existing document visualization techniques cannot effectively display how the documents are relevant to users' interest, making it difficult for users to identify the most pertinent information. Third, existing embedding generation and visualization processes suffer from a lack of interpretability, making it difficult to understand, trust and use the result for decision-making. In this paper, we present a novel visual analytics pipeline for user-driven document representation and iterative information seeking (VADIS). VADIS introduces a prompt-based attention model (PAM) that generates dynamic document embedding and document relevance adjusted to the user's query. To effectively visualize these two pieces of information, we design a new document map that leverages a circular grid layout to display documents based on both their relevance to the query and the semantic similarity. Additionally, to improve the interpretability, we introduce a corpus-level attention visualization method to improve the user's understanding of the model focus and to enable the users to identify potential oversight. This visualization, in turn, empowers users to refine, update and introduce new queries, thereby facilitating a dynamic and iterative information-seeking experience. We evaluated VADIS quantitatively and qualitatively on a real-world dataset of biomedical research papers to demonstrate its effectiveness.
Yamei Tu, Po-Yin Yen, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2024 KG-PRE-view: Democratizing a TVCG Knowledge Graph through Visual Explorations
abstract
IEEE Transactions on Visualization and Computer Graphics (TVCG) publishes cutting-edge research in the fields of visualization, computer graphics, and virtual and augmented realities. Within the TVCG ecosystem, different stakeholders make decisions based on available information related to TVCG almost on a daily basis. The decisions involve various tasks such as the retrieval of research ideas and trends, the invitation of peer reviewers, and the selection of editorial board members, just to name a few. To make well-informed decisions in these contexts, a data-driven approach is necessary. However, the current IEEE digital library only provides access to individual papers. Transforming this wealth of data into valuable insights is a daunting task, requiring specialized expertise and effort in tasks such as data crawling, cleaning, analysis, and visualizations. To address the needs of the community in facilitating more efficient and transparent decision-making, we construct and publicly release a TVCG knowledge graph (TVCG-KG). TVCG-KG is a structured representation of heterogeneous information, including the metadata of each publication such as author, affiliation, title, and semantic information such as method, task, data. Despite the widespread use of KGs in various downstream applications, a noticeable gap exists in the visualization literature regarding the full exploitation of the rich semantics embedded within KGs. While it might seem intuitive to just employ interactive graph-based visualization for KGs, we propose that knowledge discovery over KG is a series of visual exploratory tasks that can benefit from using multiple visualization techniques and designs. We conducted an evaluation of TVCG-KG quality and demonstrated its practical utility through several real-world cases. Our data and code are accessible via the following URL: https://github.com/yasmineTYM/TVCG-KG.git.
Yamei Tu, Han-Wei Shen
PacificVis1
2024 DocFlow: A Visual Analytics System for Question-Based Document Retrieval and Categorization
abstract
A systematic review (SR) is essential with up-to-date research evidence to support clinical decisions and practices. However, the growing literature volume makes it challenging for SR reviewers and clinicians to discover useful information efficiently. Many human-in-the-loop information retrieval approaches (HIR) have been proposed to rank documents semantically similar to users' queries and provide interactive visualizations to facilitate document retrieval. Given that the queries are mainly composed of keywords and keyphrases retrieving documents that are semantically similar to a query does not necessarily respond to the clinician's need. Clinicians still have to review many documents to find the solution. The problem motivates us to develop a visual analytics system, DocFlow, to facilitate information-seeking. One of the features of our DocFlow is accepting natural language questions. The detailed description enables retrieving documents that can answer users' questions. Additionally, clinicians often categorize documents based on their backgrounds and with different purposes (e.g., populations, treatments). Since the criteria are unknown and cannot be pre-defined in advance, existing methods can only achieve categorization by considering the entire information in documents. In contrast, by locating answers in each document, our DocFlow can intelligently categorize documents based on users' questions. The second feature of our DocFlow is a flexible interface where users can arrange a sequence of questions to customize their rules for document retrieval and categorization. The two features of this visual analytics system support a flexible information-seeking process. The case studies and the feedback from domain experts demonstrate the usefulness and effectiveness of our DocFlow.
Yamei Tu, Yu-Shuen Wang, Po-Yin Yen, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2024 PhraseMap: Attention-Based Keyphrases Recommendation for Information Seeking
abstract
Many Information Retrieval (IR) approaches have been proposed to extract relevant information from a large corpus. Among these methods, phrase-based retrieval methods have been proven to capture more concrete and concise information than word-based and paragraph-based methods. However, due to the complex relationship among phrases and a lack of proper visual guidance, achieving user-driven interactive information-seeking and retrieval remains challenging. In this study, we present a visual analytic approach for users to seek information from an extensive collection of documents efficiently. The main component of our approach is a PhraseMap, where nodes and edges represent the extracted keyphrases and their relationships, respectively, from a large corpus. To build the PhraseMap, we extract keyphrases from each document and link the phrases according to word attention determined using modern language models, i.e., BERT. As can be imagined, the graph is complex due to the extensive volume of information and the massive amount of relationships. Therefore, we develop a navigation algorithm to facilitate information seeking. It includes (1) a question-answering (QA) model to identify phrases related to users' queries and (2) updating relevant phrases based on users' feedback. To better present the PhraseMap, we introduce a resource-controlled self-organizing map (RC-SOM) to evenly and regularly display phrases on grid cells while expecting phrases with similar semantics to stay close in the visualization. To evaluate our approach, we conducted case studies with three domain experts in diverse literature. The results and feedback demonstrate its effectiveness, usability, and intelligence.
Yamei Tu, Yu-Shuen Wang, Po-Yin Yen, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.1
2023 SDRQuerier: A Visual Querying Framework for Cross-National Survey Data Recycling
abstract
Public opinion surveys constitute a widespread, powerful tool to study peoples' attitudes and behaviors from comparative perspectives. However, even global surveys can have limited geographic and temporal coverage, which can hinder the production of comprehensive knowledge. To expand the scope of comparison, social scientists turn to ex-post harmonization of variables from datasets that cover similar topics but in different populations and/or at different times. These harmonized datasets can be analyzed as a single source and accessed through various data portals. However, the Survey Data Recycling (SDR) research project has identified three challenges faced by social scientists when using data portals: the lack of capability to explore data in-depth or query data based on customized needs, the difficulty in efficiently identifying related data for studies, and the incapability to evaluate theoretical models using sliced data. To address these issues, the SDR research project has developed the SDRQuerier, which is applied to the harmonized SDR database. The SDRQuerier includes a BERT-based model that allows for customized data queries through research questions or keywords (Query-by-Question), a visual design that helps users determine the availability of harmonized data for a given research question (Query-by-Condition), and the ability to reveal the underlying relational patterns among substantive and methodological variables in the database (Query-by-Relation), aiding in the rigorous evaluation or improvement of regression models. Case studies with multiple social scientists have demonstrated the usefulness and effectiveness of the SDRQuerier in addressing daily challenges.
Yamei Tu, Olga Li, Junpeng Wang 0001, Han-Wei Shen, Przemek Powalko, Irina Tomescu-Dubrow, Kazimierz M. Slomczynski, Spyros Blanas, J. Craig Jenkins
IEEE Trans. Vis. Comput. Graph.1
2021 KeywordMap: Attention-based Visual Exploration for Keyword Analysis
abstract
With the high growth rate of text data, extracting meaningful information from a large corpus becomes increasingly difficult. Keyword extraction and analysis is a common approach to tackle the problem, but it is non-trivial to identify important words in the text and represent the multifaceted properties of those words effectively. Traditional topic modeling based keyword analysis algorithms require hyper-parameters which are often difficult to tune without enough prior knowledge. In addition, the relationships among the keywords are often difficult to obtain. In this paper, we utilize the attention scores extracted from Transformer-based language models to capture word relationships. We propose a domain-driven attention tuning method, guiding the attention to learn domain-specific word relationships. From the attention, we build a keyword network and propose a novel algorithm, Attention-based Word Influence (AWI), to compute how influential each word is in the network. An interactive visual analytics system, KeywordMap, is developed to support multi-level analysis of keywords and keyword relationships through coordinated views. We measure the quality of keywords captured by our AWI algorithm quantitatively. We also evaluate the usefulness and effectiveness of KeywordMap through case studies.
Yamei Tu, Jiayi Xu 0001, Han-Wei Shen
PacificVis1
2021 USEVis: Visual analytics of attention-based neural embedding in information retrieval
abstract
Neural attention-based encoders, which effectively attend sentence tokens to their associated context without being restricted by long-term distance or dependency, have demonstrated outstanding performance in embedding sentences into meaningful representations (embeddings). The Universal Sentence Encoder (USE) is one of the most well-recognized deep neural network (DNN) based solutions, which is facilitated with an attention-driven transformer architecture and has been pre-trained on a large number of sentences from the Internet. Besides the fact that USE has been widely used in many downstream applications, including information retrieval (IR), interpreting its complicated internal working mechanism remains challenging. In this work, we present a visual analytics solution towards addressing this challenge. Specifically, focused on semantics and syntactics (concepts and relations) that are critical to domain clinical IR, we designed and developed a visual analytics system, i.e., USEVis. The system investigates the power of USE in effectively extracting sentences’ semantics and syntactics through exploring and interpreting how linguistic properties are captured by attentions. Furthermore, by thoroughly examining and comparing the inherent patterns of these attentions, we are able to exploit attentions to retrieve sentences/documents that have similar semantics or are closely related to a given clinical problem in IR. By collaborating with domain experts, we demonstrate use cases with inspiring findings to validate the contribution of our work and the effectiveness of our system.
Xiaonan Ji, Yamei Tu, Junpeng Wang 0001, Han-Wei Shen, Po-Yin Yen
Vis. Informatics2