Yafeng Lu

dblp:154/1918 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
2since 2021 · last 2021
0000-0002-3978-9550ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
7 papers
Visualization and visual analytics · 97% Image and video processing · 3%
Software engineering, system software, and programming languages
2 papers
Software testing · 88% Software maintenance and evolution · 12%
Databases, data mining, and information retrieval
2 papers
Database system architecture and tuning · 33% Query processing and optimization · 33% Spatial and temporal data management · 17%
Artificial intelligence
2 papers
Trustworthy machine learning · 82% Learning paradigms · 18%

Topics — the 18 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.512021
OoDAnalyzer: Interactive Analysis of Out-of-Distribution Samples · IEEE Trans. Vis. Comput. Graph. 2021
Visualization and visual analytics
graph visualization
0.512021
Same Stats, Different Graphs: Exploring the Space of Graphs in Terms of Graph Properties · IEEE Trans. Vis. Comput. Graph. 2021
Visualization and visual analytics › visual analytics
visual analysis
0.512021
OoDAnalyzer: Interactive Analysis of Out-of-Distribution Samples · IEEE Trans. Vis. Comput. Graph. 2021
Visualization and visual analytics
visual analytics
0.512021
Same Stats, Different Graphs: Exploring the Space of Graphs in Terms of Graph Properties · IEEE Trans. Vis. Comput. Graph. 2021
Software testing
regression testing
0.522016
An extensive study of static regression test selection in modern software evolution · SIGSOFT FSE 2016
How does regression test prioritization perform in real-world software evolution? · ICSE 2016
Visualization and visual analytics
uncertainty visualization
0.412020
Exploring the Sensitivity of Choropleths under Attribute Uncertainty · IEEE Trans. Vis. Comput. Graph. 2020
Visualization and visual analytics
crowdsourced annotation
0.412019
An Interactive Method to Improve Crowdsourced Annotations · IEEE Trans. Vis. Comput. Graph. 2019
Visualization and visual analytics
spatiotemporal visualization
0.412019
A Visual Analytics Framework for Spatiotemporal Trade Network Analysis · IEEE Trans. Vis. Comput. Graph. 2019
Software testing › regression testing
regression test selection
0.212016
An extensive study of static regression test selection in modern software evolution · SIGSOFT FSE 2016
Software testing › regression testing
test case prioritization
0.212016
How does regression test prioritization perform in real-world software evolution? · ICSE 2016
Software maintenance and evolution
software evolution
0.122016
An extensive study of static regression test selection in modern software evolution · SIGSOFT FSE 2016
How does regression test prioritization perform in real-world software evolution? · ICSE 2016
Graph algorithms and graph theory
random graph generation
0.112021
Same Stats, Different Graphs: Exploring the Space of Graphs in Terms of Graph Properties · IEEE Trans. Vis. Comput. Graph. 2021
Data mining
clustering
0.112020
Exploring the Sensitivity of Choropleths under Attribute Uncertainty · IEEE Trans. Vis. Comput. Graph. 2020
Spatial and temporal data management
spatial analysis
0.112020
Exploring the Sensitivity of Choropleths under Attribute Uncertainty · IEEE Trans. Vis. Comput. Graph. 2020
Machine learning › Learning paradigms › weakly supervised learning
learning from crowds
0.112019
An Interactive Method to Improve Crowdsourced Annotations · IEEE Trans. Vis. Comput. Graph. 2019
Computational social science and digital humanities
political science
0.112019
A Visual Analytics Framework for Spatiotemporal Trade Network Analysis · IEEE Trans. Vis. Comput. Graph. 2019
Image and video processing › pattern detection
anomaly detection
0.112019
A Visual Analytics Framework for Spatiotemporal Trade Network Analysis · IEEE Trans. Vis. Comput. Graph. 2019
Software testing › regression testing
test suite augmentation
0.112016
How does regression test prioritization perform in real-world software evolution? · ICSE 2016

Methods — techniques the papers use, named apart from their topics

sampling · 1.0random graph generation · 1.0kNN-based grid layout · 1.0spatial autocorrelation · 0.9simulation-based analysis · 0.9parallel coordinates · 0.9constrained projection · 0.8confusion visualization · 0.8entity extraction · 0.6query planning · 0.5mapreduce · 0.5graph processing · 0.5deep ensembles · 0.5deep ensemble · 0.5scatterplot visualization · 0.4network analytics · 0.4multi-view framework · 0.4anomaly detection · 0.4
YearPublicationVenuePosition
2021 Same Stats, Different Graphs: Exploring the Space of Graphs in Terms of Graph Properties
abstract
Data analysts commonly utilize statistics to summarize large datasets. While it is often sufficient to explore only the summary statistics of a dataset (e.g., min/mean/max), Anscombe's Quartet demonstrates how such statistics can be misleading. We consider a similar problem in the context of graph mining. To study the relationships between different graph properties, we examine low-order non-isomorphic graphs and provide a simple visual analytics system to explore correlations across multiple graph properties. However, for larger graphs, studying the entire space quickly becomes intractable. We use different random graph generation methods to further look into the distribution of graph properties for higher order graphs and investigate the impact of various sampling methodologies. We also describe a method for generating many graphs that are identical over a number of graph properties and statistics yet are clearly different and identifiably distinct.
Utkarsh Soni, Yafeng Lu, Vahan Huroyan, Ross Maciejewski, Stephen G. Kobourov
IEEE Trans. Vis. Comput. Graph.3
2021 OoDAnalyzer: Interactive Analysis of Out-of-Distribution Samples
abstract
One major cause of performance degradation in predictive models is that the test samples are not well covered by the training data. Such not well-represented samples are called OoD samples. In this article, we propose OoDAnalyzer, a visual analysis approach for interactively identifying OoD samples and explaining them in context. Our approach integrates an ensemble OoD detection method and a grid-based visualization. The detection method is improved from deep ensembles by combining more features with algorithms in the same family. To better analyze and understand the OoD samples in context, we have developed a novelkNN-based grid layout algorithm motivated by Hall's theorem. The algorithm approximates the optimal layout and has O(kN2)O(kN2) time complexity, faster than the grid layout algorithm with overall best performance but O(N3)O(N3) time complexity. Quantitative evaluation and case studies were performed on several datasets to demonstrate the effectiveness and usefulness of OoDAnalyzer.
Changjian Chen, Jun Yuan 0003, Yafeng Lu, Yang Liu 0014, Hang Su 0006, Songtao Yuan, Shixia Liu
IEEE Trans. Vis. Comput. Graph.3
2020 Exploring the Sensitivity of Choropleths under Attribute Uncertainty
abstract
The choropleth map is an essential tool for spatial data analysis. However, the underlying attribute values of a spatial unit greatly influence the statistical analyses and map classification procedures when generating a choropleth map. If the attribute values incorporate a range of uncertainty, a critical task is determining how much the uncertainty impacts both the map visualization and the statistical analysis. In this paper, we present a visual analytics system that enhances our understanding of the impact of attribute uncertainty on data visualization and statistical analyses of these data. Our system consists of a parallel coordinates-based uncertainty specification view, an impact river and impact matrix visualization for region-based and simulation-based analysis, and a dual-choropleth map and t-SNE plot for visualizing the changes in classification and spatial autocorrelation over the range of uncertainty in the attribute values. We demonstrate our system through three use cases illustrating the impact of attribute uncertainty in geographic analysis.
Zhaosong Huang, Yafeng Lu, Elizabeth A. Mack, Wei Chen 0001, Ross Maciejewski
IEEE Trans. Vis. Comput. Graph.2
2019 An Interactive Method to Improve Crowdsourced Annotations
abstract
In order to effectively infer correct labels from noisy crowdsourced annotations, learning-from-crowds models have introduced expert validation. However, little research has been done on facilitating the validation procedure. In this paper, we propose an interactive method to assist experts in verifying uncertain instance labels and unreliable workers. Given the instance labels and worker reliability inferred from a learning-from-crowds model, candidate instances and workers are selected for expert validation. The influence of verified results is propagated to relevant instances and workers through the learning-from-crowds model. To facilitate the validation of annotations, we have developed a confusion visualization to indicate the confusing classes for further exploration, a constrained projection method to show the uncertain labels in context, and a scatter-plot-based visualization to illustrate worker reliability. The three visualizations are tightly integrated with the learning-from-crowds model to provide an iterative and progressive environment for data validation. Two case studies were conducted that demonstrate our approach offers an efficient method for validating and improving crowdsourced annotations.
Shixia Liu, Changjian Chen, Yafeng Lu, Fang-Xin Ou-Yang, Bin Wang 0021
IEEE Trans. Vis. Comput. Graph.3
2019 A Visual Analytics Framework for Spatiotemporal Trade Network Analysis
abstract
Economic globalization is increasing connectedness among regions of the world, creating complex interdependencies within various supply chains. Recent studies have indicated that changes and disruptions within such networks can serve as indicators for increased risks of violence and armed conflicts. This is especially true of countries that may not be able to compete for scarce commodities during supply shocks. Thus, network-induced vulnerability to supply disruption is typically exported from wealthier populations to disadvantaged populations. As such, researchers and stakeholders concerned with supply chains, political science, environmental studies, etc. need tools to explore the complex dynamics within global trade networks and how the structure of these networks relates to regional instability. However, the multivariate, spatiotemporal nature of the network structure creates a bottleneck in the extraction and analysis of correlations and anomalies for exploratory data analysis and hypothesis generation. Working closely with experts in political science and sustainability, we have developed a highly coordinated, multi-view framework that utilizes anomaly detection, network analytics, and spatiotemporal visualization methods for exploring the relationship between global trade networks and regional instability. Requirements for analysis and initial research questions to be investigated are elicited from domain experts, and a variety of visual encoding techniques for rapid assessment of analysis and correlations between trade goods, network patterns, and time series signatures are explored. We demonstrate the application of our framework through case studies focusing on armed conflicts in Africa, regional instability measures, and their relationship to international global trade.
Yafeng Lu, Shade T. Shutters, Michael Steptoe, Feng Wang 0012, Steven Landis, Ross Maciejewski
IEEE Trans. Vis. Comput. Graph.2
2018 Same Stats, Different Graphs - (Graph Statistics and Why We Need Graph Drawings)
Utkarsh Soni, Yafeng Lu, Ross Maciejewski, Stephen G. Kobourov
GD3
2018 The Perception of Graph Properties in Graph Layouts
abstract
Abstract When looking at drawings of graphs, questions about graph density, community structures, local clustering and other graph properties may be of critical importance for analysis. While graph layout algorithms have focused on minimizing edge crossing, symmetry, and other such layout properties, there is not much known about how these algorithms relate to a user's ability to perceive graph properties for a given graph layout. In this study, we apply previously established methodologies for perceptual analysis to identify which graph drawing layout will help the user best perceive a particular graph property. We conduct a large scale (n = 588) crowdsourced experiment to investigate whether the perception of two graph properties (graph density and average local clustering coefficient) can be modeled using Weber's law. We study three graph layout algorithms from three representative classes (Force Directed ‐ FD, Circular, and Multi‐Dimensional Scaling ‐ MDS), and the results of this experiment establish the precision of judgment for these graph layouts and properties. Our findings demonstrate that the perception of graph density can be modeled with Weber's law. Furthermore, the perception of the average clustering coefficient can be modeled as an inverse of Weber's law, and the MDS layout showed a significantly different precision of judgment than the FD layout.
Utkarsh Soni, Yafeng Lu, Brett Hansen, Helen C. Purchase, Stephen G. Kobourov, Ross Maciejewski
Comput. Graph. Forum2
2018 A Visual Analytics Framework for Identifying Topic Drivers in Media Events
abstract
Media data has been the subject of large scale analysis with applications of text mining being used to provide overviews of media themes and information flows. Such information extracted from media articles has also shown its contextual value of being integrated with other data, such as criminal records and stock market pricing. In this work, we explore linking textual media data with curated secondary textual data sources through user-guided semantic lexical matching for identifying relationships and data links. In this manner, critical information can be identified and used to annotate media timelines in order to provide a more detailed overview of events that may be driving media topics and frames. These linked events are further analyzed through an application of causality modeling to model temporal drivers between the data series. Such causal links are then annotated through automatic entity extraction which enables the analyst to explore persons, locations, and organizations that may be pertinent to the media topic of interest. To demonstrate the proposed framework, two media datasets and an armed conflict event dataset are explored.
Yafeng Lu, Steven Landis, Ross Maciejewski
IEEE Trans. Vis. Comput. Graph.1
2017 The State-of-the-Art in Predictive Visual Analytics
abstract
Abstract Predictive analytics embraces an extensive range of techniques including statistical modeling, machine learning, and data mining and is applied in business intelligence, public health, disaster management and response, and many other fields. To date, visualization has been broadly used to support tasks in the predictive analytics pipeline. Primary uses have been in data cleaning, exploratory analysis, and diagnostics. For example, scatterplots and bar charts are used to illustrate class distributions and responses. More recently, extensive visual analytics systems for feature selection, incremental learning, and various prediction tasks have been proposed to support the growing use of complex models, agent‐specific optimization, and comprehensive model comparison and result exploration. Such work is being driven by advances in interactive machine learning and the desire of end‐users to understand and engage with the modeling process. In this state‐of‐the‐art report, we catalogue recent advances in the visualization community for supporting predictive analytics. First, we define the scope of predictive analytics discussed in this article and describe how visual analytics can support predictive analytics tasks in a predictive visual analytics (PVA) pipeline. We then survey the literature and categorize the research with respect to the proposed PVA pipeline. Systems and techniques are evaluated in terms of their supported interactions, and interactions specific to predictive analytics are discussed. We end this report with a discussion of challenges and opportunities for future research in predictive visual analytics.
Yafeng Lu, Rolando Garcia, Brett Hansen, Michael Gleicher, Ross Maciejewski
Comput. Graph. Forum1
2016 SQL-SA for big data discovery polymorphic and parallelizable SQL user-defined scalar and aggregate infrastructure in Teradata Aster 6.20
abstract
There is increasing demand to integrate big data analytic systems using SQL. Given the vast ecosystem of SQL applications, enabling SQL capabilities allows big data platforms to expose their analytic potential to a wide variety of end users, accelerating discovery processes and providing significant business value. Most existing big data frameworks are based on one particular programming model such as MapReduce or Graph. However, data scientists are often forced to manually create adhoc data pipelines to connect various big data tools and platforms to serve their analytic needs. When the analytic tasks change, these data pipelines may be costly to modify and maintain. In this paper we present SQL-SA, a polymorphic and parallelizable SQL scalar and aggregate infrastructure in Aster 6.20. This infrastructure extends Aster 6's MapReduce and Graph capabilities to support polymorphic user-defined scalar and aggregate functions using flexible SQL syntax. The implementation enhances main Aster components including query syntax, API, planning and execution extensively. Integrating these new user-defined scalar and aggregate functions with Aster MapReduce and Graph functions, Aster 6.20 enables data scientists to integrate diverse programming models in a single SQL statement. The statement is automatically converted to an optimal data pipeline and executed in parallel. Using a real world business problem and data, Aster 6.20 demonstrates a significant performance advantage (25%+) over Hadoop Pig and Hive.
Robert M. Wehrmeister, James Shau, Abhirup Chakraborty, Daley Alex, Awny Al Omari, Feven Atnafu, Jeff Davis, Litao Deng, Deepak Jaiswal, Chittaranjan Keswani, Yafeng Lu, Tom Reyes, Kashif Siddiqui, David E. Simmen, Devendra Vidhani, Daniel Yu
ICDE12
2016 How does regression test prioritization perform in real-world software evolution?
abstract
In recent years, researchers have intensively investigated various topics in test prioritization, which aims to re-order tests to increase the rate of fault detection during regression testing. While the main research focus in test prioritization is on proposing novel prioritization techniques and evaluating on more and larger subject systems, little effort has been put on investigating the threats to validity in existing work on test prioritization. One main threat to validity is that existing work mainly evaluates prioritization techniques based on simple artificial changes on the source code and tests. For example, the changes in the source code usually include only seeded program faults, whereas the test suite is usually not augmented at all. On the contrary, in real-world software development, software systems usually undergo various changes on the source code and test suite augmentation. Therefore, it is not clear whether the conclusions drawn by existing work in test prioritization from the artificial changes are still valid for real-world software evolution. In this paper, we present the first empirical study to investigate this important threat to validity in test prioritization. We reimplemented 24 variant techniques of both the traditional and time-aware test prioritization, and investigated the impacts of software evolution on those techniques based on the version history of 8 real-world Java programs from GitHub. The results show that for both traditional and time-aware test prioritization, test suite augmentation significantly hampers their effectiveness, whereas source code changes alone do not influence their effectiveness much.
Yafeng Lu, Yiling Lou, Shiyang Cheng 0002, Lingming Zhang 0001, Dan Hao 0001, Yangfan Zhou 0002, Lu Zhang 0023
ICSE1
2016 An extensive study of static regression test selection in modern software evolution
abstract
Regression test selection (RTS) aims to reduce regression testing time by only re-running the tests affected by code changes. Prior research on RTS can be broadly split into dy namic and static techniques. A recently developed dynamic RTS technique called Ekstazi is gaining some adoption in practice, and its evaluation shows that selecting tests at a coarser, class-level granularity provides better results than selecting tests at a finer, method-level granularity. As dynamic RTS is gaining adoption, it is timely to also evaluate static RTS techniques, some of which were proposed over three decades ago but not extensively evaluated on modern software projects.
Owolabi Legunsen, Farah Hariri, August Shi, Yafeng Lu, Lingming Zhang 0001, Darko Marinov
SIGSOFT FSE4
2016 Exploring Evolving Media Discourse Through Event Cueing
abstract
Online news, microblogs and other media documents all contain valuable insight regarding events and responses to events. Underlying these documents is the concept of framing, a process in which communicators act (consciously or unconsciously) to construct a point of view that encourages facts to be interpreted by others in a particular manner. As media discourse evolves, how topics and documents are framed can undergo change, shifting the discussion to different viewpoints or rhetoric. What causes these shifts can be difficult to determine directly; however, by linking secondary datasets and enabling visual exploration, we can enhance the hypothesis generation process. In this paper, we present a visual analytics framework for event cueing using media data. As discourse develops over time, our framework applies a time series intervention model which tests to see if the level of framing is different before or after a given date. If the model indicates that the times before and after are statistically significantly different, this cues an analyst to explore related datasets to help enhance their understanding of what (if any) events may have triggered these changes in discourse. Our framework consists of entity extraction and sentiment analysis as lenses for data exploration and uses two different models for intervention analysis. To demonstrate the usage of our framework, we present a case study on exploring potential relationships between climate change framing and conflicts in Africa.
Yafeng Lu, Michael Steptoe, Sarah E. Burke, Jiun-Yi Tsai, Hasan Davulcu, Douglas C. Montgomery, Steven R. Corman, Ross Maciejewski
IEEE Trans. Vis. Comput. Graph.1
2015 Half a Century of Practice: Who Is Still Storing Plaintext Passwords?
Erick Bauman, Yafeng Lu, Zhiqiang Lin 0001
ISPEC2