VLDB 2026 Research / reviewers in the wild / expert
Yifan Hu 0001
dblp:92/498-1
· DBLP profile ↗
56ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0003-2017-924XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 6 since 2021Theory of computation · 13 · 2 first-authorDatabases, data management, data science and information retrieval · 11 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beauty in the Eye of AI: Aligning LLMs and Vision Models with Human Aesthetics in Network VisualizationabstractAbstract Network visualization has traditionally relied on heuristic metrics, such as stress, under the assumption that optimizing them leads to aesthetic and informative layouts. However, no single metric consistently produces the most effective results. A data‐driven alternative is to learn from human preferences, where labelers select their favored visualization among multiple layouts of the same graphs. These human‐preference labels can then be used to train a generative model that approximates human aesthetic preferences. However, obtaining human labels at scale is costly and time‐consuming. As a result, this generative approach has so far been tested only with machine‐labeled data [WYHS24]. In this paper, we explore the use of large language models (LLMs) and vision models (VMs) as proxies for human judgment. Through a carefully designed user study involving 27 participants, we curated a large set of human preference labels. We used this data both to better understand human preferences and to bootstrap LLM/VM labelers. We show that prompt engineering that combines few‐shot examples and diverse input formats, such as image embeddings, significantly improves LLM‐‐human alignment, and additional filtering by the confidence score of the LLM pushes the alignment to human‐‐human levels. Furthermore, we demonstrate that carefully trained VMs can achieve VM‐human alignment at a level comparable to that between human labelers. Our results suggest that AI can feasibly serve as a scalable proxy for human labelers. Han-Wei Shen, Yifan Hu 0001 |
Comput. Graph. Forum | 5 |
| 2025 | Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and VerificationabstractRecent advancements in large language models (LLMs) have been fueled by large-scale training corpora drawn from diverse sources such as websites, news articles, and books.These datasets often contain explicit user information, such as person names, addresses, that LLMs may unintentionally reproduce in their generated outputs.Beyond such explicit content, LLMs can also leak identity-revealing cues through implicit signals such as distinctive writing styles, raising significant concerns about authorship privacy.There are three major automated tasks in authorship privacy, namely authorship obfuscation (AO), authorship mimicking (AM), and authorship verification (AV).Prior research has studied AO, AM, and AV independently.However, their interplays remain under-explored, which leaves a major research gap, especially in the era of LLMs, where they are profoundly shaping how we curate and share user-generated content, and the distinction between machine-generated and human-authored text is also increasingly blurred.This work then presents the first unified framework for analyzing the dynamic relationships among LLM-enabled AO, AM, and AV in the context of authorship privacy.We quantify how they interact with each other to transform human-authored text, examining effects at a single point in time and iteratively over time.We also examine the role of demographic metadata, such as gender, academic background, in modulating their performances, inter-task dynamics, and privacy risks. Tuc Nguyen, Yifan Hu 0001, Thai Le |
EMNLP | 2 |
| 2024 | SmartGD: A GAN-Based Graph Drawing Framework for Diverse Aesthetic GoalsabstractWhile a multitude of studies have been conducted on graph drawing, many existing methods only focus on optimizing a single aesthetic aspect of graph layouts, which can lead to sub-optimal results. There are a few existing methods that have attempted to develop a flexible solution for optimizing different aesthetic aspects measured by different aesthetic criteria. Furthermore, thanks to the significant advance in deep learning techniques, several deep learning-based layout methods were proposed recently. These methods have demonstrated the advantages of deep learning approaches for graph drawing. However, none of these existing methods can be directly applied to optimizing non-differentiable criteria without special accommodation. In this work, we propose a novel Generative Adversarial Network (GAN) based deep learning framework for graph drawing, called SmartGD, which can optimize different quantitative aesthetic goals, regardless of their differentiability. To demonstrate the effectiveness and efficiency of SmartGD, we conducted experiments on minimizing stress, minimizing edge crossing, maximizing crossing angle, maximizing shape-based metrics, and a combination of multiple aesthetics. Compared with several popular graph drawing algorithms, the experimental results show that SmartGD achieves good performance both quantitatively and qualitatively. Kevin Yen, Yifan Hu 0001, Han-Wei Shen |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | CrypText: Database and Interactive Toolkit of Human-Written Text Perturbations in the WildabstractUser-generated textual contents on the Internet are often noisy, erroneous, and not in correct grammar. In fact, some online users choose to express their opinions online through carefully perturbed texts, especially in controversial topics (e.g., politics, vaccine mandate) or abusive contexts (e.g., cyberbullying, hate-speech). However, to the best of our knowledge, there is no framework that explores these online "human-written" perturbations (as opposed to algorithm-generated perturbations). Therefore, we introduce an interactive system called CrypText. CrypText is a data-intensive application that provides the users with a database and several tools to extract and interact with human-written perturbations. Specifically, CrypText helps look up, perturb, and normalize (i.e., de-perturb) texts. CrypText also provides an interactive interface to monitor and analyze text perturbations online. The demo is available at: https://lethaiq.github.io/anthro. Thai Le, Yiran Ye, Yifan Hu 0001, Dongwon Lee 0001 |
ICDE | 3 |
| 2023 | MGEL: Multigrained Representation Analysis and Ensemble Learning for Text ModerationabstractIn this work, we describe our efforts in addressing two typical challenges involved in the popular text classification methods when they are applied to text moderation: the representation of multibyte characters and word obfuscations. Specifically, a multihot byte-level scheme is developed to significantly reduce the dimension of one-hot character-level encoding caused by the multiplicity of instance-scarce non-ASCII characters. In addition, we introduce a simple yet effective weighting approach for fusing n-gram features to empower the classical logistic regression. Surprisingly, it outperforms well-tuned representative neural networks greatly. As a continual effort toward text moderation, we endeavor to analyze the current state-of-the-art (SOTA) algorithm bidirectional encoder representations from transformers (BERT), which works well in context understanding but performs poorly on intentional word obfuscations. To resolve this crux, we then develop an enhanced variant and remedy this drawback by integrating byte and character decomposition. It advances the SOTA performance on the largest abusive language datasets as demonstrated by our comprehensive experiments. Our work offers a feasible and effective framework to tackle word obfuscations. Fei Tan 0002, Changwei Hu, Yifan Hu 0001, Kevin Yen, Zhi Wei 0001, Aasish Pappu, Se Rim Park, Keqian Li |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Editorial: Guest Editors' Introduction: Special Section on IEEE PacificVis 2023abstractThis special section of theIEEE Transactions on Visualization and Computer Graphics (IEEE TVCG)presents the five most highly rated papers from the 2023 IEEE Pacific Visualization Symposium (IEEE PacificVis), hosted in Seoul, Korea from April 18 to Apr 21, 2023. IEEE PacificVis, sponsored by the IEEE Visualization and Graphics Technical Committee (VGTC), aims to foster greater exchange between visualization researchers and practitioners, especially in the Asia-Pacific region. This forum has grown to be a truly international event, attracting submissions and attendees from many countries, not only in the Asia-Pacific but also in Europe, America, and beyond. Thus, IEEE PacificVis is serving the additional purpose of sharing the latest advances in the field of visualization with researchers and practitioners in the region and, also, introducing research developments from the region to the broader international visualization research community. Jaegul Choo, Timo Ropinski, Yifan Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | UNICON: A UNIform CONstraint Based Graph Layout FrameworkabstractWe propose UNICON, a UNIform CONstraint based graph layout framework that supports both soft and hard constraints. We extend the stress model to accommodate soft constraints by incorporating them in the objective functions, optimized by stochastic gradient descent. For hard constraints, such as inequalities or equalities in the layout space, we utilize a gradient projection method to satisfy them. A visualization prototype system is implemented based on this framework for the user to interactively add or remove constraints to generate the desired layouts. We demonstrate the efficiency, quality, and flexibility of the framework and the system on a number of datasets with a wide range of user-defined constraints. Jiacheng Yu, Yifan Hu 0001, Xiaoru Yuan |
PacificVis | 2 |
| 2021 | Hadoop-MTA: a system for Multi Data-center Trillion Concepts Auto-ML atop HadoopabstractThe ever-growing computation capability distributed infrastructure brings tremendous opportunities for mining and analysis of data that was impossible otherwise. Meanwhile, the inherent computation model of distributed system also brings unique and non-trivial challenges for traditional Auto-ML, including the explosion of data dimensions, the expected absence of features, and the heterogeneity of information. This is especially the case in modern Internet enterprises, where data in the scale of trillions are stored in multiple data centers, and the discovery of subtle signals could incur significant impact in revenue and welfare. How can we best harness the large scale distributed machine learning, but without keeping engineers constantly in the loop? In this work, we present Hadoop-MTA, a system for Multi Data-center, Trillion Concepts, Auto-ML on top of the Hadoop distributed computation environment that leverages sparsity aware heterogeneous knowledge graph representation and dimensionality agnostic parallel learning. Through multiple large scale experiments, we find that Hadoop-MTA significantly output-performs competitive state of the art distributed learning algorithms and scales well to trillion scale data-sets. Our model is rolled out to Hadoop serving infrastructure in Yahoo covering billions of unique identities and shows improvements 129.5% accuracy and 106.5 % weighted F1-score (more than 2x) on key targeting use cases. Keqian Li, Yifan Hu 0001, Manisha Verma, Fei Tan 0002, Changwei Hu, Tejaswi Kasturi, Kevin Yen |
IEEE BigData | 2 |
| 2021 | BAN: Large Scale Brand ANonymization for Creative Recommendation via Label Light AdaptationabstractOne of the primary component in ads creative recommendation system is the brand anonymization that removes brand-specific information from ad text for legal compliance and providing ready to use template for the advertisers to customize and consume. In our previous work [1] on ads creative recommendation system, the anonymization is done via a block list created solely based on manual reviewing, which is expensive and limits in the scale of the deployment of the ads recommendation. In this work we investigate a large scale, automated approach for brand anonymization. Such a problem presents many unique and non-trivial challenges, including the domain specificity of the brand entities, the fine-granularity requirements of structured output, the tight constraint of the limited contexts, the high level of grammatical noise in the advertisement data, and the heterogeneity of information required to perform anonymization. We propose a transformer model that leverage implicit knowledge together with a label-light adaptation procedure for this task. Our model is rolled out to ads systems in Yahoo that cover billions of impression traffic per month and improved previous production system by 68.3% F1-score on token level prediction and 61.6% on ad level prediction. Keqian Li, Kevin Yen, Shaunak Mishra, Yifan Hu 0001, Changwei Hu, Manisha Verma |
IEEE BigData | 4 |
| 2021 | TSI: An Ad Text Strength Indicator using Text-to-CTR and Semantic-Ad-SimilarityabstractComing up with effective ad text is a time consuming process, and particularly challenging for small businesses with limited advertising experience. When an inexperienced advertiser onboards with a poorly written ad text, the ad platform has the opportunity to detect low performing ad text, and provide improvement suggestions. To realize this opportunity, we propose an ad text strength indicator (TSI) which: (i) predicts the click-through-rate (CTR) for an input ad text, (ii) fetches similar existing ads to create a neighborhood around the input ad, (iii) and compares the predicted CTRs in the neighborhood to declare whether the input ad is strong or weak. In addition, as suggestions for ad text improvement, TSI shows anonymized versions of superior ads (higher predicted CTR) in the neighborhood. For (i), we propose a BERT based text-to-CTR model trained on impressions and clicks associated with an ad text. For (ii), we propose a sentence-BERT based semantic-ad-similarity model trained using weak labels from ad campaign setup data. Offline experiments demonstrate that our BERT based text-to-CTR model achieves a significant lift in CTR prediction AUC for cold start (new) advertisers compared to bag-of-words based baselines. In addition, our semantic-textual-similarity model for similar ads retrieval achieves a [email protected] of 0.93 (for retrieving ads from the same product category); this is significantly higher compared to unsupervised TF-IDF, word2vec, and sentence-BERT baselines. Finally, we share promising online results from advertisers in the Yahoo (Verizon Media) ad platform where a variant of TSI was implemented with sub-second end-to-end latency. Shaunak Mishra, Changwei Hu, Manisha Verma, Kevin Yen, Yifan Hu 0001, Maxim Sviridenko |
CIKM | 5 |
| 2021 | BERT-Beta: A Proactive Probabilistic Approach to Text ModerationabstractText moderation for user generated content, which helps to promote healthy interaction among users, has been widely studied and many machine learning models have been proposed.In this work, we explore an alternative perspective by augmenting reactive reviews with proactive forecasting.Specifically, we propose a new concept text toxicity propensity to characterize the extent to which a text tends to attract toxic comments.Beta regression is then introduced to do the probabilistic modeling, which is demonstrated to function well in comprehensive experiments.We also propose an explanation method to communicate the model decision clearly.Both propensity scoring and interpretation benefit text moderation in a novel manner.Finally, the proposed scaling mechanism for the linear model offers useful insights beyond this work. Fei Tan 0002, Yifan Hu 0001, Kevin Yen, Changwei Hu |
EMNLP (1) | 2 |
| 2021 | What's in a name? - gender classification of names with character based machine learning models
Yifan Hu 0001, Changwei Hu, Thanh Tran 0005, Tejaswi Kasturi, Elizabeth Joseph, Matt Gillingham |
Data Min. Knowl. Discov. | 1 |
| 2021 | UrbanMotion: Visual Analysis of Metropolitan-Scale Sparse TrajectoriesabstractVisualizing massive scale human movement in cities plays an important role in solving many of the problems that modern cities face (e.g., traffic optimization, business site configuration). In this article, we study a big mobile location dataset that covers millions of city residents, but is temporally sparse on the trajectory of individual user. Mapping sparse trajectories to illustrate population movement poses several challenges from both analysis and visualization perspectives. In the literature, there are a few techniques designed for sparse trajectory visualization; yet they do not consider trajectories collected from mobile apps that possess long-tailed sparsity with record intervals as long as hours. This article introduces UrbanMotion, a visual analytics system that extends the original wind map design by supporting map-matched local movements, multi-directional population flows, and population distributions. Effective methods are proposed to extract and aggregate population movements from dense parts of the trajectories leveraging their long-tailed sparsity. Both characteristic and anomalous patterns are discovered and visualized. We conducted three case studies, one comparative experiment, and collected expert feedback in the application domains of commuting analysis, event detection, and business site configuration. The study result demonstrates the significance and effectiveness of our system in helping to complete key analytics tasks for urban users. Lei Shi 0002, Congcong Huang, Meijun Liu, Tao Jiang 0054, Zhihao Tan, Yifan Hu 0001, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2020 | Repulsive Attention: Rethinking Multi-head Attention as Bayesian InferenceabstractBang An, Jie Lyu, Zhenyi Wang, Chunyuan Li, Changwei Hu, Fei Tan, Ruiyi Zhang, Yifan Hu, Changyou Chen. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Bang An 0001, Jie Lyu 0004, Zhenyi Wang 0001, Chunyuan Li, Changwei Hu, Fei Tan 0002, Ruiyi Zhang 0002, Yifan Hu 0001, Changyou Chen |
EMNLP (1) | 8 |
| 2020 | TNT: Text Normalization based Pre-training of Transformers for Content ModerationabstractIn this work, we present a new language pre-training model TNT (Text Normalization based pre-training of Transformers) for content moderation.Inspired by the masking strategy and text normalization, TNT is developed to learn language representation by training transformers to reconstruct text from four operation types typically seen in text manipulation: substitution, transposition, deletion, and insertion.Furthermore, the normalization involves the prediction of both operation types and token labels, enabling TNT to learn from more challenging tasks than the standard task of masked word recovery.As a result, the experiments demonstrate that TNT outperforms strong baselines on the hate speech classification task.Additional text normalization experiments and case studies show that TNT is a new potential approach to misspelling correction. Fei Tan 0002, Yifan Hu 0001, Changwei Hu, Keqian Li, Kevin Yen |
EMNLP (1) | 2 |
| 2020 | HABERTOR: An Efficient and Effective Deep Hatespeech DetectorabstractWe present our HABERTOR model for detecting hatespeech in large scale user-generated content.Inspired by the recent success of the BERT model, we propose several modifications to BERT to enhance the performance on the downstream hatespeech classification task.HABERTOR inherits BERT's architecture, but is different in four aspects: (i) it generates its own vocabularies and is pre-trained from the scratch using the largest scale hatespeech dataset; (ii) it consists of Quaternionbased factorized components, resulting in a much smaller number of parameters, faster training and inferencing, as well as less memory usage; (iii) it uses our proposed multisource ensemble heads with a pooling layer for separate input sources, to further enhance its effectiveness; and (iv) it uses a regularized adversarial training with our proposed finegrained and adaptive noise magnitude to enhance its robustness.Through experiments on the large-scale real-world hatespeech dataset with 1.4M annotated comments, we show that HABERTOR works better than 15 state-ofthe-art hatespeech detection methods, including fine-tuning Language Models.In particular, comparing with BERT, our HABERTOR is 4∼5 times faster in the training/inferencing phase, uses less than 1/3 of the memory, and has better performance, even though we pretrain it by using less than 1% of the number of words.Our generalizability analysis shows that HABERTOR transfers well to other unseen hatespeech datasets and is a more efficient and effective alternative to BERT for the hatespeech classification. Thanh Tran 0005, Yifan Hu 0001, Changwei Hu, Kevin Yen, Fei Tan 0002, Kyumin Lee, Se Rim Park |
EMNLP (1) | 2 |
| 2020 | Eiffel: Evolutionary Flow Map for Influence Graph VisualizationabstractThe visualization of evolutionary influence graphs is important for performing many real-life tasks such as citation analysis and social influence analysis. The main challenges include how to summarize large-scale, complex, and time-evolving influence graphs, and how to design effective visual metaphors and dynamic representation methods to illustrate influence patterns over time. In this work, we present Eiffel, an integrated visual analytics system that applies triple summarizations on evolutionary influence graphs in the nodal, relational, and temporal dimensions. In numerical experiments, Eiffel summarization results outperformed those of traditional clustering algorithms with respect to the influence-flow-based objective. Moreover, a flow map representation is proposed and adapted to the case of influence graph summarization, which supports two modes of evolutionary visualization (i.e., flip-book and movie) to expedite the analysis of influence graph dynamics. We conducted two controlled user experiments to evaluate our technique on influence graph summarization and visualization respectively. We also showcased the system in the evolutionary influence analysis of two typical scenarios, the citation influence of scientific papers and the social influence of emerging online events. The evaluation results demonstrate the value of Eiffel in the visual analysis of evolutionary influence graphs. Lei Shi 0002, Yifan Hu 0001, Hanghang Tong, Chaoli Wang 0001, Tong Yang 0003, Deyun Wang, Shuo Liang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | OnionGraph: Hierarchical topology+attribute multivariate network visualizationabstractHierarchical abstraction is a scalable strategy to deal with large networks. Existing visualization methods have allowed to aggregate the network nodes into hierarchies based on the node attributes or network topology, each of which has its own advantage. Very few previous system has the capability to enjoy the best of both worlds. This paper presents OnionGraph, an integrated framework for the exploratory visual analysis of heterogeneous multivariate networks. OnionGraph allows nodes to be aggregated based on either node attributes, topology, or a hierarchical combination of both. These aggregations can be split, merged and filtered under the focus+context interaction model, or automatically traversed by the information-theoretic navigation method. Node aggregations that contain subsets of nodes are displayed by the onion metaphor, indicating the level and details of the abstraction. We have evaluated the OnionGraph tool in three real-world cases. Performance experiments demonstrate that on a commodity desktop, our method can scale to million-node networks while preserving the interactivity for analysis. Lei Shi 0002, Qi Liao 0002, Hanghang Tong, Yifan Hu 0001, Chaoli Wang 0001, Chuang Lin 0002, Weihong Qian |
Vis. Informatics | 4 |
| 2019 | ORIGIN: Non-Rigid Network AlignmentabstractNetwork alignment is a fundamental task in many high-impact applications. Most of the existing approaches either explicitly or implicitly consider the alignment matrix as a linear transformation to map one network to another, and might overlook the complicated alignment relationship across networks. On the other hand, node representation learning based alignment methods are hampered by the incomparability among the node representations of different networks. In this paper, we propose a unified semi-supervised deep model (ORIGIN) that simultaneously finds the non-rigid network alignment and learns node representations in multiple networks in a mutually beneficial way. The key idea is to learn node representations by the effective graph convolutional networks, which subsequently enable us to formulate network alignment as a point set alignment problem. The proposed method offers two distinctive advantages. First (node representations), unlike the existing graph convolutional networks that aggregate the node information within a single network, we can effectively aggregate the auxiliary information from multiple sources, achieving far-reaching node representations. Second (network alignment), guided by the high-quality node representations, our proposed non-rigid point set alignment approach overcomes the bottleneck of the linear transformation assumption. We conduct extensive experiments that demonstrate the proposed non-rigid alignment method is (1) effective, outperforming both the state-of-the-art linear transformation-based methods and node representation based methods, and (2) efficient, with a comparable computational time between the proposed multi-network representation learning component and its single-network counterpart. Hanghang Tong, Jiejun Xu, Yifan Hu 0001, Ross Maciejewski |
IEEE BigData | 4 |
| 2019 | Adaptive Feature Redundancy MinimizationabstractMost existing feature selection methods select the top-ranked features according to certain criterion. However, without considering the redundancy among the features, the selected ones are frequently highly correlated with each other, which is detrimental to the performance. To tackle this problem, we propose a framework regarding adaptive redundancy minimization (ARM) for the feature selection. Unlike other feature selection methods, the proposed model has the following merits: (1) The redundancy matrix is adaptively constructed instead of presetting it as the priori information. (2) The proposed model could pick out the discriminative and non-redundant features via minimizing the global redundancy of the features. (3) ARM can reduce the redundancy of the features from both supervised and unsupervised perspectives. Rui Zhang 0017, Hanghang Tong, Yifan Hu 0001 |
CIKM | 3 |
| 2019 | A Deep Structural Model for Analyzing Correlated Multivariate Time SeriesabstractMultivariate time series are routinely encountered in real-world applications, and in many cases, these time series are strongly correlated. In this paper, we present a deep learning structural time series model which can (i) handle correlated multivariate time series input, and (ii) forecast the targeted temporal sequence by explicitly learning/extracting the trend, seasonality, and event components. The trend is learned via a 1D and 2D temporal CNN and LSTM hierarchical neural net. The CNN-LSTM architecture can (i) seamlessly leverage the dependency among multiple correlated time series in a natural way, (ii) extract the weighted differencing feature for better trend learning, and (iii) memorize the long-term sequential pattern. The seasonality component is approximated via a non-liner function of a set of Fourier terms, and the event components are learned by a simple linear function of regressor encoding the event dates. We compare our model with several state-of-the-art methods through a comprehensive set of experiments on a variety of time series data sets, such as forecasts of Amazon AWS Simple Storage Service (S3) and Elastic Compute Cloud (EC2) billings, and the closing prices for corporate stocks in the same category. Changwei Hu, Yifan Hu 0001, Sungyong Seo |
ICMLA | 2 |
| 2019 | Large-Scale Gender/Age Prediction of Tumblr UsersabstractTumblr, as a leading content provider and social media, attracts 371 million monthly visits, 280 million blogs and 53.3 million daily posts. The popularity of Tumblr provides great opportunities for advertisers to promote their products through sponsored posts. However, it is a challenging task to target specific demographic groups for ads, since Tumblr does not require user information like gender and ages during their registration. Hence, to promote ad targeting, it is essential to predict user's demography using rich content such as posts, images and social connections. In this paper, we propose graph based and deep learning models for age and gender predictions, which take into account user activities and content features. For graph based models, we come up with two approaches, network embedding and label propagation, to generate connection features as well as directly infer user's demography. For deep learning models, we leverage convolutional neural network (CNN) and multilayer perceptron (MLP) to prediction users' age and gender. Experimental results on real Tumblr daily dataset, with hundreds of millions of active users and billions of following relations, demonstrate that our approaches significantly outperform the baseline model, by improving the accuracy relatively by 81% for age, and the AUC and accuracy by 5% for gender. Changwei Hu, Yifan Hu 0001, Tejaswi Kasturi, Shanmugam Ramasamy, Matt Gillingham, Keith Yamamoto |
ICMLA | 3 |
| 2019 | A Coloring Algorithm for Disambiguating Graph and Map DrawingsabstractDrawings of non-planar graphs always result in edge crossings. When there are many edges crossing at small angles, it is often difficult to follow these edges, because of the multiple visual paths resulted from the crossings that slow down eye movements. In this paper we propose an algorithm that disambiguates the edges with automatic selection of distinctive colors. Our proposed algorithm computes a near optimal color assignment of a dual collision graph, using a novel branch-and-bound procedure applied to a space decomposition of the color gamut. We give examples demonstrating this approach in real world graphs and maps, as well as a user study to establish its effectiveness and limitations. Yifan Hu 0001, Lei Shi 0002 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2018 | HARP: Hierarchical Representation Learning for NetworksabstractWe present HARP, a novel method for learning low dimensional embeddings of a graph’s nodes which preserves higher-order structural features. Our proposed method achieves this by compressing the input graph prior to embedding it, effectively avoiding troublesome embedding configurations (i.e. local minima) which can pose problems to non-convex optimization. HARP works by finding a smaller graph which approximates the global structure of its input. This simplified graph is used to learn a set of initial representations, which serve as good initializations for learning representations in the original, detailed graph. We inductively extend this idea, by decomposing a graph in a series of levels, and then embed the hierarchy of graphs from the coarsest one to the original graph. HARP is a general meta-strategy to improve all of the state-of-the-art neural algorithms for embedding graphs, including DeepWalk, LINE, and Node2vec. Indeed, we demonstrate that applying HARP’s hierarchical paradigm yields improved implementations for all three of these methods, as evaluated on classification tasks on real-world graphs such as DBLP, BlogCatalog, and CiteSeer, where we achieve a performance gain over the original implementations by up to 14% Macro F1. Haochen Chen, Bryan Perozzi, Yifan Hu 0001, Steven Skiena |
AAAI | 3 |
| 2018 | CorrelatedMultiples: Spatially Coherent Small Multiples With Constrained Multi-Dimensional ScalingabstractAbstract Displaying small multiples is a popular method for visually summarizing and comparing multiple facets of a complex data set. If the correlations between the data are not considered when displaying the multiples, searching and comparing specific items become more difficult since a sequential scan of the display is often required. To address this issue, we introduce CorrelatedMultiples, a spatially coherent visualization based on small multiples, where the items are placed so that the distances reflect their dissimilarities. We propose a constrained multi‐dimensional scaling (CMDS) solver that preserves spatial proximity while forcing the items to remain within a fixed region. We evaluate the effectiveness of our approach by comparing CMDS with other competing methods through a controlled user study and a quantitative study, and demonstrate the usefulness of CorrelatedMultiples for visual search and comparison in three real‐world case studies. Yifan Hu 0001, Stephen C. North, Han-Wei Shen |
Comput. Graph. Forum | 2 |
| 2017 | Nationality Classification Using Name EmbeddingsabstractNationality identification unlocks important demographic information, with many applications in biomedical and sociological research. Existing name-based nationality classifiers use name substrings as features and are trained on small, unrepresentative sets of labeled names, typically extracted from Wikipedia. As a result, these methods achieve limited performance and cannot support fine-grained classification. Junting Ye, Shuchu Han, Yifan Hu 0001, Baris Coskun, Meizhu Liu, Hong Qin 0001, Steven Skiena |
CIKM | 3 |
| 2016 | There is More to Streamgraphs than Movies: Better Aesthetics via Ordering and LassoingabstractAbstract Streamgraphs were popularized in 2008 when The New York Times used them to visualize box office revenues for 7500 movies over 21 years. The aesthetics of a streamgraph is affected by three components: the ordering of the layers, the shape of the lowest curve of the drawing, known as the baseline, and the labels for the layers. As of today, the ordering and baseline computation algorithms proposed in the paper of Byron and Wattenberg are still considered the state of the art. However, their ordering algorithm exploits statistical properties of the movie revenue data that may not hold in other data . In addition, the baseline optimization is based on a definition of visual energy that in some cases results in considerable amount of visual distortion. We offer an ordering algorithm that works well regardless of the properties of the input data , and propose a 1‐norm based definition of visual energy and the associated solution method that overcomes the limitation of the original baseline optimization procedure. Furthermore, we propose an efficient layer labeling algorithm that scales linearly to the data size in place of the brute‐force algorithm adopted by Byron and Wattenberg. We demonstrate the advantage of our algorithms over existing techniques on a number of real world data sets. Marco Di Bartolomeo, Yifan Hu 0001 |
Comput. Graph. Forum | 2 |
| 2014 | Hierarchical Focus+Context Heterogeneous Network VisualizationabstractAggregation is a scalable strategy for dealing with large network data. Existing network visualizations have allowed nodes to be aggregated based on node attributes or network topology, each of which has its own advantages. However, very few previous systems have the capability to enjoy the best of both worlds. This paper presents OnionGraph, an integrated framework for exploratory visual analysis of large heterogeneous networks. OnionGraph allows nodes to be aggregated based on either node attributes, topology, or a mixture of both. Subsets of nodes can be flexibly split and merged under the hierarchical focus+context interaction model, supporting sophisticated analysis of the network data. Node aggregations that contain subsets of nodes are displayed with multiple concentric circles, or the onion metaphor, indicating how many levels of abstraction they contain. We have evaluated the OnionGraph tool in two real-world cases. Performance experiments demonstrate that on a commodity desktop, OnionGraph can scale to million-node networks while preserving the interactivity for analysis. Lei Shi 0002, Qi Liao 0002, Hanghang Tong, Yifan Hu 0001, Chuang Lin 0002 |
PacificVis | 4 |
| 2014 | MapSets: Visualizing Embedded and Clustered Graphs
Alon Efrat, Yifan Hu 0001, Stephen G. Kobourov, Sergey Pupyrev |
GD | 2 |
| 2014 | A Coloring Algorithm for Disambiguating Graph and Map Drawings
Yifan Hu 0001, Lei Shi 0002 |
GD | 1 |
| 2014 | MicroEye: Visual summary of microblogsphere from the eye of celebritiesabstractMicro-blogging is a new kind of online broadcast medium rising in latest years which is called weibo in China. Every day, weibo users are fed with a large number of incurious micro-blogs. They may also pay close attention to what celebrities concerned about, but they would rather roughly know the summary of micro-blogs the celebrities, however, they can't add attentions to all the celebrities' friends. In this work, we provide weibo users with an overview of the relevant micro-blogs and offer them a visual summary of the micro-blogs that appeal to famous people. Our goal is to help users select their most interested micro-blogs from a large number of micro-blogs. Yifan Hu 0001 |
ICCCN | 2 |
| 2014 | How to Display Group Information on Node-Link Diagrams: An EvaluationabstractWe present the results of evaluating four techniques for displaying group or cluster information overlaid on node-link diagrams: node coloring, GMap, BubbleSets, and LineSets. The contributions of the paper are three fold. First, we present quantitative results and statistical analyses of data from an online study in which approximately 800 subjects performed 10 types of group and network tasks in the four evaluated visualizations. Specifically, we show that BubbleSets is the best alternative for tasks involving group membership assessment; that visually encoding group information over basic node-link diagrams incurs an accuracy penalty of about 25 percent in solving network tasks; and that GMap's use of prominent group labels improves memorability. We also show that GMap's visual metaphor can be slightly altered to outperform BubbleSets in group membership assessment. Second, we discuss visual characteristics that can explain the observed quantitative differences in the four visualizations and suggest design recommendations. This discussion is supported by a small scale eye-tracking study and previous results from the visualization literature. Third, we present an easily extensible user study methodology. Radu Jianu, Adrian Rusu, Yifan Hu 0001, Douglas Taggart |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2013 | CompactMap: A mental map preserving visual interface for streaming text dataabstractAs text streams become increasingly available from social media such as Facebook and Twitter, visual analysis of streaming text data is playing an important role in most business sectors. A fundamental challenge in visualizing a large amount of streaming text data is to preserve the user's mental map to enable tracking dynamic changes in topics, while simultaneously utilizing the display space efficiently. In this paper, we present CompactMap, an online visual interface that packs text clusters efficiently, with stable updates to maintain the user's mental map. It achieves spatiotemporally coherent layouts by dynamically matching clusters across time, and removing cluster overlaps according to spatial proximity and constraints. We developed a visual search engine based on CompactMaps for exploring a large amount of text streams in details on demand. We demonstrate the effectiveness of our approach in a controlled user study compared with a competing method. Yifan Hu 0001, Stephen C. North, Han-Wei Shen |
IEEE BigData | 2 |
| 2013 | COAST: A Convex Optimization Approach to Stress-Based Embedding
Emden R. Gansner, Yifan Hu 0001, Shankar Krishnan |
GD | 2 |
| 2013 | A Maxent-Stress Model for Graph LayoutabstractIn some applications of graph visualization, input edges have associated target lengths. Dealing with these lengths is a challenge, especially for large graphs. Stress models are often employed in this situation. However, the traditional full stress model is not scalable due to its reliance on an initial all-pairs shortest path calculation. A number of fast approximation algorithms have been proposed. While they work well for some graphs, the results are less satisfactory on graphs of intrinsically high dimension, because some nodes may be placed too close together, or even share the same position. We propose a solution, called the maxent-stress model, which applies the principle of maximum entropy to cope with the extra degrees of freedom. We describe a force-augmented stress majorization algorithm that solves the maxent-stress model. Numerical results show that the algorithm scales well, and provides acceptable layouts for large, nonrigid graphs. This also has potential applications to scalable algorithms for statistical multidimensional scaling (MDS) with variable distances. Emden R. Gansner, Yifan Hu 0001, Stephen C. North |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | A maxent-stress model for graph layoutabstractIn some applications of graph visualization, input edges have associated target lengths. Dealing with these lengths is a challenge, especially for large graphs. Stress models are often employed in this situation. However, the traditional full stress model is not scalable due to its reliance on an initial all-pairs shortest path calculation. A number of fast approximation algorithms have been proposed. While they work well for some graphs, the results are less satisfactory on graphs of intrinsically high dimension, because nodes overlap unnecessarily. We propose a solution, called the maxent-stress model, which applies the principle of maximum entropy to cope with the extra degrees of freedom. We describe a force-augmented stress majorization algorithm that solves the maxent-stress model. Numerical results show that the algorithm scales well, and provides acceptable layouts for large, non-rigid graphs. This also has potential applications to scalable algorithms for statistical multidimensional scaling (MDS) with variable distances. Emden R. Gansner, Yifan Hu 0001, Stephen C. North |
PacificVis | 2 |
| 2012 | Embedding, clustering and coloring for dynamic mapsabstractWe describe a practical approach for visualizing multiple relationships defined on the same dataset using a geographic map metaphor, where clusters of nodes form countries and neighboring countries correspond to nearby clusters. Our aim is to provide a visualization that allows us to compare two or more such maps (showing an evolving dynamic process, or obtained using different relationships). In the case where we are considering multiple relationships, e.g., different similarity metrics, we also provide an interactive tool to visually explore the effect of combining two or more such relationships. Our method ensures good readability and mental map preservation, based on dynamic node placement with node stability, dynamic clustering with cluster stability, and dynamic coloring with color stability. Yifan Hu 0001, Stephen G. Kobourov, Sankar Veeramoni |
PacificVis | 1 |
| 2012 | Visualizing Streaming Text Data with Dynamic Graphs and Maps
Emden R. Gansner, Yifan Hu 0001, Stephen C. North |
GD | 2 |
| 2012 | Optimal Polygonal Representation of Planar Graphs
Christian A. Duncan, Emden R. Gansner, Yifan Hu 0001, Michael Kaufmann 0001, Stephen G. Kobourov |
Algorithmica | 3 |
| 2012 | Drawing Large Graphs by Low-Rank Stress MajorizationabstractAbstract Optimizing a stress model is a natural technique for drawing graphs: one seeks an embedding into Rd which best preserves the induced graph metric. Current approaches to solving the stress model for a graph with |𝒱| nodes and |ɛ| edges require the full all‐pairs shortest paths (APSP) matrix, which takes O(|𝒱|2 log |ɛ|+|𝒱‖ɛ|) time and O(|𝒱|2) space. We propose a novel algorithm based on a low‐rank approximation to the required matrices. The crux of our technique is an observation that it is possible to approximate the full APSP matrix, even when only a small subset of its entries are known. Our algorithm takes time O(k|𝒱|+|𝒱|log|𝒱|+|ɛ|) per iteration with a preprocessing time of O(k3+ k(|ɛ|+|𝒱| log |𝒱|) + k2|𝒱|) and memory usage of O(k|𝒱|), where a user‐defined parameter k trades off quality of approximation with running time and space. We give experimental results which show, to the best of our knowledge, the largest (albeit approximate) full stress model based layouts to date. Marc Khoury, Yifan Hu 0001, Shankar Krishnan, Carlos Scheidegger |
Comput. Graph. Forum | 2 |
| 2012 | Visualizing Dynamic Data with MapsabstractMaps offer a familiar way to present geographic data (continents, countries), and additional information (topography, geology), can be displayed with the help of contours and heat-map overlays. In this paper, we consider visualizing large-scale dynamic relational data by taking advantage of the geographic map metaphor. We describe a map-based visualization system which uses animation to convey dynamics in large data sets, and which aims to preserve the viewer's mental map while also offering readable views at all times. Our system is fully functional and has been used to visualize user traffic on the Internet radio station last.fm, as well as TV-viewing patterns from an IPTV service. All map images in this paper are available in high-resolution at [1] as are several movies illustrating the dynamic visualization. Daisuke Mashima, Stephen G. Kobourov, Yifan Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2012 | Intelligent Graph Layout Using Many Users' InputabstractIn this paper, we propose a new strategy for graph drawing utilizing layouts of many sub-graphs supplied by a large group of people in a crowd sourcing manner. We developed an algorithm based on Laplacian constrained distance embedding to merge subgraphs submitted by different users, while attempting to maintain the topological information of the individual input layouts. To facilitate collection of layouts from many people, a light-weight interactive system has been designed to enable convenient dynamic viewing, modification and traversing between layouts. Compared with other existing graph layout algorithms, our approach can achieve more aesthetic and meaningful layouts with high user preference. Xiaoru Yuan, Limei Che, Yifan Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2011 | Multilevel agglomerative edge bundling for visualizing large graphsabstractGraphs are often used to encapsulate relationships between objects. Node-link diagrams, commonly used to visualize graphs, suffer from visual clutter on large graphs. Edge bundling is an effective technique for alleviating clutter and revealing high-level edge patterns. Previous methods for general graph layouts either require a control mesh to guide the bundling process, which can introduce high variation in curvature along the bundles, or all-to-all force and compatibility calculations, which is not scalable. We propose a multilevel agglomerative edge bundling method based on a principled approach of minimizing ink needed to represent edges, with additional constraints on the curvature of the resulting splines. The proposed method is much faster than previous ones, able to bundle hundreds of thousands of edges in seconds, and one million edges in a few minutes. Emden R. Gansner, Yifan Hu 0001, Stephen C. North, Carlos Scheidegger |
PacificVis | 2 |
| 2011 | Visualizing dynamic data with mapsabstractMaps offer a familiar way to present geographic data (continents, countries), and additional information (topography, geology), can be displayed with the help of contours and heat-map overlays. In this paper we consider visualizing large-scale dynamic relational data by taking advantage of the geographic map metaphor. We describe a system that visualizes user traffic on the Internet radio station last.fm and address challenges in mental map preservation, as well as issues in animated map-based visualization. Daisuke Mashima, Stephen G. Kobourov, Yifan Hu 0001 |
PacificVis | 3 |
| 2011 | The university of Florida sparse matrix collectionabstractWe describe the University of Florida Sparse Matrix Collection, a large and actively growing set of sparse matrices that arise in real applications. The Collection is widely used by the numerical linear algebra community for the development and performance evaluation of sparse matrix algorithms. It allows for robust and repeatable experiments: robust because performance results with artificially generated matrices can be misleading, and repeatable because matrices are curated and made publicly available in many formats. Its matrices cover a wide spectrum of domains, include those arising from problems with underlying 2D or 3D geometry (as structural engineering, computational fluid dynamics, model reduction, electromagnetics, semiconductor devices, thermodynamics, materials, acoustics, computer graphics/vision, robotics/kinematics, and other discretizations) and those that typically do not have such geometry (optimization, circuit simulation, economic and financial modeling, theoretical and quantum chemistry, chemical process simulation, mathematics and statistics, power networks, and other networks and graphs). We provide software for accessing and managing the Collection, from MATLAB™, Mathematica™, Fortran, and C, as well as an online search capability. Graph visualization of the matrices is provided, and a new multilevel coarsening scheme is proposed to facilitate this task. Timothy A. Davis 0001, Yifan Hu 0001 |
ACM Trans. Math. Softw. | 2 |
| 2010 | GMap: Visualizing graphs and clusters as mapsabstractInformation visualization is essential in making sense out of large data sets. Often, high-dimensional data are visualized as a collection of points in 2-dimensional space through dimensionality reduction techniques. However, these traditional methods often do not capture well the underlying structural information, clustering, and neighborhoods. In this paper, we describe GMap, a practical algorithm for visualizing relational data with geographic-like maps. We illustrate the effectiveness of this approach with examples from several domains. Emden R. Gansner, Yifan Hu 0001, Stephen G. Kobourov |
PacificVis | 2 |
| 2010 | On Touching Triangle Graphs
Emden R. Gansner, Yifan Hu 0001, Stephen G. Kobourov |
GD | 2 |
| 2010 | On Maximum Differential Graph Coloring
Yifan Hu 0001, Stephen G. Kobourov, Sankar Veeramoni |
GD | 1 |
| 2010 | Optimal Polygonal Representation of Planar Graphs
Emden R. Gansner, Yifan Hu 0001, Michael Kaufmann 0001, Stephen G. Kobourov |
LATIN | 2 |
| 2009 | Extending the spring-electrical model to overcome warping effectsabstractThe spring-electrical model based force directed algorithm is widely used for drawing undirected graphs, and sophisticated implementations can be very efficient for visualizing large graphs. However, our practical experience shows that in many cases, layout quality suffers as a result of non-uniform vertex density. This gives rise to warping effects in that vertices on the outskirt of the drawing are often closer to each other than those near the center, and branches in a tree-like graph tend to cling together. In this paper we propose algorithms that overcome these effects. The algorithms combine the efficiency and good global structure of the spring-electrical model, with the flexibility of the Kamada-Kawai stress model of in specifying the ideal edge length, and are very effective in overcoming the warping effects. Yifan Hu 0001, Yehuda Koren |
PacificVis | 1 |
| 2009 | GMap: Drawing Graphs as Maps
Emden R. Gansner, Yifan Hu 0001, Stephen G. Kobourov |
GD | 2 |
| 2009 | Putting recommendations on the map: visualizing clusters and relationsabstractFor users, recommendations can sometimes seem odd or counterintuitive. Visualizing recommendations can remove some of this mystery, showing how a recommendation is grouped with other choices. A drawing can also lead a user's eye to other options. Traditional 2D-embeddings of points can be used to create a basic layout, but these methods, by themselves, do not illustrate clusters and neighborhoods very well. In this paper, we propose the use of geographic maps to enhance the definition of clusters and neighborhoods, and consider the effectiveness of this approach in visualizing similarities and recommendations arising from TV shows. Emden R. Gansner, Yifan Hu 0001, Stephen G. Kobourov, Chris Volinsky |
RecSys | 2 |
| 2008 | Efficient Node Overlap Removal Using a Proximity Stress Model
Emden R. Gansner, Yifan Hu 0001 |
GD | 2 |
| 2008 | Collaborative Filtering for Implicit Feedback DatasetsabstractA common task of recommender systems is to improve customer experience through personalized recommendations based on prior implicit feedback. These systems passively track different sorts of user behavior, such as purchase history, watching habits and browsing activity, in order to model user preferences. Unlike the much more extensively researched explicit feedback, we do not have any direct input from the users regarding their preferences. In particular, we lack substantial evidence on which products consumer dislike. In this work we identify unique properties of implicit feedback datasets. We propose treating the data as indication of positive and negative preference associated with vastly varying confidence levels. This leads to a factor model which is especially tailored for implicit feedback recommenders. We also suggest a scalable optimization procedure, which scales linearly with the data size. The algorithm is used successfully within a recommender system for television shows. It compares favorably with well tuned implementations of other known methods. In addition, we offer a novel way to give explanations to recommendations given by this factor model. Yifan Hu 0001, Yehuda Koren, Chris Volinsky |
ICDM | 1 |
| 2007 | A numerical evaluation of sparse direct solvers for the solution of large sparse symmetric linear systems of equationsabstractIn recent years a number of solvers for the direct solution of large sparse symmetric linear systems of equations have been developed. These include solvers that are designed for the solution of positive definite systems as well as those that are principally intended for solving indefinite problems. In this study, we use performance profiles as a tool for evaluating and comparing the performance of serial sparse direct solvers on an extensive set of symmetric test problems taken from a range of practical applications. Nicholas I. M. Gould, Jennifer A. Scott, Yifan Hu 0001 |
ACM Trans. Math. Softw. | 3 |
| 2007 | Experiences of sparse direct symmetric solversabstractWe recently carried out an extensive comparison of the performance of state-of-the-art sparse direct solvers for the numerical solution of symmetric linear systems of equations. Some of these solvers were written primarily as research codes while others have been developed for commercial use. Our experiences of using the different packages to solve a wide range of problems arising from real applications were mixed. In this paper, we highlight some of these experiences with the aim of providing advice to both software developers and users of sparse direct solvers. We discuss key features that a direct solver should offer and conclude that while performance is an essential factor to consider when choosing a code, there are other features that a user should also consider looking for that vary significantly between packages. Jennifer A. Scott, Yifan Hu 0001 |
ACM Trans. Math. Softw. | 2 |