Fan Guo 0006

dblp:55/6662-6 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
0since 2021 · last 2014
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 3 first-authorArtificial intelligence and machine learning · 6 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Information retrieval · 68% Data mining · 32%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 92% Computational social science and digital humanities · 8%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 50% High-performance computing · 50%
Artificial intelligence
1 paper
Graph learning · 50% Probabilistic and Bayesian machine learning · 50%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 15 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › user behavior › search behavior
click model
0.222009
Click chain model in web search · WWW 2009
Efficient multiple-click models in web search · WSDM 2009
Information retrieval › ranking
relevance estimation
0.222009
Click chain model in web search · WWW 2009
BBM: bayesian browsing model from petabyte-scale data · KDD 2009
Data mining
multimodal data mining
0.112010
QMAS: Querying, Mining and Summarization of Multi-modal Databases · ICDM 2010
Data mining › anomaly detection
outlier detection
0.112010
QMAS: Querying, Mining and Summarization of Multi-modal Databases · ICDM 2010
Information retrieval › user behavior
click log analysis
0.112009
BBM: bayesian browsing model from petabyte-scale data · KDD 2009
Information retrieval › user behavior › search behavior › click model
position bias
0.112009
Efficient multiple-click models in web search · WSDM 2009
Multimedia analysis and retrieval › cross-modal retrieval
multimodal image search
0.112008
C-DEM: a multi-modal query system for Drosophila Embryo databases · Proc. VLDB Endow. 2008
Machine learning › Graph learning
dynamic graph
0.112007
Recovering temporally rewiring networks: a model-based approach · ICML 2007
Bioinformatics and computational biology › gene expression analysis › microarray data analysis
array CGH analysis
0.112007
GIMscan: A New Statistical Method for Analyzing Whole-Genome Array CGH Data · RECOMB 2007
Bioinformatics and computational biology
cancer genomics
0.112007
GIMscan: A New Statistical Method for Analyzing Whole-Genome Array CGH Data · RECOMB 2007
Bioinformatics and computational biology › cancer genomics
copy number analysis
0.112007
GIMscan: A New Statistical Method for Analyzing Whole-Genome Array CGH Data · RECOMB 2007
Data mining
clustering
0.012010
QMAS: Querying, Mining and Summarization of Multi-modal Databases · ICDM 2010
Information retrieval › ranking › result ranking
search result ranking
0.012009
Click chain model in web search · WWW 2009
Data mining
time series analysis
0.012008
Cut-and-stitch: efficient parallel learning of linear dynamical systems on smps · KDD 2008
Computational social science and digital humanities
social network analysis
0.012007
Recovering temporally rewiring networks: a model-based approach · ICML 2007

Methods — techniques the papers use, named apart from their topics

expectation-maximization · 0.3bayesian inference · 0.2wavelet feature extraction · 0.2random walk with restart · 0.2principal component analysis · 0.2OpenMP · 0.2time series · 0.1hidden temporal exponential random graph model · 0.1summarization · 0.1importance sampling · 0.1statistical methods · 0.1
YearPublicationVenuePosition
2014 QuMinS: Fast and scalable querying, mining and summarizing multi-modal databases
Robson L. F. Cordeiro, Fan Guo 0006, Donna S. Haverkamp, James H. Horne, Ellen K. Hughes, Gunhee Kim, Luciana A. S. Romani, Priscila P. Coltri, Tamires T. Souza, Agma J. M. Traina, Caetano Traina Jr., Christos Faloutsos
Inf. Sci.2
2011 MultiAspectForensics: Pattern Mining on Large-Scale Heterogeneous Networks with Tensor Analysis
abstract
Modern applications such as web knowledge base, network traffic monitoring and online social networks have made available an unprecedented amount of network data with rich types of interactions carrying multiple attributes, for instance, port number and time tick in the case of network traffic. The design of algorithms to leverage this structured relationship with the power of computing to assist researchers and practitioners for better understanding, exploration and navigation of this space of information has become a challenging, albeit rewarding, topic in social network analysis and data mining. The constantly growing scale and enriching genres of network data always demand higher levels of efficiency, robustness and generalizability where existing approaches with successes on small, homogeneous network data are likely to fall short. We introduce MultiAspectForensics, a handy tool to automatically detect and visualize novel sub graph patterns within a local community of nodes in a heterogenous network, such as a set of vertices that form a dense bipartite graph whose edges share exactly the same set of attributes. We apply the proposed method on three data sets from distinct application domains, present empirical results and discuss insights derived from these patterns discovered. Our algorithm, built on scalable tensor analysis procedures, captures spectral properties of network data and reveals informative signals for subsequent domain-specific study and investigation, such as suspicious port-scanning activities in the scenario of cyber-security monitoring.
Koji Maruhashi, Fan Guo 0006, Christos Faloutsos
ASONAM2
2010 QMAS: Querying, Mining and Summarization of Multi-modal Databases
abstract
Given a large collection of images, very few of which have labels, how can we guess the labels of the remaining majority, and how can we spot those images that need brand new labels, different from the existing ones? Current automatic labeling techniques usually scale super linearly with the data size, and/or they fail when only a tiny amount of labeled data is provided. In this paper, we propose QMAS (Querying, Mining And Summarization of Multi-modal Databases), a fast solution to the following problems: (i) low-labor labeling (L3) – given a collection of images, very few of which are labeled with keywords, find the most suitable labels for the remaining ones, and (ii) mining and attention routing – in the same setting, find clusters, the top-NO outlier images, and the top-NR representative images. We report experiments on real satellite images, two large sets (1.5GB and 2.25GB) of proprietary images and a smaller set (17MB) of public images. We show that QMAS scales linearly with the data size, being up to 40 times faster than top competitors (GCap), obtaining better or equal accuracy. In contrast to other methods, QMAS does low-labor labeling (L3), that is, it works even with tiny initial label sets. It also solves both presented problems and spots tiles that potentially require new labels.
Robson L. F. Cordeiro, Fan Guo 0006, Donna S. Haverkamp, James H. Horne, Ellen K. Hughes, Gunhee Kim, Agma J. M. Traina, Caetano Traina Jr., Christos Faloutsos
ICDM2
2010 Bayesian Browsing Model: Exact Inference of Document Relevance from Petabyte-Scale Data
abstract
A fundamental challenge in utilizing Web search click data is to infer user-perceived relevance from the search log. Not only is the inference a difficult problem involving statistical reasonings but the bulky size, together with the ever-increasing nature, of the log data imposes extra requirements on scalability. In this paper, we propose the Bayesian Browsing Model (BBM), which performs exact inference of the document relevance, only requires a single pass of the data (i.e., the optimal scalability), and is shown effective. We present two sets of experiments to evaluate the model effectiveness and scalability. On the first set of over 50 million search instances of 1.1 million distinct queries, BBM outperforms the state-of-the-art competitor by 29.2% in log-likelihood while being 57 times faster. On the second click log set, spanning a quarter of petabyte, we showcase the scalability of BBM: we implemented it on a commercial MapReduce cluster, and it took only 3 hours to compute the relevance for 1.15 billion distinct query-URL pairs.
Chao Liu 0001, Fan Guo 0006, Christos Faloutsos
ACM Trans. Knowl. Discov. Data2
2009 BBM: bayesian browsing model from petabyte-scale data
abstract
Given a quarter of petabyte click log data, how can we estimate the relevance of each URL for a given query? In this paper, we propose the Bayesian Browsing Model (BBM), a new modeling technique with following advantages: (a) it does exact inference; (b) it is single-pass and parallelizable; (c) it is effective.
Chao Liu 0001, Fan Guo 0006, Christos Faloutsos
KDD2
2009 Efficient multiple-click models in web search
abstract
Many tasks that leverage web search users' implicit feedback rely on a proper and unbiased interpretation of user clicks. Previous eye-tracking experiments and studies on explaining position-bias of user clicks provide a spectrum of hypotheses and models on how an average user examines and possibly clicks web documents returned by a search engine with respect to the submitted query. In this paper, we attempt to close the gap between previous work, which studied how to model a single click, and the reality that multiple clicks on web documents in a single result page are not uncommon. Specifically, we present two multiple-click models: the independent click model (ICM) which is reformulated from previous work, and the dependent click model (DCM) which takes into consideration dependencies between multiple clicks. Both models can be efficiently learned with linear time and space complexities. More importantly, they can be incrementally updated as new click logs flow in. These are well-demanded properties in reality.
Fan Guo 0006, Chao Liu 0001, Yi Min Wang
WSDM1
2009 Click chain model in web search
abstract
Given a terabyte click log, can we build an efficient and effective click model? It is commonly believed that web search click logs are a gold mine for search business, because they reflect users' preference over web documents presented by the search engine. Click models provide a principled approach to inferring user-perceived relevance of web documents, which can be leveraged in numerous applications in search businesses. Due to the huge volume of click data, scalability is a must.We present the click chain model (CCM), which is based on a solid, Bayesian framework. It is both scalable and incremental, perfectly meeting the computational challenges imposed by the voluminous click logs that constantly grow. We conduct an extensive experimental study on a data set containing 8.8 million query sessions obtained in July 2008 from a commercial search engine. CCM consistently outperforms two state-of-the-art competitors in a number of metrics, with over 9.7% better log-likelihood, over 6.2% better click perplexity and much more robust (up to 30%) prediction of the first and the last clicked position.
Fan Guo 0006, Chao Liu 0001, Anitha Kannan, Tom Minka, Michael J. Taylor 0001, Yi Min Wang, Christos Faloutsos
WWW1
2008 Cut-and-stitch: efficient parallel learning of linear dynamical systems on smps
abstract
Multi-core processors with ever increasing number of cores per chip are becoming prevalent in modern parallel computing. Our goal is to make use of the multi-core as well as multi-processor architectures to speed up data mining algorithms. Specifically, we present a parallel algorithm for approximate learning of Linear Dynamical Systems (LDS), also known as Kalman Filters (KF). LDSs are widely used in time series analysis such as motion capture modeling, visual tracking etc. We propose Cut-And-Stitch (CAS), a novel method to handle the data dependencies from the chain structure of hidden variables in LDS, so as to parallelize the EM-based parameter learning algorithm. We implement the algorithm using OpenMP on both a supercomputer and a quad-core commercial desktop. The experimental results show that parallel algorithms using Cut-And-Stitch achieve comparable accuracy and almost linear speedups over the serial version. In addition, Cut-And-Stitch can be generalized to other models with similar linear structures such as Hidden Markov Models (HMM) and Switching Kalman Filters (SKF).
Lei Li 0005, Wenjie Fu 0002, Fan Guo 0006, Todd C. Mowry, Christos Faloutsos
KDD3
2008 C-DEM: a multi-modal query system for Drosophila Embryo databases
abstract
The amount of biological data publicly available has experienced an exponential growth as the technology advances. Online databases are now playing an important role as information repositories as well as easily accessible platforms for researchers to communicate and contribute. Recent research projects in image bioinformatics produce a number of databases of images, which visualize the spatial expression pattern of a gene (eg. "fj"), and most of which also have one or several annotation keywords (eg., "embryonic hindgut"). C-DEM is an online system for Drosophila (= fruit-fly) Embryo images Mining. It supports queries from all three modalities to all three, namely, (a) genes, (b) images of gene expression, and (c) annotation keywords of the images. Thus, it can find images that are similar to a given image, and/or related to the desirable annotation keywords, and/or related to specific genes. Typical queries are what are most suitable keywords to assign to image insitu28465.jpg or find images that are related to gene "fj", and to the keyword "embryonic hindgut" . C-DEM uses state-of-the-art feature extraction methods for images (wavelets and principal component analysis). It envisions the whole database as a tri-partite graph (one type for each modality), and it uses fast and flexible proximity measures, namely, random walk with restarts (RWR). In addition to flexible querying, C-DEM allows for navigation: the user can click on the results of an earlier query (image thumbnails and/or keywords and/or genes), and the system will report the most related images (and keywords, and genes). The demo is on a real Drosophila Embryo database, with 10,204 images, 2,969 distinct genes, and 113 annotation keywords. The query response time is below one second on a commodity desktop.
Fan Guo 0006, Lei Li 0005, Christos Faloutsos, Eric P. Xing
Proc. VLDB Endow.1
2007 Recovering temporally rewiring networks: a model-based approach
abstract
A plausible representation of relational information among entities in dynamic systems such as a living cell or a social community is a stochastic network which is topologically rewiring and semantically evolving over time. While there is a rich literature on modeling static or temporally invariant networks, much less has been done toward modeling the dynamic processes underlying rewiring networks, and on recovering such networks when they are not observable. We present a class of hidden temporal exponential random graph models (htERGMs) to study the yet unexplored topic of modeling and recovering temporally rewiring networks from time series of node attributes such as activities of social actors or expression levels of genes. We show that one can reliably infer the latent time-specific topologies of the evolving networks from the observation. We report empirical results on both synthetic data and a Drosophila lifecycle gene expression data set, in comparison with a static counterpart of htERGM.
Fan Guo 0006, Steve Hanneke, Wenjie Fu 0002, Eric P. Xing
ICML1
2007 GIMscan: A New Statistical Method for Analyzing Whole-Genome Array CGH Data
Yanxin Shi, Fan Guo 0006, Wei Wu 0023, Eric P. Xing
RECOMB2