VLDB 2026 Research / reviewers in the wild / expert
Jian Guo 0002
dblp:96/2596-2
· DBLP profile ↗
7ranked-venue papers
3as first author
0since 2021 · last 2011
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorArtificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Data mining · 54% Information retrieval · 26% Web and social media mining · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
1 paper |
Information theory · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › text mining › topic modeling
dynamic topic model |
0.1 | 1 | 2011 | Tracking trends: incorporating term volume into temporal topic models · KDD 2011 |
Data mining › text mining
topic modeling |
0.1 | 1 | 2011 | Tracking trends: incorporating term volume into temporal topic models · KDD 2011 |
Web and social media mining › user behavior analysis
voting behavior |
0.1 | 1 | 2011 | User reputation in a comment rating environment · KDD 2011 |
Information retrieval › ranking › multi-objective ranking
diversity-aware ranking |
0.1 | 1 | 2010 | DivRank: the interplay of prestige and diversity in information networks · KDD 2010 |
Information retrieval
ranking |
0.1 | 1 | 2010 | DivRank: the interplay of prestige and diversity in information networks · KDD 2010 |
Bioinformatics and computational biology › protein function prediction
protein subcellular localization prediction |
0.1 | 1 | 2006 | TSSub: eukaryotic protein subcellular localization by extracting features from profiles · Bioinform. 2006 |
Web and social media mining
user-generated content |
0.0 | 1 | 2011 | User reputation in a comment rating environment · KDD 2011 |
Information theory › probability theory › stochastic processes
state-space models |
0.0 | 1 | 2011 | Tracking trends: incorporating term volume into temporal topic models · KDD 2011 |
Data mining › structured data mining
graph mining |
0.0 | 1 | 2010 | DivRank: the interplay of prestige and diversity in information networks · KDD 2010 |
Data mining › structured data mining › graph mining
information network analysis |
0.0 | 1 | 2010 | DivRank: the interplay of prestige and diversity in information networks · KDD 2010 |
Methods — techniques the papers use, named apart from their topics
supervised learning · 0.2autoregressive model · 0.2text feature prediction · 0.1tensor factorization · 0.1state-space model · 0.1state space model · 0.1bias smoothing · 0.1random walk · 0.1prestige and diversity optimization · 0.1support vector machine · 0.1probabilistic neural network · 0.1feature fusion · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | User reputation in a comment rating environmentabstractReputable users are valuable assets of a web site. We focus on user reputation in a comment rating environment, where users make comments about content items and rate the comments of one another. Intuitively, a reputable user posts high quality comments and is highly rated by the user community. To our surprise, we find that the quality of a comment judged editorially is almost uncorrelated with the ratings that it receives, but can be predicted using standard text features, achieving accuracy as high as the agreement between two editors! However, extracting a pure reputation signal from ratings is difficult because of data sparseness and several confounding factors in users' voting behavior. To address these issues, we propose a novel bias-smoothed tensor model and empirically show that our model significantly outperforms a number of alternatives based on Yahoo! News, Yahoo! Buzz and Epinions datasets. Bee-Chung Chen, Jian Guo 0002, Belle L. Tseng, Jie Yang 0015 |
KDD | 2 |
| 2011 | Tracking trends: incorporating term volume into temporal topic modelsabstractText corpora with documents from a range of time epochs are natural and ubiquitous in many fields, such as research papers, newspaper articles and a variety of types of recently emerged social media. People not only would like to know what kind of topics can be found from these data sources but also wish to understand the temporal dynamics of these topics and predict certain properties of terms or documents in the future. Topic models are usually utilized to find latent topics from text collections, and recently have been applied to temporal text corpora. However, most proposed models are general purpose models to which no real tasks are explicitly associated. Therefore, current models may be difficult to apply in real-world applications, such as the problems of tracking trends and predicting popularity of keywords. In this paper, we introduce a real-world task, tracking trends of terms, to which temporal topic models can be applied. Rather than building a general-purpose model, we propose a new type of topic model that incorporates the volume of terms into the temporal dynamics of topics and optimizes estimates of term volumes. In existing models, trends are either latent variables or not considered at all which limits the potential for practical use of trend information. In contrast, we combine state-space models with term volumes with a supervised learning model, enabling us to effectively predict the volume in the future, even without new documents. In addition, it is straightforward to obtain the volume of latent topics as a by-product of our model, demonstrating the superiority of utilizing temporal topic models over traditional time-series tools (e.g., autoregressive models) to tackle this kind of problem. The proposed model can be further extended with arbitrary word-level features which are evolving over time. We present the results of applying the model to two datasets with long time periods and show its effectiveness over non-trivial baselines. Liangjie Hong, Dawei Yin 0001, Jian Guo 0002, Brian D. Davison 0001 |
KDD | 3 |
| 2010 | DivRank: the interplay of prestige and diversity in information networksabstractInformation networks are widely used to characterize the relationships between data items such as text documents. Many important retrieval and mining tasks rely on ranking the data items based on their centrality or prestige in the network. Beyond prestige, diversity has been recognized as a crucial objective in ranking, aiming at providing a non-redundant and high coverage piece of information in the top ranked results. Nevertheless, existing network-based ranking approaches either disregard the concern of diversity, or handle it with non-optimized heuristics, usually based on greedy vertex selection. Qiaozhu Mei, Jian Guo 0002, Dragomir R. Radev |
KDD | 2 |
| 2008 | PairProSVM: Protein Subcellular Localization Based on Local Pairwise Profile Alignment and SVMabstractThe subcellular locations of proteins are important functional annotations. An effective and reliable subcellular localization method is necessary for proteomics research. This paper introduces a new method---PairProSVM---to automatically predict the subcellular locations of proteins. The profiles of all protein sequences in the training set are constructed by PSI-BLAST and the pairwise profile-alignment scores are used to form feature vectors for training a support vector machine (SVM) classifier. It was found that PairProSVM outperforms the methods that are based on sequence alignment and amino-acid compositions even if most of the homologous sequences have been removed. This paper also demonstrates that the performance of PairProSVM is sensitive (and somewhat proportional) to the degree of its kernel matrix meeting the Mercer's condition. PairProSVM was evaluated on Reinhardt and Hubbard's, Huang and Li's, and Gardy et al.'s protein datasets. The overall accuracies on these three datasets reach 99.3\\%, 76.5\\%, and 91.9\\%, respectively, which are higher than or comparable to those obtained by sequence alignment and by the methods compared in this paper. Man-Wai Mak, Jian Guo 0002, Sun-Yuan Kung |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2006 | TSSub: eukaryotic protein subcellular localization by extracting features from profilesabstractUNLABELLED: This paper introduces a new subcellular localization system (TSSub) for eukaryotic proteins. This system extracts features from both profiles and amino acid sequences. Four different features are extracted from profiles by four probabilistic neural network (PNN) classifiers, respectively (the amino acid composition from whole profiles; the amino acid composition from the N-terminus of profiles; the dipeptide composition from whole profiles and the amino acid composition from fragments of profiles). In addition, a support vector machine (SVM) classifier is added to implement the residue-couple feature extracted from amino acid sequences. The results from the five classifiers are fused by an additional SVM classifier. The overall accuracies of this TSSub reach 93.0 and 77.4% on Reinhardt and Hubbard's eukaryotic protein dataset and Huang and Li's eukaryotic protein dataset, respectively. The comparison with existing methods results shows TSSub provides better prediction performance than existing methods. AVAILABILITY: The web server is available from http://166.111.24.5/webtools/TSSub/index.html. Jian Guo 0002, Yuanlie Lin |
Bioinform. | 1 |
| 2005 | A novel method for protein subcellular localization: Combining residue-couple model and SVM
Jian Guo 0002, Yuanlie Lin, Zhirong Sun |
APBC | 1 |
| 2004 | A Novel Method for Protein Subcellular Localization Based on Boosting and Probabilistic Neural Network
Jian Guo 0002, Yuanlie Lin, Zhirong Sun |
APBC | 1 |