VLDB 2026 Research / reviewers in the wild / expert
Tian Guo 0002
dblp:55/3523-2
· DBLP profile ↗
12ranked-venue papers
9as first author
0since 2021 · last 2020
0000-0002-0365-0063ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-authorDatabases, data management, data science and information retrieval · 9 · 8 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Graph learning · 27% Deep learning architectures and training · 24% Trustworthy machine learning · 24% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
directed graph |
0.4 | 1 | 2020 | Low-dimensional statistical manifold embedding of directed graphs · ICLR 2020 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
low-dimensional embedding |
0.4 | 1 | 2020 | Low-dimensional statistical manifold embedding of directed graphs · ICLR 2020 |
Machine learning › Graph learning
network embedding |
0.4 | 1 | 2020 | Low-dimensional statistical manifold embedding of directed graphs · ICLR 2020 |
Machine learning › Trustworthy machine learning › interpretability › attention analysis
attention-based explanation |
0.4 | 1 | 2019 | Exploring interpretable LSTM neural networks over multi-variable data · ICML 2019 |
Machine learning › Trustworthy machine learning
interpretability |
0.4 | 1 | 2019 | Exploring interpretable LSTM neural networks over multi-variable data · ICML 2019 |
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM |
0.4 | 1 | 2019 | Exploring interpretable LSTM neural networks over multi-variable data · ICML 2019 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.4 | 1 | 2019 | Exploring interpretable LSTM neural networks over multi-variable data · ICML 2019 |
Data mining › time series analysis
time series forecasting |
0.3 | 1 | 2018 | Bitcoin Volatility Forecasting with a Glimpse into Buy and Sell Orders · ICDM 2018 |
Machine learning › Time series and sequential data › time series analysis
time series learning |
0.3 | 1 | 2017 | Hybrid Neural Networks for Learning the Trend in Time Series · IJCAI 2017 |
Machine learning › Time series and sequential data
time series analysis |
0.1 | 1 | 2019 | Exploring interpretable LSTM neural networks over multi-variable data · ICML 2019 |
Methods — techniques the papers use, named apart from their topics
temporal mixture model · 0.7rolling incremental learning · 0.7order book feature analysis · 0.7statistical manifold embedding · 0.4variable-wise hidden states · 0.4mixture attention · 0.4LSTM · 0.4long short-term memory · 0.3feature fusion · 0.3convolutional neural network · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Low-dimensional statistical manifold embedding of directed graphs
Thorben Funke, Tian Guo 0002, Alen Lancic, Nino Antulov-Fantulin |
ICLR | 2 |
| 2019 | Exploring interpretable LSTM neural networks over multi-variable dataabstractFor recurrent neural networks trained on time series with target and exogenous variables, in addition to accurate prediction, it is also desired to provide interpretable insights into the data. In this paper, we explore the structure of LSTM recurrent neural networks to learn variable-wise hidden states, with the aim to capture different dynamics in multi-variable time series and distinguish the contribution of variables to the prediction. With these variable-wise hidden states, a mixture attention mechanism is proposed to model the generative process of the target. Then we develop associated training methods to jointly learn network parameters, variable and temporal importance w.r.t the prediction of the target variable. Extensive experiments on real datasets demonstrate enhanced prediction performance by capturing the dynamics of different variables. Meanwhile, we evaluate the interpretation results both qualitatively and quantitatively. It exhibits the prospect as an end-to-end framework for both forecasting and knowledge extraction over multi-variable data. Tian Guo 0002, Tao Lin 0004, Nino Antulov-Fantulin |
ICML | 1 |
| 2018 | Bitcoin Volatility Forecasting with a Glimpse into Buy and Sell OrdersabstractBitcoin is one of the most prominent decentralized digital cryptocurrencies. Ability to understand which factors drive the fluctuations of the Bitcoin price and to what extent they are predictable is interesting both from the theoretical and practical perspective. In this paper, we study the problem of the Bitcoin short-term volatility forecasting based on volatility history and order book data. Order book, consisting of buy and sell orders over time, reflects the intention of the market and is closely related to the evolution of volatility. We propose temporal mixture models capable of adaptively exploiting both volatility history and order book features. By leveraging rolling and incremental learning and evaluation procedures, we demonstrate the prediction performance of our model as well as studying the robustness, in comparison to a variety of statistical and machine learning baselines. Meanwhile, our temporal mixture model enables to decipher the time-varying effect of order book features on volatility. It demonstrates the prospect of our temporal mixture model as an interpretable forecasting framework over heterogeneous Bitcoin data. Tian Guo 0002, Albert Bifet, Nino Antulov-Fantulin |
ICDM | 1 |
| 2017 | Hybrid Neural Networks for Learning the Trend in Time SeriesabstractThe trend of time series characterizes the intermediate upward and downward behaviour of time series. Learning and forecasting the trend in time series data play an important role in many real applications, ranging from resource allocation in data centers, load schedule in smart grid, and so on. Inspired by the recent successes of neural networks, in this paper we propose TreNet, a novel end-to-end hybrid neural network to learn local and global contextual features for predicting the trend of time series. TreNet leverages convolutional neural networks (CNNs) to extract salient features from local raw data of time series. Meanwhile, considering the long-range dependency existing in the sequence of historical trends of time series, TreNet uses a long-short term memory recurrent neural network (LSTM) to capture such dependency. Then, a feature fusion layer is to learn joint representation for predicting the trend. TreNet demonstrates its effectiveness by outperforming CNN, LSTM, the cascade of CNN and LSTM, Hidden Markov Model based method and various kernel based baselines on real datasets. Tao Lin 0004, Tian Guo 0002, Karl Aberer |
IJCAI | 2 |
| 2016 | Data Summarization with Social ContextsabstractWhile social data is being widely used in various applications such as sentiment analysis and trend prediction, its sheer size also presents great challenges for storing, sharing and processing such data. These challenges can be addressed by data summarization which transforms the original dataset into a smaller, yet still useful, subset. Existing methods find such subsets with objective functions based on data properties such as representativeness or informativeness but do not exploit social contexts, which are distinct characteristics of social data. Further, till date very little work has focused on topic preserving data summarization, despite the abundant work on topic modeling. This is a challenging task for two reasons. First, since topic model is based on latent variables, existing methods are not well-suited to capture latent topics. Second, it is difficult to find such social contexts that provide valuable information for building effective topic-preserving summarization model. To tackle these challenges, in this paper, we focus on exploiting social contexts to summarize social data while preserving topics in the original dataset. We take Twitter data as a case study. Through analyzing Twitter data, we discover two social contexts which are important for topic generation and dissemination, namely (i) CrowdExp topic score that captures the influence of both the crowd and the expert users in Twitter and (ii) Retweet topic score that captures the influence of Twitter users' actions. We conduct extensive experiments on two real-world Twitter datasets using two applications. The experimental results show that, by leveraging social contexts, our proposed solution can enhance topic-preserving data summarization and improve application performance by up to 18%. Hao Zhuang 0002, Rameez Rahman, Xia Ben Hu, Tian Guo 0002, Pan Hui 0001, Karl Aberer |
CIKM | 4 |
| 2016 | Robust Online Time Series Prediction with Recurrent Neural NetworksabstractTime series forecasting for streaming data plays an important role in many real applications, ranging from IoT systems, cyber-networks, to industrial systems and healthcare. However the real data is often complicated with anomalies and change points, which can lead the learned models deviating from the underlying patterns of the time series, especially in the context of online learning mode. In this paper we present an adaptive gradient learning method for recurrent neural networks (RNN) to forecast streaming time series in the presence of anomalies and change points. We explore the local features of time series to automatically weight the gradients of the loss of the newly available observations with distributional properties of the data in real time. We perform extensive experimental analysis on both synthetic and real datasets to evaluate the performance of the proposed method. Tian Guo 0002, Zhao Xu 0001, Xin Yao 0001, Karl Aberer, Koichi Funaya |
DSAA | 1 |
| 2016 | A Distributed Mining Framework for Influence in Evolving Entities
Tian Guo 0002, Karl Aberer |
EDBT | 1 |
| 2016 | Efficient Distributed Decision Trees for Robust Regression
Tian Guo 0002, Konstantin Kutzkov, Mohamed Ahmed 0001, Jean-Paul Calbimonte, Karl Aberer |
ECML/PKDD (2) | 1 |
| 2015 | SigCO: Mining significant correlations via a distributed real-time computation engineabstractThe dramatic rise of time-series data produced in a variety of contexts, such as stock markets, mobile sensing, sensor networks, data centre monitoring, etc., has fuelled the development of large-scale distributed real-time computation systems (e.g., Apache Storm, Samza, Spark Streaming, S4, etc.). However, it is still unclear how certain time series mining tasks could be performed using such new emerging systems. In this paper, we focus on the task of efficiently discovering statistically significant correlations among a large number of time series via a distributed realtime computation engine. We propose a framework referred to as SigCO. In SigCO, we put forward a novel partition-aware data shuffling, which is able to adaptively shuffle time series data only to the relevant nodes of the distributed real-time computation engine. On the other hand, in SigCO we design a δ-hypercube structure based correlation computation approach which is capable of pruning unnecessary correlation computations. Finally, our extensive experimental evaluations on real and synthetic datasets establish that SigCO outperforms the baseline approaches in terms of diverse performance metrics. Tian Guo 0002, Jean-Paul Calbimonte, Hao Zhuang 0002, Karl Aberer |
IEEE BigData | 1 |
| 2015 | Fast Distributed Correlation Discovery Over Streaming Time-Series DataabstractThe dramatic rise of time-series data in a variety of contexts, such as social networks, mobile sensing, data centre monitoring, etc., has fuelled interest in obtaining real-time insights from such data using distributed stream processing systems. One such extremely valuable insight is the discovery of correlations in real-time from large-scale time-series data. A key challenge in discovering correlations is that the number of time-series pairs that have to be analyzed grows quadratically in the number of time-series, giving rise to a quadratic increase in both computation cost and communication cost between the cluster nodes in a distributed environment. To tackle the challenge, we propose a framework called AEGIS. AEGIS exploits well-established statistical properties to dramatically prune the number of time-series pairs that have to be evaluated for detecting interesting correlations. Our extensive experimental evaluations on real and synthetic datasets establish the efficacy of AEGIS over baselines. Tian Guo 0002, Saket Sathe 0001, Karl Aberer |
CIKM | 1 |
| 2014 | Online Indexing and Distributed Querying Model-View Sensor Data in the Cloud
Tian Guo 0002, Thanasis G. Papaioannou, Hao Zhuang 0002, Karl Aberer |
DASFAA (1) | 1 |
| 2013 | Model-view sensor data management in the cloudabstractInfinite nature of sensor data poses a serious challenge for query processing even in a cloud infrastructure. Model-based sensor data approximation reduces the amount of data for query processing, but all modeled segments need to be scanned, in the worst case. In this paper, we propose an innovative index for modeled segments in key-value stores, namely KVI-index. KVI-index has an in-memory tree component and a secondary structure materialized in the key-value store that maps the tree nodes to the modeled data segments. Then, we introduce a KVI-index-Scan-MapReduce hybrid approach to perform efficient query processing. As proved by a series of experiments in a real private cloud infrastructure, our approach outperforms in query response time and index updating efficiency both Hadoop-based parallel processing of the raw sensor data and multiple alternative indexing approaches of model-view data. Tian Guo 0002, Thanasis G. Papaioannou, Karl Aberer |
IEEE BigData | 1 |