VLDB 2026 Research / reviewers in the wild / expert
Alexander Patrick Mathews
dblp:169/3191
· DBLP profile ↗
10ranked-venue papers
4as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Vision and language · 34% Deep learning architectures and training · 21% Probabilistic and Bayesian machine learning · 12% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 50% Web and social media mining · 22% Data mining · 22% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% |
Topics — the 27 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
image captioning |
1.2 | 4 | 2020 | Transform and Tell: Entity-Aware News Image Captioning · CVPR 2020 SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018 SentiCap: Generating Image Descriptions with Sentiments · AAAI 2016 |
Machine learning › Deep learning architectures and training › neural operator
fourier neural operator |
0.7 | 1 | 2023 | Factorized Fourier Neural Operators · ICLR 2023 |
Machine learning › Deep learning architectures and training
neural operator |
0.7 | 1 | 2023 | Factorized Fourier Neural Operators · ICLR 2023 |
Computational science and engineering
partial differential equation solver |
0.7 | 1 | 2023 | Factorized Fourier Neural Operators · ICLR 2023 |
Computer vision › Vision and language › image captioning › controllable image captioning
stylized image captioning |
0.5 | 2 | 2018 | SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018 Captioning Images Using Different Styles · ACM Multimedia 2015 |
Machine learning › Graph learning › dynamic graph learning
dynamic graph modeling |
0.5 | 1 | 2021 | Radflow: A Recurrent, Aggregated, and Decomposable Model for Networks of Time Series · WWW 2021 |
Machine learning › Graph learning
network embedding |
0.5 | 1 | 2021 | Radflow: A Recurrent, Aggregated, and Decomposable Model for Networks of Time Series · WWW 2021 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
neural point process |
0.5 | 1 | 2021 | UNIPoint: Universally Approximating Point Processes Intensities · AAAI 2021 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process |
0.5 | 1 | 2021 | UNIPoint: Universally Approximating Point Processes Intensities · AAAI 2021 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.5 | 1 | 2021 | UNIPoint: Universally Approximating Point Processes Intensities · AAAI 2021 |
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation |
0.5 | 1 | 2021 | UNIPoint: Universally Approximating Point Processes Intensities · AAAI 2021 |
Data mining › time series analysis
time series forecasting |
0.5 | 1 | 2021 | Radflow: A Recurrent, Aggregated, and Decomposable Model for Networks of Time Series · WWW 2021 |
Visualization and visual analytics
time series visualization |
0.5 | 1 | 2021 | AttentionFlow: Visualising Influence in Networks of Time Series · WSDM 2021 |
Computer vision › Vision and language › image captioning › fine-grained image captioning
entity-aware captioning |
0.4 | 1 | 2020 | Transform and Tell: Entity-Aware News Image Captioning · CVPR 2020 |
Computer vision › Vision and language › image captioning
news image captioning |
0.4 | 1 | 2020 | Transform and Tell: Entity-Aware News Image Captioning · CVPR 2020 |
Information retrieval › text summarization
comparative summarization |
0.4 | 1 | 2019 | Comparative Document Summarisation via Classification · AAAI 2019 |
Information retrieval › text summarization
extractive summarization |
0.4 | 1 | 2019 | Comparative Document Summarisation via Classification · AAAI 2019 |
Information retrieval
text summarization |
0.4 | 1 | 2019 | Comparative Document Summarisation via Classification · AAAI 2019 |
Machine learning › Generative modeling
style transfer |
0.3 | 1 | 2018 | SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018 |
Natural language and speech › Language models and text generation › controllable text generation
text style transfer |
0.3 | 1 | 2018 | SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018 |
Robotics › Motion planning and robot control › robot control › contact control › contact task control › robot force control
adaptive impedance control |
0.2 | 1 | 2016 | SentiCap: Generating Image Descriptions with Sentiments · AAAI 2016 |
Natural language and speech › Language models and text generation › text generation › neural text generation
recurrent neural network text generation |
0.2 | 1 | 2016 | SentiCap: Generating Image Descriptions with Sentiments · AAAI 2016 |
Computer vision › Vision and language
cross-modal attention |
0.1 | 1 | 2020 | Transform and Tell: Entity-Aware News Image Captioning · CVPR 2020 |
Machine learning and data management
data selection |
0.1 | 1 | 2019 | Comparative Document Summarisation via Classification · AAAI 2019 |
Computer vision › Vision and language
multimodal representation |
0.1 | 1 | 2018 | SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic representation |
0.1 | 1 | 2018 | SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018 |
Computer vision › Vision and language › image captioning
diverse image captioning |
0.1 | 1 | 2015 | Captioning Images Using Different Styles · ACM Multimedia 2015 |
Methods — techniques the papers use, named apart from their topics
tree ring encoding · 1.5recurrent neural network · 1.5line chart · 1.5ego network visualization · 1.5multi-head attention · 1.4tensor factorization · 1.3fourier transform · 1.3multi-layer decomposition · 1.0stone-weierstrass theorem · 0.5basis functions · 0.5transformer language model · 0.4image and face embeddings · 0.4byte-pair encoding · 0.4maximum mean discrepancy · 0.4gradient-based optimisation · 0.4binary classification · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Factorized Fourier Neural Operators
Alasdair Tran, Alexander Patrick Mathews, Lexing Xie, Cheng Soon Ong |
ICLR | 2 |
| 2021 | UNIPoint: Universally Approximating Point Processes IntensitiesabstractPoint processes are a useful mathematical tool for describing events over time, and so there are many recent approaches for representing and learning them. One notable open question is how to precisely describe the flexibility of point process models and whether there exists a general model that can represent all point processes. Our work bridges this gap. Focusing on the widely used event intensity function representation of point processes, we provide a proof that a class of learnable functions can universally approximate any valid intensity function. The proof connects the well known Stone-Weierstrass Theorem for function approximation, the uniform density of non-negative continuous functions using a transfer functions, the formulation of the parameters of a piece-wise continuous functions as a dynamic system, and a recurrent neural network implementation for capturing the dynamics. Using these insights, we design and implement UNIPoint, a novel neural point process model, using recurrent neural networks to parameterise sums of basis function upon each event. Evaluations on synthetic and real world datasets show that this simpler representation performs better than Hawkes process variants and more complex neural network-based approaches. We expect this result will provide a practical basis for selecting and tuning models, as well as furthering theoretical work on representational complexity and learnability. Alexander Soen, Alexander Patrick Mathews, Daniel Grixti-Cheng, Lexing Xie |
AAAI | 2 |
| 2021 | AttentionFlow: Visualising Influence in Networks of Time SeriesabstractThe collective attention on online items such as web pages, search terms, and videos reflects trends that are of social, cultural, and economic interest. Moreover, attention trends of different items exhibit mutual influence via mechanisms such as hyperlinks or recommendations. Many visualisation tools exist for time series, network evolution, or network influence; however, few systems connect all three. In this work, we present AttentionFlow, a new system to visualise networks of time series and the dynamic influence they have on one another. Centred around an ego node, our system simultaneously presents the time series on each node using two visual encodings: a tree ring for an overview and a line chart for details. AttentionFlow supports interactions such as overlaying time series of influence, and filtering neighbours by time or flux. We demonstrate AttentionFlow using two real-world datasets, VevoMusic and WikiTraffic. We show that attention spikes in songs can be explained by external events such as major awards, or changes in the network such as the release of a new song. Separate case studies also demonstrate how an artist's influence changes over their career, and that correlated Wikipedia traffic is driven by cultural interests. More broadly, AttentionFlow can be generalised to visualise networks of time series on physical infrastructures such as road networks, or natural phenomena such as weather and geological measurements. Minjeong Shin, Alasdair Tran, Alexander Patrick Mathews, Georgiana Lyall, Lexing Xie |
WSDM | 4 |
| 2021 | Radflow: A Recurrent, Aggregated, and Decomposable Model for Networks of Time SeriesabstractWe propose a new model for networks of time series that influence each other. Graph structures among time series are found in diverse domains, such as web traffic influenced by hyperlinks, product sales influenced by recommendation, or urban transport volume influenced by road networks and weather. There has been recent progress in graph modeling and in time series forecasting, respectively, but an expressive and scalable approach for a network of series does not yet exist. We introduce Radflow, a novel model that embodies three key ideas: a recurrent neural network to obtain node embeddings that depend on time, the aggregation of the flow of influence from neighboring nodes with multi-head attention, and the multi-layer decomposition of time series. Radflow naturally takes into account dynamic networks where nodes and edges change over time, and it can be used for prediction and data imputation tasks. On real-world datasets ranging from a few hundred to a few hundred thousand nodes, we observe that Radflow variants are the best performing model across a wide range of settings. The recurrent component in Radflow also outperforms N-BEATS, the state-of-the-art time series model. We show that Radflow can learn different trends and seasonal patterns, that it is robust to missing nodes and edges, and that correlated temporal patterns among network neighbors reflect influence strength. We curate WikiTraffic, the largest dynamic network of time series with 366K nodes and 22M time-dependent links spanning five years. This dataset provides an open benchmark for developing models in this area, with applications that include optimizing resources for the web. More broadly, Radflow has the potential to improve forecasts in correlated time series networks such as the stock market, and impute missing measurements in geographically dispersed networks of natural phenomena. Alasdair Tran, Alexander Patrick Mathews, Cheng Soon Ong, Lexing Xie |
WWW | 2 |
| 2020 | Transform and Tell: Entity-Aware News Image CaptioningabstractWe propose an end-to-end model which generates captions for images embedded in news articles. News images present two key challenges: they rely on real-world knowledge, especially about named entities; and they typically have linguistically rich captions that include uncommon words. We address the first challenge by associating words in the caption with faces and objects in the image, via a multi-modal, multi-head attention mechanism. We tackle the second challenge with a state-of-the-art transformer language model that uses byte-pair-encoding to generate captions as a sequence of word parts. On the GoodNews dataset, our model outperforms the previous state of the art by a factor of four in CIDEr score (13 to 54). This performance gain comes from a unique combination of language models, word representation, image embeddings, face embeddings, object embeddings, and improvements in neural network design. We also introduce the NYTimes800k dataset which is 70% larger than GoodNews, has higher article quality, and includes the locations of images within articles as an additional contextual cue. Alasdair Tran, Alexander Patrick Mathews, Lexing Xie |
CVPR | 2 |
| 2019 | Comparative Document Summarisation via ClassificationabstractThispaperconsidersextractivesummarisationinacomparative setting: given two or more document groups (e.g., separated by publication time), the goal is to select a small number of documents that are representative of each group, and also maximally distinguishable from other groups. We formulate a set of new objective functions for this problem that connect recent literature on document summarisation, interpretable machine learning, and data subset selection. In particular, by casting the problem as a binary classification amongst different groups, we derive objectives based on the notion of maximum mean discrepancy, as well as a simple yet effective gradient-based optimisation strategy. Our new formulation allows scalable evaluations of comparative summarisation as a classification task, both automatically and via crowd-sourcing. To this end, we evaluate comparative summarisation methods on a newly curated collection of controversial news topics over 13months.Weobserve thatgradient-based optimisationoutperforms discrete and baseline approaches in 15 out of 24 different automatic evaluation settings. In crowd-sourced evaluations, summaries from gradient optimisation elicit 7% more accurate classification from human workers than discrete optimisation. Our result contrasts with recent literature on submodular data subset selection that favours discrete optimisation. We posit that our formulation of comparative summarisation will prove useful in a diverse range of use cases such as comparing content sources, authors, related topics, or distinct view points. Umanga Bista, Alexander Patrick Mathews, Minjeong Shin, Aditya Krishna Menon, Lexing Xie |
AAAI | 2 |
| 2018 | SemStyle: Learning to Generate Stylised Image Captions Using Unaligned TextabstractLinguistic style is an essential part of written communication, with the power to affect both clarity and attractiveness. With recent advances in vision and language, we can start to tackle the problem of generating image captions that are both visually grounded and appropriately styled. Existing approaches either require styled training captions aligned to images or generate captions with low relevance. We develop a model that learns to generate visually relevant styled captions from a large corpus of styled text without aligned images. The core idea of this model, called SemStyle, is to separate semantics and style. One key component is a novel and concise semantic term representation generated using natural language processing techniques and frame semantics. In addition, we develop a unified language model that decodes sentences with diverse word choices and syntax for different styles. Evaluations, both automatic and manual, show captions from SemStyle preserve image semantics, are descriptive, and are style shifted. More broadly, this work provides possibilities to learn richer image descriptions from the plethora of linguistic data available on the web. Alexander Patrick Mathews, Lexing Xie, Xuming He 0001 |
CVPR | 1 |
| 2016 | SentiCap: Generating Image Descriptions with SentimentsabstractThe recent progress on image recognition and language modeling is making automatic description of image content a reality. However, stylized, non-factual aspects of the written description are missing from the current systems. One such style is descriptions with emotions, which is commonplace in everyday communication, and influences decision-making and interpersonal relationships. We design a system to describe an image with emotions, and present a model that automatically generates captions with positive or negative sentiments. We propose a novel switching recurrent neural network with word-level regularization, which is able to produce emotional image captions using only 2000+ training sentences containing sentiments. We evaluate the captions with different automatic and crowd-sourcing metrics. Our model compares favourably in common quality metrics for image captioning. In 84.6% of cases the generated positive captions were judged as being at least as descriptive as the factual captions. Of these positive captions 88% were confirmed by the crowd-sourced workers as having the appropriate sentiment. Alexander Patrick Mathews, Lexing Xie, Xuming He 0001 |
AAAI | 1 |
| 2015 | Captioning Images Using Different StylesabstractI develop techniques that can be used to incorporate stylistic objectives into existing image captioning systems. Style is generally a very tricky concept to define, thus I concentrate on two specific components of style. First I develop a technique for predicting how people will name visual objects. I demonstrate that this technique could be used to generate captions with human like naming conventions. Full details are available in a recent publication. Second I outline a system for generating sentences which express a strong positive or negative sentiment. Finally I present two possible future directions which are aimed at modelling style more generally. These are learning to imitate an individuals captioning style and generating a diverse set of captions for a single image. Alexander Patrick Mathews |
ACM Multimedia | 1 |
| 2015 | Choosing Basic-Level Concept Names Using Visual and Language ContextabstractWe study basic-level categories for describing visual concepts, and empirically observe context-dependant basic level names across thousands of concepts. We propose methods for predicting basic-level names using a series of classification and ranking tasks, producing the first large scale catalogue of basic-level names for hundreds of thousands of images depicting thousands of visual concepts. We also demonstrate the usefulness of our method with a picture-to-word task, showing strong improvement over recent work by Ordonez et al, by modeling of both visual and language context. Our study suggests that a model for naming visual concepts is an important part of any automatic image/video captioning and visual story-telling system. Alexander Patrick Mathews, Lexing Xie, Xuming He 0001 |
WACV | 1 |