Alexander Patrick Mathews

dblp:169/3191 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Vision and language · 34% Deep learning architectures and training · 21% Probabilistic and Bayesian machine learning · 12%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 50% Web and social media mining · 22% Data mining · 22%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
image captioning
1.242020
Transform and Tell: Entity-Aware News Image Captioning · CVPR 2020
SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018
SentiCap: Generating Image Descriptions with Sentiments · AAAI 2016
Machine learning › Deep learning architectures and training › neural operator
fourier neural operator
0.712023
Factorized Fourier Neural Operators · ICLR 2023
Machine learning › Deep learning architectures and training
neural operator
0.712023
Factorized Fourier Neural Operators · ICLR 2023
Computational science and engineering
partial differential equation solver
0.712023
Factorized Fourier Neural Operators · ICLR 2023
Computer vision › Vision and language › image captioning › controllable image captioning
stylized image captioning
0.522018
SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018
Captioning Images Using Different Styles · ACM Multimedia 2015
Machine learning › Graph learning › dynamic graph learning
dynamic graph modeling
0.512021
Radflow: A Recurrent, Aggregated, and Decomposable Model for Networks of Time Series · WWW 2021
Machine learning › Graph learning
network embedding
0.512021
Radflow: A Recurrent, Aggregated, and Decomposable Model for Networks of Time Series · WWW 2021
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
neural point process
0.512021
UNIPoint: Universally Approximating Point Processes Intensities · AAAI 2021
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process
0.512021
UNIPoint: Universally Approximating Point Processes Intensities · AAAI 2021
Machine learning › Deep learning architectures and training
recurrent neural network
0.512021
UNIPoint: Universally Approximating Point Processes Intensities · AAAI 2021
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation
0.512021
UNIPoint: Universally Approximating Point Processes Intensities · AAAI 2021
Data mining › time series analysis
time series forecasting
0.512021
Radflow: A Recurrent, Aggregated, and Decomposable Model for Networks of Time Series · WWW 2021
Visualization and visual analytics
time series visualization
0.512021
AttentionFlow: Visualising Influence in Networks of Time Series · WSDM 2021
Computer vision › Vision and language › image captioning › fine-grained image captioning
entity-aware captioning
0.412020
Transform and Tell: Entity-Aware News Image Captioning · CVPR 2020
Computer vision › Vision and language › image captioning
news image captioning
0.412020
Transform and Tell: Entity-Aware News Image Captioning · CVPR 2020
Information retrieval › text summarization
comparative summarization
0.412019
Comparative Document Summarisation via Classification · AAAI 2019
Information retrieval › text summarization
extractive summarization
0.412019
Comparative Document Summarisation via Classification · AAAI 2019
Information retrieval
text summarization
0.412019
Comparative Document Summarisation via Classification · AAAI 2019
Machine learning › Generative modeling
style transfer
0.312018
SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018
Natural language and speech › Language models and text generation › controllable text generation
text style transfer
0.312018
SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018
Robotics › Motion planning and robot control › robot control › contact control › contact task control › robot force control
adaptive impedance control
0.212016
SentiCap: Generating Image Descriptions with Sentiments · AAAI 2016
Natural language and speech › Language models and text generation › text generation › neural text generation
recurrent neural network text generation
0.212016
SentiCap: Generating Image Descriptions with Sentiments · AAAI 2016
Computer vision › Vision and language
cross-modal attention
0.112020
Transform and Tell: Entity-Aware News Image Captioning · CVPR 2020
Machine learning and data management
data selection
0.112019
Comparative Document Summarisation via Classification · AAAI 2019
Computer vision › Vision and language
multimodal representation
0.112018
SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic representation
0.112018
SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text · CVPR 2018
Computer vision › Vision and language › image captioning
diverse image captioning
0.112015
Captioning Images Using Different Styles · ACM Multimedia 2015

Methods — techniques the papers use, named apart from their topics

tree ring encoding · 1.5recurrent neural network · 1.5line chart · 1.5ego network visualization · 1.5multi-head attention · 1.4tensor factorization · 1.3fourier transform · 1.3multi-layer decomposition · 1.0stone-weierstrass theorem · 0.5basis functions · 0.5transformer language model · 0.4image and face embeddings · 0.4byte-pair encoding · 0.4maximum mean discrepancy · 0.4gradient-based optimisation · 0.4binary classification · 0.4
YearPublicationVenuePosition
2023 Factorized Fourier Neural Operators
Alasdair Tran, Alexander Patrick Mathews, Lexing Xie, Cheng Soon Ong
ICLR2
2021 UNIPoint: Universally Approximating Point Processes Intensities
abstract
Point processes are a useful mathematical tool for describing events over time, and so there are many recent approaches for representing and learning them. One notable open question is how to precisely describe the flexibility of point process models and whether there exists a general model that can represent all point processes. Our work bridges this gap. Focusing on the widely used event intensity function representation of point processes, we provide a proof that a class of learnable functions can universally approximate any valid intensity function. The proof connects the well known Stone-Weierstrass Theorem for function approximation, the uniform density of non-negative continuous functions using a transfer functions, the formulation of the parameters of a piece-wise continuous functions as a dynamic system, and a recurrent neural network implementation for capturing the dynamics. Using these insights, we design and implement UNIPoint, a novel neural point process model, using recurrent neural networks to parameterise sums of basis function upon each event. Evaluations on synthetic and real world datasets show that this simpler representation performs better than Hawkes process variants and more complex neural network-based approaches. We expect this result will provide a practical basis for selecting and tuning models, as well as furthering theoretical work on representational complexity and learnability.
Alexander Soen, Alexander Patrick Mathews, Daniel Grixti-Cheng, Lexing Xie
AAAI2
2021 AttentionFlow: Visualising Influence in Networks of Time Series
abstract
The collective attention on online items such as web pages, search terms, and videos reflects trends that are of social, cultural, and economic interest. Moreover, attention trends of different items exhibit mutual influence via mechanisms such as hyperlinks or recommendations. Many visualisation tools exist for time series, network evolution, or network influence; however, few systems connect all three. In this work, we present AttentionFlow, a new system to visualise networks of time series and the dynamic influence they have on one another. Centred around an ego node, our system simultaneously presents the time series on each node using two visual encodings: a tree ring for an overview and a line chart for details. AttentionFlow supports interactions such as overlaying time series of influence, and filtering neighbours by time or flux. We demonstrate AttentionFlow using two real-world datasets, VevoMusic and WikiTraffic. We show that attention spikes in songs can be explained by external events such as major awards, or changes in the network such as the release of a new song. Separate case studies also demonstrate how an artist's influence changes over their career, and that correlated Wikipedia traffic is driven by cultural interests. More broadly, AttentionFlow can be generalised to visualise networks of time series on physical infrastructures such as road networks, or natural phenomena such as weather and geological measurements.
Minjeong Shin, Alasdair Tran, Alexander Patrick Mathews, Georgiana Lyall, Lexing Xie
WSDM4
2021 Radflow: A Recurrent, Aggregated, and Decomposable Model for Networks of Time Series
abstract
We propose a new model for networks of time series that influence each other. Graph structures among time series are found in diverse domains, such as web traffic influenced by hyperlinks, product sales influenced by recommendation, or urban transport volume influenced by road networks and weather. There has been recent progress in graph modeling and in time series forecasting, respectively, but an expressive and scalable approach for a network of series does not yet exist. We introduce Radflow, a novel model that embodies three key ideas: a recurrent neural network to obtain node embeddings that depend on time, the aggregation of the flow of influence from neighboring nodes with multi-head attention, and the multi-layer decomposition of time series. Radflow naturally takes into account dynamic networks where nodes and edges change over time, and it can be used for prediction and data imputation tasks. On real-world datasets ranging from a few hundred to a few hundred thousand nodes, we observe that Radflow variants are the best performing model across a wide range of settings. The recurrent component in Radflow also outperforms N-BEATS, the state-of-the-art time series model. We show that Radflow can learn different trends and seasonal patterns, that it is robust to missing nodes and edges, and that correlated temporal patterns among network neighbors reflect influence strength. We curate WikiTraffic, the largest dynamic network of time series with 366K nodes and 22M time-dependent links spanning five years. This dataset provides an open benchmark for developing models in this area, with applications that include optimizing resources for the web. More broadly, Radflow has the potential to improve forecasts in correlated time series networks such as the stock market, and impute missing measurements in geographically dispersed networks of natural phenomena.
Alasdair Tran, Alexander Patrick Mathews, Cheng Soon Ong, Lexing Xie
WWW2
2020 Transform and Tell: Entity-Aware News Image Captioning
abstract
We propose an end-to-end model which generates captions for images embedded in news articles. News images present two key challenges: they rely on real-world knowledge, especially about named entities; and they typically have linguistically rich captions that include uncommon words. We address the first challenge by associating words in the caption with faces and objects in the image, via a multi-modal, multi-head attention mechanism. We tackle the second challenge with a state-of-the-art transformer language model that uses byte-pair-encoding to generate captions as a sequence of word parts. On the GoodNews dataset, our model outperforms the previous state of the art by a factor of four in CIDEr score (13 to 54). This performance gain comes from a unique combination of language models, word representation, image embeddings, face embeddings, object embeddings, and improvements in neural network design. We also introduce the NYTimes800k dataset which is 70% larger than GoodNews, has higher article quality, and includes the locations of images within articles as an additional contextual cue.
Alasdair Tran, Alexander Patrick Mathews, Lexing Xie
CVPR2
2019 Comparative Document Summarisation via Classification
abstract
Thispaperconsidersextractivesummarisationinacomparative setting: given two or more document groups (e.g., separated by publication time), the goal is to select a small number of documents that are representative of each group, and also maximally distinguishable from other groups. We formulate a set of new objective functions for this problem that connect recent literature on document summarisation, interpretable machine learning, and data subset selection. In particular, by casting the problem as a binary classification amongst different groups, we derive objectives based on the notion of maximum mean discrepancy, as well as a simple yet effective gradient-based optimisation strategy. Our new formulation allows scalable evaluations of comparative summarisation as a classification task, both automatically and via crowd-sourcing. To this end, we evaluate comparative summarisation methods on a newly curated collection of controversial news topics over 13months.Weobserve thatgradient-based optimisationoutperforms discrete and baseline approaches in 15 out of 24 different automatic evaluation settings. In crowd-sourced evaluations, summaries from gradient optimisation elicit 7% more accurate classification from human workers than discrete optimisation. Our result contrasts with recent literature on submodular data subset selection that favours discrete optimisation. We posit that our formulation of comparative summarisation will prove useful in a diverse range of use cases such as comparing content sources, authors, related topics, or distinct view points.
Umanga Bista, Alexander Patrick Mathews, Minjeong Shin, Aditya Krishna Menon, Lexing Xie
AAAI2
2018 SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text
abstract
Linguistic style is an essential part of written communication, with the power to affect both clarity and attractiveness. With recent advances in vision and language, we can start to tackle the problem of generating image captions that are both visually grounded and appropriately styled. Existing approaches either require styled training captions aligned to images or generate captions with low relevance. We develop a model that learns to generate visually relevant styled captions from a large corpus of styled text without aligned images. The core idea of this model, called SemStyle, is to separate semantics and style. One key component is a novel and concise semantic term representation generated using natural language processing techniques and frame semantics. In addition, we develop a unified language model that decodes sentences with diverse word choices and syntax for different styles. Evaluations, both automatic and manual, show captions from SemStyle preserve image semantics, are descriptive, and are style shifted. More broadly, this work provides possibilities to learn richer image descriptions from the plethora of linguistic data available on the web.
Alexander Patrick Mathews, Lexing Xie, Xuming He 0001
CVPR1
2016 SentiCap: Generating Image Descriptions with Sentiments
abstract
The recent progress on image recognition and language modeling is making automatic description of image content a reality. However, stylized, non-factual aspects of the written description are missing from the current systems. One such style is descriptions with emotions, which is commonplace in everyday communication, and influences decision-making and interpersonal relationships. We design a system to describe an image with emotions, and present a model that automatically generates captions with positive or negative sentiments. We propose a novel switching recurrent neural network with word-level regularization, which is able to produce emotional image captions using only 2000+ training sentences containing sentiments. We evaluate the captions with different automatic and crowd-sourcing metrics. Our model compares favourably in common quality metrics for image captioning. In 84.6% of cases the generated positive captions were judged as being at least as descriptive as the factual captions. Of these positive captions 88% were confirmed by the crowd-sourced workers as having the appropriate sentiment.
Alexander Patrick Mathews, Lexing Xie, Xuming He 0001
AAAI1
2015 Captioning Images Using Different Styles
abstract
I develop techniques that can be used to incorporate stylistic objectives into existing image captioning systems. Style is generally a very tricky concept to define, thus I concentrate on two specific components of style. First I develop a technique for predicting how people will name visual objects. I demonstrate that this technique could be used to generate captions with human like naming conventions. Full details are available in a recent publication. Second I outline a system for generating sentences which express a strong positive or negative sentiment. Finally I present two possible future directions which are aimed at modelling style more generally. These are learning to imitate an individuals captioning style and generating a diverse set of captions for a single image.
Alexander Patrick Mathews
ACM Multimedia1
2015 Choosing Basic-Level Concept Names Using Visual and Language Context
abstract
We study basic-level categories for describing visual concepts, and empirically observe context-dependant basic level names across thousands of concepts. We propose methods for predicting basic-level names using a series of classification and ranking tasks, producing the first large scale catalogue of basic-level names for hundreds of thousands of images depicting thousands of visual concepts. We also demonstrate the usefulness of our method with a picture-to-word task, showing strong improvement over recent work by Ordonez et al, by modeling of both visual and language context. Our study suggests that a model for naming visual concepts is an important part of any automatic image/video captioning and visual story-telling system.
Alexander Patrick Mathews, Lexing Xie, Xuming He 0001
WACV1