Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Alexander Long

dblp:156/9630 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0003-4300-9454ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 59% Optimization for machine learning · 13% Reinforcement learning · 10%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 100%
Theoretical computer science
1 paper
Information theory · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
2.632025
Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training · NeurIPS 2025
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism · NeurIPS 2025
Nesterov Method for Asynchronous Pipeline Parallel Optimization · ICML 2025
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
1.722025
Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training · NeurIPS 2025
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism · NeurIPS 2025
Machine learning › Optimization for machine learning › gradient-based optimization
accelerated gradient methods
0.912025
Nesterov Method for Asynchronous Pipeline Parallel Optimization · ICML 2025
Machine learning › Efficient and distributed learning › model compression › neural network compression
activation compression
0.912025
Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training · NeurIPS 2025
Machine learning › Efficient and distributed learning › distributed training
asynchronous optimization
0.912025
Nesterov Method for Asynchronous Pipeline Parallel Optimization · ICML 2025
Natural language and speech › Language models and text generation › language modeling
long-context language modeling
0.912025
Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training · NeurIPS 2025
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.912025
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism · NeurIPS 2025
Machine learning › Optimization for machine learning › gradient-based optimization › accelerated gradient methods
nesterov acceleration
0.912025
Nesterov Method for Asynchronous Pipeline Parallel Optimization · ICML 2025
Machine learning › Efficient and distributed learning › distributed training › model parallelism
pipeline parallelism
0.912025
Nesterov Method for Asynchronous Pipeline Parallel Optimization · ICML 2025
Geometric modeling and processing
implicit neural representation
0.812024
A sampling theory perspective on activations for implicit neural representations · ICML 2024
Information theory › signal processing
sampling theory
0.812024
A sampling theory perspective on activations for implicit neural representations · ICML 2024
Computer vision › Image recognition and object detection
image classification
0.612022
Retrieval Augmented Classification for Long-Tail Visual Recognition · CVPR 2022
Machine learning › Learning paradigms
long-tailed recognition
0.612022
Retrieval Augmented Classification for Long-Tail Visual Recognition · CVPR 2022
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.612022
Fast and Data Efficient Reinforcement Learning from Pixels via Non-parametric Value Approximation · AAAI 2022
Machine learning › Reinforcement learning
value-based reinforcement learning
0.612022
Fast and Data Efficient Reinforcement Learning from Pixels via Non-parametric Value Approximation · AAAI 2022
Information retrieval › retrieval augmentation
retrieval-augmented classification
0.612022
Retrieval Augmented Classification for Long-Tail Visual Recognition · CVPR 2022
Natural language and speech › Language models and text generation
large language model training
0.312025
Nesterov Method for Asynchronous Pipeline Parallel Optimization · ICML 2025
Machine learning › Reinforcement learning › deep reinforcement learning › visual reinforcement learning
pixel-based reinforcement learning
0.212022
Fast and Data Efficient Reinforcement Learning from Pixels via Non-parametric Value Approximation · AAAI 2022
Machine learning › Transfer learning and domain adaptation › pre-training and adaptation
pretrained model reuse
0.212022
Retrieval Augmented Classification for Long-Tail Visual Recognition · CVPR 2022

Methods — techniques the papers use, named apart from their topics

sampling theory · 1.5dynamical systems analysis · 1.5external memory · 1.1transformer network · 0.9nesterov accelerated gradient · 0.9mixtures of subspaces · 0.9low-rank reparameterization · 0.9low-dimensional subspace compression · 0.9convergence analysis · 0.9retrieval module · 0.6proximity graph lookup · 0.6non-parametric memory · 0.6non-parametric approximation · 0.6monte carlo methods · 0.6
YearPublicationVenuePosition
2025 Nesterov Method for Asynchronous Pipeline Parallel Optimization
abstract
Pipeline Parallelism (PP) enables large neural network training on small, interconnected devices by splitting the model into multiple stages. To maximize pipeline utilization, asynchronous optimization is appealing as it offers 100% pipeline utilization by construction. However, it is inherently challenging as the weights and gradients are no longer synchronized, leading to *stale (or delayed) gradients*. To alleviate this, we introduce a variant of Nesterov Accelerated Gradient (NAG) for asynchronous optimization in PP. Specifically, we modify the look-ahead step in NAG to effectively address the staleness in gradients. We theoretically prove that our approach converges at a sublinear rate in the presence of fixed delay in gradients. Our experiments on large-scale language modelling tasks using decoder-only architectures with up to **1B parameters**, demonstrate that our approach significantly outperforms existing asynchronous methods, even surpassing the synchronous baseline.
Thalaiyasingam Ajanthan, Sameera Ramasinghe, Gil Avraham, Alexander Long
ICML5
2025 Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization
abstract
We consider a decentralized setup in which the participants collaboratively train and serve a large neural network, and where each participant only processes a subset of the model. In this setup, we explore the possibility of unmaterializable weights, where a full weight set is never available to any one participant. We introduce Unextractable Protocol Models (UPMs): a training and inference framework that leverages the sharded model setup to ensure model shards (i.e.,, subsets) held by participants are incompatible at different time steps. UPMs periodically inject time-varying, random, invertible transforms at participant boundaries; preserving the overall network function yet rendering cross-time assemblies incoherent. On Qwen-2.5-0.5B and Llama-3.2-1B, 10 000 transforms leave FP32 perplexity unchanged ($\Delta$PPL$< 0.01$; Jensen–Shannon drift $<4 \times 10^{-5}$), and we show how to control growth for lower precision datatypes. Applying a transform every 30s adds 3% latency, 0.1% bandwidth, and 10% GPU-memory overhead at inference, while training overhead falls to 1.6% time and < 1% memory. We consider several attacks, showing that the requirements of direct attacks are impractical and easy to defend against, and that gradient-based fine-tuning of stitched partitions consumes $\geq 60\%$ of the tokens required to train from scratch. By enabling models to be collaboratively trained yet not extracted, UPMs make it practical to embed programmatic incentive mechanisms in community-driven decentralized training.
Alexander Long, Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Gil Avraham, Violetta Shevchenko, Hadi M. Dolatabadi, Sameera Ramasinghe
NeurIPS1
2025 Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism
abstract
Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to communication bottlenecks. While existing compression techniques are effective in data-parallel, they do not extend to model parallelism. Unlike data-parallel training, where weight gradients are exchanged, model-parallel requires compressing activations and activation gradients as they propagate through layers, accumulating compression errors. We propose a novel compression algorithm that compresses both forward and backward passes, enabling up to 99% compression with no convergence degradation with negligible memory/compute overhead. By leveraging a recursive structure in transformer networks, we predefine a low-dimensional subspace to confine the activations and gradients, allowing full reconstruction in subsequent layers. Our method achieves up to 100x improvement in communication efficiency and enables training billion-parameter-scale models over low-end GPUs connected via consumer-grade internet speeds as low as 80Mbps, matching the convergence of centralized datacenter systems with 100Gbps connections with model parallel.
Sameera Ramasinghe, Thalaiyasingam Ajanthan, Gil Avraham, Alexander Long
NeurIPS5
2025 Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training
abstract
Pretraining language models with extended context windows enhances their ability to leverage rich information during generation. Existing methods split input sequences into chunks, broadcast them across multiple devices, and compute attention block by block which incurs significant communication overhead. While feasible in high-speed clusters, these methods are impractical for decentralized training over low-bandwidth connections. We propose a compression method for communication-efficient context parallelism in decentralized settings, achieving a remarkable compression rate of over 95% with negligible overhead and no loss in convergence. Our key insight is to exploit the intrinsic low-rank structure of activation outputs by dynamically constraining them to learned mixtures of subspaces via efficient reparameterizations. We demonstrate scaling billion-parameter decentralized models to context lengths exceeding 100K tokens on networks as slow as 300Mbps, matching the wall-clock convergence speed of centralized models on 100Gbps interconnects.
Sameera Ramasinghe, Thalaiyasingam Ajanthan, Hadi M. Dolatabadi, Gil Avraham, Violetta Shevchenko, Chamin Hewa Koneputugodage, Alexander Long
NeurIPS8
2024 A sampling theory perspective on activations for implicit neural representations
abstract
Implicit Neural Representations (INRs) have gained popularity for encoding signals as compact, differentiable entities. While commonly using techniques like Fourier positional encodings or non-traditional activation functions (e.g., Gaussian, sinusoid, or wavelets) to capture high-frequency content, their properties lack exploration within a unified theoretical framework. Addressing this gap, we conduct a comprehensive analysis of these activations from a sampling theory perspective. Our investigation reveals that, especially in shallow INRs, $\mathrm{sinc}$ activations—previously unused in conjunction with INRs—are theoretically optimal for signal encoding. Additionally, we establish a connection between dynamical systems and INRs, leveraging sampling theory to bridge these two paradigms.
Hemanth Saratchandran, Sameera Ramasinghe, Violetta Shevchenko, Alexander Long, Simon Lucey
ICML4
2024 LipAT: Beyond Style Transfer for Controllable Neural Simulation of Lipstick using Cosmetic Attributes
abstract
Lipstick virtual try-on (VTO) experiences have become widespread across the e-commerce sector and assist users in eliminating the guesswork of shopping online. However, such experiences still lack in both realism and accuracy. In this work, we propose LipAT, a neural framework that blends the strengths of Physics-Based Rendering (PBR) and Neural Style Transfer (NST) approaches to directly apply lipstick onto face images given lipstick attributes (e.g., colour, finish type). LipAT consists of a physics aware neural lipstick application module (LAM) to apply lipstick on face images given its attributes and Lipstick Refiner Module (LRM) to improve the realism by refining the imperfections. Unlike the NST approaches, LipAT allows precise and controllable lipstick attribute preservation, without requiring crude approximations and inference of various intertwined environment factors (e.g., scene lighting, face structure etc) involved in image generation that is required for accurate PBR. We propose an experimental framework with quantitative metrics to evaluate different desirable aspects of the lipstick attribute driven try-on alongside user studies to further validate our findings. Our results show that LipAT considerably outperforms fully-automated PBR approaches in preserving realism and the NST approaches in preserving various lipstick attributes such as finish types.
Amila Silva, Olga Moskvyak, Alexander Long, Ravi Garg, Stephen Gould, Gil Avraham, Anton van den Hengel
WACV3
2022 Fast and Data Efficient Reinforcement Learning from Pixels via Non-parametric Value Approximation
abstract
We present Nonparametric Approximation of Inter-Trace returns (NAIT), a Reinforcement Learning algorithm for discrete action, pixel-based environments that is both highly sample and computation efficient. NAIT is a lazy-learning approach with an update that is equivalent to episodic Monte-Carlo on episode completion, but that allows the stable incorporation of rewards while an episode is ongoing. We make use of a fixed domain-agnostic representation, simple distance based exploration and a proximity graph-based lookup to facilitate extremely fast execution. We empirically evaluate NAIT on both the 26 and 57 game variants of ATARI100k where, despite its simplicity, it achieves competitive performance in the online setting with greater than 100x speedup in wall-time.
Alexander Long, Alan Blair 0001, Herke van Hoof
AAAI1
2022 Retrieval Augmented Classification for Long-Tail Visual Recognition
abstract
We introduce Retrieval Augmented Classification (RAC), a generic approach to augmenting standard image classification pipelines with an explicit retrieval module. RAC consists of a standard base image encoder fused with a parallel retrieval branch that queries a non-parametric external memory of pre-encoded images and associated text snippets. We apply RAC to the problem of long-tail classification and demonstrate a significant improvement over previous state-of-the-art on Places365-LT and iNaturalist-2018 (14.5% and 6.7% respectively), despite using only the training datasets themselves as the external information source. We demonstrate that RAC's retrieval module, without prompting, learns a high level of accuracy on tail classes. This, in turn, frees the base encoder to focus on common classes, and improve its performance thereon. RAC represents an alternative approach to utilizing large, pretrained models without requiring fine-tuning, as well as a first step towards more effectively making use of external memory within common computer vision architectures.
Alexander Long, Wei Yin 0006, Thalaiyasingam Ajanthan, Pulak Purkait, Ravi Garg, Alan Blair 0001, Chunhua Shen, Anton van den Hengel
CVPR1
2021 Spatial Dimensions of Algorithmic Transparency: A Summary
abstract
Spatial data brings an important dimension to AI’s quest for algorithmic transparency. For example, data driven computer-aided policy-decisions use measures of segregation (e.g., dissimilarity index) or income-inequality (e.g., Gini index), and these measures are affected by space partitioning choice. This may lead policymakers to underestimate the level of inequality or segregation within a region. The problem stems from the fact that many segregation based analyses use aggregated census data but do not report result sensitivity to choice of spatial partitioning (e.g., census block, tract). Beyond the well-known Modifiable Areal Unit Problem, this paper shows (via mathematical proofs as well as case studies with census data and census based synthetic micro-population data) that values of many measures (e.g., Gini index, dissimilarity index) diminish monotonically with increasing spatial-unit size in a hierarchical space partitioning (e.g., block, block-group, tract), however the ranking based on spatially aggregated measures remain sensitive to the scale of spatial partitions (e.g., block, block group). This paper highlights the need for social scientists to report how rankings of inequality are affected by the choice of spatial partitions.
Jayant Gupta, Alexander Long, Corey Kewei Xu, Shashi Shekhar 0001
SSTD2
2014 SEEM: a scalable visualization for comparing multiple large sets of attributes for malware analysis
abstract
Recently, the number of observed malware samples has rapidly increased, expanding the workload for malware analysts. Most of these samples are not truly unique, but are related through shared attributes. Identifying these attributes can enable analysts to reuse analysis and reduce their workload. Visualizing malware attributes as sets could enable analysts to better understand the similarities and differences between malware. However, existing set visualizations have difficulty displaying hundreds of sets with thousands of elements, and are not designed to compare different types of elements between sets, such as the imported DLLs and callback domains across malware samples. Such analysis might help analysts, for example, to understand if a group of malware samples are behaviorally different or merely changing where they send data.
Robert Gove, Joshua Saxe, Sigfried Gold, Alexander Long, Giacomo Bergamo
VizSEC4
2014 Detecting malware samples with similar image sets
abstract
This paper proposes a method for identifying and visualizing similarity relationships between malware samples based on their embedded graphical assets (such as desktop icons and button skins). We argue that analyzing such relationships has practical merit for a number of reasons. For example, we find that malware desktop icons are often used to trick users into running malware programs, so identifying groups of related malware samples based on these visual features can highlight themes in the social engineering tactics of today's malware authors. Also, when malware samples share rare images, these image sharing relationships may indicate that the samples were generated or deployed by the same adversaries.
Alexander Long, Joshua Saxe, Robert Gove
VizSEC1