Angela W. Li

dblp:371/4251 · also Wu Angela Li · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-4523-3401ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Static Analysis for Efficient Streaming Tokenization
abstract
Tokenization, also referred to as lexing or scanning, is the computational task of partitioning an input text into a sequence of substrings called tokens. Tokenization is one of the first stages of program compilation, it is used in natural language processing, and it is also useful for processing unstructured text or semi-structured data such as JSON, CSV, and XML. A tokenizer is typically specified as a list of regular expressions, which is called a tokenization grammar. Each regular expression describes a class of tokens (e.g., integer, floating-point number, variable identifier, string literal). The semantics of tokenization employs the longest match policy to disambiguate among the possible choices. This policy says that we should prefer a longer token over a shorter one. It is also known as the maximal munch policy.
Angela W. Li, Yudi Yang, Konstantinos Mamouras
ASPLOS (2)1
2026 An Efficient Algorithm for Streaming BPE Tokenization
abstract
Tokenization is an essential text preprocessing step in almost all large language models (LLMs), and byte-pair encoding (BPE) is a popular tokenization method used by models such as GPT, GPT-2, and RoBERTa. Since LLMs have many applications that require the fast processing of large amounts of data (e.g., real-time document summarization and analysis), offline tokenization algorithms may have high latency or prohibitive memory requirements. In this paper, we study BPE tokenization with a focus on providing a streaming implementation. A BPE tokenizer is specified with an ordered list of token merge rules, where each rule describes the merging of two adjacent tokens. We introduce the concept of delay for a list of BPE merge rules, which corresponds to the amount of lookahead needed before tokens can be finalized. We view BPE tokenization as a sequence-to-sequence transduction and show how to obtain a bound on delay. This bound enables streaming tokenization with a low memory footprint. We propose a novel streaming algorithm that uses a small amount of memory (independent of the input text) and has linear time complexity in the length of the input text. Our experimental evaluation shows that our algorithm performs well in comparison to existing BPE tokenizers.
Konstantinos Mamouras, Angela W. Li, Yudi Yang
Proc. ACM Program. Lang.2
2026 Revisiting sparse error correction: Model analysis and new algorithms
Angela W. Li
Signal Process.3
2025 Verified and Efficient Matching of Regular Expressions with Lookaround
abstract
Regular expressions can be extended with lookarounds for contextual matching. This paper discusses a Coq formalization of the theory of regular expressions with lookarounds. We provide an efficient and purely functional algorithm for matching expressions with lookarounds and verify its correctness. The algorithm runs in time linear in both the size of the regular expression as well as the input string. Our experimental results provide empirical support to our complexity analysis. To the best of our knowledge, this is the first formalization of a linear-time matching algorithm for regular expressions with lookarounds.
Agnishom Chattopadhyay, Angela W. Li, Konstantinos Mamouras
CPP2
2025 Efficient Algorithms for the Uniform Tokenization Problem
abstract
Tokenization (also known as scanning or lexing) is a computational task that has applications in the lexical analysis of programs during compilation and in data extraction and analysis for unstructured or semistructured data (e.g., data represented using the JSON and CSV data formats). We propose two algorithms for the tokenization problem that have linear time complexity (in the length of the input text) without using large amounts of memory. We also show that an optimized version of one of these algorithms performs well compared to prior approaches on practical tokenization workloads.
Angela W. Li, Konstantinos Mamouras
Proc. ACM Program. Lang.1
2025 Streaming Validation of JSON Documents Against Schemas
Alexis Le Glaunec, Angela W. Li, Konstantinos Mamouras
Proc. VLDB Endow.2
2024 A Computational Study on Sentence-based Next Speaker Prediction in Multiparty Conversations
abstract
In this paper we present a computational study to quantitatively examine the task of predicting the next speaker in multi-party conversations using machine learning models. To accomplish this, we create features that accurately represent information relevant to speaker changes in such conversations. We utilize sentence-based models, rather than the widely-used InterPausal Unit (IPU)-based models, and extend the definition of verbal backchanneling to include additional reactions that signify listeners’ attention or interest. Through extensive experiments with various machine learning models and inputs, we show that our sentence-based models outperform existing IPU-based models, with the best model achieving 61.39% accuracy. Our study provides design implications and recommendations for the development of virtual agents or humanoid robots with interactive social interaction capabilities.
Meng-Chen Lee, Angela W. Li, Zhigang Deng 0001
IVA2
2024 Static Analysis for Checking the Disambiguation Robustness of Regular Expressions
abstract
Regular expressions are commonly used for finding and extracting matches from sequence data. Due to the inherent ambiguity of regular expressions, a disambiguation policy must be considered for the match extraction problem, in order to uniquely determine the desired match out of the possibly many matches. The most common disambiguation policies are the POSIX policy and the greedy (PCRE) policy. The POSIX policy chooses the longest match out of the leftmost ones. The greedy policy chooses a leftmost match and further disambiguates using a greedy interpretation of Kleene iteration to match as many times as possible. The choice of disambiguation policy can affect the output of match extraction, which can be an issue for reusing regular expressions across regex engines. In this paper, we introduce and study the notion of disambiguation robustness for regular expressions. A regular expression is robust if its extraction semantics is indifferent to whether the POSIX or greedy disambiguation policy is chosen. This gives rise to a decision problem for regular expressions, which we prove to be PSPACE-complete. We propose a static analysis algorithm for checking the (non-)robustness of regular expressions and two performance optimizations. We have implemented the proposed algorithms and we have shown experimentally that they are practical for analyzing large datasets of regular expressions derived from various application domains.
Konstantinos Mamouras, Alexis Le Glaunec, Angela W. Li, Agnishom Chattopadhyay
Proc. ACM Program. Lang.3