Kareem Ahmed

dblp:188/6144 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
13since 2021 · last 2025
0000-0002-1252-8625ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Controllable Generation via Locally Constrained Resampling
abstract
Autoregressive models have demonstrated an unprecedented ability at modeling the intricacies of natural language. However, they continue to struggle with generating complex outputs that adhere to logical constraints. Sampling from a fully-independent distribution subject to a constraint is hard. Sampling from an autoregressive distribution subject to a constraint is doubly hard: We have to contend not only with the hardness of the constraint but also the distribution's lack of structure. We propose a tractable probabilistic approach that performs Bayesian conditioning to draw samples subject to a constraint. By factoring in information about the entire sequence, our approach offers better contextual awareness during constrained generation compared to current greedy approaches. Starting from a model sample, we induce a local, factorized distribution which we can tractably condition on the constraint. To generate samples that satisfy the constraint, we sample from the conditional distribution, correct for biases in the sample weights, and resample. The resulting samples closely approximate the target distribution and are guaranteed to satisfy the constraints. We evaluate our approach on several tasks, including LLM detoxification and solving Sudoku puzzles. We show that by disallowing a list of toxic expressions our approach is able to steer the model's outputs away from toxic generations, outperforming similar approaches to detoxification. We also show that our approach achieves a perfect accuracy on Sudoku, compared to less than $50\%$ for GPT4-o and Gemini 1.5.
Kareem Ahmed, Kai-Wei Chang 0001, Guy Van den Broeck
ICLR1
2024 Where is the signal in tokenization space?
abstract
Large Language Models (LLMs) are typically shipped with tokenizers that deterministically encode text into so-called canonical token sequences, to which the LLMs assign probability values.One common assumption is that the probability of a piece of text is the probability of its canonical token sequence.However, the tokenization of a string is not unique: e.g., the Llama2 tokenizer encodes Tokens as [Tok,ens], but [Tok,en,s] also represents the same text.In this paper, we study noncanonical tokenizations.We prove that, given a string, it is computationally hard to find the most likely tokenization for an autoregressive LLM, as well as to compute the marginal probability over all possible tokenizations.We then show how the marginal is, in most cases, indistinguishable from the canonical probability.Surprisingly, we then empirically demonstrate the existence of a significant amount of signal hidden within tokenization space.Notably, by simply aggregating the probabilities of noncanonical tokenizations, we achieve improvements across a range of LLM evaluation benchmarks for a variety of architectures, including transformers and state space models.
Renato Lui Geh, Honghua Zhang, Kareem Ahmed, Benjie Wang 0001, Guy Van den Broeck
EMNLP3
2024 Probabilistically Rewired Message-Passing Neural Networks
abstract
Message-passing graph neural networks (MPNNs) emerged as powerful tools for processing graph-structured input. However, they operate on a fixed input graph structure, ignoring potential noise and missing information. Furthermore, their local aggregation mechanism can lead to problems such as over-squashing and limited expressive power in capturing relevant graph structures. Existing solutions to these challenges have primarily relied on heuristic methods, often disregarding the underlying data distribution. Hence, devising principled approaches for learning to infer graph structures relevant to the given prediction task remains an open challenge. In this work, leveraging recent progress in exact and differentiable k-subset sampling, we devise probabilistically rewired MPNNs (PR-MPNNs), which learn to add relevant edges while omitting less beneficial ones. For the first time, our theoretical analysis explores how PR-MPNNs enhance expressive power, and we identify precise conditions under which they outperform purely randomized approaches. Empirically, we demonstrate that our approach effectively mitigates issues like over-squashing and under-reaching. In addition, on established real-world datasets, our method exhibits competitive or superior predictive performance compared to traditional MPNN models and recent graph transformer architectures.
Chendi Qian, Andrei Manolache, Kareem Ahmed, Zhe Zeng 0001, Guy Van den Broeck, Mathias Niepert, Christopher Morris 0001
ICLR3
2024 Scaling Tractable Probabilistic Circuits: A Systems Perspective
abstract
Probabilistic Circuits (PCs) are a general framework for tractable deep generative models, which support exact and efficient probabilistic inference on their learned distributions. Recent modeling and training advancements have enabled their application to complex real-world tasks. However, the time and memory inefficiency of existing PC implementations hinders further scaling up. This paper proposes PyJuice, a general GPU implementation design for PCs that improves prior art in several regards. Specifically, PyJuice is 1-2 orders of magnitude faster than existing systems (including very recent ones) at training large-scale PCs. Moreover, PyJuice consumes 2-5x less GPU memory, which enables us to train larger models. At the core of our system is a compilation process that converts a PC into a compact representation amenable to efficient block-based parallelization, which significantly reduces IO and makes it possible to leverage Tensor Cores available in modern GPUs. Empirically, PyJuice can be used to improve state-of-the-art PCs trained on image (e.g., ImageNet32) and language (e.g., WikiText, CommonGen) datasets. We further establish a new set of baselines on natural image and language datasets by benchmarking existing PC structures but with much larger sizes and more training epochs, with the hope of incentivizing future research. Code is available at https://github.com/Tractables/pyjuice.
Anji Liu, Kareem Ahmed, Guy Van den Broeck
ICML2
2024 Snake species classification using deep learning techniques
abstract
Abstract Incorrect snake identification from the observable visual traits is a major reason of death resulting from snake bites. The classification of snake species has a significant role in determining the appropriate treatment without any delay, the delay may cause dangerous complications or lead to the death of the victim. The difficulty of classifying snakes by human lies in the variations of snake pattern based on geographic variation and age, the intraclass variance is high for specific classes and the interclass variance is low among others, and there may be two remarkably similar types in shape, with one being toxic and the other not. The limitation of the experts’ number in the herpetology and their geographical distribution leads us to the importance of using deep learning in the snake species classification. A model to classify snake species accately is proposed in this study. It is divided into two main processes, detecting the salient object by applying Salient Object Detection (SOD) model based on VGG16 architecture is the first process, the presence of snakes in places with a complex background led to the necessity of separating the salient object, then the classification model is applied with use of image augmentations parameters which improved the results. Four CNN models were used in the classification process including VGG16, ResNet50, MobileNetV2, and DenseNet121. Different experiments on 5,10,16,20, 22, and 45 number of classes and different models were conducted, and the model achieved unprecedented results. The results indicated that the VGG16, DenseNet121, and MobileNetV2 have achieved superior results in the same order from highest to lowest accuracy. The best accuracy is achieved using VGG16 architecture with accuracy 97.09% when using 45 number of classes.
Kareem Ahmed, Mai A. Gad, Amal Elsayed Aboutabl
Multim. Tools Appl.1
2023 Semantic Strengthening of Neuro-Symbolic Learning
abstract
Numerous neuro-symbolic approaches have recently been proposed typically with the goal of adding symbolic knowledge to the output layer of a neural network. Ideally, such losses maximize the probability that the neural network’s predictions satisfy the underlying domain. Unfortunately, this type of probabilistic inference is often computationally infeasible. Neuro-symbolic approaches therefore commonly resort to fuzzy approximations of this probabilistic objective, sacrificing sound probabilistic semantics, or to sampling which is very seldom feasible. We approach the problem by first assuming the constraint decomposes conditioned on the features learned by the network. We iteratively strengthen our approximation, restoring the dependence between the constraints most responsible for degrading the quality of the approximation. This corresponds to computing the mutual information between pairs of constraints conditioned on the network’s learned features, and may be construed as a measure of how well aligned the gradients of two distributions are. We show how to compute this efficiently for tractable circuits. We test our approach on three tasks: predicting a minimum-cost path in Warcraft, predicting a minimum-cost perfect matching, and solving Sudoku puzzles, observing that it improves upon the baselines while sidestepping intractability.
Kareem Ahmed, Kai-Wei Chang 0001, Guy Van den Broeck
AISTATS1
2023 SIMPLE: A Gradient Estimator for k-Subset Sampling
Kareem Ahmed, Zhe Zeng 0001, Mathias Niepert, Guy Van den Broeck
ICLR1
2023 A Pseudo-Semantic Loss for Autoregressive Models with Logical Constraints
abstract
Neuro-symbolic AI bridges the gap between purely symbolic and neural approaches to learning. This often requires maximizing the likelihood of a symbolic constraint w.r.t the neural network's output distribution. Such output distributions are typically assumed to be fully-factorized. This limits the applicability of neuro-symbolic learning to the more expressive auto-regressive distributions, e.g., transformers. Under such distributions, computing the likelihood of even simple constraints is #P-hard. Instead of attempting to enforce the constraint on the entire likelihood distribution, we propose to do so on a random, local approximation thereof. More precisely, we approximate the likelihood of the constraint with the pseudolikelihood of the constraint centered around a model sample. Our approach is factorizable, allowing us to reuse solutions to sub-problems---a main tenet for the efficient computation of neuro-symbolic losses. It also provides a local, high fidelity approximation of the likelihood: it exhibits low entropy and KL-divergence around the model sample. We tested our approach on Sudoku and shortest-path prediction cast as auto-regressive generation, and observe that we greatly improve upon the base model's ability to predict logically-consistent outputs. We also tested our approach on the task of detoxifying large language models. We observe that using a simple constraint disallowing a list of toxic words, we are able to steer the model's outputs away from toxic generations, achieving SoTA compared to previous approaches.
Kareem Ahmed, Kai-Wei Chang 0001, Guy Van den Broeck
NeurIPS1
2023 A Unified Approach to Count-Based Weakly Supervised Learning
abstract
High-quality labels are often very scarce, whereas unlabeled data with inferred weak labels occurs more naturally. In many cases, these weak labels dictate the frequency of each respective class over a set of instances. In this paper, we develop a unified approach to learning from such weakly-labeled data, which we call *count-based weakly-supervised learning*. At the heart of our approach is the ability to compute the probability of exactly $k$ out of $n$ outputs being set to true. This computation is differentiable, exact, and efficient. Building upon the previous computation, we derive a *count loss* penalizing the model for deviations in its distribution from an arithmetic constraint defined over label counts.
Vinay Shukla, Zhe Zeng 0001, Kareem Ahmed, Guy Van den Broeck
NeurIPS3
2022 PYLON: A PyTorch Framework for Learning with Constraints
abstract
Deep learning excels at learning task information from large amounts of data, but struggles with learning from declarative high-level knowledge that can be more succinctly expressed directly. In this work, we introduce PYLON, a neuro-symbolic training framework that builds on PyTorch to augment procedurally trained models with declaratively specified knowledge. PYLON lets users programmatically specify constraints as Python functions and compiles them into a differentiable loss, thus training predictive models that fit the data whilst satisfying the specified constraints. PYLON includes both exact as well as approximate compilers to efficiently compute the loss, employing fuzzy logic, sampling methods, and circuits, ensuring scalability even to complex models and constraints. Crucially, a guiding principle in designing PYLON is the ease with which any existing deep learning codebase can be extended to learn from constraints in a few lines code: a function that expresses the constraint, and a single line to compile it into a loss. Our demo comprises of models in NLP, computer vision, logical games, and knowledge graphs that can be interactively trained using constraints as supervision.
Kareem Ahmed, Tao Li 0039, Thy Ton, Quan Guo, Kai-Wei Chang 0001, Parisa Kordjamshidi, Vivek Srikumar, Guy Van den Broeck, Sameer Singh 0001
AAAI1
2022 Semantic Probabilistic Layers for Neuro-Symbolic Learning
abstract
We design a predictive layer for structured-output prediction (SOP) that can be plugged into any neural network guaranteeing its predictions are consistent with a set of predefined symbolic constraints. Our Semantic Probabilistic Layer (SPL) can model intricate correlations, and hard constraints, over a structured output space all while being amenable to end-to-end learning via maximum likelihood.SPLs combine exact probabilistic inference with logical reasoning in a clean and modular way, learning complex distributions and restricting their support to solutions of the constraint. As such, they can faithfully, and efficiently, model complex SOP tasks beyond the reach of alternative neuro-symbolic approaches. We empirically demonstrate that SPLs outperform these competitors in terms of accuracy on challenging SOP tasks such as hierarchical multi-label classification, pathfinding and preference learning, while retaining perfect constraint satisfaction.
Kareem Ahmed, Stefano Teso, Kai-Wei Chang 0001, Guy Van den Broeck, Antonio Vergari
NeurIPS1
2022 Neuro-symbolic entropy regularization
abstract
In structured output prediction, the goal is to jointly predict several output variables that together encode a structured object – a path in a graph, an entity-relation triple, or an ordering of objects. Such a large output space makes learning hard and requires vast amounts of labeled data. Different approaches leverage alternate sources of supervision. One approach – entropy regularization – posits that decision boundaries should lie in low-probability regions. It extracts supervision from unlabeled examples, but remains agnostic to the structure of the output space. Conversely, neuro-symbolic approaches exploit the knowledge that not every prediction corresponds to a valid structure in the output space. Yet, they do not further restrict the learned output distribution.This paper introduces a framework that unifies both approaches. We propose a loss, neuro-symbolic entropy regularization, that encourages the model to confidently predict a valid object. It is obtained by restricting entropy regularization to the distribution over only the valid structures. This loss can be computed efficiently when the output constraint is expressed as a tractable logic circuit. Moreover, it seamlessly integrates with other neuro-symbolic losses that eliminate invalid predictions. We demonstrate the efficacy of our approach on a series of semi-supervised and fully-supervised structured-prediction experiments, where it leads to models whose predictions are more accurate as well as more likely to be valid.
Kareem Ahmed, Kai-Wei Chang 0001, Guy Van den Broeck
UAI1
2022 Performance evaluation of salient object detection techniques
abstract
Abstract Recently, the detection and segmentation of salient objects that attract the attention of human visual in images is determined by using salient object detection (SOD) techniques. As an essential computer vision problem, SOD has increasingly attracted the researchers’ interest over the years. While a lot of SOD models and applications have been proposed, there is still a lack of deep understanding of the issues and achievements. A comprehensive study on the recent techniques of SOD is provided in this paper. Precisely, this paper presents a review of SOD techniques from various perspectives. Various image segmentation techniques are presented such as segmentation based on machine learning or deep learning, the second perspective concentrates on classifying them into supervised and unsupervised learning techniques and the last one based on manual approach, semi-automatic approach, and fully automatic approach and so on. Then, the paper presents a summarization of datasets used for SOD. Finally, analyses of SOD models and comparison results are presented.
Kareem Ahmed, Mai A. Gad, Amal Elsayed Aboutabl
Multim. Tools Appl.1
2018 Action recognition using fast HOG3D of integral videos and Smith-Waterman partial matching
abstract
Recognising human activity from video stream has become one of the most interesting applications in computer vision. In this study, a novel hybrid technique for human action recognition is proposed based on fast HOG3D of integral videos and Smith–Waterman partial shape matching of the fused frame. The proposed technique is divided into two main stages, the first stage extracts a set of foreground snippets from the input video, and extracts the histogram of 3D gradient orientations from the spatio‐temporal volumetric data; and the second stage fuses a set of key frames from current snippet and extracts the contours from the fused frame. Non‐linear support vector machine (SVM) decision trees are used to classify HOG3D features into one of fixed action categories. On the other hand, Smith–Waterman partial shape matching is used to compare between the contour of the fused frame and the stored template contour of specified action. The results from SVM and Smith–Waterman partial shape matching are then combined. The experimental results show that combining non‐linear SVM decision trees of HOG3D features and Smith–Waterman partial shape matching of fused contours improved the accuracy of the classification model while maintaining efficiency in time elapsed for training.
Ibrahim M. El-Henawy, Kareem Ahmed, Hamdi Mahmoud
IET Image Process.2