Mahdi Karami

dblp:90/394 · DBLP profile ↗
← Back
11ranked-venue papers
9as first author
5since 2021 · last 2026
0000-0002-5096-6204ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 5 since 2021Computer networks · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 35% Graph learning · 20% Representation and self-supervised learning · 17%
Computer networks
1 paper
Physical-layer communications · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
sequence modeling
1.622025
Best of Both Worlds: Advantages of Hybrid Graph Sequence Models · ICML 2025
Orchid: Flexible and Data-Dependent Convolution for Sequence Modeling · NeurIPS 2024
Machine learning › Graph learning
graph neural network
0.912025
Best of Both Worlds: Advantages of Hybrid Graph Sequence Models · ICML 2025
Machine learning › Deep learning architectures and training
hybrid architecture
0.912025
Best of Both Worlds: Advantages of Hybrid Graph Sequence Models · ICML 2025
Machine learning › Deep learning architectures and training
attention mechanism
0.812024
Orchid: Flexible and Data-Dependent Convolution for Sequence Modeling · NeurIPS 2024
Machine learning › Generative modeling
autoregressive model
0.812024
HiGen: Hierarchical Graph Generative Networks · ICLR 2024
Machine learning › Graph learning
graph generation
0.812024
HiGen: Hierarchical Graph Generative Networks · ICLR 2024
Machine learning › Representation and self-supervised learning
multi-view learning
0.512021
Deep Probabilistic Canonical Correlation Analysis · AAAI 2021
Machine learning › Representation and self-supervised learning › multi-view learning › canonical correlation analysis
probabilistic canonical correlation analysis
0.512021
Deep Probabilistic Canonical Correlation Analysis · AAAI 2021
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.512021
Deep Probabilistic Canonical Correlation Analysis · AAAI 2021
Machine learning › Generative modeling › normalizing flow
invertible convolution
0.412019
Invertible Convolutional Flow · NeurIPS 2019
Machine learning › Generative modeling
normalizing flow
0.412019
Invertible Convolutional Flow · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.312017
Multi-view Matrix Factorization for Linear Dynamical System Estimation · NIPS 2017
Machine learning › Representation and self-supervised learning
matrix factorization
0.312017
Multi-view Matrix Factorization for Linear Dynamical System Estimation · NIPS 2017
Machine learning › Representation and self-supervised learning › matrix factorization
multi-view matrix factorization
0.312017
Multi-view Matrix Factorization for Linear Dynamical System Estimation · NIPS 2017
Machine learning › Graph learning
graph representation learning
0.312025
Best of Both Worlds: Advantages of Hybrid Graph Sequence Models · ICML 2025
Physical-layer communications › modulation
adaptive modulation
0.212013
Pilot Symbol Parameter Optimization Based on Imperfect Channel State Prediction for OFDM Systems · IEEE Trans. Commun. 2013
Physical-layer communications
channel estimation
0.212013
Pilot Symbol Parameter Optimization Based on Imperfect Channel State Prediction for OFDM Systems · IEEE Trans. Commun. 2013
Physical-layer communications › channel estimation
training signal design
0.212013
Pilot Symbol Parameter Optimization Based on Imperfect Channel State Prediction for OFDM Systems · IEEE Trans. Commun. 2013
Physical-layer communications › channel estimation
channel prediction
0.012013
Pilot Symbol Parameter Optimization Based on Imperfect Channel State Prediction for OFDM Systems · IEEE Trans. Commun. 2013

Methods — techniques the papers use, named apart from their topics

transformer · 0.9tokenization · 0.9linear RNN · 0.9hierarchical affinity clustering · 0.9shift equivariance · 0.8recursive factorization · 0.8multinomial distribution · 0.8conditioning network · 0.8autoregressive generation · 0.8deep generative network · 0.5minimum mean square error estimation · 0.2
YearPublicationVenuePosition
2026 FedBNR: A Fully Global Federated Gaussian Process
abstract
Uncertainty estimation plays a key role in many practical areas such as simulation and parameter optimization. Gaussian process (GP) is one popular model that provides naturally well-calibrated uncertainty estimates. However, it is challenging to learn a global GP posterior under the federated learning (FL) framework. In FL, clients’ private data should not be shared, but merging local kernels directly leads to privacy leakage. Previous works that consider federated GPs avoid it and focus on the personalized setting. This sacrifices information from other clients that can be exploited to benefit generalization. We present Federated Bayesian Neural Regression (FedBNR) that learns a global federated GP while respecting clients’ privacy. We incorporate deep kernel learning and random features by defining a unifying random kernel (URK). URK enables a principled approach of learning a global posterior as if all client data is centralized. Experiments conducted on real world regression datasets show statistically significant improvements.
Haolin Yu, Kaiyang Guo, Mahdi Karami, Pascal Poupart
Mach. Learn.3
2025 Best of Both Worlds: Advantages of Hybrid Graph Sequence Models
abstract
Modern sequence models (e.g., Transformers and linear RNNs) emerged as dominant backbones of recent deep learning frameworks, mainly due to their efficiency, representational power, and/or ability to capture long-range dependencies. Recently, adopting these sequence models for graph-structured data has gained popularity as the alternative to Message Passing Neural Networks (MPNNs). There is, however, a lack of a common foundation about what constitutes a good graph sequence model, and a mathematical description of the benefits and deficiencies in adopting different sequence models for learning on graphs. To this end, we introduce the Graph Sequence Model (GSM), a unifying framework for applying sequence models to graph data. The GSM framework allows us to understand, evaluate, and compare the power of different sequence model backbones in graph tasks. Building on this insight, we propose GSM++, a fast hybrid model that hierarchically tokenizes the graph using Hierarchical Affinity Clustering (HAC) and then encodes these sequences via a hybrid architecture. The theoretical and experimental findings confirm that GSM++ outperforms baseline models on most benchmarks.
Ali Behrouz, Ali Parviz, Mahdi Karami, Clayton Sanford, Bryan Perozzi, Vahab S. Mirrokni
ICML3
2024 HiGen: Hierarchical Graph Generative Networks
abstract
Most real-world graphs exhibit a hierarchical structure, which is often overlooked by existing graph generation methods. To address this limitation, we propose a novel graph generative network that captures the hierarchical nature of graphs and successively generates the graph sub-structures in a coarse-to-fine fashion. At each level of hierarchy, this model generates communities in parallel, followed by the prediction of cross-edges between communities using separate neural networks. This modular approach enables scalable graph generation for large and complex graphs. Moreover, we model the output distribution of edges in the hierarchical graph with a multinomial distribution and derive a recursive factorization for this distribution. This enables us to generate community graphs with integer-valued edge weights in an autoregressive manner. Empirical studies demonstrate the effectiveness and scalability of our proposed generative model, achieving state-of-the-art performance in terms of graph quality across various benchmark datasets. Code available at https://github.com/Karami-m/HiGen_main.
Mahdi Karami
ICLR1
2024 Orchid: Flexible and Data-Dependent Convolution for Sequence Modeling
abstract
In the rapidly evolving field of deep learning, the demand for models that are both expressive and computationally efficient has never been more critical. This paper introduces Orchid, a novel architecture designed to address the quadratic complexity of traditional attention mechanisms without compromising the ability to capture long-range dependencies and in-context learning. At the core of this architecture lies a new data-dependent global convolution layer, which contextually adapts its kernel conditioned on input sequence using a dedicated conditioning neural network. We design two simple conditioning networks that maintain shift equivariance in our data-dependent convolution operation. The dynamic nature of the proposed convolution kernel grants Orchid high expressivity while maintaining quasilinear scalability for long sequences. We evaluate the proposed model across multiple domains, including language modeling and image classification, to highlight its performance and generality. Our experiments demonstrate that this architecture not only outperforms traditional attention-based architectures such as BERT and Vision Transformers with smaller model sizes, but also extends the feasible sequence length beyond the limitations of the dense attention layers. This achievement represents a significant step towards more efficient and scalable deep learning models for sequence modeling.
Mahdi Karami, Ali Ghodsi 0001
NeurIPS1
2021 Deep Probabilistic Canonical Correlation Analysis
abstract
We propose a deep generative framework for multi-view learning based on a probabilistic interpretation of canonical correlation analysis (CCA). The model combines a linear multi-view layer in the latent space with deep generative networks as observation models, to decompose the variability in multiple views into a shared latent representation that describes the common underlying sources of variation and a set of viewspecific components. To approximate the posterior distribution of the latent multi-view layer, an efficient variational inference procedure is developed based on the solution of probabilistic CCA. The model is then generalized to an arbitrary number of views. An empirical analysis confirms that the proposed deep multi-view model can discover subtle relationships between multiple views and recover rich representations.
Mahdi Karami, Dale Schuurmans
AAAI1
2019 Invertible Convolutional Flow
abstract
Normalizing flows can be used to construct high quality generative probabilistic models, but training and sample generation require repeated evaluation of Jacobian determinants and function inverses. To make such computations feasible, current approaches employ highly constrained architectures that produce diagonal, triangular, or low rank Jacobian matrices. As an alternative, we investigate a set of novel normalizing flows based on the circular and symmetric convolutions. We show that these transforms admit efficient Jacobian determinant computation and inverse mapping (deconvolution) in O(N log N) time. Additionally, element-wise multiplication, widely used in normalizing flow architectures, can be combined with these transforms to increase modeling flexibility. We further propose an analytic approach to designing nonlinear elementwise bijectors that induce special properties in the intermediate layers, by implicitly introducing specific regularizers in the loss. We show that these transforms allow more effective normalizing flow models to be developed for generative image models.
Mahdi Karami, Dale Schuurmans, Jascha Sohl-Dickstein, Laurent Dinh, Daniel Duckworth
NeurIPS1
2017 Multi-view Matrix Factorization for Linear Dynamical System Estimation
abstract
We consider maximum likelihood estimation of linear dynamical systems with generalized-linear observation models. Maximum likelihood is typically considered to be hard in this setting since latent states and transition parameters must be inferred jointly. Given that expectation-maximization does not scale and is prone to local minima, moment-matching approaches from the subspace identification literature have become standard, despite known statistical efficiency issues. In this paper, we instead reconsider likelihood maximization and develop an optimization based strategy for recovering the latent states and transition parameters. Key to the approach is a two-view reformulation of maximum likelihood estimation for linear dynamical systems that enables the use of global optimization algorithms for matrix factorization. We show that the proposed estimation strategy outperforms widely-used identification algorithms such as subspace identification methods, both in terms of accuracy and runtime.
Mahdi Karami, Martha White, Dale Schuurmans, Csaba Szepesvári
NIPS1
2013 Pilot Symbol Parameter Optimization Based on Imperfect Channel State Prediction for OFDM Systems
abstract
The optimization of pilot symbol parameters can improve the spectral efficiency of adaptive modulation for orthogonal frequency division multiplexing (OFDM) systems, since pilot symbols impose an overhead on the system consuming power and bandwidth. An optimal pilot symbol assisted adaptive modulation (PSAAM) scheme for OFDM systems is proposed that maximizes spectral efficiency by adapting the power and constellation size of each subcarrier based on employing imperfect channel state information (CSI) at the transmitter. The pilot symbol power and spacing is also optimized in this scheme. A suboptimum scheme that decreases computational complexity without perceivable loss in performance is also presented. The optimality of minimum mean square error (MMSE) channel prediction for OFDM systems expressed in terms of a lower bound on spectral efficiency is approached. It is proved that the rectangular pilot pattern with equi-spaced and equal power pilot tones achieves the minimum MSE of the channel prediction in addition to having the advantage of simplifying PSAAM design. Numerical results show the importance of optimal pilot parameter adjustment for rapidly fading channels.
Mahdi Karami, Ali Olfat, Norman C. Beaulieu
IEEE Trans. Commun.1
2012 Channel adaptive power allocation and pilot optimization for OFDM systems
abstract
An adaptive resource allocation scheme is designed for OFDM that maximizes the average transmission rate by adapting the power of the data symbols while optimizing the pilot symbol spacing and power. Assuming imperfect channel state information (CSI) at the transmitter, the exact solution obtained for optimum power loading across data subcarriers can not be expressed in terms of elementary functions. A useful approximation, that has the simple form of a water-filling function, is derived for determining the optimum power loading function that is based on imperfect CSI. The merits of the proposed adaptive resource allocation schemes are corroborated by numerical results.
Mahdi Karami, Norman C. Beaulieu
GLOBECOM1
2010 Pilot Symbol Assisted Adaptive Modulation for OFDM Systems with Imperfect Channel State Information
abstract
Adaptive modulation exhibits the advantage of great spectral efficiency for high data rate wireless transmission, where the transmission power and rate of subchannels in orthogonal frequency division multiplexing (OFDM) systems are matched to the channel state information (CSI). CSI can be acquired at the receiver by MMSE channel estimation based on pilot symbol assisted modulation (PSAM), which is a widely used technique, and fed back to the transmitter. While more accurate channel estimation will result in more reliable transmission, pilot symbols don''t carry information and they impose an overhead on the system consuming power and bandwidth. An optimal adaptive PSAM scheme, that maximizes spectral efficiency by optimizing pilot power and spacing simultaneously is proposed. A rectangular pilot pattern is adopted to simplify the adaptive scheme.
Mahdi Karami, Ali Olfat, Norman C. Beaulieu
GLOBECOM1
2008 Fast Blind Adaptive Channel Shortening Using Signal Subspace
abstract
In cyclic-prefixed communication systems, if the delay spread of the channel is longer than the cyclic prefix (CP) a channel-shortening equalizer (CSE) can be used to restore the desired operation of such systems. Since in time-varying environment we are interested in fast adaptive equalizer with tracking capability, the aim of this paper is to propose RLS-type algorithm for channel shortening. In this paper, we first propose an RLS-type algorithm to estimate the eigenvector corresponding to the smallest eigenvalue of a matrix and based on this algorithm we develop an RLS-type blind channel shortener. We also, based on PAST algorithm, propose an RLS-type update rule to shorten the channel under MMSE criterion. Simulations show the speed advantage of the proposed algorithms.
Mahdi Karami, Ali Olfat
VTC Spring1