Rana Ali Amjad

dblp:126/4861 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0002-0007-4697ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Theory of computation · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorComputer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence
abstract
Providing Language Models (LMs) with relevant evidence in the context (either via retrieval or user-provided) can significantly improve their ability to provide better-grounded responses. However, recent studies have found that LMs often struggle to fully comprehend and utilize key evidence from the context, especially when it contains noise and irrelevant information—an issue common in real-world scenarios.To address this, we propose SelfElicit, an inference-time approach that helps LMs focus on key contextual evidence through self-guided explicit highlighting.By leveraging the inherent evidence-finding capabilities of LMs using the attention scores of deeper layers, our method automatically identifies and emphasizes key evidence within the input context, facilitating more accurate and grounded responses without additional training or iterative prompting.We demonstrate that SelfElicit brings consistent and significant improvement on multiple evidence-based QA tasks for various LM families while maintaining computational efficiency.Our code and documentation are available at https://github.com/ZhiningLiu1998/SelfElicit.
Zhining Liu 0002, Rana Ali Amjad, Ravinarayana Adkathimar, Tianxin Wei, Hanghang Tong
ACL (1)2
2022 Invertible Low-Divergence Coding
abstract
Several applications in communication, control, and learning require approximating target distributions to within small informational divergence. The additional requirement of invertibility usually leads to using encoders that are one-to-one mappings, also known as distribution matchers. However, even the best one-to-one encoders have divergences that grow logarithmically with the block length. To overcome this limitation, an encoder is proposed that has an invertible one-to-many mapping and a low-rate random number generator (RNG). Two algorithms are developed to design the mapping by assigning strings in either a most-likely first or least-likely first order. Both algorithms give information rates approaching the entropy of the target distribution with exponentially decreasing divergence and with vanishing RNG rate in the block length.
Patrick Schulte, Rana Ali Amjad, Thomas Wiegart, Gerhard Kramer
IEEE Trans. Inf. Theory2
2022 Understanding Neural Networks and Individual Neuron Importance via Information-Ordered Cumulative Ablation
abstract
In this work, we investigate the use of three information-theoretic quantities-entropy, mutual information with the class variable, and a class selectivity measure based on Kullback-Leibler (KL) divergence-to understand and study the behavior of already trained fully connected feedforward neural networks (NNs). We analyze the connection between these information-theoretic quantities and classification performance on the test set by cumulatively ablating neurons in networks trained on MNIST, FashionMNIST, and CIFAR-10. Our results parallel those recently published by Morcos et al., indicating that class selectivity is not a good indicator for classification performance. However, looking at individual layers separately, both mutual information and class selectivity are positively correlated with classification performance, at least for networks with ReLU activation functions. We provide explanations for this phenomenon and conclude that it is ill-advised to compare the proposed information-theoretic quantities across layers. Furthermore, we show that cumulative ablation of neurons with ascending or descending information-theoretic quantities can be used to formulate hypotheses regarding the joint behavior of multiple neurons, such as redundancy and synergy, with comparably low computational cost. We also draw connections to the information bottleneck theory for NNs.
Rana Ali Amjad, Kairen Liu, Bernhard C. Geiger
IEEE Trans. Neural Networks Learn. Syst.1
2021 Neural Augmentation of Kalman Filter with Hypernetwork for Channel Tracking
abstract
We propose Hypernetwork Kalman Filter (HKF) for tracking applications with multiple different dynamics. The HKF combines generalization power of Kalman filters with expressive power of neural networks. Instead of keeping a bank of Kalman filters and choosing one based on approximating the actual dynamics, HKF adapts itself to each dynamics based on the observed sequence. Through extensive experiments on CDL-B channel model, we show that the HKF can be used for tracking the channel over a wide range of Doppler values, matching Kalman filter performance with genie Doppler information. At high Doppler values, it achieves around 2dB gain over genie Kalman filter. The HKF generalizes well to unseen Doppler, SNR values and pilot patterns unlike LSTM, which suffers from severe performance degradation.
Kumar Pratik, Rana Ali Amjad, Arash Behboodi, Joseph B. Soriaga, Max Welling
GLOBECOM2
2020 Up or Down? Adaptive Rounding for Post-Training Quantization
abstract
When quantizing neural networks, assigning each floating-point weight to its nearest fixed-point value is the predominant approach. We find that, perhaps surprisingly, this is not the best we can do. In this paper, we propose AdaRound, a better weight-rounding mechanism for post-training quantization that adapts to the data and the task loss. AdaRound is fast, does not require fine-tuning of the network, and only uses a small amount of unlabelled data. We start by theoretically analyzing the rounding problem for a pre-trained neural network. By approximating the task loss with a Taylor series expansion, the rounding task is posed as a quadratic unconstrained binary optimization problem. We simplify this to a layer-wise local loss and propose to optimize this loss with a soft relaxation. AdaRound not only outperforms rounding-to-nearest by a significant margin but also establishes a new state-of-the-art for post-training quantization on several networks and tasks. Without fine-tuning, we can quantize the weights of Resnet18 and Resnet50 to 4 bits while staying within an accuracy loss of 1%.
Markus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos, Tijmen Blankevoort
ICML2
2020 Bayesian Bits: Unifying Quantization and Pruning
abstract
We introduce Bayesian Bits, a practical method for joint mixed precision quantization and pruning through gradient based optimization. Bayesian Bits employs a novel decomposition of the quantization operation, which sequentially considers doubling the bit width. At each new bit width, the residual error between the full precision value and the previously rounded value is quantized. We then decide whether or not to add this quantized residual error for a higher effective bit width and lower quantization noise. By starting with a power-of-two bit width, this decomposition will always produce hardware-friendly configurations, and through an additional 0-bit option, serves as a unified view of pruning and quantization. Bayesian Bits then introduces learnable stochastic gates, which collectively control the bit width of the given tensor. As a result, we can obtain low bit solutions by performing approximate inference over the gates, with prior distributions that encourage most of them to be switched off. We experimentally validate our proposed method on several benchmark datasets and show that we can learn pruned, mixed precision networks that provide a better trade-off between accuracy and efficiency than their static bit width equivalents.
Mart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad, Ying Wang 0051, Tijmen Blankevoort, Max Welling
NeurIPS4
2020 Learning Representations for Neural Network-Based Classification Using the Information Bottleneck Principle
abstract
In this theory paper, we investigate training deep neural networks (DNNs) for classification via minimizing the information bottleneck (IB) functional. We show that the resulting optimization problem suffers from two severe issues: First, for deterministic DNNs, either the IB functional is infinite for almost all values of network parameters, making the optimization problem ill-posed, or it is piecewise constant, hence not admitting gradient-based optimization methods. Second, the invariance of the IB functional under bijections prevents it from capturing properties of the learned representation that are desirable for classification, such as robustness and simplicity. We argue that these issues are partly resolved for stochastic DNNs, DNNs that include a (hard or soft) decision rule, or by replacing the IB functional with related, but more well-behaved cost functions. We conclude that recent successes reported about training DNNs using the IB framework must be attributed to such solutions. As a side effect, our results indicate limitations of the IB framework for the analysis of DNNs. We also note that rather than trying to repair the inherent problems in the IB functional, a better approach may be to design regularizers on latent representation enforcing the desired properties directly.
Rana Ali Amjad, Bernhard C. Geiger
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Co-Clustering via Information-Theoretic Markov Aggregation
abstract
We present an information-theoretic cost function for co-clustering, i.e., for simultaneous clustering of two sets based on similarities between their elements. By constructing a simple random walk on the corresponding bipartite graph, our cost function is derived from a recently proposed generalized framework for information-theoretic Markov chain aggregation. The goal of our cost function is to minimize relevant information loss, hence it connects to the information bottleneck formalism. Moreover, via the connection to Markov aggregation, our cost function is not ad hoc, but inherits its justification from the operational qualities associated with the corresponding Markov aggregation problem. We furthermore show that, for appropriate parameter settings, our cost function is identical to well-known approaches from the literature, such as “Information-Theoretic Co-Clustering” by Dhillon et al. Hence, understanding the influence of this parameter admits a deeper understanding of the relationship between previously proposed information-theoretic cost functions. We highlight some strengths and weaknesses of the cost function for different parameters. We also illustrate the performance of our cost function, optimized with a simple sequential heuristic, on several synthetic and real-world data sets, including the Newsgroup20 and the MovieLens100k data sets.
Clemens Blöchl, Rana Ali Amjad, Bernhard C. Geiger
IEEE Trans. Knowl. Data Eng.2
2018 Information Rates and Error Exponents for Probabilistic Amplitude Shaping
abstract
Probabilistic Amplitude Shaping (PAS) is a codedmodulation scheme in which the encoder is a concatenation of a distribution matcher with a systematic Forward Error Correction (FEC) code. For reduced computational complexity the decoder can be chosen as a concatenation of a mismatched FEC decoder and dematcher. This work studies the theoretic limits of PAS. The classical joint source-channel coding (JSCC) setup is modified to include systematic FEC and the mismatched FEC decoder. At each step error exponents and achievable rates for the corresponding setup are derived.
Rana Ali Amjad
ITW1
2015 Channel resolvability codes based on concatenation and sparse linear encoding
abstract
A concatenation of two encoders is used to construct channel resolvability codes. The code of the first encoder has large minimum distance and the second encoder is linear and has a sparse generator matrix. If the first encoder has encoding complexity O(n) or O(n log n), where n is the length of the codewords, an overall encoding complexity O(n log n) can be achieved. One can tune the sparsity to trade off the complexity of the second encoder against the minimum distance requirement of the first code, and to trade off the complexity of one of the encoders and the informational divergence scaling.
Rana Ali Amjad, Gerhard Kramer
ISIT1
2014 Informational divergence and entropy rate on rooted trees with probabilities
abstract
Rooted trees with probabilities are used to analyze properties of variable length codes. A bound is derived on the difference between the entropy rates of such codes and memoryless sources. The bound is in terms of normalized informational divergence and is used to derive converses for exact random number generation, resolution coding, and distribution matching.
Georg Böcherer, Rana Ali Amjad
ISIT2
2013 Fixed-to-variable length distribution matching
abstract
Fixed-to-variable length (f2v) matchers are used to reversibly transform an input sequence of independent and uniformly distributed bits into an output sequence of bits that are (approximately) independent and distributed according to a target distribution. The degree of approximation is measured by the informational divergence between the output distribution and the target distribution. An algorithm is developed that efficiently finds optimal f2v codes. It is shown that by encoding the input bits blockwise, the informational divergence per bit approaches zero as the block length approaches infinity. A relation to data compression by Tunstall coding is established.
Rana Ali Amjad, Georg Böcherer
ISIT1
2013 Fixed-to-variable length resolution coding for target distributions
abstract
The number of random bits required to approximate a target distribution in terms of un-normalized informational divergence is considered. It is shown that for a variable-to-variable length encoder, this number is lower bounded by the entropy of the target distribution. A fixed-to-variable length encoder is constructed using M-type quantization and Tunstall coding. It is shown that the encoder achieves in the limit an un-normalized informational divergence of zero with the number of random bits per generated symbol equal to the entropy of the target distribution. Numerical results show that the proposed encoder significantly outperforms the optimal block-to-block encoder in the finite length regime.
Georg Böcherer, Rana Ali Amjad
ITW2
2013 Average throughput maximization for energy harvesting transmitters with causal energy arrival information
abstract
We consider an energy harvesting node which transmits data using the energy it harvests from the environment. In the simple scenario of point-to-point communication, the performance of the node in terms of throughput is greatly influenced by the transmission strategy it employs and its knowledge about the energy arriving process. We assume in this work that the transmitting node does not have non-causal information about the energy to be harvested in future, but only has available the statistics of the energy arriving process which is stationary. The practical considerations that the energy storage capacity of the node is limited and there is additional energy consumption within the circuitry of the node are taken into account by our system model. Viewing the system as a finite-state Markov decision process, we optimize the transmission policy the node employs as a function of the energy storage state after an energy arrival, by using the policy-iteration algorithm. The asymptotic performance of the system in terms of average throughput is studied within the established theoretical and algorithmic framework, under the several transmission strategies we propose. Simulation results indicate the advantage of each strategy with respect to a certain range of values that the system parameters may take.
Qing Bai, Rana Ali Amjad, Josef A. Nossek
WCNC2