Nasik Muhammad Nafi

dblp:271/9440 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-8492-9217ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Deep learning architectures and training · 42% Reinforcement learning · 32% Efficient and distributed learning · 21%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention
0.912025
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
large-scale distributed training
0.912025
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.912025
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025
High-performance computing › supercomputing
exascale computing
0.912025
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025
Machine learning › Reinforcement learning
off-policy reinforcement learning
0.712023
Multi-Horizon Learning in Procedurally-Generated Environments for Off-Policy Reinforcement Learning (Student Abstract) · AAAI 2023
Machine learning › Reinforcement learning
value-based reinforcement learning
0.712023
Multi-Horizon Learning in Procedurally-Generated Environments for Off-Policy Reinforcement Learning (Student Abstract) · AAAI 2023
Environmental and earth informatics › climate science
climate downscaling
0.312025
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025
Machine learning › Learning theory
generalization
0.212023
Multi-Horizon Learning in Procedurally-Generated Environments for Off-Policy Reinforcement Learning (Student Abstract) · AAAI 2023

Methods — techniques the papers use, named apart from their topics

tile-wise sequence scaling · 2.6residual learning · 2.6bayesian regularization · 2.6rainbow · 0.7hyperbolic discounting · 0.7advantage-based action selection · 0.7
YearPublicationVenuePosition
2025 A Minimalist Approach to Augmentation-based Self-supervised Representation Learning for On-policy Reinforcement Learning
Nasik Muhammad Nafi, William H. Hsu
AAMAS1
2025 ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling
abstract
Sparse observations and coarse-resolution climate models limit effective regional decision-making, underscoring the need for robust downscaling. However, existing AI methods struggle with generalization across variables and geographies and are constrained by the quadratic complexity of Vision Transformer (ViT) self-attention. We introduce ORBIT-2, a scalable foundation model for global, hyper-resolution climate downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with R2 scores in range of 0.98–0.99 against observation data.
Xiao Wang 0004, Jong-Youl Choi, Takuya Kurihana, Isaac Lyngaas, Hong-Jun Yoon, Xi Xiao 0003, David Pugmire, Nasik Muhammad Nafi, Aristeidis Tsaris, Ashwin M. Aji, Maliha Hossain, Mohamed Wahib, Dali Wang, Peter E. Thornton, Prasanna Balaprakash, Moetasim Ashfaq, Dan Lu 0001
SC9
2024 Analyzing the Sensitivity to Policy-Value Decoupling in Deep Reinforcement Learning Generalization
abstract
The existence of policy-value representation asymmetry negatively affects the generalization capability of traditional actor-critic architectures that use a shared representation of policy and value. To address this representation asymmetry, fully decoupled/separated networks for policy and value have been proposed, though they come with increased computational overhead. Recent research has suggested that partial separation of networks results in similar generalization performance with reduced computational costs. Thus, the questions arise: Do we really need two separate networks? Is there any particular scenario where only full separation works? Does increasing the degree of separation in a partially separated network improve generalization? To answer these questions, we present the first comprehensive study of the generalization performance of four different extents of decoupling of the policy and value networks, namely: fully shared, early separation, late separation, and full separation on the challenging RL generalization benchmark Procgen, a suite of 16 procedurally-generated environments, and the Crafter benchmark. Interestingly, we observe that early separation does not produce the expected generalization. Our findings suggest that, unless there is a distinct or explicit predetermined source of value estimation, partial late separation is an effective strategy for capturing necessary policy-value representation asymmetry and obtaining competitive generalization in unseen scenarios.
Nasik Muhammad Nafi, Raja Farrukh Ali, William H. Hsu
IJCNN1
2023 Multi-Horizon Learning in Procedurally-Generated Environments for Off-Policy Reinforcement Learning (Student Abstract)
abstract
Value estimates at multiple timescales can help create advanced discounting functions and allow agents to form more effective predictive models of their environment. In this work, we investigate learning over multiple horizons concurrently for off-policy reinforcement learning by using an advantage-based action selection method and introducing architectural improvements. Our proposed agent learns over multiple horizons simultaneously, while using either exponential or hyperbolic discounting functions. We implement our approach on Rainbow, a value-based off-policy algorithm, and test on Procgen, a collection of procedurally-generated environments, to demonstrate the effectiveness of this approach, specifically to evaluate the agent's performance in previously unseen scenarios.
Raja Farrukh Ali, Kevin Duong, Nasik Muhammad Nafi, William H. Hsu
AAAI3
2023 Relevant Instance Segmentation in American Football Practice Images to Aid Risky Tackle Detection
abstract
This paper addresses the problem of relevant region segmentation, as a pretext task for a defined multi-object scene classification task, with a specialized application to risky tackle detection from American football practice videos. The downstream task of classifying each frame from such a video as depicting a risky tackle or not depends on the interaction between the tackle-performing player and the target dummy. In both automated and manual approaches, if these two objects can not be differentiated from other objects as part of the analysis, false positive and false negative scene misclassification errors may result, to the detriment of both precision and recall. While player detection appears to be a simple human detection task, specific poses and occlusion due to the dummy make the instance segmentation task particularly challenging in the case of American football practice videos. In this paper, we present a new annotated dataset of tackle practice images and for the first time demonstrate instance segmentation in American football practice images leveraging the new dataset. Further, we show that the Cascade Mask R-CNN based segmentation approach is more suitable for the problem than another popular segmentation model, simple Mask R-CNNs, by characterizing the inherent difficulty of the task and comparing experimental results.
Nasik Muhammad Nafi, Ashley Rediger, Scott Dietrich, William H. Hsu
ICMLA1
2023 Policy Optimization with Augmented Value Targets for Generalization in Reinforcement Learning
abstract
Our work aims to improve the generalization performance of a reinforcement learning (RL) agent in unseen environment variations. The value function used in RL agents is frequently overfitted, leading to poor generalization performance. In this work, we argue that the task completion time is highly impacted by the varying environmental conditions, thus resulting in variation in episode lengths, and consequently, the value estimation. Therefore, learning from a limited variation of the environments, the agent gets biased to the value estimates that correspond to the observed episode lengths. To this end, we introduce Augmented Value Targets (AVaTar), which generates multiple value function targets considering the possibility of episode length variation and optimizes the value function with the average of these targets. We demonstrate that optimizing the average of the augmented targets is computationally more feasible than independently leveraging those pseudotargets. Evaluations on the Procgen and Crafter benchmark show that our proposed approach is effective in generalizing the value estimates over unseen contexts and significantly outperforms the standard policy gradient algorithm Proximal Policy Optimization (PPO). Furthermore, comparison and integration with the recent generalizationspecific approach UCB-DrAC indicate that AVaTar outperforms UCB-DrAC in most of the environments from Procgen.
Nasik Muhammad Nafi, Giovanni Poggi-Corradini, William H. Hsu
IJCNN1
2022 Attention-based Partial Decoupling of Policy and Value for Generalization in Reinforcement Learning
abstract
In this work, we introduce Attention-based Partially Decoupled Actor-Critic (APDAC), an actor-critic architecture for generalization in reinforcement learning, which partially separates the policy and the value functions. To learn directly from images, traditional actor-critic architectures use a shared network to represent the policy and value functions. While a shared representation allows parameter and feature sharing, it can also lead to overfitting that catastrophically damages generalization performance. On the other hand, two separate networks for policy and value can help to avoid overfitting and reduce the generalization gap, but at the cost of added complexity both in terms of architecture design and computation time. APDAC is a hybrid architecture that builds upon the combined strengths of both architectures by sharing initial layer blocks of the network and separating the later ones for policy and value. APDAC incorporates an attention mechanism to enable robust representation learning. We present meaningful visualization of the policy and value that explains the perception of the trained agent. Our empirical analysis, including an ablation study, shows that APDAC significantly outperforms the standard PPO baseline on the challenging RL generalization benchmark Procgen and achieves performance that is competitive with the recent state-of-the-art method (IDAAC) while using fewer convolutional layers and requiring less computational time. Our code is available at https://github.com/nasiknafi/apdac.
Nasik Muhammad Nafi, Creighton Glasscock, William H. Hsu
ICMLA1
2021 eGAN: Unsupervised Approach to Class Imbalance Using Transfer Learning
Ademola Okerinde, William H. Hsu, Tom Theis, Nasik Muhammad Nafi, Lior Shamir
CAIP (1)4