Momin Haider

dblp:391/5642 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 61% Reinforcement learning · 39%
Computer networks
1 paper
Routing and switching · 77% Network performance modeling · 23%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
decoding
0.812024
Fast Best-of-N Decoding via Speculative Rejection · NeurIPS 2024
Natural language and speech › Language models and text generation › alignment
inference-time alignment
0.812024
Fast Best-of-N Decoding via Speculative Rejection · NeurIPS 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
NetworkGym: Reinforcement Learning Environments for Multi-Access Traffic Management in Network Simulation · NeurIPS 2024
Routing and switching
load sharing
0.812024
NetworkGym: Reinforcement Learning Environments for Multi-Access Traffic Management in Network Simulation · NeurIPS 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.212024
Fast Best-of-N Decoding via Speculative Rejection · NeurIPS 2024
Network performance modeling
network simulation
0.212024
NetworkGym: Reinforcement Learning Environments for Multi-Access Traffic Management in Network Simulation · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

value-function pessimism · 1.5TD3+BC · 1.5CQL · 1.5speculative rejection · 0.8reward model · 0.8
YearPublicationVenuePosition
2024 NetworkGym: Reinforcement Learning Environments for Multi-Access Traffic Management in Network Simulation
abstract
Mobile devices such as smartphones, laptops, and tablets can often connect to multiple access networks (e.g., Wi-Fi, LTE, and 5G) simultaneously.Recent advancements facilitate seamless integration of these connections below the transport layer, enhancing the experience for apps that lack inherent multi-path support.This optimization hinges on dynamically determining the traffic distribution across networks for each device, a process referred to as multi-access traffic splitting.This paper introduces NetworkGym, a high-fidelity network environment simulator that facilitates generating multiple network traffic flows and multi-access traffic splitting.This simulator facilitates training and evaluating different RL-based solutions for the multi-access traffic splitting problem.Our initial explorations demonstrate that the majority of existing state-of-the-art offline RL algorithms (e.g. CQL) fail to outperform certain hand-crafted heuristic policies on average.This illustrates the urgent need to evaluate offline RL algorithms against a broader range of benchmarks, rather than relying solely on popular ones such as D4RL.We also propose an extension to the TD3+BC algorithm, named Pessimistic TD3 (PTD3), and demonstrate that it outperforms many state-of-the-art offline RL algorithms.PTD3's behavioral constraint mechanism, which relies on value-function pessimism, is theoretically motivated and relatively simple to implement.We open source our code and offline datasets at github.com/hmomin/networkgym.
Momin Haider, Ming Yin 0003, Menglei Zhang, Arpit Gupta, Yu-Xiang Wang 0003
NeurIPS1
2024 Fast Best-of-N Decoding via Speculative Rejection
abstract
The safe and effective deployment of Large Language Models (LLMs) involves a critical step called alignment, which ensures that the model's responses are in accordance with human preferences. Prevalent alignment techniques, such as DPO, PPO and their variants, align LLMs by changing the pre-trained model weights during a phase called post-training. While predominant, these post-training methods add substantial complexity before LLMs can be deployed. Inference-time alignment methods avoid the complex post-training step and instead bias the generation towards responses that are aligned with human preferences. The best-known inference-time alignment method, called Best-of-N, is as effective as the state-of-the-art post-training procedures. Unfortunately, Best-of-N requires vastly more resources at inference time than standard decoding strategies, which makes it computationally not viable. In this work, we introduce Speculative Rejection, a computationally-viable inference-time alignment algorithm. It generates high-scoring responses according to a given reward model, like Best-of-N does, while being between 16 to 32 times more computationally efficient.
Hanshi Sun, Momin Haider, Huitao Yang, Jiahao Qiu, Ming Yin 0003, Mengdi Wang 0001, Peter L. Bartlett, Andrea Zanette
NeurIPS2