EDBT 2026 Demo / reviewers in the wild / expert
Momin Haider
dblp:391/5642
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 61% Reinforcement learning · 39% | |
| Computer networks
1 paper |
Routing and switching · 77% Network performance modeling · 23% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
decoding |
0.8 | 1 | 2024 | Fast Best-of-N Decoding via Speculative Rejection · NeurIPS 2024 |
Natural language and speech › Language models and text generation › alignment
inference-time alignment |
0.8 | 1 | 2024 | Fast Best-of-N Decoding via Speculative Rejection · NeurIPS 2024 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.8 | 1 | 2024 | NetworkGym: Reinforcement Learning Environments for Multi-Access Traffic Management in Network Simulation · NeurIPS 2024 |
Routing and switching
load sharing |
0.8 | 1 | 2024 | NetworkGym: Reinforcement Learning Environments for Multi-Access Traffic Management in Network Simulation · NeurIPS 2024 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.2 | 1 | 2024 | Fast Best-of-N Decoding via Speculative Rejection · NeurIPS 2024 |
Network performance modeling
network simulation |
0.2 | 1 | 2024 | NetworkGym: Reinforcement Learning Environments for Multi-Access Traffic Management in Network Simulation · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
value-function pessimism · 1.5TD3+BC · 1.5CQL · 1.5speculative rejection · 0.8reward model · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | NetworkGym: Reinforcement Learning Environments for Multi-Access Traffic Management in Network SimulationabstractMobile devices such as smartphones, laptops, and tablets can often connect to multiple access networks (e.g., Wi-Fi, LTE, and 5G) simultaneously.Recent advancements facilitate seamless integration of these connections below the transport layer, enhancing the experience for apps that lack inherent multi-path support.This optimization hinges on dynamically determining the traffic distribution across networks for each device, a process referred to as multi-access traffic splitting.This paper introduces NetworkGym, a high-fidelity network environment simulator that facilitates generating multiple network traffic flows and multi-access traffic splitting.This simulator facilitates training and evaluating different RL-based solutions for the multi-access traffic splitting problem.Our initial explorations demonstrate that the majority of existing state-of-the-art offline RL algorithms (e.g. CQL) fail to outperform certain hand-crafted heuristic policies on average.This illustrates the urgent need to evaluate offline RL algorithms against a broader range of benchmarks, rather than relying solely on popular ones such as D4RL.We also propose an extension to the TD3+BC algorithm, named Pessimistic TD3 (PTD3), and demonstrate that it outperforms many state-of-the-art offline RL algorithms.PTD3's behavioral constraint mechanism, which relies on value-function pessimism, is theoretically motivated and relatively simple to implement.We open source our code and offline datasets at github.com/hmomin/networkgym. Momin Haider, Ming Yin 0003, Menglei Zhang, Arpit Gupta, Yu-Xiang Wang 0003 |
NeurIPS | 1 |
| 2024 | Fast Best-of-N Decoding via Speculative RejectionabstractThe safe and effective deployment of Large Language Models (LLMs) involves a critical step called alignment, which ensures that the model's responses are in accordance with human preferences. Prevalent alignment techniques, such as DPO, PPO and their variants, align LLMs by changing the pre-trained model weights during a phase called post-training. While predominant, these post-training methods add substantial complexity before LLMs can be deployed. Inference-time alignment methods avoid the complex post-training step and instead bias the generation towards responses that are aligned with human preferences. The best-known inference-time alignment method, called Best-of-N, is as effective as the state-of-the-art post-training procedures. Unfortunately, Best-of-N requires vastly more resources at inference time than standard decoding strategies, which makes it computationally not viable. In this work, we introduce Speculative Rejection, a computationally-viable inference-time alignment algorithm. It generates high-scoring responses according to a given reward model, like Best-of-N does, while being between 16 to 32 times more computationally efficient. Hanshi Sun, Momin Haider, Huitao Yang, Jiahao Qiu, Ming Yin 0003, Mengdi Wang 0001, Peter L. Bartlett, Andrea Zanette |
NeurIPS | 2 |