EDBT 2026 Demo / reviewers in the wild / expert
Bilgehan Sel
dblp:335/1479
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0001-8701-6539ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 34% Reinforcement learning · 31% Planning, search and constraint satisfaction · 18% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model reasoning |
2.5 | 3 | 2025 | LLMs Can Reason Faster Only If We Let Them · ICML 2025 LLMs Can Plan Only If We Tell Them · ICLR 2025 Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models · ICML 2024 |
Machine learning › Reinforcement learning
safe reinforcement learning |
2.3 | 3 | 2025 | Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation · AAAI 2024 A CMDP-within-online framework for Meta-Safe Reinforcement Learning · ICLR 2023 |
Machine learning › Trustworthy machine learning › adversarial machine learning
adversarial defense |
0.9 | 1 | 2025 | Reinforcement Learning with Backtracking Feedback · NeurIPS 2025 |
Machine learning › Reinforcement learning › multi-objective reinforcement learning
constrained multi-objective reinforcement learning |
0.9 | 1 | 2025 | Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Natural language and speech › Language models and text generation
large language model fine-tuning |
0.9 | 1 | 2025 | Reinforcement Learning with Backtracking Feedback · NeurIPS 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
LLM-based planning |
0.9 | 1 | 2025 | LLMs Can Plan Only If We Tell Them · ICLR 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
long-horizon planning |
0.9 | 1 | 2025 | LLMs Can Plan Only If We Tell Them · ICLR 2025 |
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization |
0.8 | 1 | 2024 | Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation · AAAI 2024 |
Machine learning › Learning theory › generalization bounds
covering number bound |
0.7 | 1 | 2023 | On Solution Functions of Optimization: Universal Approximation and Covering Number Bounds · AAAI 2023 |
Mathematical optimization › continuous optimization
convex optimization |
0.7 | 1 | 2023 | On Solution Functions of Optimization: Universal Approximation and Covering Number Bounds · AAAI 2023 |
Machine learning › Optimization for machine learning
multi-objective optimization |
0.3 | 1 | 2025 | Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › planning evaluation
planning benchmarks |
0.3 | 1 | 2025 | LLMs Can Plan Only If We Tell Them · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning |
0.3 | 1 | 2025 | Reinforcement Learning with Backtracking Feedback · NeurIPS 2025 |
Machine learning › Optimization for machine learning › multi-objective optimization
gradient manipulation |
0.2 | 1 | 2024 | Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation · AAAI 2024 |
Machine learning › Optimization for machine learning
constrained optimization |
0.2 | 1 | 2023 | A CMDP-within-online framework for Meta-Safe Reinforcement Learning · ICLR 2023 |
Mathematical optimization › continuous optimization
linear and quadratic programming |
0.2 | 1 | 2023 | On Solution Functions of Optimization: Universal Approximation and Covering Number Bounds · AAAI 2023 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 1.7reinforcement learning · 1.7reward model · 0.9primal-based optimization · 0.9natural policy gradient · 0.9chain-of-thought prompting · 0.9algorithm-of-thoughts · 0.9soft switching policy optimization · 0.8gradient manipulation · 0.8convergence analysis · 0.8universal approximation · 0.7tame geometry · 0.7rate-distortion theory · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLMs Can Plan Only If We Tell ThemabstractLarge language models (LLMs) have demonstrated significant capabilities in natural language processing and reasoning, yet their effectiveness in autonomous planning has been under debate. While existing studies have utilized LLMs with external feedback mechanisms or in controlled environments for planning, these approaches often involve substantial computational and development resources due to the requirement for careful design and iterative backprompting. Moreover, even the most advanced LLMs like GPT-4 struggle to match human performance on standard planning benchmarks, such as the Blocksworld, without additional support. This paper investigates whether LLMs can independently generate long-horizon plans that rival human baselines. Our novel enhancements to Algorithm-of-Thoughts (AoT), which we dub AoT+, help achieve state-of-the-art results in planning benchmarks out-competing prior methods and human baselines all autonomously. Bilgehan Sel, Ruoxi Jia 0001, Ming Jin 0002 |
ICLR | 1 |
| 2025 | LLMs Can Reason Faster Only If We Let ThemabstractLarge language models (LLMs) are making inroads into classical AI problems such as automated planning, yet key shortcomings continue to hamper their integration. Chain-of-Thought (CoT) struggles in complex multi-step reasoning, and Tree-of-Thoughts requires multiple queries that increase computational overhead. Recently, Algorithm-of-Thoughts (AoT) have shown promise using in-context examples, at the cost of significantly longer solutions compared to CoT. Aimed at bridging the solution length gap between CoT and AoT, this paper introduces AoT-O3, which combines supervised finetuning on AoT-style plans with a reinforcement learning (RL) framework designed to reduce solution length. The RL component uses a reward model that favors concise, valid solutions while maintaining planning accuracy. Empirical evaluations indicate that AoT-O3 shortens solution length by up to 80\% compared to baseline AoT while maintaining or surpassing prior performance. These findings suggest a promising pathway for more efficient, scalable LLM-based planning. Bilgehan Sel, Lifu Huang, Naren Ramakrishnan, Ruoxi Jia 0001, Ming Jin 0002 |
ICML | 1 |
| 2025 | Reinforcement Learning with Backtracking FeedbackabstractAddressing the critical need for robust safety in Large Language Models (LLMs), particularly against adversarial attacks and in-distribution errors, we introduce Reinforcement Learning with Backtracking Feedback (RLBF). This framework advances upon prior methods, such as BSAFE, by primarily leveraging a Reinforcement Learning (RL) stage where models learn to dynamically correct their own generation errors. Through RL with critic feedback on the model's live outputs, LLMs are trained to identify and recover from their actual, emergent safety violations by emitting an efficient "backtrack by x tokens" signal, then continuing generation autoregressively. This RL process is crucial for instilling resilience against sophisticated adversarial strategies, including middle filling, Greedy Coordinate Gradient (GCG) attacks, and decoding parameter manipulations. To further support the acquisition of this backtracking capability, we also propose an enhanced Supervised Fine-Tuning (SFT) data generation strategy (BSAFE+). This method improves upon previous data creation techniques by injecting violations into coherent, originally safe text, providing more effective initial training for the backtracking mechanism. Comprehensive empirical evaluations demonstrate that RLBF significantly reduces attack success rates across diverse benchmarks and model scales, achieving superior safety outcomes while critically preserving foundational model utility. Bilgehan Sel, Vaishakh Keshava, Phillip Wallis, Lukas Rutishauser, Ming Jin 0002, Dingcheng Li |
NeurIPS | 1 |
| 2025 | Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement LearningabstractIn numerous reinforcement learning (RL) problems involving safety-critical systems, a key challenge lies in balancing multiple objectives while simultaneously meeting all stringent safety constraints. To tackle this issue, we propose a primal-based framework that orchestrates policy optimization between multi-objective learning and constraint adherence. Our method employs a novel natural policy gradient manipulation method to optimize multiple RL objectives and overcome conflicting gradients between different objectives, since the simple weighted average gradient direction may not be beneficial for specific objectives due to misaligned gradients of different objectives. When there is a violation of a hard constraint, our algorithm steps in to rectify the policy to minimize this violation. Particularly, We establish theoretical convergence and constraint violation guarantees, and our proposed method also outperforms prior state-of-the-art methods on challenging safe multi-objective RL tasks. Shangding Gu, Bilgehan Sel, Yuhao Ding, Lu Wang 0029, Qingwei Lin, Alois C. Knoll, Ming Jin 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient ManipulationabstractEnsuring the safety of Reinforcement Learning (RL) is crucial for its deployment in real-world applications. Nevertheless, managing the trade-off between reward and safety during exploration presents a significant challenge. Improving reward performance through policy adjustments may adversely affect safety performance. In this study, we aim to address this conflicting relation by leveraging the theory of gradient manipulation. Initially, we analyze the conflict between reward and safety gradients. Subsequently, we tackle the balance between reward and safety optimization by proposing a soft switching policy optimization method, for which we provide convergence analysis. Based on our theoretical examination, we provide a safe RL framework to overcome the aforementioned challenge, and we develop a Safety-MuJoCo Benchmark to assess the performance of safe RL algorithms. Finally, we evaluate the effectiveness of our method on the Safety-MuJoCo Benchmark and a popular safe benchmark, Omnisafe. Experimental results demonstrate that our algorithms outperform several state-of-the-art baselines in terms of balancing reward and safety optimization. Shangding Gu, Bilgehan Sel, Yuhao Ding, Lu Wang 0029, Qingwei Lin, Ming Jin 0002, Alois C. Knoll |
AAAI | 2 |
| 2024 | Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMsabstractBilgehan Sel, Priya Shanmugasundaram, Mohammad Kachuee, Kun Zhou, Ruoxi Jia, Ming Jin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Bilgehan Sel, Priya Shanmugasundaram, Mohammad Kachuee, Ruoxi Jia 0001, Ming Jin 0002 |
ACL (1) | 1 |
| 2024 | Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language ModelsabstractCurrent literature, aiming to surpass the "Chain-of-Thought" approach, often resorts to external modi operandi involving halting, modifying, and then resuming the generation process to boost Large Language Models’ (LLMs) reasoning capacities. Due to their myopic perspective, they escalate the number of query requests, leading to increased costs, memory, and computational overheads. Addressing this, we propose the Algorithm of Thoughts—a novel strategy that propels LLMs through algorithmic reasoning pathways. By employing algorithmic examples fully in-context, this overarching view of the whole process exploits the innate recurrence dynamics of LLMs, expanding their idea exploration with merely one or a few queries. Our technique outperforms earlier single-query methods and even more recent multi-query strategies that employ an extensive tree search algorithms while using significantly fewer tokens. Intriguingly, our results suggest that instructing an LLM using an algorithm can lead to performance surpassing that of the algorithm itself, hinting at LLM’s inherent ability to weave its intuition into optimized searches. We probe into the underpinnings of our method’s efficacy and its nuances in application. The code and related content can be found in: https://algorithm-of-thoughts.github.io Bilgehan Sel, Ahmad Al-Tawaha, Vanshaj Khattar, Ruoxi Jia 0001, Ming Jin 0002 |
ICML | 1 |
| 2023 | On Solution Functions of Optimization: Universal Approximation and Covering Number BoundsabstractWe study the expressibility and learnability of solution functions of convex optimization and their multi-layer architectural extension. The main results are: (1) the class of solution functions of linear programming (LP) and quadratic programming (QP) is a universal approximant for the smooth model class or some restricted Sobolev space, and we characterize the rate-distortion, (2) the approximation power is investigated through a viewpoint of regression error, where information about the target function is provided in terms of data observations, (3) compositionality in the form of deep architecture with optimization as a layer is shown to reconstruct some basic functions used in numerical analysis without error, which implies that (4) a substantial reduction in rate-distortion can be achieved with a universal network architecture, and (5) we discuss the statistical bounds of empirical covering numbers for LP/QP, as well as a generic optimization problem (possibly nonconvex) by exploiting tame geometry. Our results provide the **first rigorous analysis of the approximation and learning-theoretic properties of solution functions** with implications for algorithmic design and performance guarantees. Ming Jin 0002, Vanshaj Khattar, Harshal Kaushik, Bilgehan Sel, Ruoxi Jia 0001 |
AAAI | 4 |
| 2023 | A CMDP-within-online framework for Meta-Safe Reinforcement Learning
Vanshaj Khattar, Yuhao Ding, Bilgehan Sel, Javad Lavaei, Ming Jin 0002 |
ICLR | 3 |