
Bahram Behzadian
Research Scientist, Meta
I study policy improvement when the model of the world or the evaluator is imperfect.
About
I work on reinforcement learning and sequential decision-making. The through-line across my research is policy improvement when the model of the world or the evaluator is imperfect: robust MDPs when dynamics are uncertain, and truncated or approximate evaluators when reward signals are unreliable.
My recent work studies policy improvement under verifier-style rewards in recursive reasoning models, with applications to RL post-training, RLVR, and reward hacking under imperfect feedback. Earlier work developed scalable algorithms for robust and risk-sensitive MDPs.
Research Interests
- Policy improvement under truncated or approximate evaluators
- Verifier-style rewards, RLVR, and reward hacking
- Robust MDPs and decision-making under uncertainty
- Model-based RL and planning under learned or imperfect models
Featured Work
UPI-TRM: Policy Improvement under Truncated Evaluators with Verifier-Style Rewards
Bahram Behzadian, Brett Daley, Gopeshh Subbaraj, and Houssam Nassif
ICLR 2026 Workshop on AI with Recursive Self-Improvement
Studies when policy improvement remains reliable under truncated evaluators and when imperfect verifier feedback creates failure modes.