Portrait of Bahram Behzadian

Bahram Behzadian

Research Scientist, Meta

I study policy improvement when the model of the world or the evaluator is imperfect.

About

I work on reinforcement learning and sequential decision-making. The through-line across my research is policy improvement when the model of the world or the evaluator is imperfect: robust MDPs when dynamics are uncertain, and truncated or approximate evaluators when reward signals are unreliable.

My recent work studies policy improvement under verifier-style rewards in recursive reasoning models, with applications to RL post-training, RLVR, and reward hacking under imperfect feedback. Earlier work developed scalable algorithms for robust and risk-sensitive MDPs.

Research Interests

  • Policy improvement under truncated or approximate evaluators
  • Verifier-style rewards, RLVR, and reward hacking
  • Robust MDPs and decision-making under uncertainty
  • Model-based RL and planning under learned or imperfect models