Prompts for How Does a Policy Improve
The human-side prompts that developed and steered this essay, in conversation order. Assistant responses and interface metadata are omitted. The two source slide decks are private; their filenames identify the material without publishing the files.
Source format
Which format do you prefer to read slides in: PDF, ODP, or PPTX? They have embedded screenshots of articles (figures etc.)
Initial request
Attached slides: Model-Based_RL.pdf and PWIL-Relative_Entropy-MPO-AWR.pdf.
I’m thinking that we could write an expository essay for the next slot on our blog, built around the attached slides, also present under notes/private/ . It would be quite technical. If nothing else, a refresher for me.
First pass
In response to a choice between a proposed structure and a complete draft:
Proposed structure and technical through-line
Mathematical connections
Would this be covering connections to ELBO and EM?
Recent research
These slides are at least 5 years old. Are there recent research highlights worth mentioning?
Length and scope
You can budget for 8,000 words, a substantial writeup. The hard limit is 11,000 words, and I’d rather not split into two articles.
Drafting
Go for it, thanks!
Generalization in the opening example
“Its usual behavior is to take the familiar route”: maybe e.g. “Its instinct is to take the route that looks familiar”? That implies generalization the former suggests revisitng an exact state.
Explicit state arguments
“For a recurring example, imagine a delivery robot deciding whether to wait” – maybe clearer to write with the function parameters? \pi_0(s,.) and Q^{\pi_0}(s,.) WDYT?
Introductory companion
The subsequent discussion led to a separate introductory article, “How Can a Reward Train a Neural Network?”. Its prompt companion records the framing and the request to introduce Bellman optimality before policy gradients.
Return distributions and conditioning on outcomes
The discussion of discounting, eligibility traces, risk, exploration, and upside-down RL is recorded in the introductory article’s prompt companion. It led to additions to both articles.
Sounds good. How do we incorporate this discussion into the articles? Just the introductory article, or do we split?
Sounds good. Let’s do that.
Relative feedback and representation learning
GRPO made me think of contrastive vs. regularized methods in representation learning. I wonder if it’s worth bringing up, but in the core article that I haven’t started reading yet? (Too heavy for the introductory one.)
Thanks! In “Which World Model”, we write “[Mattick] agrees with LeCun that the contrastive instruction to pull similar examples together and push others apart insufficiently constrains a useful representation.” Maybe we could refer back to it saying sth like that RL has internalized that lesson?