Rules that pull against each other
Three rules, each defensible on its own, combined into a trap where inventing something was the only legal move left.
Three rules in a policy at the same time, each of them reasonable on its own.
make exactly one change per pass a mandate to act you may not stop while a value is vague an exit closed you may not touch a protected value an option removed
Once the real improvements run out, deletion and invention are the only legal actions left. Not because the model is badly behaved, but because every honest move has been ruled out while it is being told it must move.
That single trap explains four pathologies that look unrelated when you meet them. Churn, where passes group things together and then split them back apart to arrive at the same content. Deletion of perfectly good decisions for the crime of being too vague to sharpen. Laundering of protected values by moving them somewhere else first. And questions invented so the pass has something to show.
Both fixes are structural rather than wording. Exactly one change becomes at most one change. And a second legal ending gets added, distinct from finished: the work can honestly end blocked, waiting on facts only a person has. Being told exactly what to sharpen, being unable to and returning nothing is an answer rather than a failure to converge.
Every rule has two ends
Add a line to stop a small model making up entries in the log. Something like: if you cannot name a change, write "no change" rather than inventing an edit. Reasonable on its own. It is also a blessed, zero effort way to give up. The next run declares the work finished on the first pass.
That line becomes the third statement in the policy pointing towards stopping, against a single line telling the model to find a move. A rule against inventing and a rule against stopping early point in opposite directions, so every policy edit needs testing against both ends rather than the one it was written for.
A second lesson is attached to that one and it is about diagnosis rather than design. The explanation in the paragraph above is the one written at the time and it was almost certainly wrong, because the offending line was not in the copy of the policy that was actually running. Establish that a change was in effect before crediting or blaming it. A plausible culprit is not an established one and the most recent change is the most seductive wrong answer available.
Never let an opinion end the run
A gate that can mark a field as needing a person will end the run if that mark is treated as terminal. Give a 36B model two vague fields with an obvious improvement sitting between them and it sets the mark on both, so the run ends on the first pass having done nothing. An opinion ended it before a single attempt.
The same shape turns up anywhere a flag is used as a filter. Everything comes back ineligible, the machinery never runs, nothing fails and the count sits unchanged. It is easy to make twice. The second time, the fix was written twenty lines above it in the same file.
A flag like that is advice about order, never about eligibility. Try anyway. Being wrong then costs one pass instead of the whole run.