Rewrites Avatar actor instructions from positive and negative trajectories.
Each round evaluates the current actor, contrasts bounded positive and negative examples, asks typed predictors for feedback and a rewritten instruction, evaluates that candidate, and retains it only when it improves the configured objective.