Versioning the prompt finally made our rollout chart mean something
We rewrote a prompt and the overall answer score rose, so the first chart looked great. Then we split results by experience level. Newcomers got clearer steps while expert customers got longer answers that buried the usefulness. A model update had also landed that afternoon, which meant the aggregate chart couldn’t tell us which change moved quality. We versioned the prompt in
