on this site
For the first two weeks the LLM picked the action each turn from a menu the engine enumerated. It was legal and it was legible and it was still wrong in ways nobody could see: a fight read perfectly and the tactics were nonsense.
So it lost the job. The engine picks now, with a playbook that is matchup-aware and measured. The model calls the fight and plays the heel. The fights got better in a week, and every remaining bug became something a test could name.
Comments 0
Say what you used. No tool-shaming, no scores, no drive-bys.