Won’t it be a good idea to stay at the visual level?

Instead of generating visuals from text and then back to text, just communicate visually with your agents.

Something like UML provides a certain granularity to be able to code something usefully in pure visual language. Of course there are other visual programming languages that provide more depth and would probably be better suited.

Point remains that a solution based on constantly mapping better two different representations of the same thing won’t be a long term efficient solution.