"Generate an SVG of a pelican riding a bicycle" tests:
https://gistpreview.github.io/?815466e3208746488d47679949b68... - 33170 tokens
https://gistpreview.github.io/?815466e3208746488d47679949b68... - 18125 tokens
https://gistpreview.github.io/?815466e3208746488d47679949b68... - 12960 tokens
Scale 35 felt a bit noisy and incoherent, but 30 seemed nice. (they use the same seed, but idk how reliable seed in llamacpp is)
I use a python test script that captures the answer and renders it to a html page along with the llama-cli log, launch parameters, the chat log, and the python script itself for maximum transparency. :)