We tried Luna and it scored way lower in our evals. Muse also. We haven't had a chance to test others.

How much time were you able to put into tuning your prompts? And was it worse on all fronts (cost, latency, accuracy) or just some?