Its 200-300 TPS. GPT is at 50-60. And its like 5% worse? Yeah that is a good tradeoff.

Are you watching/waiting on your agents? I care zero about t/sec, quality is far more important that quantity or latency, but they run in the background and I check in from time to time