Homa has been around for a while. Here's the 2018 paper.[1]
The core idea: When a message arrives at the sender’s transport module, Homa divides the message into two parts: an initial unscheduled portion (the first RTTbytes bytes), followed by a scheduled portion. The sender transmits the unscheduled bytes immediately, using one or more DATA packets. The scheduled bytes are not transmitted until requested explicitly by the receiver using GRANT packets.
So it sends blind for short requests, then needs a go-ahead from the receiver. That's reasonable when the main application is a remote procedure call. It's reminiscent of QNX's networking protocol, which is also single packet message request/response but can also handle arbitrarily long messages.
What makes this work today is that per-packet processing overhead in hardware switches is low vs. per-byte overhead. In early software driven switches, per-packet overhead tended to dominate, and sending small packets was very inefficient. In modern hardware switches, where FPGAs are doing the processing, the per-packet overhead is low enough that small packets are not inefficient.
It's amusing that web stuff is so bloated today that any transaction under 1MB is considered "small". So this is not a suitable protocol for open web use.
[1] https://people.csail.mit.edu/alizadeh/papers/homa-sigcomm18....
The packet sizes should be the same for TCP or Homa; the difference is in sender-based vs. receiver-based congestion control.
Upvote for the QNX mention! I loved using QNX for RT control stuff.
Also, thanks for the clear summary and context.