Airgapped LLM inferrence server can't serve their output tokens, right?
They can expose just their inference port, possible via some supervisor. The inference consumer can also be air gapped. This kind of segmentation is increasingly common for high value services.
Then it’s not air gapped…
not at a high bitrate
They can expose just their inference port, possible via some supervisor. The inference consumer can also be air gapped. This kind of segmentation is increasingly common for high value services.
Then it’s not air gapped…
not at a high bitrate