There are some very interesting latent compaction approaches like this[1] for when you can control the whole inference stack. i.e in on-device and datacenter inference.
There are some very interesting latent compaction approaches like this[1] for when you can control the whole inference stack. i.e in on-device and datacenter inference.