Tldr, strata is a waste of time.

at least for vision? Are there similar comparisons for language? It seems like vision is often an afterthought when it comes to bootstrapping these newer inference engines