> Building local multimodal search with this would be amazing.

(lawyer here) — I’m curious: for what you’re describing, wouldn’t the machine need to be constantly running/updating the embeddings to take updated and new files into account? If so, how would that computation load compare to, say, Spotlight constantly updating its index?

The embedding model stays loaded in memory. It is used for turning your search keywords into embeddings.

The index you’re thinking of is made once per file.. then you compare and search in embeddings. Add/modify files = asynchronous updating or adding corresponding embeddings using the model in memory onto wherever you persist those embeddings(say SQLite)..

Also how you turn a file into one or more items is a separate question and would likely need tuning to circumstances, usually big documents are split (chunked) sometimes at paragraph or even more granularly. Where to optimally chunk alone is not easy. This also allows you to then search for a specific part of the document, at the expense of not taking the wide context into account, but embeddings usually struggle with too many tokens anyway.