But what if the model you're using doesn't have image processing capabilities?

You can set different compaction strategy, currently "Summarize in place and keep the current session", "Generate handoff and continue in a new session", "Drop heavy content in place, recover via artifact", Snapcompact as mentioned

I dunno, I didn't read in-depth. Hopefully you don't gotta zoom in with human eyeballs.