When I built rem, I spent significant effort getting screenshot -> ocr + screenshot -> ffmpeg loop energy efficient, but it definitely is more expensive than accessibility API.
You also save a lot of disk space and writes to disk.
That being said, you lose the cool swipe to go back in time and search through history and visually see, features.
And situations where accessibility isn't supported.
And as others have mentioned, built in ocr is definitely better than tesseract.