Sure, object recognition is System 1, but Jev's own use cases list things like security incident triage, invoice approval, agent escalation, support actions, etc.
Those are generally the kind of tasks that require "System 2" in humans.
To be clear, I think the whole "System 1 vs System 2" framing is a pretty limiting way to think about AI (and thinking in general).
They are "System 2" in humans, because we haven't evolved the necessary wetware circuitry for it to be "System 1".
But with models, we can train them to answer such questions without verbal reasoning.