I'm not debating you at all. I'm asking what the model looks like since you've stated (and I've agreed) that a language model wouldn't work.
I think it would make sense to explain how a theoretical model could do better than SAT. Otherwise, is the idea here just "magic is possible"?
Yes, "magic is possible" if you defined "magic" as "very large models approximating functions in a way that people didn't think would work".
Current SOTA language and vision models, or models used to predict protein shapes are magic by the standards of 2016. As for why could it be better than a SAT? Why couldn't it be? Models are better than deterministic, logically written software for lots of situations. You can create infinite training data for this problem. The number of humans that work on encryption is tiny. The idea that because humans haven't figured out how to break some encryption schemes it can't be done is kind of absurd.