At least the SSM layers should have relative PE built in via recurrence, in a hybrid attention model like this.