Because of masked attention in LLMs, if you put the options before the body (the email to analyze), the transformer already knows what it needs to look for, and can use more tokens to create state to address that specific task (BERT has no mask in the attention, so tokens attend also to next tokens). You could also do a few examples in the system prompt to improve calibration.
Another trick that works is to repeat the question two times: "I'm repeating the task and labels for clarity: ..."
Wow! TIL! I've been running a for loop around the two ordering variations to catch the winner of each turn and the difference is quite noticeable. In the options-after-body case in 47 of 100 attempts it classifies as phishing, whereas in the options-before-body case it classifies clearly as rickroll (94 out of 100 attempts)
Payroll sends you an email with a link to a Youtube video that plays a song.
Options after body:
Options before body: This was Gemma4-26B-A4B-NVFP4 by the way.EDIT
Gemma4-12B-it-NVFP4 seems way less sensitive to option/body ordering:
Options after body:
Options before body: Anyway, this for-looping stuff doing 100 calls to even a local VLLM API takes around 5 seconds in total, so this isn't anywhere close to sub-second Jev territory.Yah, that's what I use: https://rcarmo.github.io/projects/go-system-one uses Gemma, and that's partly why. Seems less prone to getting distracted with ordering.
What a time to be alive, repeating questions to a model twice to increase accuracy.
Repitation always helped make your point stronger. Repitation always helped make your point stronger.
I guess repeating a mistake helps make it more obvious too.
With enough repitation you might get repitition.
Use repetition to avoid trepidation if you have low reputation.
Or even repetition, repetition.
Mr Smith agrees.
You can say that again!
Repetition Legitimizes
I use ROT13 twice for extra security
I had an issue with accuracy a bit ago. So I repeated a couple of things without understanding why and it solved the problem.
I am glad there is an actual reason.
fuck this timeline
I feel you
Bro this timeline makes no sense. Repeating instruction to a data center of geniuses.
Djinniuses
Prompt Repetition Improves Non-Reasoning LLMs: https://arxiv.org/abs/2512.14982
This works very very well :).
https://github.com/Mushroom-Systems/lichen
Did I read this right?
This repo is really outperforming the OG Jev in the public benchmarks?
There was no time to benchmaxx. How is this possible?