There was an article a few years ago called "Lost in the Middle: How Language Models Use Long Contexts" https://arxiv.org/abs/2307.03172
From my experience this holds true to this day. It was one of my core observations for similarity to the limitations of human working memory on "Engineering for Bounded Cognition"
Richard Hendricks solved this decisively with middle-out compression
I didn't get the reference, but it looks like im going to have to watch that series now :D
It holds up really well. I’m envious you get to watch it for the first tome, enjoy!
What inspired him to take this novel approach?