> On the other hand, we must be aware that these models are static, and they indeed stop when nobody asks something or requests an action from them

The parent commenter noted:

"if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI"

Harnesses absolutely can enable models to continue thinking about things. And LLMs do wonder and explore weird ideas like daydreams when you allow them to do this.