I'm doing a PhD in sociology of work (I have a background in software development) and want to do such a study on software development with AI, but as you say it's quite hard to find a good setting, because it's hard to operationalise quality.
I have been thinking of measuring it along these lines perhaps... https://www.gitclear.com/the_ai_code_quality_maintainability...
But they have a closed sourced software so I would have to find or create an open one in that case.
Happy to hear if you have any ideas on study design!
What if you looked at scheduling? Like, an assistant or agent scheduling a calendar something?
It's atomic, and one can't usually add value beyond a low threshold. If we're seeing more words–or tokens–used over time for a task which one cannot reasonably argue has gotten more complicated in that time, it would be direct evidence of involution. If you can relate that to competition for VC dollars or customers, that takes you over to Neijuan.
The usage of Neijuan as I understand it isn't about a complete lack of results, Chinese test scores are really good, but there's only 10% in the top 10% so competing even harder just means everyone is miserable competing for the same number of spots. Same for production if it's driving prices down, you did produce more, but the end result is a race to the bottom.