It's not so clear that AI is an amplifier.. The paper has some fascinating analysis on this topic:
"At the other end of the distribution, AI students who spend more than 65 minutes on their homework receive homework and exam scores similar to those of non-AI students, suggesting that these students do not use generative AI for homework assignments. However, this group consists entirely of students who adopted generative AI no more than Öve months. Six months after adoption, no AI student spends more than 65 minutes completing their homework (see Figure A5). This is consistent with the gradual process of learning how to use AI tools. It also suggests that AI crowds out the highest level of e§ort."
"Interestingly, in the range of 50-65 minutes, the median and the interquartile range of exam scores of AI and non-AI students are similar. This implies that, in the range where AI students and non-AI students have overlapping homework times, students who spend the same amount of time completing homework on average receive similar exam scores."
"This pattern shows that students who spend the same amount of time on homework learn similarly, with or without generative AI. In other words, generative AI reduces time spent learning for the majority of AI students but not learning efficiency for those who spend the same time studying as the non-AI students."
Personally in my own work it's pretty obvious how it's an amplifier. You waste far less time trying to find an answer to something that confuses you.
You might say "it's good to learn research skills" and that's true to an extent, but tutors have always made people better students. And AI is a tutor you can message at any time, day or night, for free.
"AI is a tutor you can message at any time, day or night, for free"
which makes it not a tutor. if it is true that tutors have always made people better students, then those tutors are definitely imposing some limits on how many answers they give you and requiring you to do some thinking. I think you could have premised your same argument by, "copying a smart kid's answers has always made people better students...."
That's why you see the split in outcomes. The fraction of kids who value self improvement are going to get smarter, and the ones who just want to slide by will fall behind.
Nothing stopping kids from asking AI to give you responses like this, it is more than capable. There is always going to be the temptation to take short cuts though.
Yes, but is it the correct answer?
I wasted several days trying to have AI teach me containers. I would have been much better off just reading the docs.
Did you initially ask the tool to read the docs?
I wouldn’t be surprised if you did and still had the same issue, but if you relied on the training data recall alone, you definitely shortchanged yourself.
I’ve had decent results with tasking a model to run pre-research and summarization as a checklist, then have a second instance of the model(s) review and then develop my learning plan learn, versus the times I just asked Claude or ChatGPT to explain something to me “from memory.”
The knowledge AI has is vast. Communicating it to you is slow and narrow baud. The chat interface itself as a medium is lacking and will need to be replaced without another medium.
Finding an answer isn't always an objective, though. I agree it often has been too much in baseline schooling.
First, there's a certain amount of baseline knowledge we'd like students to possess. Without a certain prerequisite amount of underlying information committed to memory, it gets far more difficult to achieve fluency in a topic.
But more than that, there's other skills we're trying to build: frustration tolerance, processing contradictory information, disciplined problem solving. You only really develop these skills through productive struggle. If you find a way to shortcut the productive struggle, students truggle.
> but tutors have always made people better students.
Sure. Bloom showed us that students taught with a combination of tutorial and mastery methods, one-on-one, outperform students in a normal classroom by roughly 2 sigma.
The paradox has always been-- why hasn't technology unlocked these gains for students in normal classrooms? If we could boost everyone's performance by this amount, it would be huge for society-- but society can't afford to teach everyone with tutorial methods.
Since the 1970s, we've invested in edtech towards trying to make this happen, but most of it has actually had net-negative effects as best as we can measure. AI, so far, looks to be much worse.
I think part of the answer is that a big part of what makes a conventional classroom work are social pressures. So far, it looks like AI (and edtech in general) does more to dismantle conventional pedagogy and to break down the social fabric of the classroom, than it has improved differentiation or unlocked this tutorial effect more broadly.
> If you find a way to shortcut the productive struggle, students truggle.
I was assuming this was a typo and was thinking about making a joke about it, but it does appear to be slang that fits the context:
https://www.urbandictionary.com/define.php?term=Truggle
> 1. the standard of perpetual intellectual failure made by an individual.
> Truggle; the standard defenintion of a person who is a failureat everything.
Was it a typo or did you actually mean this?
the difference in the vast majority of cases is that people go to a tutor with an intentional stance of cultivating a particular kind of practice, not merely to get an answer to a question. It's not technically impossible to do this with a chatbot but far less likely given that, unlike an even half decent tutor, a chatbot will never cultivate that attitude in you. The elimination of that friction is exactly why people talk to bots.
The good thing about a tutor is that he or she isn't always available and knows they won't be in the future, so they instill good habits and independence in you.
> a chatbot will never cultivate that attitude in you
Hence the whole "AI is basically an amplifier of bad and good" argument made earlier. The ones who do this because they want to understand, don't need to cultivate this at all, it naturally happens with chatbots. They don't just ask for the answer to a question, but then dig into why it's like that and what not. But the ones that don't care, now have to do even less to get the fast answer without understanding.
[dead]
Right I was not clear enough about separating my opinion from the findings.
It would be good to see the effect of access to AI during preparation on those who previously achieved top 10% points in exams of similar topics before. I suspect they would benefit further.
My apologies for coming off as over-enthusiastic, I am currently obsessed with this study. Here is another quote:
"The negative learning effects are larger for students with higher initial achievement. The differences in the estimated full (6-10 month average) effects are substantial, with a 50% gap between the most negative effect (-24 percent) for the highest tercile and the least negative (-16 percent) for the lowest tercile. "
Not top 10% as you asked, but the closest to what you asked. My working hypothesis is that top performance is highly correlated with willingness to work hard, and AI decreases the motivation to work hard.
Interesting. Top academic performance is mostly correlated with conscientiousness (willingness to work hard and keeping track of things) and intelligence. And I'd add motivation and interest to that too.
I think if you take a physics class where the student is intelligent and intrinsically motivated through their own interest (I admit this is rare) then AI probably helps.
I think if they are intelligent, intrinsically motivated, and willing to pursue knowledge beyond what the class requires, it probably helps.
That last part is key. Intrinsic motivation doesn't mean you pursue it outside normal bounds. My kid loves soccer, its her second favorite thing in the world, she has an absurdly high tolerance for physical discomfort while playing, but she doesn't play it at home. There's other things she's rather do, such as play with her toys.
When you move the bar to something even less interesting to most kids like science, you're going to have a pretty huge falloff. You're basically selecting for kids who choose to do it in their spare time. I know a lot of smart kids (I run a boyscout troop, my wife a girlscout troop, both with lots of high achievers), and none of them do this.
well most university exams are designed to measure how much you study. so we didn't really need a study to tell us, "Exams continue to measure what they are designed to measure."
they're not designed to measure general aptitude, or function as admissions criteria, or screen for job applications, or any other numerous things they are used for.
there can be many questions of pedagogy. one of them is, what do our exams measure and how do we use them? professors who say, "My exam is designed to measure who studies, not be used for all these other purposes that they are actually used for" - I don't buy it. It's the same as late night comedians saying they are not responsible for solutions, even when spending 90% of their air time making political jokes.
THIS is the pedagogical issue, that pedagogy has NEVER caught up with the scope of responsibilities. This is acute in STEM - I mean, the humanities departments are generally pretty well run, all things considered, in this regard. Generative AI is accelerating that pre-existing crisis.
The fact that exam scores are correlated with how much you study is not the same as exams only reflect how much you study. Two students who study the same amount could have very different exam scores. The reason that there still is a strong correlation between exam score and time of study is because if all other things being equal, students who studies more have higher exam scores.
i'm not saying they reflect how much you study. they reflect a lot of things, including that. but you ask the people who write the tests, they're going to say, how much you know or how much you study, but nonetheless, they are limited. i agree with you. that's part of my point.
let's imagine a different study. we instead compare AI-users and non-users on a Wechsler (IQ-adjacent) test.
overall, it would be surprising if AI usage impacted your Wechsler scores. someone has done this study and the impact is quite quite small. BUT. do we care? We don't use Wechsler scores for admissions, we don't use them for jobs, we don't use them for... are you getting it now? A Wechsler family test is measuring something real, just like a university exam measures something. But what do we USE them for? Wechsler and a typical university exam are, in some senses, EQUALLY vague in terms of their fitness for purpose for answering a question like, "should we hire this guy?"
Like there is an association between IQ and earnings but it is actually surprisingly small! There is an association with math education and earnings and it is also surprisingly small. And consider how many people get by just fine without using a single piece of math education once they have finished school - like what if maximizing your earnings isn't all that it is about? Are you getting it now?
The issue isn't the AI usage. I can find tests that are immune to AI usage. The issue is using tests for things that they are not designed for. We pick and choose, for some subtle but nonetheless pervasive cultural reasons, which tests we use for which purpose, and very frequently, not because they are calibrated for the chosen purpose. This is coming from someone who scores very well on all these tests, and have kids, so I have a very strong incentive to buy into the status quo, and I'm telling you: academic testing has been fucked up for a long, long time.
Education is a complex topic. I don't claim to understand it, nor do I think we can sort this out in a hacker news thread. What I reject is the simplistic claims like "exams don't mean anything". There are certainly exams that are badly designed, but there are also exams that are well designed. A well designed exam can look very badly to different people, based on what they know and where they come from. It's a bad idea to think exam scores as the single metric of education quality, it is equally bad to reject them completely, because chances are any alternative measurement people come up with are going to be worse.
You're right that academic exams also measure something along the lines of instruction following / obedience / willingness to jump over hoops for no good reason etc. and that's often a good signal for most kinds of jobs.
I just started teaching undergrad CS courses after ~20 years of various non-academic jobs and it took me about two semesters to realize that almost nothing matters except how I assess students.
Two weeks of ADHD-fueled research later, I concluded that academia is actively resistant to implementing assessment reform because it would expose the utter pointlessness of most of what happens in university classrooms.
The reality is that we have no idea what most university exams measure because they are ad hoc, written by amateurs (yes, most professors are untrained in pedagogical methods) with zero psychometric validity analysis.
I never studied in university and yet I acheived good exam scores. If you understand a topic and ave a reasonable memory and ability to apply my our understanding not much studying is required in my experience. If you don't understand the big picture then you got to laboriously keep track of and manipulate a bunch of disparate pieces.
> well most university exams are designed to measure how much you study.
Huh? They're designed to measure how much you know. They can't see how much you study, nor would they have reason to be interested.
How much you know at the moment of the exam. Not how much you truly know.
at least in my experience in university - i didn't really ask this question, since it is obvious to me, but some students have asked it during lecture, or some instructors have volunteered the answer ahead of time - if you ask how to perform better on the exam, usually the instructors say, "here's what you should study." they never say, "know more." the thing i am talking about is consistent with the paper. really, your takeaway should be, exams can't see how much you know!
Because "know more" isn't actionable. Knowing more is achieved by studying but not necessarily more time spent studying, but well spent effort. Staring at the page for hours and saying "I don't understand" doesn't help. Solve exercise problems, explain the material to fellow students, discuss it with them, make mind maps, bullet point summaries, work through derivations step by step, etc. There are many techniques.
At the end of the day though what matters is what you know. Furthermore, if it's a serious subject, it shouldn't matter whether you learned it from this teacher or from another school and teacher, as long as your knowledge is correct. Knowing the idiosyncracies of this particular teacher should not factor into the grade. A serious subject can be learned on one continent and examined on another. Bullshit courses are all about learning pet peeves and hobby horses of a particular teacher.
You become a person who knows more, by studying more about the things you need to know.
I would usually say something along the lines of “everything we covered is on the table” or “everything we covered since the last exam is on the table” depending on the nature of the test. That’s the same message as “know more” but I think it sounds politer.
Self reports of AI use is worthless
Social stigma against just juices people to stick to socially desirable answers; no doubt a whole bunch who self reported as non-AI users actually used AI
Sounds like the opposite of the social stigma in tech companies, where most people are under pressure to self-report as AI users when they are actually just using normal methods to get the same result (and spending the time saved however they like) (nothing against genAI for making work take more time, it's still good for business as long as the extra work can be moved around)
Depends who's using it. Like many tools the force multiplier depends on the operator.
I'm confident it's an amplifier for people who know how learning works and already do a lot of it, successfully. However the level of "learning fluency" I'm talking about isn't reached for many until late college or grad school, and sometimes not at all. So I'm not surprised by the quoted results for 12-18 year olds.
> Depends who's using it. Like many tools the force multiplier depends on the operator.
How would you prove/disprove this assumption without falling into a True Scotsman fallacy?
It's a fair question. In theory you could do an experiment where the subjects were grad students or professors, or top performing college students. Give them some fixed amount of time to understand some new subject on which they'll be tested, only 1 group has access to an LLM with appropriate context for the learning task, etc. That's just one idea...
And people who want to learn. Most teenagers lack agency in their studies. They aren't in high school because they love it but because they have no choice.
Sometimes things are just common sense pure and simple.
> Sometimes things are just common sense pure and simple.
The discussion is about "AI", so common sense is out the window. These people's professional reputations depend on addict-level "AI" usage remaining socially acceptable.