Surprisingly I agree with you. My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue.
A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.
I think the key question is more effective at what?
I see plenty of anecdotal evidence that models have been trained fantastically well—and getting better—at writing to trigger the right neurons in the human population to produce “This is interesting/informative/correct” responses in bulk.
Could their ability to produce those responses run far ahead of their ability to actually achieve the last in reality? Sure seems plausible, and then where are we?
Many people are bad at writing, including scientists. Improvements could be articulating a concept in a way that the reader will understand it clearly.
My hypothesis is that an LLM’s ability to concoct prose that will convince even an expert of the validity of an idea is largely independent of the LLM’s ability to validate the idea itself, or whether the idea is correct in the first place. Especially if it’s being prompted to do the former, not the latter.
My anecdotal evidence is the LLM-generated, inchoate technical dross that is routinely upvoted onto the hn front page. Much of it isn’t even coherent enough to be wrong, but the readership here finds it interesting!
I agree prose and validation are different. I dont know anyone who contests this.
> if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue.
That's a big "If".
If a research is good, the author still has to clear all the hurdles in publishing. "Writing your own paper" is just one more hurdle.
> A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.
That's just a different way of saying "if the majority of CS papers are crap, lets just accept this reality".
So, go on, publish away all your AI-induced "research", but the bar is slowly going to be raised anyway to reject that. That's how science always worked - when a bar is not sufficient to exclude the crap, it is raised.
I reviewed for ACL (big NLP conf) and another similar conference recently. Lots of crap (though there always has been). One superficially well-written paper was likely AI-plagiarized (i.e. the AI sampled an idea from prior work and rewrote it) and got desk rejected for it.
Reviewers were also totally unengaged. Of 20 reviews I read (from my reviewers or from reviewers on the same papers), maybe 2 were mediocre, and the rest were crap (though likely not AI).
The notion that science will somehow benefit from this is about as stupid an idea as you can have. Science relies on skepticism. AIs are not skeptical, and many folks are submitting papers because they stand to gain something, not because they are motivated to do good research or develop new understanding. Fields are being inundated with garbage that is maximally indistinguishable from real work (that's the training objective for LLMs). This in turn maximizes the cost of identifying bad work.
This is the same enshittification process that we see everywhere else. You get spam phone calls because there is no reason for a spammer not to call you. "Researchers" are submitting spam papers because there is no cost to doing so with some possible gain. Absent intervention, this eventually drives the community value of the network to zero (or potentially negative, if friction costs to switching are high).
And yet the quality of work at ACL is still higher than NeurIPS, and that's with ACL this year having over 40% of all posters looking nearly identical due to everyone claude coding their posters...
I mean, sure. But what an absolutely insane predicate. "Not reducing the quality of the output substantially" is (as an AI might say...) load bearing there.
Other problems include: Signal to Noise Ratio going through the roof.
Yeah I was careful on purpose with my statement. :D
Is it? I guess look -- I'm an academic, I've read piles of articles such as these. Signal to noise, already not great.
Yes, I feel like there's room to improve things, I just strongly doubt that "using AI to detect AI" is a particularly useful thing to do here.
> if it makes communicating research more effective, while not reducing the quality of the output substantially
That’s a big if. We all know that’s not what’s happening.
> My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue.
That's a big if. ArXiv is not peer reviewed and LLMs basically interpolate and extrapolate text, which makes them essentially fluff generators. Even in the most charitable interpretation, LLMs enable those with nothing to say to say nothing while meeting surface-level style guides.
We can't know if real science is happening in the background but I'd wager that the majority of these papers is not complete slop but real findings with AI generated text used to communicate it. If it was just straight slop I would be really worried.
> We can't know if real science is happening in the background but I'd wager that the majority of these papers is not complete slop but real findings with AI generated text used to communicate it. If it was just straight slop I would be really worried.
Why don't you read them and see? The ones I looked at were clear slop.
> (...) but I'd wager that the majority of these papers is not complete slop (...)
That's a huge assumption, and one that goes against the whole notion of using LLMs to generate text. AI slop is by far the norm.