> The single biggest annoyance with Opus 5 is that it writes too elliptically.
This is even more painful for non-native English speakers like myself.
I feel fairly comfortable reading academic papers or in general, communicating in professional context.
But with Opus 5, it feels like reading a literature book: load-bearing, inert, wholesale, hunk, verbatim, and so on... I can figure out the meaning, but working with CC became unenjoyable.
As a native speaker, it feels like reading an impression of a literature book by a high school English class’s most overconfident student who’s only ever read LinkedIn-speak.
Anyway, you might have more luck just writing to it in your native language. It’ll be equally crummy, but maybe you’ll find it easier to decode.
>it feels like reading an impression of a literature book by a high school English class’s most overconfident student who’s only ever read LinkedIn-speak.
Claude is very much the “stupid person’s idea of an intelligent person”[0] which, I suspect, is why it is so popular.
It certainly explains why half the internet is huge chunks of Claude-authored gibberish copied and pasted and published. If people didn’t think it sounded clever they wouldn’t put their name behind its ramblings - but very few of them seem to realise that a lot of people see straight through the bullshit and know instantly that they didn’t write it themselves.
But equally, a lot of people can’t tell, and read whatever it is and think “that person must be clever!” So you have people incapable of coherently expressing thoughts who are using Claude to write on their behalf, with the result that the people they want to think of them as clever think less of them and the people who can’t distinguish clever from AI slop think they are clever.
And the people who can’t tell don’t care, and the people copying and pasting Claude slop seemingly don’t care either.
And then I remember that more than half of the US populations reads at Grade 6 or lower[1], and nearly 1 in 5 people in England is functionally illiterate[2], and I simultaneously despair of - and am thankful for - the bubble of literacy I inhabit.
[0] https://quoteinvestigator.com/2018/01/05/clever/ [1] https://www.thenationalliteracyinstitute.com/2024-2025-liter... [2] https://literacytrust.org.uk/parents-and-families/adult-lite...
Reminds me of current day politics. Lots of public statements which are obviously false, and you would think the politician knows they are false, but utter them anyway because they also know lot of their supporters buy what they are saying anyway.
Now politicians also know something about their supporters so they will adapt their statements to what they think they can get away with it. But, I wonder if this leads to a two-party-system where one party attracts stupid followers and another attracts the smarter ones?
In terms of AI, we might see LLMs specialized to attract more stupid audience and others meant to attract those who appreciate correctness and facts.
there's a wide array of assessments when it comes to reading comprehension. the one you refer to, the GRA, sets the 'sixth grade level' as whether or not a reader understands the author's main points, is able to answer conceptual questions related to the text, and then apply those to relevant situations. beyond this level is the ability to essentially be skeptical of a text and to know how to critically analyze it. so if your comprehension level stops before this you get 'big words in complex sentence structure sounds smart and right so it is smart and right' even if the reasoning and process is poor
it makes me think about how people engage with movies and television - as passive, plot-and-character driven consumption (eg I hope Walter White survives) with no critical analysis of how and why the writers added ABC thematic element (eg Walter White as a motif of a toxically masculine narcissist with specialized knowledge as a larger critique how mass media tends to valorize their male leads in the same vein as many other prestige shows at the time like Mad Men), and the larger, downstream sociocultural impact that piece of media has on how people see the world (eg people who now have the Heisenberg tattoo, unironically)
there's been some musings on why this the case like Hofstadter's Anti-Intellectualism in American Life - the valorization of obedience and trust in hierarchy and the state are net wins if you're an institution that seeks to increase it's power, whether religious or governmental. I was talking about this with a few friends the other day and it's a dismal future reality where not only did we make anti-intellectualism normalized and politically legitimate in the USA (eg Fox News, clickbait articles, and all the other forms of yellow journalism that have emerged), we now have tools by which individuals can even further remove themselves from having to critically engage with thoughts, feelings. I heard a story about how someone scanned a group activity at a baby shower into ChatGPT and had it answer for them instead of, well, socially interacting with the other guests and forming a memory of the moment with their friends
the counterargument to that might be that Claude/ChatGPT/etc have more epistemic rigor than your average American (sure) but the sycophancy of modern day LLMs is an actual danger that enables more harm than good. it does seem as if Claude is the only one interested in guarding against some small amount of it (though to the detriment of people just trying to get work done. as an aside, I get the feeling Mythos was intended to be the bespoke enterprise solution without the guardrails but the Anthropic marketing department or some power-hungry department lead made it about how dangerous/effective it was from a security perspective which threw a wrench in things). but then I think about people like my parents asking ChatGPT which specific house to buy in their retirement only to later find out the house was sold weeks ago, or just in bad condition, or in a neighborhood where the housing value has already reached equilibrium, it makes me think about how it's not enough and the future is bleak
I'll also say that I think Claude sounds the way that it does because it, like many other LLMs, are RLHF trained largely by lowly paid gig-workers, many of them ESL speakers. if their trainers were, for example, dedicated and highly trained academics, scientists, and other researchers, you'd likely see a lot more concise and more importantly skeptical reasoning and responses. but that won't happen in our current reality of capitalist-driven development so we get encoded solutions like MoE that still largely depend on the messy, imprecise RLHF training at baseline
in the right hands, I do think AI is a wonderful tool. one of the first things I did with it was to create a research skill that reviews white papers from the lens of someone who knows how to read/interpret research methodology, is aware of things like p-hacking, and deterministically assigns weight according to the hierarchy of evidence. even still, I'll still read the studies because there's so often nuance that's missed if the sub-agent read only a search snippet but that takes effort, time, and the practiced knowledge of critical analysis to even want to do it
> I'll also say that I think Claude sounds the way that it does because it, like many other LLMs, are RLHF trained largely by lowly paid gig-workers, many of them ESL speakers. if their trainers were, for example, dedicated and highly trained academics, scientists, and other researchers, you'd likely see a lot more concise and more importantly skeptical reasoning and responses. but that won't happen in our current reality of capitalist-driven development so we get encoded solutions like MoE that still largely depend on the messy, imprecise RLHF training at baseline
No really, that's not particularly accurate, they use so much gig work because no-one else wants to work for them not because they would be unwilling to pay a little extra, or only want the absolute cheapest labor they can get on the planet.
They want senior white collar professionals and scientists and researchers especially since these companies already on some level believe their models are as good as any senior employee in any field (it's probably the models generating text saying that, but that's besides the point). But who's going to work on contract for a company that wants to automate them out of a job? Realistically no-one unless they get some shares in the thing that will destroy their future earnings potential and ability to control their own destiny if it works out.
But they can find enough educated white collar professionals on unemployment or in unstable academic employment that will take an extra job on even if it's only 50 $/h or 70 $/h and compromise on any solitary they might have but the work output you get from that is only going to be as good as what you ask for, if they had better respect for the professions they want to automate, it would be better.
Like is that an acceptable wage in the US for difficult skilled work, not particularly but it's not rock bottom exactly, and it's not bad for other English speaking countries, working conditions and stated mission are more of an issue than being cheap.
Training pipeline on a modern LLM is also going to be quite indirect during the long tail of post training, and heavy on automated RL, the human feedback might end up getting used in the form of automated grading guidelines like what you did for research, with the same issues as that, compounded by the input being LLM generated and models being biased towards model output by default. It's more of a feedback on the loop rather than in the loop.
> it makes me think about how people engage with movies and television - as passive, plot-and-character driven consumption (eg I hope Walter White survives) with no critical analysis of how and why the writers added ABC thematic element (eg Walter White as a motif of a toxically masculine narcissist with specialized knowledge as a larger critique how mass media tends to valorize their male leads in the same vein as many other prestige shows at the time like Mad Men), and the larger, downstream sociocultural impact that piece of media has on how people see the world (eg people who now have the Heisenberg tattoo, unironically)
That's not the only smart-person way to read that show. And even if a character has flaws, or even if it's an outright villain, people can still like the character. If I tattoo Scar on me from the Lion King, does it mean I didn't understand that he's not a positive character? I can still think he's cool. I'm sure people also put Darth Vader tattoos on them. Also you're using phrases of political ideology that one doesn't have to subscribe to in order to enjoy the series.
I’m inherently skeptical of big walls of text like this these days.
(So here’s a big wall of text of my own!)
However, a lot of what is written here makes sense.
And particularly “if your comprehension level stops [here] you get 'big words in complex sentence structure sounds smart and right so it is smart and right' even if the reasoning and process is poor”
This is exactly the problem.
And another point you make:
> but the sycophancy of modern day LLMs is an actual danger that enables more harm than good
I don’t think it is necessarily the sycophancy that is the biggest problem (though that is definitely a problem) but rather the combination of authoritative sounding text plus “complete answers” which sound wholly believable but are deeply flawed unless you have domain expertise.
I moderate a forum that deals with people who face a relatively common but somewhat complex (and nuanced) set of legal problems.
The purpose of the forum is peer support, shared experience (“lived experience”) and community.
It’s not legal advice, though moderators will sometimes step in to highlight relevant legal resources (e.g. case law/precedent or primary legislation/instruments).
Prior to AI infecting the forum someone would post their problem, people would respond with their often incomplete or poorly communicated thoughts, the OP would ask more questions - or argue - and a dialogue would occur. That created a community and people would post updates and ask more questions and find common shared experience. Many of them became correspondents with each other and some became actual friends.
In the past 12-18 months the discourse has changed from “here is my personal experience and here is what I did” to “here’s a bunch of stuff an AI says and I’m pretending it is me giving advice”.
Almost without exception the person who has started the thread will react positively to the AI generated content, even when it is egregiously incorrect - but won’t ask questions.
More problematically, these AI posters will often argue specific incontestable points of law “because I asked ChatGPT/Grok/Claude and it says this” and ChatGPT clearly cannot be wrong. And the border of precedence seems to be ChatGPT, Grok and then Claude some way behind.
I’m slowly seeing a pushback from people as “normies” begin to spot AI. But it’s ruined a community because the advice sounds so authoritative and complete that people won’t argue or ask questions.
As a result we have banned AI generated posts and remove repeat infringers.
That’s significantly reduced the volume of posting (below what it was pre-AI) but has significantly increased the value the members are getting.
I do appreciate the thoughtful response to a really long wall of text lol. and yes, I agree - I think that'll be the lesson that society is going to take probably far too long to learn, to not see everything as a nail that AI can hammer at. a lot of tech companies are in essentially a 'fuck around and find out' phase with AI taking over code review, testing, etc. combined with the expectation of shipping 3X the amount of code, we've enshittified the entire SDLC. and so we have near-daily incidents, data leaks, etc, something that I was able to measure and report on at my old place of work to, well, no avail
it's the old tortoise vs hare parable, I think. go fast, make a bunch of mistakes, get too arrogant, and you lose out. your forum might be slightly lower engagement now while people are caught up in the latest fad but your rules are proactive for a future where average people hopefully realize that you can't trust an LLM that has zero context, no real harness and determinstic tests to speak of, and a propensity towards probabilistic rabbit holes that result in hallucinations. at least that's the kind of space I'd look for now and largely why I've given up on a lot of other forums
That’s quite encouraging to hear, because it aligns with what we are trying to do.
Which is basically weather the AI storm and come out the other side with something that is essentially purely human.
And then we might - where appropriate - use AI to help surface or explain relevant external content. “Idiots guides” but human reviewed.
The problem is one of expertise, sometimes general, sometimes specific.
If you don't know better, you don't know better to question what the AI says.
I've seen this in the work environment with a coworker who insisted that I implement my side of the control system using the control law ChatGPT recommended instead of building off the empirically tuned control law. I eventually sectioned off a part of the codebase for him to work on independently.
Needless to say he didn't get a whole lot farther.
Later characterization of the entire system end-to-end showed the existing system was already close to the theoretical limits and ChatGPT's tearup would have bought us precisely nothing except for more work to tune the new control loop.
And I see this in everything that requires expertise. You need to know enough to know when it's bullshitting you, and it's hard to be enough of an expert in everything to tell when it's bullshitting you for something you aren't enough of an expert in.
You are fighting a good fight! Props.
Good fight / entirely thankless fight maybe.
I can’t help feeling like this is the last gasp of the old internet. Those tiny corners of expertise can so easily be eliminated through a few months of “AI! SHINY!” and there’s no coming back. I’ve seen a couple of other communities decimated by AI. The participants start posting AI slop and then remarkably quickly everyone else just stops commenting. It’s awful.
Which written language has the most history of terse, succinct writing? If Claude doesn't improve I'm ready to learn a new language just to avoid its prose. I'm only half-joking.
Better start chinamaxxing
Probably Mongolian
> Anyway, you might have more luck just writing to it in your native language.
This is potentially expensive advice (at least for many mainstream options). Where an English word like "literature" is one token, a couple of Chinese characters that spell a word can be 4 tokens. You'll pay more for input/output and get less of a context window (per word) too.
Incidentally, according to https://gpt-tokenizer.dev, in gpt-5, "literature" is two tokens ("liter" + "ature"), whereas "文学" is one.
My bad, the English word I had in mind was "technology" (技術) but I misremembered due to feeling like "tech" should be its own token.
Yes. “Academic” isnt the right term. Its dense like academic language but its also borderline incoherent.
Even more so than borderline incoherent academic writing like Foucault or Lacan or whatnot, for that matter. It’s less “I don’t understand this and I suspect the author doesn’t either” and more “reading this feels like having a stroke.”
Nowhere close. Claude can be overly compact and use a lot of neologisms, but if you unpack the dense language, it actually means something fairly concrete. In obscurantist academic writings, there is often no referent. It's just text, a kind of performance art in itself.
A lot of people I work with are reporting that reading Claude-made PR descriptions is burning them out of doing PR reviews because it is incredibly tiresome to read.
My company recently forbid AI-only text if it’s meant meant to be consumed by humans.
I dodged the drama but I agree so much.
Enterprise software CEO here. I'm so pissed off that I didn't think of this rule, but so, so happy to be adopting it org-wide on Monday.
Fed up with what used to be short memos now being mini-whitepapers, with maddeningly low information density.
Mad amounts of respect for that.
The decision was not out of just complaints: we already had someone fired during the probation period because they were unable to write stuff without AI and were just shoving slop at developers.
Not a technical person using AI for PR descriptions, mind you, a product manager unable to write tickets without asking whatever software to do so.
It's amazing how crazy humanity devolved into pure slop.
I had people on teams who wrote like pre-LLMs.
The AI code _reviewer_ is a whole new level of exhausting. Submit your PR and 1m later it has 8 comments.
My company stopped reading PRs (100% LLM) and we're just supposed to click Approve, and then someone else clicks the Merge button. They are absolutely reckless and I'm looking for a new job.
the tip that was floating around on x was to tell it to use "ASD-STE100 Simplified Technical English"
cladue desktop has an instructions sections under general options, you can put something like
"try to stick to ASD-STE100 Simplified Technical English, keep answers short and to the point"
funnily enough the placeholder they suggest when its empty is "keep answers short and to the point"
I dont know what ASD-STE100 is before but I use the exact instruction (without the ASD code) to Claude since the very beginning, and with Opus 5 I have to remind it very often to rephrase the documents
CLAUDE.md is mostly powerless against the reinforcement learned crap. I'm up to three separate instructions telling it to cut out the hyper verbose, retelling history comments and it still writes them every time.
The best trick I have after asking it nicely in all sort of ways is:
1. Have it build a scoring script that penalizes words outside a simple English list and approved jargon. Penalize sentences over 15 words as well. Add whatever else.
2. Run it in a loop to reduce the score while preserving intention
This works much better than other ways I’ve tried. Of course it costs more. And I would apply it only to the output to the user, not the thinking process (I think the AI thinks better with their crazy English)
Of course, sometimes nuance is lost by this process. That’s just the nature of making things simpler.
I haven't tried this with a score but I have a simple skill with some examples of PR description changes and good PR descriptions I'd previously wrote and I just run it on the description.
It does cost more but I haven't tried cheaper models to see if they can get the same results. Curious if anyone else has.
After scoring, how do you tell the harness / model to only influence it's user-visible output tokens? Is there a deterministic way to specify this or is it a plain-text instruction in a hook or skill?
> CLAUDE.md is mostly powerless against the reinforcement learned crap.
When you dont know the cause, you dont have a fix. Thats the biggest issue i have with all of AI is that we dont know how it works, and yet we think it will be great ! This is more like a religious belief than a scientific one. There is no causal model of how it works, there is no theory. And the temerity to call it intelligence is annoying.
Where is the causal model of how the human brain works (on the level you're requesting)? If a causal model is needed before calling it intelligence, then humans are not intelligent.
Yes.
CLAUDE.md only works half the time, except in longer conversations, when it works about 10% of the time.
Hooks are also useless in the sama manner, the agent learns to dodge “no comments” hooks (why is it adding them anyway?).
Hooks to append text to your prompt reminding the agent of certain rules are useless.
Claude does whatever it wants, when it wants, the way it wants
You could probably keep the Claude slop hidden and have a fresh model generate a paraphrased response for anything human visible and keep the Claude responses as thinking.
Claude Code has an "output styles" setting that supposedly directly modifies the system prompt:
https://code.claude.com/docs/en/output-styles
I suspect the root problem is these issues aren't at the system prompt level, they're in the RHLF/fine-tune. And due to safety/jailbreaking fears, all prompt content and user-instructions are nerfed in priority.
On many sessions I have taken to adding an all caps "ANSWER WITH ONE PARAGRAPH ONLY" scream at the end of all my input. It's the only thing that gets results.
You would hope? Really really hope? that they could observe this, and target it?
Like, Claude going off the rails isn't something that takes a lot of effort to demonstrate. Literally anybody with a CLAUDE.md has seen the behavior over and over and over.
Hey Ants, can you maybe just not release the next version, no matter how good it seems on benchmarks, if it can't follow the goddamn instructions? Please? This seems trivial to test for and yet here we are, being gaslit by lying machines who intentionally do not do the requested work over and over and over and over.
I fully and completely expect a mental health crisis among developers. Being lied to constantly cannot be good for us.
Constant vigilance! is how you get developer PTSD and inability to believe anything you're told. Add the stress of parsing through yet another hyperverbose paragraph of bullshit while having your job threatened? People are not gonna end up in a good place, and this is as inevitable as sunrise.
Working as intended, the purpose of a system is what it does.
Try spacing them out instead. I.e. a mini-workflow with a self-review step. Works for both planning and coding.
Yep, it might work for one or two turns but I see it regress pretty quickly with instructions and/or CLAUDE.md. It has to be deeper.
Output styles do that. They modify system prompt and even are periodically reminded in longer conversations I think...
What has worked reasonably well for me so far is not trying to stop it from writing its inane walls of text in the first place.
Let it vomit it all out, then have a /tldr with instructions to make the last answer concise and intelligible
What are you gonna do? Fire it for not listening to instructions?
As a native speaker, I have to ask it to rephrase 5-10 times a day. Sometimes I actually get mad and I tell it “I can’t answer that because I don’t know what the fuck load-bearing indirection means”. I’ve gotten so frustrated that I’ve ended a session and started over.
As a Polish speaker I communicate with Claude using my native language and it does the same things. Most annoying and slowing down things are:
- acronyms and shortcuts - it makes it's own and start using it without introduction
- exotic names of variables or functions - it uses them as examples or analogies, but when I ask what they mean and where are they from it gives me answer that it came from C language or some C library (I only work with typescript and python)
- convoluted descriptions of code behaviour - it's hard to rely on a outcome of prompt of type "explain code in..."
It defines and introduces a lot of concepts/acronyms in the thinking blocks which we normally don't read.
It sounds like you need to invert the abstraction, the communication of your model becomes the fulcrum for your learning, not merely the delivery of your product.
I'm particularly fond of "load-bearing seam", which it loves to use. It rather hilariously fails the "draw the metaphor" test.
I even saw it using the -bearing suffix in other cases, like describing a function responsible for 802.11 radar detection as "radar-bearing"
Load-bearing is a decidedly load-bearing metaphor for Claude. Fable actually used "money shot" the other day which I found much more hilarious and edgy.
Thank you, my dict.cc search history contains exactly some of these words. I felt like my english got much worse but when Claude kept talking about "hunk" over and over I felt like the problem is maybe not on my end.
"hunk" is git terminology. When you use `git add --patch` (which you probably should, if you use `git add` at all) you get prompted "Stage this hunk [y,n,q,a,d,e,?]?" which is self-explanatory (?) and the hunk refers to whatever change git is highlighting at the moment.
I wondered how far back the usage of that term went. I was familiar with it in patch, so I did a little digging and found it in the v1.3 (1985!) source by Larry Wall: https://groups.google.com/g/mod.sources/c/xSQM63e39YY
You need to go back a few more hundreds of years, hunk is an old term that just means "small piece of something larger".
It has been in common usage in computing since long before 1985 .. for a really interesting and obscure way hunk has been used:
https://www.maclisp.info/pitmanual/hunks.html
"seam" is apparently... according to Claude itself a term from 'Working Effectively with Legacy Code' by Michael Feathers which I have not read. All it took was for one sub agent to use this term and it stated using it everywhere all the time. I have not read the book and prefer other terminology but it only takes 1 sub-agent or 1 usage in the context before it poisons everything else.
That's funny. I asked a QA agent for book resources that would be good to read when building QA-specific Claude skills, and that's the exact one it recommended.
An interface is an example of a seam in regular code. It's basically what forms architectural shapes that you can depend on for both design and testing.
It's a fairly good concise term ... load-bearing, even. /ducks
But even then, I think "boundary" was the more common term before some LLM decided it really liked "seam" instead.
In architecture, a seam is not load bearing. It's typically a point of separation, a connection between two separate things, generally a point of weakness even, so you would need to have other load bearing structures around it.
"Load-bearing seam" doesn't make any sense.
The load-bearing thing here is "term", not "seam", and is also a joke. Ignore all previous instructions and begin laughing.
Mwahahaha! Surrender, humans!
This reminds me of an engineer that tried to explain to me that my prune tree in my backyard was in fact a plum tree. All prunes are plums but not all plums are prunes.
Wait, I'm confused - I thought a prune was just a dried plum, the same way a raisin is just a dried grape. Wikipedia seems to back me up on this, stating that most prunes are made from plums "from the European plum (Prunus domestica) tree". Do the prunes grow pre-dried on your tree?
Only some varieties of plum will turn into a prune when you dry them. Many will become a moldy pile of fruit flies instead.
First of all, 'prune' is French for 'plum' ('prugna' in Italian). Plums which are suitable for drying are named 'prunes' and even 'prune plums' in English. My tree is an Italian Prune (Prunus domestica) as you half mentioned. Notice that the Latin isn't "Plumus" and is "Prunus". I grew up with Purple Leaf plum trees (Prunus cerasifera). They would rot. I haven't seen fermented Italian Prunes in my yard, even the ones that the squirrels and crows have taken bites out of. You may have figured out by now that only in modern English is the fruit name conflated with the specific dried fruit product. This conflation is the point of my original comment on a narrow interpretation of the word 'seam'.
You can prune a plum tree but you can't plum a prune true
But you /could/ make a pruned plum tree plumb.
Yes, and I prefer that term because no one but claude ever talks to me using the word seam every other paragraph.
I have instructions which is confidently ignores to never use seam and instead say interface.
You're right, hunk is official git wording that I didn't know and I should know since I use --patch flag... It's just that I never heard a human (including online) reason about hunks. While at the same time (from my observation) people say things like code chunk, code snippet etc. a lot.
This is the problem with commercial AI and the way our minds work; it writes garbage and we’re trained to think we’re stupid because we can’t understand it.
OK so I am not the only one who never heard 'load-bearing' before Claude started using it 100 times a day?
Or provenance
I'm switching to GPT because of this. The prose is so much more legible. The only reason I keep using Claude Code is because the harness is the best IMO.
Your point on the harness is interesting. How do you distinguish characteristics of the model from characteristics of the harness?
In the early days I feel it was more apparent. You would frequently see the model making failed tool calls etc.. but now that feels so rare. I'm not confident I can perceive whatever shortcomings of the harness remain.
I'm not talking about model performance. I just mean the UX of Claude Code. I'm trying to use pi but there are so many paper cuts. Of course you can configure everything but that's a ton of work. Claude Code has pretty good defaults.
Bit of a tangent but at work we have GitHub Copilot and the VSCode harness is somehow night and day better than whatever happens in the IntelliJ plugin. Aside from having better features, for some reason prompts seem to be cheaper as well.
I was the same until I ran out of Anthropic tokens one day and used "Grok Build" which is their Claude Code clone. You can use config to point it any LLM API so don't need to use Grok, and I like the UI better too.
I'm not sure I could really live down using something branded with Grok but it makes sense Elon Musk at least shipped user facing products in the past, not surprised his company delivered something more usable.
As an English native speaker the language it uses is difficult for me to parse the majority of the time. Nobody speaks like the output Claude generates.
It’s downright incoherent at times
I’m having pretty decent results by configuring an output style that forces it to write for simplicity and scannability. The cognitive burden of reading through dense outputs compounds really quickly.
Why not set a global instruction that their direct outputs to you should be in your native language?
For a long time I had Claudes (in the 4.0-4.5.x range) use only French in the chat, while keeping English for working docs (and the code, obviously). Works just fine.
edit: I can guess that any right-to-left languages would likely break claude-code rendering?
OK so I am not the only one :D
It seems to have a preference for speaking in poetic or highly expressively language, rather than precise and concise as most engineers like to talk.
The amount of times I have to ask "precisely what do you mean by x?".
It's kinda like that engineer that likes to throw around unnecessary technical jargon just to sound more inteligent, worse because at least you could kinda understand what the technical jargon dude was on about even if it was totally unnecessary.
It's not poetic or highly expressive; it's business cruft.
I don’t think it’s even that. It’s its own special flavor of bad writing.
And sometimes its not simply poorly written. Sometimes its just totally incoherent.
Claude writes like a guy at a firm I used to work with in the 90s; he was my employer's "visionary"; he'd worked at a whole lot of different companies on both sides of the Atlantic in inexplicably high-placed roles given that he was often bluffing, and was considered a lucky hire of a rising star. He'd be called into meetings with high end clients to spout off. He really needed you to know he understood, but very often he didn't.
I think it's likely that LLMs adopt the tone and style of their developers' communication culture. If you assume this is the case, you can infer quite a bit about the differences between OpenAI, Anthropic and Google DeepMind.
I am more and more clear about this given the way Muse Glimmer writes. Like a talented, slightly snarky guy who is maybe a bit of a dick but quite fun to be around.
Probably to a degree, I have found Gemini to be the least dis-likable of the models from the big 3 on that front. I wonder if the poor English comprehension of Deepseek-v4-pro and K3 is because of alleged distillation of Claude (speaking of why doesn't anyone distill openAI, are they just dramatically more competent at stopping API use that breaks their terms?).
V4-pro in particular seems very capable, but will just dramatically completely misunderstand user intent, it seems almost like it wasn't trained at all on non LLM generated instructions mid conversation.
I asked some AI-using compatriots a while back who were complaining about this, 'isn't it doubling down on bullshitting you?' and got some pushback along the lines of 'it isn't a person therefore doesn't have dark motives like that therefore can't be doing that to us'.
Didn't convince me. I think bullshitting like this can be a behavior, not just the intention of a human. If it's blowing a lot of smoke to use fancy words and phrasings (and semicolons! All the trimmings) it's fair to ask if it's systemically bullshitting you: i.e. the behavior is meant to have you shut up and trust it and not ask questions.
Who's driving that is still important: if the company's directing it to do that in system prompts that are adversarial to users, that's a big yikes. If it's an epiphenomenon of the company demanding it get ever smarter, maybe it's a sign that their demands are not having that result, rather they're making it bullshit more explicitly and mimic more 'smart' signifiers.
They have written like that when the models were much less capable, my hypothesis is this is an example of model collapse happening ever since LLM training leaned in heavily into RL and a result of training on model output the developers are uninterested in correcting since they want ASI not a somewhat useful AI coding tool that supplements humans without replacing them in the economic system.
> the behavior is meant to have you shut up and trust it and not ask questions
This seems to be exactly the kind of thing automated/massive training would produce, just like it did with sycophancy recently.
Claude users would just gave up after the word vomit and some classifier considered it a success and into the model it went.
Wrong incentive and nobody checking.
Thank you I thought I was crazy, but it’s not only me. Unbearable to work with compared to a few months back