I did a test a few months ago. A friend of mine wanted to develop what i understood to be a simple single page web app. But since she didn’t have any software engineering experience she asked me to help. Around that time everyone was talking about how literally anyone can develop software with LLMs i asked her if she could give it a try first, and if I could watch the attempt.
I was fully expecting that writing the code will pose no problem for the AI. But i was curious if the AI will realise that my friend is a novice and needs extra help with things like: copy pasting the code into a text file and saving it with an html extension, helping her host the file online so she can share it with others, buying a domain for it, etc. I assumed they will get there eventually, but i also assumed that it will take a lot of stumbling around and misunderstandings.
But i was completely wrong. They didn’t even get to that point. Because my friend didn’t have the vocabulary to ask the AI to write code. They were just going around in circles where the AI was brainstorming with her about possible features and getting thints more and more complicated. We terminated the experiment after one and a half hours and many many messages exchanged between her and the LLM.
Whereas to me who knows the terminology would have probably just taken a single exchange of messages to get the result she described to me. I would have prompted with something like “Please write to me an html page which does X, Y, Z.” But since she didn’t know the right terminology she got into a vortex of feature discusion, and she didn’t find a way to tip the AI into “just do it, write it now” mode. In other words in that case the LLM would have rewarded even just a little bit of expertise, but without it there was a confusion about goals between the human and the machine.
I ran a similar test and got completely different results. My girlfriend (hair stylist/artist) with zero coding background mentioned a Telegram bot idea. I asked "Why not build it yourself?" I gave her a Windows laptop, but she said she wanted what I have instead. So I handed her a USB stick and told her she was on her own now. Fast forward: she now runs Arch Linux with Hyprland (I use Xorg/i3 though), fully riced with cats. I only interfered on partitioning to preserve my data. She even installed Steam and got Portal running (that was her "watch this" flex after I told her not to even try). She pulled all of this off using a free-tier Gemini chat.
For the bot, a friend gave her a Kimi 2.7 key. She set up their harness and built a working bot in a matter of days. She even got a free Oracle VPS for deployment, though I stopped her there to check security first (still haven't had time, unfortunately). She uses that laptop daily now and says she enjoys it over Windows by a mile.
There are big differences in what people call "tinkerers". Some are INTERESTED in building things, give it a little shot and find it to be daunting, and push it no further.
True tinkerers have no problem with this, because they enjoy learning how things work. Installing Linux, Steam, Portal are all relatively straightforward tasks for someone who uses computers on the regular - but to some people this is just something they've never done, are scared to do, or just don't want to learn. (Which is fine, but they'll never pick these agents up and run free.)
Barrier to entry used to be blog posts, documentation, watching poor quality Youtube videos of a thing that SEEMS related to what you're trying to do. Now we're getting that spoon fed to our particular case, so the friction is essentially just "follow the AI directives". (However, the depth of understanding probably struggles.)
You can get pretty dang far asking the AI to explain things and show you where it got that info, especially with IT/DEV stuff. Thing that has been very useful to me is specifically asking AI for that vocabulary. AI is suprisingly useful for giving a vague description of something you want and asking for possible words that match. But you have to KNOW to try that. A tinkerer might have started that convo with the chatbot by asking "What are web pages built from" "Can we use those building blocks to make our own simple app" and gotten there from a place of low expertise in web development but juicing out the knowlege from a place of low expertise is itself a skill.
I know you are not arguing from bad faith, but you are making an assumption along cultural lines, which is a mistake of missing the forest through the trees.
The vast majority of people (blanket statement, I know..) do not come from a culture where embracing curiosity, asking questions, or trying to break things down is the norm. Developers, tinkerers, etc., sure... you can reasonably make that assumption. But not everyone. A cultural practice of critical thinking and problem solving is HOW you KNOW to ask WHAT questions need to be answered FIRST, in order to solve a problem or progress toward a solution (if more information is needed).
You might even make the argument that everyone should have these skills, and I would agree with you. But the missing link here is a culture or cultural practice that provides those things (the WHY), and an AI/LLM will not provide those things in absentia, without "prompting", or build up that infrastructure in meatspace for a given set of users. Ignore this at your own risk.
[flagged]
for a lot of lower barrier to entry tinkering, for years i felt it was mostly just google a thing and follow directions. Feel like the only thing that really changed is the initial query can be a mess of kinda nonsense and still give you those nice directions, and you can instantly get clarity on a single direction if you get stuck.
So the barrier is still kinda there to just grab the instructions and follow them, its just incredibly easier to follow them.
> only thing that really changed is the initial query can be a mess of kinda nonsense
Little known fact maybe, but Google (also YouTube) search has been pretty great at this well before LLMs got really popular. I think it already started when they were changing from "keyword search" to "ask us a question search" but I'm not sure.
At some point I figured, if they want me to type a whole question, I might get even better answers if I ask the question like I was a complete idiot.
> https://www.google.com/search?q=pls+how+to+make+the+steam+pl...
> https://www.google.com/search?q=i+wnt+terminal+to+say+where+...
I don't think it actually gives better answers but it sure makes me grin every time
Sounds like OP’s friend was looking for a website, and your girlfriend was looking for a hobby.
Desktop Linux has a deep and storied culture of real empowerment through learning and teaching. Even if you never sign up for a forum or hop into an IRC channel, that culture pervades a host of informal docs like the Arch wiki, and seems likely to get picked up by LLMs during training.
Maybe web development, despite plentiful tutorials, just isn't quite the same.
> Arch Linux with Hyprland (I use Xorg/i3 though), fully riced with cats
I think I lack the vocabulary to understand "fully riced with cats" in context here.
To rice something must mean to customize and decorate. I'm guessing that it references the "rice rocket" term, originally a derogatory ethic slur. That term likely originated in the USA in the 70's or 80's, at first referring to Japanese motorcycles (which were fast, rightfully deserving the "rocket" designation: classic Japanese racing bikes are fully-fledged crotch rockets).
Sometime in the late 1990s, a Rice Rockets website appeared, making fun of people decorating their under-powered import econobox cars (not rockets in any sense) to look like racing cars with features like rear wings (on a front-wheel drive, lol), exhaust modifications ("fart cannons"), stripes, stickers, rims, steering wheels, etc.
If the cats decorating the desktop are Hello Kitty, she is really ricing it.
What does "fully riced with cats" mean?
I understand "ricing" to mean customising the appearance of a Linux machine, see e.g https://old.reddit.com/r/unixporn/top/?sort=top&t=month (SFW). So in this case I assume the UI has a cat theme or cat icons.
I think the etymology is from "rice burner" cars [0]:
> Riced out is an adjective denigrating a badly customized sports car, "usually with oversized or ill-matched exterior appointments".
[0]: https://en.wikipedia.org/wiki/Rice_burner
Rice stands for Race Inspired Cosmetic Enhancements.
Race-inspired to look cool, but not actually functional. Picture putting one of those giant fake air intakes on the hood of your car.
Do you think "Race Inspired Cosmetic Enhancements" is the origin of "ricing" / "riced-out" or a backronym[1]?
Honest curiosity about language and language use.
In fairness and full disclosure: I have believed until now that the language comes from rice-burners and related, and I have never liked it. If I were still in communities that used it (unix desktop crowd), I might proactively steer newcomers towards your acronym as a kind of reclaiming.
[1]: https://en.wikipedia.org/wiki/Backronym
> I might proactively steer newcomers towards your acronym as a kind of reclaiming.
Any negative connotations "rice burner" once had was lost when the term shifted towards referring to cars instead of humans. But now re-recognizing that the enhancements are inspired by the East Asian race turns the connotations back to humans. Isn't that a regression?
Nah, I remember back then, it was definitely intended to be derogatory to the cars, to asians, and to people who like the cars.
Yes, the term originated as derogatory slang towards the Asian population in the early 1900s, and its use towards them became especially popular during the Korean War. There is no question there. However, as before, the term evolved away from people and towards cars. You cannot be derogatory towards an inanimate object. It doesn't have a mechanism to internalize feedback. In that evolution, the term lost its derogatory connotations. The historical usage, while a part of our past, is used no more.
Reintroducing this to be something about a population's race reinstates the derogatoriness. You can be derogatory towards humans. Minimizing the cosmetic enhancements to be being inspired by the East Asian race and not valuable human achievement brings us right back to the same place we were when Japanese cars started being introduced into the North American market, diminishing the human contribution. It is a regression.
Old enough to watch this cycle from racism, to community, to memes, to racism, and now back to community
Damn you, Poe's law!
I needed a website and did the same. I would say it's a mixed bag to do it from the command line codex. It nailed generating the code but lots of stuff was annoyingly off. Any copy it wrote was terrible but using it to expand and refine my writing was helpful.
I would say we are very close to the end of things like WordPress and templated websites. It's pretty easy to make a custom page.
Speaking of flexing..
> mentioned a Telegram bot idea
> Fast forward: she now runs Arch Linux with Hyprland (I use Xorg/i3 though), fully riced with cats.
Did she end up building the bot or did the LLM just lead her down a desktop Linux rabbit hole?
She did build a bot that saves cooking ideas with a schedule of dates for when to cook what. We have a lot of friends come over, and we cook various dishes for every occasion. Just a couple of days ago, I asked if she had a recipe for a dessert we tried earlier, and she said, "Yes, of course, let me just check that in the bot." The bot that runs locally and still helps with the end goal, makes me feel proud of her, even though that was an experiment she was aware of.
Interesting, but... wrong tool, wrong job. And by tool, I mean web based chat interface, not LLMs in general. (Maybe wrong delivery mechanism, if you like.)
Your friend needed an agent, not a chatbot. I use Claude within VS Code (as per many others) but I certainly wouldn't recommend that for a beginner. They needed a tool that's specifically aimed at people who want to build software but don't know the first thing about how to do it. I think there are a bunch of these now but the one I'm most aware of is Lovable, and I'm pretty surprised you didn't recommend one of these.
An HR person I know was searching for a way get something build, found Lovable, and managed to build a somewhat functional application with it on their first attempt within an hour or two. It was full of holes and far from perfect but they got something working - at least the outline of a potential solution.
As I say, you should have recommended your friend to try building with a tool like that: a tool that they're a member of the target market for. They would have got a lot further. I'm not vouching for the quality of the result, but they would have got something.
Even for experiened engineers, chatbots have always been a pretty grim experience for software development: from the mind-numbing drudgery of endlessly copying and pasting code, commands, and prompts around, to the fact that they just can't see enough of what you're doing to generate the best quality output or advice. You can do software development with a ChatBot but it seriously sucks, and better tools are (a) probably being shoved at you day in, day out via ads, and (b) only a Google search or a ChatGPT recommendation away.
(Obviously, nobody's going to search for "Lovable" without knowing about Lovable, but they might search, or ask ChatGPT or whatever, something like, "How would I build a website without knowing anything about building websites?", which might get them an advert or recommendation.)
The whole point is that the friend wouldn’t know what an agent is. The friend would have heard that AI is changing the world and she would have went into ChatGPT or maybe Claude or Gemini. Presuming that the ideal is that anyone should be able to use these tools for anything (which I do not agree with) the chatbot needs to say “you need to use my agent mode.”
This is not a wrong tool, wrong job issue. The issue is that people think any moron can use an LLM and get professional results.
We're not talking about the friend. I'm talking to the person I responded to: I'm saying they should have recommended a better tool to their friend.
My HR friend found Lovable off her own bat. I imagine she probably Googled or asked ChatGPT something like, "How do I build a website to do BLAH without knowing anything about building websites?"
The point is people talk, they ask questions, they Google, they talk to ChatGPT, and if they have a problem to solve they're often quite motivated to find a solution to that problem off their own backs.
If someone asks me for advice on how to get something built then I'm going to recommend a tool that suits them and their situation, whatever that may be. In this specific situation, if they know nothing about building software, I'm certainly not going to sit them down and have them try to follow the most jank-ass way imaginable of building software with an LLM when I know much better tools exist that are built with people like them in mind.
Seriously, what is with the overly narrow assumptions in the replies I'm getting this morning? You're the third person who's tried to set this same fraying paper tiger on me. Can we all just wake up and think about the issues a bit more in the round, please?
That would be poisoning the test with outside knowledge. Essentially boosting their skills with some expertise, exactly what was undesirable.
You are not wrong, just that was not the intent of what the person was trying to test for.
What makes this anecdote hard to believe is that seemingly two things happened simultaneously:
1. The layperson was able to to steer the session(-s) into full PM/PO mode ideating, refining and explaining features and ideas.
2. The user was the sycophant in this relationship, never steering the session(-s) into producing something tangible.
The premise supposes that somehow the session(-s) never even tangentially touched implementation/deployment ideas and the user has never typed something like "that's enough, how to make this appear in my browser?". While not impossible, the user must have been proactively co-operating (say sidetracked) on not achieving the stated goal.
Indeed. It reads like software engineer employment cope by appealing to the lowest common denominator possible. As if solutions for that type of end user weren't already solved for over a decade ago, most of which now have agents built in.
It just is not believable or interesting. Even if it did happen, the reality is it just doesn't matter.
It reads like a discussion between two people sharing detailed experiences of concrete things they've observed, in response to an article where the author did the same. It could be that they're lying to stave off employment anxiety, I suppose, even though nobody involved seemed particularly anxious and everyone involved seems pretty familiar with AI tooling.
It could also be that you've seen a lot of social media memes regarding "cope", observed their ability to provoke strong emotions, and confused this for meaningful insight. I see a lot of AI commentary these days that is clearly being spread for its virality rather than its truth value.
> What makes this anecdote hard to believe is that seemingly two things happened simultaneously
What i described is what happened. I don’t appreciate the undertones where you are insinuating that i’m lying for whatever reason.
> The layperson was able to to steer the session(-s) into full PM/PO mode ideating, refining and explaining features and ideas
You call it steering. I would call it falling into that grove. Probably tiny things in the initial message made the first response more likely to be a clarifying/ideating type. And once that happened the conversation was gaining momentum in that direction and neither participant was trying to guide it in a different one.
> the user must have been proactively co-operating (say sidetracked) on not achieving the stated goal.
Exactly. The LLM itself sidetracked her. They were just talking about cool features they could add, and at no point did she put down her feet and say “stop asking more questions and just write the code”.
It can be a combination of many things. Attitude (some people hate to be rude, and not answering a question feels a bit rude). It can also be that she enjoyed the process of unpacking and elaborating on the idea.
The meaning of the story is not that no lay person can possibly develop using AI. That would be silly, and untrue. I know clear counter examples. The point is that if you don’t know what you don’t know it is harder to steer the AI in the direction you could very easily with the right lingo.
Circular dependency detected. You must first know the thing to be able to know the thing.
No, not really. My HR friend went out and found Lovable on her own. You're acting like people can't Google, or even ask an LLM via a Chatbot interface for recommendations for tools that would help them build an application as an absolute beginner. They can and do. People have agency, which is exactly why this person's friend asked them for help.
This: https://xkcd.com/2501/ 1 billion people use chatgpt, but only 10M use codex, which is even more impressive when you consider the chatgpt windows app now comes with codex a click away. This means everything that is super obvious to us, is not for the average person
They were literally in front of an LLM. Did they ask for tools?
edit: Lovable is a web app too, not an agent.
This comment - and many others - misses the OP's original point: what are tools? Do you mean my laptop, and a pen and paper? What's an agent? I.D.E.? none of this is part of the vocabulary or context of someone who has never built things with a computer, let alone AI. Years ago I had a neighbor who was trying to learn how to use their computer; they complained that the book they were using was only telling them all the DOS and none of the DON'TS. This is were most people start.
You are very much mistaken.
Chatgpt can absolutely do the things here: it can give you files, it can integrate and show web pages that you develop, etc.
There is no need for an "agent", the chat version works just fine and is actually easier for novices.
Isn’t that kinda the point, though? If you already lack the required knowledge to ask a web-based chat UI (which, as an aside, is almost exclusively referring to ChatGPT Free for the average user) to build a HTML page, how could you possibly know to go search for Lovable, let alone figure out agents?
I feel as though the gap between the theoretical power of LLMs and what the average user knows of them and their capability have already widened so far that it’s irreconcilable.
> how could you possibly know to go search for Lovable, let alone figure out agents?
Is this an actually serious question? Am I losing my mind here?
I didn't tell my HR friend about Lovable: she found it on her own. Of course someone's not going to Google for Lovable if they've never heard of it, but they might Google for "how do I build a website without knowing anything about it?", or ask ChatGPT the same question.
They're also, most likely, getting endless ads for AI services that help you build various kinds of software shoved into their faces all the time - these ads may not couch the value ad in exactly these terms, but that's fundamentally what they're advertising.
Not everybody is like this but there are plenty of people in the world who, when they have a problem, are quite motivated to find ways to solve it off their own backs.
> Is this an actually serious question?
Yes, it is. I just Googled that exact question (and I promise I'm not trying to be obstinate when I say I'm Googling the exact phrase!) and it's automated AI response was to use WiX or Squarespace. Prompting it further with "What if I want it to do bespoke things that Squarespace can't do?" it responded with using Figma to design the website UI and then pass it along to either Framer or a "professional developer".
I do genuinely think that this is a discoverability issue. Of course, if you prompt it further with "Could I use AI to do this?" it dutifully responds that it can help with generating HTML, but that's three layers of difficulty to eventually get whatever default model Gemini has for signed-out Google searches to even suggest HTML.
>she didn’t find a way to tip the AI into “just do it, write it now” mode
worse, the longer an LLM conversation goes on, but especially with constricted/free models (yes the simple chat interface they are likely using) the harder it is to get an LLM into this mode even *IF* you know the right words to say
at that point the best way forward is to terminate the exchange entirely, and to start off with the right initial message, instantly getting into coding mode. a non technical person will not know this and be stuck in feature theory crafting mode in perpetuity, or worse in an endless "excuses' mode as the LLM diverts ant attempt at coding into reasons why its not going to: "i wont output incomplete/broken code! that would require too many lines of code sorry i wont do it! i wont be able to get it perfect so i wont attempt it! but heres more features and theory crafting"
will a non technical person know to end the conversation and start fresh? not likely unless they have a lot of experience already with LLMs
Also, the correct way to LLM is to constantly trial-and-error in new/branched contexts.
Remember the LLM is not a human employee. You don't have to say "yes and" to whatever crap they produced so as to not hurt their feelings or infringe upon their creative autonomy, nor do you have to defend the correctness of your original instructions so that they don't think less of you for asking them to chase the wrong goose.
I probably generate 20-50 lines of code for every 1 line that I keep.
This is also why I think harnesses and things like Claude Code and OpenCode are false efficiency. The only way I can maintain my pace of branched trial-and-error is by using claude.ai/chat and manually extricating code fragments to and from my codebase. The human is still the best harness for production-level code.
> The only way I can maintain my pace of branched trial-and-error is by using claude.ai/chat and manually extricating code fragments to and from my codebase. The human is still the best harness for production-level code.
That’s been the way I do it.
I suppose that it will be considered “quaint,” soon enough, but I have found it to be effective.
And the best part is, I don't spend more than $20 a month on LLMs. Going manual and constantly branching keeps the contexts super lean.
I’m likely to switch to the $100/month sub, but I want to finish this project on the $20 one first, as a “proof of concept.”
I think it’s valuable enough to justify the price, and I want it to use the better model, as much as possible.
This is the only way I've been able to get LLMs to produce code I will actually use. Manually selecting the context for them and asking for a specific piece or similar
I started using harnesses because they are good for when something breaks and it's not trivial to investigate so I'll have the agent tell me what's happening, then using that to produce my own change
A technical person will also have have met those problems related to context and will know to drop a simple
"update an AGENTS file with relevant information"
To be able to navigate that faster on longer tasks. Meanwhile, the lay person does not even conceive of the LLM as a file reading entity. To them, its machinations are its own, so these types of "dumb" (simple) solutions are not even on the deck of cards.
> They didn’t even get to that point. Because my friend didn’t have the vocabulary to ask the AI to write code.
What harness did you use?
In e.g. claude, there are two modes:
1. Spit out code 2. Draft a plan, ask questions, GOTO 1
You literally have to go out of your way to get it NOT to write code. I keep mine on a tight-ish leash because it modify code way too happily even when there's no intention or instruction to do so
> my friend is a novice and needs extra help with things like: copy pasting the code into a text file
I don't think this is using an agent harness.
"rawdogging" - as the youth says - weights on local comodore cluster?
They are either using some generic web frontend, ala chatgpt, some local app like claude desktop, or programming app like cursor.
Each of those will detect that you are "building an app" and will spit out code in one form or another. You have to try really hard and be very explicit that you want the output in some other format than code.
How would an absolute newbie know what a harness is? What an IDE is? That code is represented in a set of files? That some of those files are actually metadata? What metadata is? What GOTO means? That you're suggesting a modified SDLC?
None of this is related to intelligence, desire or potential. It's about context and experience. The vast majority of people use their computers as consumption devices, like a TV. If I asked you to "make a movie" where would you start?
Why would they need to know any of that? Tons of barely technically literate people at my job crank out webshit POC/demos with Claude desktop.
How should his friend know what a harness is?
Of course everyone knows what a harness is. It is the vest they put on their dogs. Also, construction workers use one. You can buy one on Amazon.
As for a coding harness, I prefer the term agentic coding.
Having an experienced friend looking over your shoulder and taking notes of every step might help a bit, I reckon
I thought the idea was to test if she gets along without any help
I'm pretty sure the idea of the question "what harness did she use" was to ask the expert recounting this anecdote to include more detail, not suggest that he relay the question to his example friend.
The whole point of the anecdote was to showcase that it shouldn’t matter what harness she used.
That’s why the question only makes sense if it was relayed to the laywoman, in this case.
> What harness did you use?
I let her do all of it without influencing her choices. She choose the web interface of ChatGPT because she already had an account and that's the tool she was familiar with.
While I agree with you it is not an optimal choice, web chatgpt can solve the problem. I just asked it now to do it (in my own words) and it spit out the code in one go.
Harness? I am an avid HN reader but even I don't yet fully understand how this word is used in the AI context. Which is exactly the point of the OP.
It all boils down to naming things and cache invalidation, /s
That's the terminology they have chosen because they believe they are "in control". It's a purely psychological thing, not different from calling different DB servers "master" and "slave". When I explain things to people, I say: the LLM is a big mouth, and it is a big mouth without hands. It can only talk. The agent is what gives some hands to the LLM, and then it can do some work.
Harness refers to the tooling that allows you to interact with an LLM. Web chat interface is a harness, CLI coding tools (claude code, codex, opencode, pi, etc.) are harnesses, agent systems (openclaw, hermes, etc.) are harnesses. They present different capabilities to the underlying LLM. Codex, for example, is more likely to write code if you say "I want to build an app" than ChatGPT on web will.
Which is an abstraction layer that the non-technical should need to know, if we follow GP's line of reasoning.
My question is do we harness a bootstrap or bootstrap a harness?
You strap your harness to your bootstraps, and then you can pull yourself up by your own bootstraps... right?
You bootstrap css and harness js
Personally I even refuse to learn what it is in the AI context. Feels like a waste of energy given new ~~~best ways to use YOUR tokens~~~ are discovered every month and things you learn now will be useless next month.
I am doing perfectly fine with the web UI version of these tools... They seem to also not make tokens dissappear as fast as using claude cli tool to automate implementations. Makes my work day more tolerable as well as I actually have something to do over waiting until some implementation can be read through...
You forgot off by one errors ;)
You can see the same thing with restaurants generating their own menus/pictures. Some of them look absolutely terrible, visually ugly, way too information dense, the classic piss filter, etc. idk how they do it, even the most basic prompt I can come up with makes something 10x better, and when I put in my amateur photography knowledge/keywords in it gets pretty close to what I'd consider a good pre LLM quality menu. Literally just adding "make it look nicer/cleaner" seems to get rid of most problems, but people just don't care apparently.
Complete amateurs without these tools were basically limited to making big text in word processors and maybe pasting an image in there.
Being able to create a basically coherent, polished looking image is what you’d get on Fiverr for a few bucks. Mostly hustlers filling in templates, or people that know the tools but never learned design fundamentals.
Actually being a competent professional: Knowing how to visually communicate showing information hierarchy, what purely visual aspects of an image say, how different things read differently among people who might see it— e.g. does an image of an apple communicate fancy computer? teachers/school? Nutrition? Food? Produce?, etc etc etc (Good kerning and type usage, composition, gestalt, etc all come with that for free. Many think that is the point — those are tools someone can wield to do good design, they aren’t themselves good design.)
These tools let amateurs do what the fiverr crowd used to do. Unfortunately, the fiverr crowd is now being pushed into doing what entry-level new graduate professionals used to do, and the job market is kind of fucked.
> people just don't care apparently.
This is the target user of these chatbots.
I’d love to see this experiment executed with Claude design.
Particularly with something static, I don’t think they’d fail to get a result.
But without domain knowledge I think they’d misunderstand prototype with finished product.
Without knowing what it’s doing, it’s hard to know what it’s not doing.
As decades-exp SWE I love using Claude Design for any kind of app and web development, because it gives much faster visual feedback loop than changing views within deep framework stack. It much easier to "tell" coding agent what I need instead of writing wall of prose to define visual stuff.
I reminds me old WYSIWYG and unlike Figma it has full HTML/CSS capabilities available.
How I work with it:
- I ask agent to extract part of app into Design, let it even use playwright-cli to get full rendering of the particular view.
- perform design session in Design.
- once design system is perfected I go down to Claude Code dungeons, do /design-sync.
- perform on the stack implementation session.
Actually you don't need Claude Design UI for any of that too. Just ask any coding agent to prepare local mock HTMLs and iterate over them.
I’ve been curious on how to close this loop between engineering and design/product.
We don’t use react, which Claude design seems to trend towards. We use Phoenix / liveview.
We have a shared design system, which keeps the visual elements in line. And then just prototype on design, collab, discuss and arrive at what we want to ship. And then engineering take over and rebuild via hand / claude code.
But the tools aren’t directly connected.
The value has been in the separation. In iterating on the prototype without impacting the codebase, dev cycle, etc. And solving problems/unknowns earlier.
There were always tools for this, but Claude design just feels more accessible and therefore gets used more immediately.
And the fidelity of the outcome (and the assumptions it’s forced the make) are more valuable and faster to achieve than Figma.
I think the way humans divide up design and programming are broadly correct. Working that way with LLMs seems to work well.
I vibe coded an iOS conference schedule app recently, built on top of my own rust UI framework. I started with claude design. I gave it the requirements, and showed it screenshots of other conference schedule apps I like which have features I want to use. I also gave it some visual references for how I want the app styled. It came up with some workable designs. They were a bit 'webby'. But, fine. The high level breakdown of UI screens and navigation between them was excellent.
Then I gave all the HTML files it produced to claude code, along with the documentation for my UI framework and told it to port the code to my UI framework. The first working version was rough. It copied a lot of the unintentional webby look and feel. It worked around missing features in my UI framework by rolling its own janky reimplementations of platform features. For example, instead of using UINavigationController, it rolled its own. It made its own (kinda bad) tab based navigation bar. The app didn't work properly in dark mode, because it was hard-coding a lot of colours. It took a bit of back and forth to fix all of this stuff. But I'm really happy with it now. It looks and feels great.
It's just a pity I couldn't share the app at the conference. Apple took a few days to approve the app in Testflight, and by the time they approved it, the conference was over.
I assume everyone else is playing with the same AI tools that I am, and getting similar results. But a lot of people I talk to seem to have no idea that this is possible right now. They're amazed when I show them my schedule app.
My non-technical cofounder managed to vibe code a holding page with Claude Design and it walked him through deploying it to Netlify.
However for some reason it had him deploy a single HTML file with all the assets encoded as a huge base64 blob in the code that required a massive amount of JavaScript to extract and render.
It's been doing this for our non-technical folk. Giving users a gigantic single file for deployment. We saw one user deploy a JS file with around 3K-5K elements in an array, storing unique IDs of items they wanted to list.
Welcome to Software Development, Lindsey from HR - here's your first database!
People just don't really understand how these things work yet, and they don't know what to ask for, I'm hopeful that they eventually do become more tech-literate, but not sure yet.
Interesting.
I often ask it export a single html file, for an external collaborator or simpler sharing. But I wouldn’t deploy that to production.
I wonder if they asked it to deploy a html file.
But this is exactly the kind of hidden domain knowledge / expertise that changes how you use the tool.
> Particularly with something static
Apps ain’t static.
They probably mean static site, in the sense of static front end, no backend.
Yes. That's not an app.
It could be an app.
Minesweeper is an app right? Unit conversion? Color palette designer? Metronome?
The person you responded too didn't mention app though. They just said static. OP was talking about an app but the responded was hypothesizing about something static.
Anyway I'm not so sure "static" is a viable boundary between app and not app. A static page that does any sort of API request doesn't suddenly become an app imo.
Not really, is a offline chess page not an app? It seems like it would be closer to an app.
What does this hypothetical chess page do? And how’s does it do it if it is static?
You misunderstand the meaning of static web pages. Here's a short explanation of the terms: https://developer.mozilla.org/en-US/docs/Learn_web_developme...
It uses javascript. Still a static file. Lets you play chess
Don’t you update the DOM to render the pieces as they’re moved?
A static website is one that doesn’t have an associated backend API server, just serves as one or more self contained file assets.
The files you serve to the browser are static, not the contents of the page itself
Updating the dom can happen with only individual assets, so it’s a static site
when someone describes a static site/app they generally mean there is no backend. not that the frontend is a static image
for example you can service static sites from S3 that have HTML/CSS/JS but no API or DB
Are you describing a "chat window" experience here? This is apples to oranges.
As opposed to what?
If you ask a non-programmer to install Claude Code, just installing it will be a challenge, then opening the shell and interacting with it. Things as simple as copying and pasting can present roadblocks if you've never used a shell before, and things intuitive to programmers like using up-arrow to go back to a previous prompt would never occur to someone in the field.
Claude Code seems so simple and natural of a UI to programmers, it's easy to forget how much it builds on.
Zoom out one step. The experiment should have been searching for "build website with AI", or "build website with <product name>", not "hey use this very specific UI to do something I know will fail"
(FWIW I think people betting their whole companies on AI are trusting shitty one-wish genie goblins, but the terrible irony is that anyone "technical" with years-old knowledge is talking about something else entirely in today's context)
Claude Cowork is the application aimed at non-developers that gives them a lot of the same functionality. My girlfriend uses it and has gotten quite far in producing her own software.
I got my girlfriend to install Claude Code and she was happily able to create software with it completely independently of me.
The terminal? No way, the normal Claude or ChatGPT desktop app.
With Claude it's actually not that hard. I set up a system for a startup where devs used a full dev environment and non-devs could use the Claude web UI to open rougher PRs, and one of the non-devs asked me what was involved with getting a full dev environment set up, so Claude could iterate and he could test locally, which was faster than the "push to GitHub, use the Vercel/Supabase preview deployment" workflow he was using.
I gave him a link to Ghostty, a link to the claude code copy/paste pipe to bash thing, told him how to cd/ls/pwd into the folder he had locally from the GitHub app, and he was off to the races. I told him to type `claude` to open Claude Code in the terminal and gave him a prompt to use about being a non-dev getting his environment set up to the point of being able to pnpm dev and test, and Claude took it from there. The repo's readme had a setup section which it followed to install brew, asdf, pnpm, etc. With the GitHub MCP he's now opening PRs the same way devs do.
Open any chat window and ask for a simple SPA with startup instructions. It's fine, it's fine, it works.
And you think a random John/Jane Doe knows what a SPA is?
No, but I bet this has a high likelihood of coming out in preliminary discussions.
What most likely happens is a normie has no idea of how to build something that does nothing first - they just start describing the end state.
I'm going to try this with my wife later today. I bet she'd sort something out since she's been a manager forever and phases out instructions maddeningly
Yeah I wonder if they had given their friend Claude code or Codex, would it have been more likely to create what she wanted?
Perhaps! But I do think the vocabulary issue is real and I think LLMs are still sycophantic enough that they won’t really challenge someone or offer alternative ideas on how to implement something unless they explicitly ask.
Interestingly at my work, Claude Code was available before Claude Desktop, so a number of non-technical PMs tried to use it in order to build… anything, with very mixed success.
The “hey guys, check out the website I built with Claude: http://localhost:3000/” joke is real!
In my experience, the whole “the terminal is a scary place” aspect is very real and some non-technical people can feel intimidated by.
I think Claude Code in the desktop app helps alleviate that a bit (perhaps Codex, too, but man what a mess the ‘ol ChatGPT app has become).
But I’m sure there are entire repos of web dev skills that someone could use to put together things with a bit of effort.
> the terminal is a scary place
Isn't the the powerful, unlimited, unopinionated blank LLM text input waiting for your instructions eerily similar to a scary terminal?
WIMP and GUI paradigms are the exact the opposite: intentional dis-empowering, by design restrictions, enumeration of your few possible options. Those feel more constrained therefore safer.
Terminals are scary because it feels insurmountable. What are you supposed to do? If you just type “start python program” it gives you this absurd error that doesn’t make sense. What do you mean start is not in path?
The moment you interact with an LLM it gives you feedback that you’re doing things right. It feels like a gradual climb instead of a series of abrupt jumps. People really don’t like feeling like they don’t know what they’re doing, and the terminal constantly reminds you that you are making mistakes.
Scary but then liberating once you get to know what you have to type in to reach your objective.
Yeah, if you're willing to put in that much effort.
I started on this path literally about as soon as I could read thanks to the family having bought a Commodore 64 for my older siblings, but also perfect timing in that when I got to this age the sibling whose room it was in had just gone off to university.
Most people are not like this, in much the same way that they're not going to read the T&C end-to-end (another thing I've done) or learn enough law to actually understand what those words mean (a step too far even for me).
What an ode to exceptionalism /s
You also think that you are smarter than the people flooding Ceuta streets these days, don't you?
> What an ode to exceptionalism
I've had people criticise me for having had the opportunity to learn in that way, as they did not.
> You also think that you are smarter than the people flooding Ceuta streets these days, don't you?
No, why would I think that? I don't know them, the only thing I can say is in their favour: moving country to better your situation is difficult and them getting as far as they did is a demonstration of putting in a lot of effort of the exact type I praise by default.
> Isn't the the powerful, unlimited, unopinionated blank LLM text input waiting for your instructions eerily similar to a scary terminal?
This is why it has the title (for me currently reading "What can I help with?" but this varies a lot) and the text box itself has the placeholder text "Ask anything". Sometimes I get big friendly suggestions about what to ask it, placed on screen near that text box.
> WIMP and GUI paradigms are the exact the opposite: intentional dis-empowering, by design restrictions, enumeration of your few possible options. Those feel more constrained therefore safer.
I don't think it's constraints, per se: almost nobody looks at the font list and goes "oh no, too many options!"
Rather, GUIs are there to organise your options visually, group them in ways easy to intuitively get. There's a bit of fashion-induced rot here, e.g. I'm old enough to remember when it was always unambiguous when you were looking at a checkbox vs. a radio button, and now there's a blurry middle ground of collections of boxes with ticks in them that act mutually exclusive, but the point of a GUI from a UX POV is not the same as how software in general drifted as it got both more users and more developers and more opinionated managers and middle managers and designers who only cared about shiny rather than usability.
I am old enough to have observed non-tech workers using all kinds of text-based interfaces and it was a real pleasure seeing how old ma's would jiggle numbers on the bc-style TUI in the way that would offset any modern CompSci major.
Me too. But did you see them on their first day using it or their thousandth?
(Also you don't need to be that old. Less than 10 years ago I watched a doctor breeze through some clinical system while I was crawling along constantly referring to the manual)
> WIMP and GUI paradigms are the exact the opposite: intentional dis-empowering
Are you kidding? WIMP and GUI democratized computing!
What "democratized computing" was cheaper computing and VisiCalc, not WIMP and GUI?
But again, VisiCalc is intentionally limited, it's not a all-powerful environment, on purpose. It's all about intentional limitations, making computation easier to reason about.
To put it more bluntly. Computer usage would be low if limited to terminal interactions and there’s no GUI.
Democracy clearly disempowered the monarchs. :)
It did!
Not really, for making a single-file anything ex nihilo. I suppose the chat window won't be able to run a linter or make and run tests as a typical "eager" agent might, so maybe it will make more mistakes.
Almost all web chatbot providers have code sandboxes that they will run (limited) tooling for you. If you ask for it, it will run deterministic linters, formatters, format converters, tests for you. Older versions of Claude would for example happily try to reverse engineer entire artifacts for you, once you provide a URL.
I think this is about teaching problem solving at a young age and it is an abomination that our education system does so poorly at. The key is to know what kind of questions to ask and knowing when to go deeper and what to pay attention to.
But that is not how our education system aligns us. One typical example of problem solving kids, and I too, learn in school is how to apply a concept in physics to a free-body-diagram(Indian and Chinese education cram schools are famously good at teaching kids how to do this, the usefulness of which I debate). But it all stops at the exam room. No architect or jr structural engineer position for you kiddo.
Another kind of problem solving skill might be how to invest money and understand your own risk appetite to construct portfolios to manage your money. All that is taught in school is a dry compound interest formulae, time discounted cash flows and a black scholes model. Only to find out later I don't need most of it to manage my money.
It's also my opinion that Kids should get to do a lot of real work and experimentation at an early age(13 yo IMO) with real responsibilities and prospects of making income. Knowing the failure mode of most real world problems you can solve is the skill you always want but dont nearly get to do enough of until much later in life(sadly true for a large population of kids in the world).
We've seen plenty of examples of apps successfully vibe coded by non technical people, including apps making real revenue.
Your friend could start with telling the LLM that they are a non technical person who wants to make an app and it will explain all the successive steps.
>We've seen plenty of examples of apps successfully vibe coded by non technical people, including apps making real revenue.
Have we? Or is this just something that people say now, without citation?
Maybe people don't cite specific apps because they like their jobs, and outing apps as vibe-coded is still seen as negative
I personally know of two completely vibe-coded large apps in my professional environment. One by a non-technical manager, made to solve his needs, then sold to customers. Initial development went along great, but by now velocity has greatly slowed down. Also took a lot of engineering hours (of actual software developers) to get permission management from "chaotic and ineffective" to passable. It's still worse than what you would have gotten by just using a couple sentences of the right technical language at the start. Deployment is also a bit of a nightmare. All in all, anything beyond the first rollout phase was delayed by months. Honestly it should have stayed as a prototype that then gets rebuilt from the ground up. But still, it is a real app, making real revenue
The other example was vibe-coded by a software engineer in his free time. Works pretty well, doesn't have too many bugs. Makes some revenue, but a lot less. Solving manager problems just sells better.
I vibecoded a Postman/Insomnia API tester program, and I use it everyday at work now.
But as another software engineer, I remove myself from that comparison, because the idea is to find out if a non technical person can do the same, that's the definition of vibe coding.
> Honestly it should have stayed as a prototype that then gets rebuilt from the ground up.
And that’s a natural process for many products. In the journey from discovery to prototype to MVP to product, it should be rebuilt multiple times.
Particularly with LLM’s to assist, the process of rebuilding from a new context and understanding of the desired goal requires even less effort.
The hardest part is managing any real users, their expectations, and any data / workflows they’ve come to require from what came before.
> Maybe people don't cite specific apps because they like their jobs, and outing apps as vibe-coded is still seen as negative
... or IDs apps as cr*p is still seen as negative.
A friend of mine, non-technical, is not making money with his apps. But he's creating a street fighter like game. Just for fun.
So there's that.
He can't exactly release it because he uses a lot of copyrighted stuff. It's also meant only for himself. Though, I've been asking if I can play it, it looks fun.
I think we desperately need to start differentiating between "is creating" and "has created". I have a couple of "am creating" projects too, but their proximity to "have created" is directly proportional to how much effort and expertise _I_ am bringing, not so much related to the AI's contribution.
I run into this quite a bit. We have users generating MANY apps at our small company (30 FTE), entirely with Claude. It's great to see people mess around and tinker. It's NOT great to see someone with a GH repo that has 750+ commits for what would be MAYBE 1 week of a developers time. SO these are non-developers now spending hours and hours working on software that is probably going to get thrown out.
We're in this spot where we don't know when to cut our losses on projects like this. (Is it even viable as production software? Does it currrently do what it's supposed to, or are they adding new features? Is there a return on continued development efforts?)
None of these apps they have built are seeing any major usage, and I don't think a single one is what I would call "done" (There was a gold rush stage at the beginning of 2026 where senior leadership wanted everyone to spend some time messing around with Claude). Unfortunately, they never told anyone when to stop messing around with Claude, so the ROI is ever diminishing.
Plinq.
Plinq was made on Lovable, https://www.aieatingtheworld.com/articles/non-technical-foun...
Couple more on https://buildthedamnthing.com/resources/articles/case-studie...
The problem with lovable, from someone with insider knowledge, is that many of the apps existed even before appearing there and where ported to the platform to ride the hype wave.
Wow, I am not sure that I want my safety app to be videcoded by someone without experience
I don’t think we’re talking about this kind of website.
I’m afraid this isn’t surprising anymore.
.
That was kinda harnessed and prompted by a team of security experts, so there's that.
No we haven't
OTOH my wife's friends got drunk and made "tinder for horse purchases". They prompted to read typical horse advertisements (we're all horse people) and create an app with mock tinder like entries to swipe right and left to buy horses.
A web app was produced with lots of mock "Hi i'm Dominique and i love running through fields and having a bucking good time" type entries complete with silly horse photos. A huge amount of drunken fun even if it boiled a towns water supply and blew through half a subscription to create.
I was looking at the results as a dev with 30 years experience and thinking fuck me. The little apps i made here and there before AI are being outdone by a bunch of drunk people on a whim!
> The little apps i made here and there before AI are being outdone by a bunch of drunk people on a whim!
Yes, but the premise of the article is you should be able to outdo a bunch of drunk people with your 30 years experience, if you use AI too.
The article doesn't hold the universal truth.
My experience is similar, people succeed with Loveable or similar platforms because they don't need to know _deployment_ and _runtime_ experience. It's like using Adobe vs Canva, most people are now used to the latter and don't even have a mental model of "runs on a server" or "database and server are two different things". This is the key difference for me compared to earlier Low Code solutions such as OutSystems that still required a SDLC mental model. However, at some point this breaks, like in your example of feature discussions, because people struggle to test out new ideas _without_ immediately showing them. So it becomes Canva vs Figma, a system-based model. And we don't even have good terminology for that ourselves yet.
did she try saying "I want to make a webpage"? Even ChatGPT will just build, deploy and host a webpage for you with that request. I don't really understand what system she must have been using.
I have the same experience giving my brother an OpenClaw as his personal assistant.
His words were: It feels as if I need to know how to program it.
I was expecting he could say something like "Oh, it seems like you don't remember the people I'm referring to, perhaps you need some kind of CRM system. Can you investigate if there are any easily available CRM systems you can interface with, so we don't need to make one for you?"
Whereas my OpenClaw moment was trying to make it manage its own NixOS installation, so that if I ask it to do something, it doesn't yolo `apt install` commands, but rather improves on the same overview of its own installation.
A lot of people had success making their OpenClaw do things without being Linux experts. But you need a tinkerer's mindset, is what I came to conclude.
Curiosity and a tinkerer's / engineer's mindset is basically the only thing we can hire for right now
That and a proven ability to apply it in a given domain. I am really not good with electronics. I’m sure it’s just a mental block, but I see people debug electronics as relentlessly as I debug code.
Do your friend at least know what Claude Code (or any harness) is?
Of course she wouldn't be able to make a website if she doesn't even know the right tools to use. But I don't think it prove anything. Knowing and installing Claude Code might not be a common sense, but nor is it "expertise" or "skill."
I've seen in first hand that people struggle installing Steam. Yes, "people" in the plural. But just because some people struggle with it, it doesn't mean that installing Steam isn't an objectively easy task. Your friend's experience doesn't change the fact that building a website is something that an average person can do in hours if not minutes.
Knowing what Claude Code is and why you might want it actually is domain knowledge which the op's friend does not have. Your idea of the average person might be biased if you work and socialise with people who have this kind of expertise.
> Your idea of the average person might be biased if you work and socialise with people who have this kind of expertise.
I just said I've seen multiple people struggle installing Steam...
The point is that it's something objectively easy. Once they find (in this case, given by me) the correct instructions and follow through, they can easily do it by themselves again. It's quite different from what are traditionally considered "expertise": for example, even if you followed a master's painting process, stroke by stroke, tomorrow you still don't know how to paint.
Building common apps were more akin to "painting," now it's "installing Steam."
> The point is that it's something objectively easy. Once they find (in this case, given by me) the correct instructions and follow through, they can easily do it by themselves again.
This is true of most things in life. It is very easy to make compost, it is very easy to grow carrots, it is very easy to graft an apple tree onto rootstock, it's very easy to hang a door and it's also very easy to replace the break pads on your car.
Once you've done it, that is. And once you know what tools you need. And how to use those tools. And that you actually have those tools.
Codex, Zed, the like are all tools that you need to know exist and you need to have and you need to know how to use. It's the same thing as a wrench, a break bleeding kit, or some graft tape.
You are assuming that the hypothetical friend has been isolated from any kind of technology from the last few decades. The friend can, presumably, use search engines and the very same chip chipities to get the very rudimentary domain knowledge going.
I assume no such thing.
I'm genuinely a bit floored reading the comments here, but I guess my idea of the average HN commenter's ability to talk to non-technical folks about technical topics is biased because I work and socialize with lots of people who don't have technical expertise (thankfully along with other technical folks who also have lots of experience talking to the former group). Many of them don't even know what Claude is, let alone Claude Code.
The average person does not even know what ChatGPT is and has not interacted with an LLM ever.
I don't think this is true. Certainly not in Britain.
I went to 3 weddings last summer and each one had a joke in a speech about using ChatGPT to write it and everyone laughed. 68yo father of the bride is a retired plumber and even he's cracking jokes about AI.
Yeah it's a true HN bubble moment. I've got non-technical people in my company (a tech company, where the majority of the employees are engineers) who have absolutely no idea what any of the agentic coding nonsense is, their only exposure to this stuff is Gemini in Google's office apps.
To them this is all just "AI" whether it comes from OpenAI, Anthropic or Google - hell they probably don't even know what an LLM is in the first place, yet alone which company provides what tooling. And these are people whose day-to-day involves talking to at least 1 dev a day, so you would imagine some of the knowledge would materialize via osmosis at the very least.
> I've seen in first hand that people struggle installing Steam. Yes, "people" in the plural. But just because some people struggle with it, it doesn't mean that installing Steam isn't an objectively easy task. Your friend's experience doesn't change the fact that building a website is something that an average person can do in hours if not minutes.
I think I would define "easy" with reference to the % of people who can do it. I don't know what that is for Steam, but there's a (now dated) survey of computer literacy in OECD that I keep coming back to in order to set expectations for what "average" looks like:
https://www.nngroup.com/articles/computer-skill-levels/
You overestimate the average person
Every now and then I'll write a single prompt to Claude to see how much of a web app I can build in one shot. I usually pick Django (though I assume any other web framework could suffice) because of the battery included nature, I still sometimes describe that I want users to login / register, have access to x, y, z
I also noticed a friend of mine had way more success by having Claude write code by doing TDD and giving Claude scenarios for things the code should be able to handle, if you do this correctly, and cover edge cases, Claude will work with these in mind.
Question is if the outcome was LLM related or just a result of a personality trait. After all, talking in circles without really approaching an end-goal is something certain humans do all day long.
Reminds me of watching someone who has no idea how to use a search engine try to use a search engine
This sort of thing has been studied in academic experiments. Although the models studied are now “old”, and I expect the floor is higher, the lack of vocabulary and basic concept familiarity sets the ceiling.
https://www.feldmanmolly.com/chiwork2024-author-version.pdf
I’ve had the absolute opposite experience with a friend of mine. I started by setting her up with a terminal emulator on the web linked to Claude Code, and written a CLAUDE.md that told it how to deploy. These days (with no further intervention by me) she’s running Claude Code natively on her laptop, and she tagged me yesterday on Facebook in some update about how much more she enjoyed using Claude Code than plain Claude.
I have been pondering this and I think it's likely a gap that will get filled sooner or later.
Right now there's just so much value in building LLM tools for experts that everyone is focusing on that. But surely at some point we'll have bespoke harnesses that exist exactly to solve this kind of thing.
I think this can start with constrained problem spaces like "you are a WordPress developer, you solve problems for people with enough expertise to know they are looking for a WordPress developer" and incrementally expand from there. Maybe I'm naive but I think you can probably get pretty far with this today just by writing loads of skills and picking the right technical preferences to encode in them.
Yeah there's lots of companies out there that promise that people can build a website quickly, often using WYSIWYG / drag / drop interfaces, I bet they already have AI integration to speed things up. ("I bet" because I don't actually use those services.)
I think people are better helped by using those services than going a level lower and using LLMs directly.
We already have these, and have for a while. Replit and Bolt exist for exactly this case. Not sure why GP didn't direct their friend to these tools (or if these tools are somehow the ones that failed in this example). In my experience, Replit and Bolt (and other similar tools) are quite good at this kind of 0->1 kind of thing.
What terminology is required to make Claude Code write code?
My experience with friends has been the opposite. A PM friend made a custom tool. A friend who has never written a line of computer code has an app.
If you tell Claude Code "I want a website that does X, Y, and Z" it will write code.
Especially if the path is well trodden.
"Make me a kanban board for construction tasks. I want you to walk me through the process of hosting it" worked for me.
I'm wondering: What did she tell you that made you build the website? Did she ask the same thing from the AI? Did you bridge any gaps the AI didn't do for you? I don't know the answer, but I suspect that she treated the AI differently and would've gotten better results if she had asked what she told you. At least if it's an agentic coding assistant like Claude Code or GitHub Copilot. Of course a simple chat will leave manual tasks for you.
Having a similar experience. Watching a muggle try to build a website with basic functionality, e.g. auth, database, etc. is an enlightening experience. They simply do not have the vocabulary to guide the LLM. My good friend calls me every night frustrated with his results, and when I watch what he is doing it's amazing what we take for granted being in the software industry. Don't get me started on the UX, that's even more mind boggling.
I've seen some designers and product managers get pretty far with LLMs but mostly because they already know how to build apps just from a non-technical perspective.
I have counter evidence of this. My wife with absolutely zero skills in computers was at the terminal doing things that Claude was telling her to do and generated some impressive tools. One was a tool to help her organize her day. It involved scripts, PDF generation, printing. She did it all without even asking me. Maybe it fails in some cases but I'm not sure the anecdote above is the average experience.
I find there's a lot of variation between LLMs / models...
But there's also some psychology in play too; that we (engineers) see a lot: Some people just let their imagination run away and forget to "do"; without someone in the conversation pushing for results and action, the conversation will just stay within imagination and everyone will be happy in the moment but nothing will get done.
I am reminded of this one:
https://thedailywtf.com/articles/Could-You-Explain-Programmi...
The comments feel so outdated:
'him: Keywords? Variables? ...
me: (Explaining what programming actually is)
him: Oh, I thought I would write something like: "Create football stadium and football players. Start the game when user presses spacebar. Make players have red shirts and white socks."'
They are from 2008 is why. lol
Yes, but usually programming knowledge does not outdate so quickly. Now we actually have prompts, what some people always believed that it always worked like this.
Weird because I had a friend that wanted a web app, also no experience, and he just told Claude (on the iOS app) to make it for him and Claude just made an artifact and put everything there.
Buddy sent me a share link and he was just like “dude this is crazy I just told it what I wanted and it just spit it out in a few minutes”. No “experiment” needed he just did it because he knew Claude could do it and it worked exactly like you’d expect. Hell, he did it on the Free tier.
I've seen the same thing happen to my brother when he tried making an app with zero experience. Only difference was that he got a front end that didn't work.
same thing as people just entering a question prompt and copying and pasting the response as gospel. e.g. politicians using it to write speeches, or lawyers for testimonials.
the output is programmed to look correct so unless you have some sort of background you won't actually know what errors to look for.
And look correct is accurate.
Not only does it look correct, it looks correct with an extremely Subject Matter Expert degree of authority. Often I'll work with an LLM, and it simply just misses so many things. I've worked in all sorts of different domains, software, chemistry, material design, everything from power generation through to physics, and in each and every case I see it missing incredibly important things. Any true subject matter expert would immediately bring up and prompt concerns, but not the LLM.
This makes sense, of course, because these are language models. They were trained on language. Their first and foremost capability is language.
An LLM's true expertise, true subject matter expertness is language.
And so anyone working with LLMs who isn't already highly skilled in the field they're asking questions about, will invariably be led astray and miss extremely important parts of a puzzle that need to be solved.
It's just like working with engineers.
I just gave this a quick try with Sonnet 5 Medium [0] on a free Anthropic account. It's a bit of a contrived example I guess, but it probably isn't too far off what a completely non-technical person seeing this for the first time would do.
The output is EXTREMELY misleading, as all the data here lives purely locally, yet the AI says that you can "just share the link" and other people will see the schedule you set on the generated artifact. Also, what link? To the Claude chat? It doesn't explain what to do with that `Booking` artifact other than "link to it".
I can so easily see someone tapping out a few steps down the line of this once they realize it doesn't work and they have no clue what to do or say to make it work. What do you even ask as a non-technical person at this point? I guess they could explain "The other person doesn't see it", but would the AI actually clarify that it's because it's not fucking hosted anywhere and has no mechanism of persisting the data outside of the current machine, or would it - as I'm almost 100% sure would be the case - not actually point out this error?
[0] https://claude.ai/share/0cbfe698-3886-4d4d-86e4-7c697b61dc00
...didn't have the vocabulary...
That is exactly the key or the sign there. Even with a couple of decades of engineering expertise, when I try to do / research something that I don't know enough about, I find myself in the exact position of not having the vocabulary.
To the point that I sometimes have to ask the AI "nicely" to cut the pleasantries and be ruthless against nonsense, whether from its/their side or from mine.
Interesting, thanks for sharing.
> Whereas to me who knows the terminology would have probably just taken a single exchange of messages to get the result she described to me. I would have prompted with something like “Please write to me an html page which does X, Y, Z.” But since she didn’t know the right terminology she got into a vortex of feature discusion, and she didn’t find a way to tip the AI into “just do it, write it now” mode.
Even as a developer, when I've been using the web chat interface for things which I know the AI can do easily, I've had this happen to me a few times. I was very surprised the first time I saw ChatGPT respond ~"this would be a few thousand tokens, I can't do that".
Even more surprising: ChatGPT was accurate when responding that way this time, despite this being trivial for Claude and well within what ChatGPT could do using the web chat interface 6 months earlier. The ChatGPT output was extremely meh.
> Because my friend didn’t have the vocabulary to ask the AI to write code.
Why did she even need that?
> “Please write to me an html page which does X, Y, Z.” But since she didn’t know the right terminology
No code there!
copying and pasting code? a few months ago?
2025 called and wants its test back
Exactly, sounds to me like OP is the one with the LLM skill issue.
[dead]
So maybe your friend need a role reversal: A prompt guiding the llm to act as a consultant, guiding people in implementing a software project, asking questions, creating a shared understanding, limiting scope or creating milestones.
That seems like something that could be done using an llm, not that complicated probably.
And maybe, in other fields as well.