Have you tried using Wispr or Willow (or any one of a thousand alternatives?)
A little odd at first but absolutely amazing for the purpose of piling context into an LLM.
Have you tried using Wispr or Willow (or any one of a thousand alternatives?)
A little odd at first but absolutely amazing for the purpose of piling context into an LLM.
It's so bizarre to me that people want to do this.
Can't you type faster than you speak? Doesn't your speaking inhibit your thinking? Aren't you self-conscious talking out loud? How are our experiences so different?
> Can't you type faster than you speak? Doesn't your speaking inhibit your thinking?
For me personally: no, I speak faster than I type; and speaking actually helps me get more ideas compared to typing.
(Not sure if that’s due to having no typing speed barrier, or maybe because speaking activates different parts of the brain.)
Once you get over the feeling of self-consciousness, it’s a great way. I even go on short walks sometimes and mumble to my phone to prepare some long prompt. Thinking works even better, when walking outside :-)
do you perchance have no inner monologue? its a physical difference in how people think that would affect typing vs talking
These are all interrelated points and the sibling comment is correct: it's a skill.
The key thing with these voice systems is that you do not need to edit anything. You can literally just stream of consciousness into them, no editing, include the backtracking, the live-revisions, etc., and it will actually all produce vastly better context for the LLM than the written thing you took even 30 seconds to edit for clarity or brevity.
I had a very similar disposition towards this idea just 6 months ago. I highly recommend trying it out. The key thing is that you do not need to edit. Just keep talking. Try it for a few weeks!
You don't need to edit anything when you are typing either. You don't even need to worry about spelling or typos.
Seems much harder to learn how to write how you never write, versus just learning to speak to a computer the same way you speak to anything else.
What do you mean how you never write? You can just write what for want to say instead of saying it out loud. It is not a special skill, it's no different than how you write a message, just that you don't hit backspace to go back and correct things.
Correct, it's no different from how you normally write, except for the ways in which it is. We agree.
When you're making a big deal out of it being "much harder" because it's "how you never write" and they're saying that's just "not hitting backspace"? No, you don't agree.
How frequently do you write in a stream of consciousness and not hit backspace?
Maybe give a ballpark estimate of characters typed per week in this manner versus characters typed where you are doing some combination of: 1) thinking about what you're writing before you write it, 2) punctuating and formatting correctly, or 3) correcting your writing output?
Ridiculous proposition. And I type correctly at 110+ wpm.
How often I do it doesn't matter because it's such a trivial thing to switch. If you're gonna "try it for a few weeks" the part of you that has to learn the typing-specific parts of that method is about 1% of the difficulty.
It's really easy to ignore typos. And the way you have to approach thinking and correcting is the same whether you're typing or voicing.
If you can't just type the way you would just talk, and you find it notably hard, it's you that's being ridiculous.
But it's literally not. You already correct yourself and revise your speech in an append-only rolling edit. You do it all day every day for decades.
Versus never writing in this way.
Have you tried the voice-based prompting, as I'm describing?
"never" schmever. I never ramble unrestrained either. Deciding not to edit at all is a skill either way. If you just want the equivalent of voice in a normal way, you remove backspace and that's it.
> Have you tried the voice-based prompting, as I'm describing?
I've never prompted a thing. I can just see your distinction is nonsense. If you think it's hard you're doing it wrong.
And wow I did that post without revising a thing. Wow.
Right. So the crux of the issue here is you have no experience with what's being discussed, while I do.
Sheesh, imagine thinking you choose not to rewind time and "edit" the speech that has already come out of your mouth lmao.
I'm doing it right now you goober! Stop calling it hard!
What experience do you insist I'm lacking in something I'm doing right now?
Maybe you lost the thread, but this is a conversation about using dictation (specifically modern dictation tools like Whisper derivatives) to prompt coding agents.
Do you think it suddenly becomes harder to avoid backspace when you're in a different text box?
The thing you're claiming is hard, writing exactly the way you would speak, is not hard.
There's also some additional benefit to going with the flow and not thinking about words much before saying them, but that's equally hard with text or voice.
You just said like two comments ago you have no experience with this. Not sure why you keep acting like you do.
I hadn't before. Then I started doing it just to show how easy it was.
It doesn't matter what text box you're typing in. The ability to type as you'd speak is easy. Without any extra delays or issues.
I hope you're not trying to argue that typing the same way into an AI prompt is harder than doing it into HN. It's just not hard in any situation. Voice isn't special.
You are struggling to hold the thread, I'm afraid. Have a good night!
I see. You've fallen into some weird pedantry to think what I'm saying isn't relevant to your argument. I hope you figure out my very simple meaning later, have a good night too!
A lot of people on the autism spectrum can have issues with speaking or speaking speed, but have no such barrier necessarily to typing speed.
I'd imagine there's a statistically large number of people that meet that criteria on this website.
I mean we don't need to do any epidemiological studies here or anything.
If someone hasn't tried it, they should. It's probably quite different from how they're expecting, might be great, and costs basically nothing. Try it for a few days and if it doesn't work in your workflow, obviously don't do it.
But I have encountered many many people who raised these exact same arguments against trying it, then tried it, and were hooked within days. Exactly 0% of people I've ever convinced to try it decided it just wasn't for them and went back to typing full-time.
[delayed]
My stream of barely-intelligible rambling comes out much faster than I can type, and it's not even close.
It's a skill like any other. You start out stuttering and second-guessing yourself, but after a while, you get better at it. And the LLM smooths out the odd mistakes better than you might think.
Voice dictation tools? Not really. Tried various dictation tools interacting on the phone, but the quality varies, and resulting prompts are very much not like I would write them.
My limiting factor is that 99% of the day I'm around people - either at work, or at home with wife and kids. There's almost no point during the day I could feel comfortable talking at an AI, and even if I stay up late, then talking risks waking the kids up.
Can't wait for some kind of subvocalization microphones to become a thing.
Get a noise machine for wife and kids bedrooms. I talk to friends late at night and my deep voice carries through walls. Works great.
I have zero desire to talk to an ai though, that was cool for about 20 minutes on my pentium 1 acer computer. Hasn’t been since. Old competent non paid Alexa was good for timers as well, the rest of the platforms a turd, nice timers though.
> that was cool for about 20 minutes on my pentium 1 acer computer. Hasn’t been since.
Oh back in the days, i.e. 20 years ago, I had a better voice control system than anything afforded by Alexa or Apple or others, using MS Speech API in its custom constrained grammar mode, plus some bootleg samples of Star Trek's computer voice + a sub-dollar microphone soldered to a long cable and hung on the side of the wardrobe.
The trick that made it work? Microsoft Speech API actually let you train voice to your own text corpus. I'd prepare all combinations of commands I want to issue, print it out, and train it over a dozen short sessions in several locations of the room and at different ambient noise levels (from silent through various genres of music playing at various loudness). End result was more reliable and had better voice-mismatch rejection than any current system I've tried.
Oh, and the real kicker? This all worked fully locally; this was before cloud was even a thing. Turns out you don't actually need cloud for reliable voice control. Nor that much processing power; PC I had then was relatively budget even for 2006.
I highly recommend getting a boom mic that you can keep right up against your mouth and literally whisper into. Take a few weeks just doing stream-of-consciousness prompting into it like this. Don't worry about editing or backtracking or revisions. It's really amazing how well the AI voice recognition → coding agent workflow works.