838
you are viewing a single comment's thread
view the rest of the comments
[-] W3dd1e@lemmy.zip 58 points 2 weeks ago

Isn’t it cheaper and quieter to just type out your prompts?

This is akin to people who have conversations on speakerphone in public places.

[-] axus@lemmy.ca 13 points 2 weeks ago

AI needs to measure the level of confidence in your voice, to calibrate its bullshit accordingly

[-] calebwill@lemmy.zip 1 points 2 weeks ago

I wonder if anyone owns a patent on this idea.

[-] mavu@discuss.tchncs.de 11 points 2 weeks ago

you assume that vibe-coders can actually touch type. or type at all.

[-] SethW@lemmy.world 4 points 2 weeks ago

I assumed they were all working from their iphones

[-] 0x0@lemmy.zip 3 points 2 weeks ago

Or be capable of critical thinking.

[-] draco_aeneus@mander.xyz 8 points 2 weeks ago

Speaking is faster than typing, I guess?

[-] lightnsfw@reddthat.com 10 points 2 weeks ago

Maybe for Eminem.

[-] AppleMango@lemmy.world 6 points 2 weeks ago

apparently audio and images are more efficient compared to text for multimodal models?

[-] chatokun@lemmy.dbzer0.com 31 points 2 weeks ago

I'm going to need significant levels of convincing. Computers have always preferred specificity and accuracy, it's half the reason I'm in my current position (MSP Escalations/level 3, half of my success at fixing issues is being extremely specific in looking up exact error messages instead of paraphrasing).

This isn't a defense of AI; on the contrary, it's my doubt that AI can read intentions/inflection/emotion better than just writing out what you actually want.

[-] qaz@lemmy.world 8 points 2 weeks ago* (last edited 2 weeks ago)

Deepseek recently published a paper in which they describe that vision tokens contain more information than text tokens and that this can be used to compress context.

We present DeepSeek-OCR as an initial investigation into the feasibility of compressing long contexts via optical 2D mapping.

Experiments show that when the number of text tokens is within 10 times that of vision tokens (i.e., a compression ratio < 10×), the model can achieve decoding (OCR) precision of 97%. Even at a compression ratio of 20×, the OCR accuracy still remains at about 60%. This shows considerable promise for research areas such as historical long-context compression and memory forgetting mechanisms in LLMs.

It reminds me of LLM caveman speak, it used to have another option to use Chinese instead of English. A language like Chinese is seemingly better at encoding information in fewer tokens and I think this is the same mechanism why OCR tokens work so well.

That said, I also doubt that voice messages are more efficient than text prompts, but it's best not to waste too much time engaging with these sorts of LinkedIn posts (and LinkedIn in general).

[-] db_null@lemmy.dbzer0.com 6 points 2 weeks ago

LLMs don't need accuracy. This just boils down to speaking being faster than typing, especially if your thought isn't fully formulated.

[-] VibeSurgeon@piefed.social 7 points 2 weeks ago

As far as I know, these workflows typically involve a transcription model to convert the audio to text, and then passing the text to the model.

[-] VinegarChunks@lemmus.org 4 points 2 weeks ago

Older people such as myself tend to hate voice-to-text I think because it was so awful in the past. And if you screwed up with it in the past it was a less understandable excuse that “I was using voice-to-text.” And because we were all forced in some way to learn to type well.

Voice to text works a little bit better now. And I think younger people know everyone else uses it and to forgive when it screws up.

[-] nightlily@leminal.space 6 points 2 weeks ago

I hate it because it’s generally intensely US-centric. Not understanding non-US accents or terms.

[-] edible_funk@lemmy.world 1 points 2 weeks ago

It worked better five years ago before they replaced the existing algorithms with AI bullshit. Keeps adding slurs to my dictionary too since they replaced keyboard prediction with modern AI and so many people use slurs Google just assumes I do too.

[-] Beehaw_Girl@beehaw.org 1 points 2 weeks ago

Voice to text is better for you now than it was in the past? I remember in 2014 voice to text worked great for me. Now it hardly ever works.

[-] Beehaw_Girl@beehaw.org 1 points 2 weeks ago

Maybe this was designed for illiterate coders.

this post was submitted on 13 Jul 2026
838 points (100.0% liked)

Programmer Humor

32559 readers
1181 users here now

Welcome to Programmer Humor!

This is a place where you can post jokes, memes, humor, etc. related to programming!

For sharing awful code theres also Programming Horror.

Rules

founded 3 years ago
MODERATORS