I wish some of my colleagues would use this for teleconferences
I just cannot tell if this is satire or not
Welcome to clown world!
Mate they could just type out the prompt instead of dictating with the vader mask.
Techbros will have workers do anything but work from home.
Isn’t it cheaper and quieter to just type out your prompts?
This is akin to people who have conversations on speakerphone in public places.
AI needs to measure the level of confidence in your voice, to calibrate its bullshit accordingly
you assume that vibe-coders can actually touch type. or type at all.
I assumed they were all working from their iphones
Speaking is faster than typing, I guess?
Maybe for Eminem.
apparently audio and images are more efficient compared to text for multimodal models?
I'm going to need significant levels of convincing. Computers have always preferred specificity and accuracy, it's half the reason I'm in my current position (MSP Escalations/level 3, half of my success at fixing issues is being extremely specific in looking up exact error messages instead of paraphrasing).
This isn't a defense of AI; on the contrary, it's my doubt that AI can read intentions/inflection/emotion better than just writing out what you actually want.
Deepseek recently published a paper in which they describe that vision tokens contain more information than text tokens and that this can be used to compress context.
We present DeepSeek-OCR as an initial investigation into the feasibility of compressing long contexts via optical 2D mapping.
Experiments show that when the number of text tokens is within 10 times that of vision tokens (i.e., a compression ratio < 10×), the model can achieve decoding (OCR) precision of 97%. Even at a compression ratio of 20×, the OCR accuracy still remains at about 60%. This shows considerable promise for research areas such as historical long-context compression and memory forgetting mechanisms in LLMs.
It reminds me of LLM caveman speak, it used to have another option to use Chinese instead of English. A language like Chinese is seemingly better at encoding information in fewer tokens and I think this is the same mechanism why OCR tokens work so well.
That said, I also doubt that voice messages are more efficient than text prompts, but it's best not to waste too much time engaging with these sorts of LinkedIn posts (and LinkedIn in general).
LLMs don't need accuracy. This just boils down to speaking being faster than typing, especially if your thought isn't fully formulated.
As far as I know, these workflows typically involve a transcription model to convert the audio to text, and then passing the text to the model.
If only there was some sort of device that would allow you to input those commands without interfering with others, or vice versa... But it's probably just a dream...
This looks uncomfortable and humiliating. Now if they were to make it in the form of a suppository...
You win best comment award on this thread. 🏆
Use a keyboard ya filthy casuals
A noise cancelling microphone isn't a terrible invention though.
I've seen those things 10 years ago when somebody was helping a deaf student in university. Used one of those things to talk to a text to speech engine during the lecture. Seemed to work quite well.
They suck, you can't get air into them without destroying the noise cancelling part so you've got to constantly be taking it off to take a breath / running out of oxygen mid sentence
You talk with your nose?
Your nose in in the mask. It's extremely hard to make a seal around your lips that doesn't include the nose
Thats true this is just the absolute worst use case scenario i can think of
Back in the day there used to be specialized equipment for this purpose called an office, of which you had your own and could close the door to
It's lazy enough to vibe code, but when you're too lazy to even type a prompt.....sheeeesh
This is fox news levels of propaganda but on the left
I love the implication that he's speaking to the AI and wants to not disturb people, but has no headphones on. so I can only assume it's responding from the speakers.
note
Yes, I know it could be responding over text instead of voice.
Which I think it is even more funny, why not just type them if is not a full audio back and forth.
If YouTube’s automated subtitles are anything to go by, it’ll randomly start thinking I’m speaking Vietnamese.
Is typing THAT hard?
When you need a nannybot? Yes.
Cap here. The product is called vocal dampener and used by singers during warmup

Every husband here is sad for all those years he didn't knew these existed. Mothers and teachers wonder if these are available in kid-sizes. Kids wonder what glue works on both this device and the teachers skin.
Why not give them private offices? If AI is increasing productivity and replacing workers they should have enough space left.
I have never seen anyone except streamers try this out in all honesty, I doubt the article is actually truthful. Even vibe coders understand how to type.
Developer to LLM: I am your father!
Before I read the text I thought this was a photo of a stenographer; they used those or similar in depositions I've given.
I got subpoenaed once and the stenographer used one of these.
Save money. Get human shaped boxes with lids that they can lay down in to save energy to increase productivity!

Next up: Humans are hardwiring themselves to the computers, with only the brain and a few select "most important" organs remaining of them.
Programmer Humor
Welcome to Programmer Humor!
This is a place where you can post jokes, memes, humor, etc. related to programming!
For sharing awful code theres also Programming Horror.
Rules
- Keep content in english
- No advertisements
- Posts must be related to programming or programmer topics