137
submitted 6 days ago* (last edited 6 days ago) by rimu@piefed.social to c/programming@programming.dev

studies show a clear trend – output is up (more code, more commits, bigger diffs), but outcomes don’t reflect that trend. If anything, the average team is taking longer to ship worse software

top 50 comments
sorted by: hot top controversial new old
[-] Digit@lemmy.today 1 points 22 hours ago

studies show a clear trend – output is up (more code, more commits, bigger diffs), but outcomes don’t reflect that trend. If anything, the average team is taking longer to ship worse software

&

burning out sooner,

suffocating high pace churn,

debugging bad code.

[-] Feyd@programming.dev 69 points 6 days ago

output is up (more code, more commits, bigger diffs)

We've known measuring output by lines of code is counterproductive for a long time.

[-] themaninblack@lemmy.world 3 points 1 day ago

I contributed -13,000 lines of code this week. Very proud

[-] jtrek@startrek.website 25 points 6 days ago

Has management known that?

A lot of problems seem to be downstream from "management are idiots and assholes"

[-] Feyd@programming.dev 7 points 6 days ago

I'm sure there were pockets but it legitimately seemed like that dragon had been slain until recently. Trying to assign more meaning to scrum points has been in vogue for a while though.

[-] jtrek@startrek.website 4 points 6 days ago

My team assigns both hours and points to tasks. I've never seen the points used for anything but they still spend time on it.

[-] Kissaki@programming.dev 3 points 5 days ago

That's... an interesting approach.

[-] dejected_warp_core@lemmy.world 2 points 6 days ago

That battle is perpetual. Lazy management sees a number that resembles a statistic, and try to use it as an easy metric for stuff it doesn't represent. Points do aggregate into velocity, which is worth measuring. But on their own, points are a proxy for estimation in $SPRINT_LENGTH days. The way to manage up is to keep making it clear that the smallest unit of estimation in Agile is the sprint length; points are used to subdivide that but only to ensure that the sprint itself is not overloaded and thus an accurate estimate.

[-] dreamkeeper@literature.cafe 3 points 5 days ago

At any minimally competent company they are aware.

However my company mostly has former engineers as engineering managers.

[-] criss_cross@lemmy.world 5 points 6 days ago

Can you retell my company that?

[-] Kissaki@programming.dev 7 points 5 days ago

For a small exorbitant consulting fee that legitimizes me, I can tell them that.

[-] ArseAssassin@sopuli.xyz 50 points 6 days ago

One study found a significant correlation between confidence in AI output and belief in the paranormal.

💀 💀 💀

[-] 87Six@lemmy.zip 8 points 6 days ago

That's HILLARIOUS

[-] MonkeMischief@lemmy.today 3 points 6 days ago

Lol that's funny. My belief in having a divinely created soul is exactly why I think humans can't be replaced by these supercharged drunken parrots. :O

[-] Senal@programming.dev 6 points 6 days ago* (last edited 6 days ago)

I am 100% genuinely not trying to start a fight, this a legitimate question to which I really would like to know the answer.

What is the difference between a divine creation and the "ghost in the machine "?

I will give context , faith is genuinely confusing to me.

I have no issue with personal faiths unless you are trying to force everyone to have the same faith as you.

I consider anything above small scale organisation of religion to be the worst thing that can happen to faith, because people are people and power corrupts.

I suppose my real question is , for a concept that prohibits proof as part of it's definition, how do you determine that one faith (or system) is better than another?

[-] arendjr@programming.dev 5 points 6 days ago

I’m not the person you responded to, but I can try to take a stab at that…

for a concept that prohibits proof as part of it's definition, how do you determine that one faith (or system) is better than another?

Objectively, you don’t. But people’s experiences go beyond the objective, and we all have our subjective experiences too. Faith is how we make sense of those.

The question then isn’t, which one is better, but whose other experiences and faiths do we relate to? We build friendships and alliances based on those. Not because they’re better, per se, but because we feel comfortable or safe with them.

From that perspective it makes a lot of sense that we value other humans, because they can make us feel understood and appreciated, and we can exchange ideas with them which in turn refines our faith. Now, some people get tricked into thinking you can do the same with machines, but well, I think you can guess where I too stand on that idea.

[-] Senal@programming.dev 2 points 3 days ago* (last edited 3 days ago)

That is a well thought out and helpful answer, that's useful to me in understanding come aspects of faith in general, thank you.

Before i continue i want to highlight the differences i see between:

- LLM (Large Language Models)
	- The trending technology that is a smaller part of a general field.
        - it's what you are interacting with using ChatGPT, Claude etc
- AI (Artifical Intelligence)
	- The general field, which includes LLM's but also Neural Networks, certain kinds of algorithms and more.
- AGI (Artificial **General** Intelligence)
	- The common scientific term for what most people think of when they hear "AI". 
        - roughly and borderline inaccurately:
              - "a machine intellect of the same type as a human, matching or exceeding common human capability"

I get that faith, mysticism etc are how some of us deal with subjective ( from our point of view ) phenomena . I'd imagine this is why gods for things like storms or seasons etc come about.

As i said , I have no issue with any of that unless it starts being forced on people, i even vaguely understand how people could come to that conclusion and i'm not pretending to have any better answers.

My original question was more about this part:

Now, some people get tricked into thinking you can do the same with machines, but well, I think you can guess where I too stand on that idea.

Are you saying that the subjective experiences with the LLM's are the trick or the feeling of understanding and appreciation with the machine ?

Some peoples experiences with LLM's (perhaps what they consider to be an AGI?) sounds like they'd fall exactly in to the type of subjective phenomena that would bring about faith as a means of processing.

That some people seem to think that that kind of faith is any less valid than a guardian deity or a wind god is where I'm struggling.

I don't understand how people are making that kind of distinction, unless it's not something based in reasoning?

I want to make clear i don't think LLM's are sapient or sentient at this point, though i can't personally disprove it and am not against being convinced otherwise should sufficient rationale be provided.

[-] arendjr@programming.dev 1 points 3 days ago

There’s a lot of levels at which I could reply to this comment, because it’s quite a philosophical topic, but let me clarify what I meant at the end of my last comment, before writing a wall of text here :)

As I said in my previous comment, faith is largely a means to make sense of our subjective experiences. We can refine our faith by talking about our experiences with other people.

So the most basic fallacy I was alluding to, from people who believe you can refine your faith by talking with an LLM is this: An LLM has no subjective experiences, because it doesn’t experience anything at all. This is one way how using the output from an LLM as a source of inspiration for your faith sets you up for psychosis. You start believing things that have no basis in reality or anyone’s genuine experiences. Rather, you may start believing in machine-generated approximations from which coherence has been lost.

But it seems your understanding was slightly different, in that people might develop their own sense of faith from experiences with an LLM itself rather than through the exchange of presumed experiences with the LLM. Do I understand correctly that you think some people attribute godlike qualities to the LLM, maybe through the presumption that it is an AGI? That would be a bit of a different discussion indeed :)

[-] Senal@programming.dev 1 points 2 days ago

So the most basic fallacy I was alluding to, from people who believe you can refine your faith by talking with an LLM is this: An LLM has no subjective experiences, because it doesn’t experience anything at all. This is one way how using the output from an LLM as a source of inspiration for your faith sets you up for psychosis. You start believing things that have no basis in reality or anyone’s genuine experiences. Rather, you may start believing in machine-generated approximations from which coherence has been lost.

Reasonable, though it could be argued they are full of subjective experiences, just all mixed together.

But i expect it's probably more about the coherence of the experience that provides the meaning in this context.

Coherence as in , one coherent explanation of a subjective experience by a person rather than a plausible sounding experience made up of many bits of experiences from a corpus of data.

But it seems your understanding was slightly different, in that people might develop their own sense of faith from experiences with an LLM itself rather than through the exchange of presumed experiences with the LLM. Do I understand correctly that you think some people attribute godlike qualities to the LLM, maybe through the presumption that it is an AGI? That would be a bit of a different discussion indeed :)

Not exactly, i think there are two different things here from my side.

The first:

I think people might attribute person like qualities to an LLM, the "ghost in the machine" as was mentioned earlier.

I wasn't trying to imply someone who knows what an AGI is supposed to be would use that knowledge to assign godlike attributes to an LLM (a "god in the machine").

Though , as a concept that would also be a valid example of "belief in the absence of proof".

I don't expect most people to know the terminology or have even an idea of what's happening behind the scenes.

The second:

The part i'm struggling with is why "faith" in something like ability for meaning to come from LLM generated text (god-like or person-like) is any less acceptable than "faith" derived from perceived meaning in seemingly random (based on the knowledge of the time) events, like thunderstorms or plagues.

They would both cover the idea of belief in something in the absence of proof, so how would someone decide one is less valid than the other.

[-] arendjr@programming.dev 1 points 2 days ago

Right, gotcha. So, a few things here:

Reasonable, though it could be argued they are full of subjective experiences, just all mixed together.

This argument wouldn’t really make sense to me, because an experience and the words describing said experience are clearly not the same. An LLM only possesses the later, but the actual experience is something conscious that an LLM does not have.

Words are approximations, and often poor ones at that. For us humans, we can use them to describe experiences so that hopefully they bring understanding in the recipient, but whether the recipient truly understands and whether they can relate to the experience, also depends on their own experiences and their empathic abilities. None of that is at work inside an LLM.

But indeed, thinking that an LLM does have such abilities would be an example of attributing person-like qualities to it.

Now, if a person does develop a sense of faith from LLM output, is that more or less valid than them developing a sense of faith through a thunderstorm? That depends on the faith, I guess.

If someone truly believes the LLM to be person-like (or godlike) that would be a nonsensical faith, as much as someone believing that firestorms are the wrath of god, for instance.

If someone believes that an LLM-generated text provided them with a critical insight that turned their life around, and therefore LLMs are divinely inspired, it’s as harmless as someone believing that a thunderbolt that struck nearby gave them the necessary wakeup call that turned their life around, and therefore thunderstorms are guided by the divine.

Of course in either case, if they start to revere LLMs/thunderstorms above other aspects of reality, it becomes problematic.

Personally, as a more spiritually-inclined person, I think we should be able to say that both thunderstorms and LLMs are awe-inspiring, and I think both can, in theory, be a vessel for the divine. But ultimately, neither are divine in and of themselves.

So isn’t there a difference between faith based on LLMs vs. faith based on firestorms? I think it really depends on the case, but there’s still an important distinction: A firestorm doesn’t pretend to be human, or doesn’t try to trick people into pretending it has humanlike qualities or experiences. So generally speaking, people tend not to develop weird faiths based on thunderstorms. But LLMs are, by design, imposters pretending to have humanlike qualities. This is why I think we tend to be dismissive of people believing in things told by LLMs, because it’s so easy to fall into the idea of ascribing real humanlike qualities to them and developing psychotic thoughts from there.

[-] MonkeMischief@lemmy.today 1 points 5 days ago

That was a very insightful answer. Well said! Thank you very much for replying. :)

I will have to contemplate a little bit, and respond to the question myself as well.

[-] Senal@programming.dev 2 points 3 days ago

i have posted a response to the above comment that might contain relevant information.

Thank you for taking the time to respond.

[-] NigelFrobisher@aussie.zone 17 points 5 days ago

I could’ve told them this, tbh.

[-] dejected_warp_core@lemmy.world 21 points 6 days ago* (last edited 6 days ago)

There are some cardinal sins within the annals of software development in the workplace. The relevant two here are:

  • Do not build a scoreboard for productivity
  • Do not equate keystrokes with effort or value rendered

These create perverse incentives that can really screw everything up, including creating permanent damage to company culture, products, and productivity. Anytime you see these things done, it's because you have (or are) lazy-ass management.

Edit: I also just learned that this is an application of Goodhart's Law.

[-] themaninblack@lemmy.world 6 points 5 days ago* (last edited 5 days ago)

Oh god a former company had the burndown chart on a big plasma screen at all times

[-] chicken@lemmy.dbzer0.com 15 points 5 days ago

LLMs cannot distinguish between recent and out-of-date information in the context, and information in the model itself, learned during training (“dominant priors”), can often “outweigh” information we give it

LLM inference is more accurate when we give them examples (demonstrations) rather than just describing what we want.

Deep neural networks, including LLMs, struggle to learn patterns with long-range dependencies, at any scale of model. They will always be “driving in fog”, with local, short-range probabilities crowding out long-range ones. In case you were wondering why they suck at the “big picture” – probabilistically, it’s a blur.

I'd guess that what all this stuff adds up to is, sustainable use of LLMs as a coding tool for nontrivial projects calls for an entirely reworked set of software development practices to conform to its limitations effectively, but the people in charge really really want and believe it to be a drop-in efficiency boost, and a big mess results. This reminds me a lot of the articles and arguments I've read over the years about low level vs high level programming languages and frameworks. Probably will play out a similar way.

[-] Simulation6@sopuli.xyz 9 points 5 days ago

LLMs can't tell the difference between code and comments in some cases. Older code bases that have had a number of hands touching it over the years have a lot of commented out code and dead methods. LLM sees that as good live code. Also there will be comments like 'this is a stupid way to do this, but I don't have time to fix it right now'. LLM has no sense of humor in this case.
And sure. I should clean all this up before hand, but who has time for that? Project has one, part time developer working on it now.

[-] ell1e@leminal.space 14 points 6 days ago* (last edited 6 days ago)

This doesn't seem to cover there is also no LLM that doesn't plagiarize, or where the training data appears to be compatible with such behavior (e.g. CC0). Now I don't know what that means legally, but morally it seems to be tossing away other project's licensing and I think for FOSS as a whole that's no good.

Also something worth reiterating: https://machinelearning.apple.com/research/illusion-of-thinking LLMs apparently can't do basic logical reasoning. Even a junior coder can do that. I'm always surprised anybody would let LLMs near their code, at all.

I recommed you reading this

It summarizes really good not only the moral, but also the legal problems of AI, vibecoding and "AI-assisted/AI-boosted" programming/engineering/development.

[-] hoshikarakitaridia@lemmy.world 15 points 6 days ago
[-] WhatAmLemmy@lemmy.world 12 points 6 days ago* (last edited 6 days ago)

It's obvious if you actually use software beyond the average literacy of a talking chimp. I use hundreds of apps across iphone, mac, and linux os's. Literally none of them have noticeably increased in quality, stability, or feature-set beyond their average between 1-5 years ago.

Mac and iphone appear to have more bugs and shittier quality control than at any other point in the last decade.

Quality software is getting harder and harder to find thanks to all the slop-abandonware being promoted by slop-content and slop-SEO on slop-enshittified search engines.

I notice far more idiocracy-grade errors in digital content, cx, business processes, product listings, etc than ever before.

Weather forecasts have gone to dogshit in the last 2 years. Even same-day forecasts can shift on a dime unpredictably. It's at the point where I check 3 apps. I never had a time where I woke up to 0% chance of rain and sunny, then looked outside to see rain. Not a sun shower. A rainy day several mm downpour.

Auto-generated subtitles are great for content that was never going to receive human attention, but they're clearly being used to replace humans. At least once a week I notice a major contextual error that completely alters the perception of the line/scene, and there's no way to submit corrections.

Art, culture, and knowledge are being actively corrupted, bastardized, and destroyed.

The future sucks, and is dumb as all fuck. Complete clown show run by criminals, rapists, pedophiles, scammers, and thieves.

[-] eah@programming.dev 2 points 6 days ago* (last edited 6 days ago)

Weather forecasts have gone to dogshit in the last 2 years. Even same-day forecasts can shift on a dime unpredictably. It’s at the point where I check 3 apps. Until this year, I never had a time where I woke up to 0% chance of rain and sunny, then looked outside to see rain. Not a sun shower. A rainy day hour-long downpour.

Supposing you're right that weather forecasts have gone to dogshit (I'm skeptical because your evidence is anecdotal), there may be reasons beside AI for that.

[-] WhatAmLemmy@lemmy.world 0 points 5 days ago* (last edited 5 days ago)

Yeah I was gonna add the weather part could be entirely explained by climate change and fascisms war on science, but forgot. Weather forecasting is actually one area where AI should far exceed what humans could ever possibly achieve without it. It is impossible for humans to process and adapt to the changing patterns across hundreds or thousands of variables interacting with each other in real time.

I'd never heard of the 5G impact, so thanks for that. I consider it unlikely unless there were a significant change in 5G usage during the last year. 5G has covered most Australian major cities since 2021, and 5g coverage is essentially non-existent outside of populated areas (90% or more of Australia).

[-] Feyd@programming.dev 1 points 6 days ago

Yes it blows my mind everyone can't see all the historically perfectly fine software starting to crumble to dust

[-] The_Decryptor@aussie.zone 1 points 5 days ago

The big one for me was rsync, maintainer started vibecoding it and the first release with those changes had a huge amount of regressions.

Yet it keeps getting held up as an example of "using AI tools properly".

[-] MagicShel@lemmy.zip 8 points 6 days ago

This aligns with my experience, largely. Of course it's still my job to maximize LLM effectiveness within my organization. Which is a delicate balancing act to protect my teams from overeager executive leadership looking for huge gains.

My own summary is that AI can be an accelerator, but the harder you lean into it, the worse outcomes will be. No matter how much code is written, you still need actual human minds to understand it and they can only handle so much volume before getting overwhelmed.

Also, if AI gives you 20% productivity gains, but that 20% goes into playing with AI trying to get more, you haven't really gained anything. Usage needs to be standardized rather than developers constantly negotiating with AI trying to coax out better outcomes.

[-] staircase@programming.dev 3 points 6 days ago* (last edited 6 days ago)

human minds ... can only handle so much volume before getting overwhelmed.

One might even consider flourishing employees as opposed to not-burned-out ones.

[-] jerakor@startrek.website 1 points 6 days ago

This is a tale older than AI. Most of the AI productivity pushes I struggle to get adopted fail not because of AI bad or its too hard to do. They fail because of a broken CI/CD pipeline. They fail because some team thinks their process is sacred and unique.

[-] faltryka@lemmy.world 7 points 6 days ago

Some of this does not line up with my lived experience pretty starkly.

Repo level markdown files with architectural guidance not working for example… I’ve found that works quite well.

Not perfectly well, but llms are designed specifically NOT to be perfect deterministic executioners. Still though, pretty well.

I have seen that in a jr engineers hands llms get to bad outcomes fast, and unintuitively (to leaders…) usage of llms in coding does not provide a path for a he engineer to upskill into a sr engineer. A sr engineer with llms though is almost always radically augmented regarding their output speed on task completion.

[-] MagicShel@lemmy.zip 6 points 6 days ago

I agree with your last paragraph. We had about 6 weeks of unlimited AI spend before the costs reached executive leadership, and in that time I saw the least experienced developers spend the most with the least to show for it.

But I will say that another factor is thinking that if you get 10% gains from a little AI, then a lot of AI will get you 100%.

But I find the article is right about repo-wide docs. At least on their own. I find having small markdowns (often in the form of skills/commands), focused on specific tasks reduces spend (especially when your execution agent is a low cost model, leaving the reasoning to dedicated agents) and gives better outcomes. Loading massive docs into every task reduces the attention to the task at hand and often confuses AI as the reasoning part of the model becomes overwhelmed and starts inferring wrong things confidently.

I suppose it heavily depends on the scale of the repo though. A large microservice with multiple upstream services it needs to call spends a lot tokens on API which is unnecessary for most tasks. And then it decides to use the wrong one.... I have stories lol.

[-] rimu@piefed.social 2 points 5 days ago

It's possible to win lots of battles but still lose the war. You can ask Trump about that :)

[-] tty5@lemmy.world 2 points 6 days ago* (last edited 6 days ago)

Repo level markdown files with architectural guidance not working for example

I've seen it become less and less effective as the size of the file(s) grew and as the codebase grew - they got increasingly more diluted or even lost in context compression. After several months of a 6 man team working on the project the rate at which they got ignored started affecting output a lot.

[-] melfie@lemmy.zip 1 points 5 days ago

Repo level markdown files with architectural guidance not working for example… I’ve found that works quite well.

Same. AGENTS.md files and the like are quite effective. Especially if you’re reviewing the code and making the LLM help you update the markdown files when it makes a mistake to prevent the same type of mistake in the future. Having concrete examples of “good” vs. “bad” to illustrate each architectural rule goes a long way.

For any feature or bug fix that is “painting with the colors already in the tray”, it makes sense to let a LLM write the code. Humans will introduce new tech and new patterns out of boredom and turn the codebase into a big Frankenstein, but the LLM will just follow the architectural guidelines indefinitely.

[-] faltryka@lemmy.world 2 points 5 days ago

Agree, I have them curated lessons.md anytime they make a mistake and have found that to be highly effective. Every now and then a lesson goes defunct and needs pruned, but I think that’s just part of the new swe skill set.

[-] JeeBaiChow@lemmy.world 7 points 6 days ago

Lol. Who knew?

[-] OpenStars@discuss.online 5 points 6 days ago

Hehehehe

LLMs struggle with negation. Telling them not to do something can often have the same effect as telling them to do it.

The future looks... unreliable.

Models may get more powerful, but not significantly more reliable. This it folks – work with what you’ve got!

We all know that true AGIs becoming smarter than humans seems inevitable, but that could be like a hundred years from now, if ever. What's unclear is what will happen two years from now, involving matters having little to do with the technology & what it is capable of and instead more to do with the economy and what jobs will be available then.

img

[-] Blurntout@lemmy.ca 3 points 6 days ago

15 - 30 years of pain followed by the end of life as we know it those who adapt may thrive but will continue to be exploited.

[-] OpenStars@discuss.online 2 points 6 days ago

Tbf to LLM manufacturers, the end of life as we know it was coming either way.

[-] LemmyBruceLeeMarvin@lemmy.ml 2 points 5 days ago

Someone's ITIL certified

this post was submitted on 13 Aug 2026
137 points (100.0% liked)

Programming

28153 readers
140 users here now

Welcome to the main community in programming.dev! Feel free to post anything relating to programming here!

Cross posting is strongly encouraged in the instance. If you feel your post or another person's post makes sense in another community cross post into it.

Hope you enjoy the instance!

Rules

Rules

  • Follow the programming.dev instance rules
  • Keep content related to programming in some way
  • If you're posting long videos try to add in some form of tldr for those who don't want to watch videos

Wormhole

Follow the wormhole through a path of communities !webdev@programming.dev



founded 3 years ago
MODERATORS