33 comments

  • frereubu 41 minutes ago
  • nsainsbury 28 minutes ago
    It's actually sad how much AI has either not improved life at all or actively made it worse from the perspective of the average person:

    Your car is still the same. Your dishwasher is still the same. The train you take to work is still the same and never comes on time. Fuel is more expensive. The roads are still congested with traffic. Your kids are (probably) doing worse at school. Food costs more. Houses cost more. Rent is higher. Buying a computer or a PS5 is more expensive. Politics is still full of mentally unstable people. The environment is getting worse. You will still die of heart disease or cancer. Food quality is worse. Your kids can't find work. Wealth inequality is accelerating.

    But hey...on the flip side...a lot of people are rapidly building software (that nobody is using), and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)

    Imagine what a "country of geniuses in a datacenter" will be able to do? Apparently...nothing.

    • trescenzi 8 minutes ago
      This is the core hubris of everyone behind the LLM craze. I do not want to deny their utility there absolutely is some there. But the inherent assumption behind the whole thing is that with a "country of geniuses" problems will melt away. But that's not the case. The reason problems like housing or famines or war aren't being solved is because structurally there's no incentives to solve them not because there's not enough geniuses out there.
      • aleph_minus_one 1 minute ago
        > The reason problems like housing or famines or war aren't being solved is because structurally there's no incentives to solve them not because there's not enough geniuses out there.

        Rather: these problems are not solved because there are not enough geniuses that are capable of convincing the masses to send the responsible politicians to hell.

    • robswc 21 minutes ago
      This is sort of my take. You could look around in 2016 and 2026 and honestly not much has changed... at least at first glance. There is a missing piece of the puzzle.

      I'm noticing this even in industry. Yea, we can build and iterate on ideas 100x faster but has this turned into anything meaningful? Not really... Hell, look at the big companies spending billions on inference... nothing shipped and things are still buggy as ever. Where is everything?

      • Rapzid 13 minutes ago
        The job market had changed that's for sure.
      • prometheus1992 4 minutes ago
        >You could look around in 2016 and 2026 and honestly not much has changed

        I find 2026 much more repulsive compared to 2016.

    • kozikow 11 minutes ago
      At least for me, personal agents have been quite useful. If you integrate them well - shopping, filling random government forms, keeping track of my personal life, optimizing training plan and nutrition, planning vacations, etc. I can do more home repairs or renovations myself with help of chat

      So basically you get personal digital assistant - very useful - but not changing my life.

      What would I need for real transformation though? Food, house cleaning, commute. Even if we had ASI tomorrow, what could it do to change my life, besides taking away my job? Design good house robot to help me with food preparation and chores... Maybe finally self driving cars launched in my area? Or it will take control, "optimize humanity survival chance" and realize we don't need as many people on this planet.

      • ptx 3 minutes ago
        Almost every time I use an AI chatbot it eventually turns out it's lying to me. I would not trust them with any of the things you mention.
    • c0rruptbytes 4 minutes ago
      LLMs could’ve been developed a lot slower without capsizing the entire supply chain and been seen a lot more positively, China is building a lot of good AI with a minuscule of the compute power and it’s fine

      there was never a need to rush other than for capitals self interest to control

      if frontier labs wanna pace the frontier, they should be honest and return the compute back to the people

    • mattnewton 20 minutes ago
      I think this is related to the structures we’ve built in society. It’s lower friction to write software than just about any of the problems you listed.

      It reminds me of the Ezra Klein quote that “OpenAl would need permits to cover its parking lot in solar panels, but it can accelerate into recursive self-improvement, as best I can tell, whenever it so chooses."

      • numbasys 2 minutes ago
        Soon: The super-intelligent AI is telling us to cover parking lots and other flat roofs in solar panels.
    • mindwok 5 minutes ago
      Whatever benefits AI has to offer on a societal level are going to take a long long time to show up, give it some more time. We had smartphones in 2007, and e.g. By way of example, Uber didn't launch ridesharing until 2013, and then that still took years to seep into the public awareness.

      By comparison, we're in the early years of AI. We're all still figuring out what this thing even means.

    • hellohello2 3 minutes ago
      Publicly-accessible AI started working 1 year ago, 2 years tops. What did you expect? It to solve war driving prices up?
    • denkmoon 5 minutes ago
      [delayed]
    • emdash 13 minutes ago
      That pretty much sums up my experience. Everything is more expensive and worse at the same time. And it feels like we're being screwed over from every direction and in new ways every day. People keep talking about this tech is supposed to do so much but everything just gets worse.
    • raz32dust 18 minutes ago
      This view is too cynical. The frontier advancements are certainly going to trickle down. Take cars for example. AI driven cars are already here. AI-designed battery systems, fuel systems, materials and 3d printing capabilities will come to cars. It will likely lead to cheaper, more efficient and more effective vehicles. If will just take time for the impact to percolate.
      • Supermancho 10 minutes ago
        Worse, the view is largely phrased as gaslighting AI. Blaming the cause of particular grievances that coincidentally have been sourced at the same time as the rise of AI, is purposeful.

        > [Your train] never comes on time. Fuel is more expensive. The roads are still congested with traffic. Your kids are (probably) doing worse at school. Food costs more. Houses cost more. Rent is higher. Buying a computer or a PS5 is more expensive. Politics is still full of mentally unstable people. The environment is getting worse. You will still die of heart disease or cancer. Food quality is worse. Your kids can't find work. Wealth inequality is accelerating.

        None of this has to do with AI.

        > But hey...on the flip side...a lot of people are rapidly building software (that nobody is using), and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)

        Plain wrong on both counts.

        > Imagine what a "country of geniuses in a datacenter" will be able to do Apparently...nothing.

        Nonsequitor, probably based on a personal beef. Datacenters, at best, are understood to be for hosting more AI - not for people to somehow live in as an arcology of discovery.

    • oblio 16 minutes ago
      https://nitter.xitter.cc/augeeidos/status/210940046382122618...

      This reply summarizes the problem from a different angle.

    • sans_souse 12 minutes ago
      I would just add that consumer products in general aren't the same - practically every business model has incorporated obsoletionism to the point it's practically cancelled out the curve of progress and evolution of tech. We could have nice things.
    • dreamcompiler 18 minutes ago
      Correction: My car spies on me and its cruise control no longer works because it slams on the brakes in the middle of the highway whenever it sees a mirage. My dishwasher broke because it got an unasked-for software upgrade.

      Things worked better before we started putting bad software in everything. As a software engineer, I know what good software looks like. The fact that almost all software is now bad -- including the software that drives AI and the software that AI generates -- makes me want to flip the table. We know how to write better software and the general public should hang us if we continue to refuse to deliver it.

    • WillowWithAWand 6 minutes ago
      The problem, as with most things, is capitalism.

      As long as the world revolves around the central focus of making as much profit as possible as quickly and easily as possible the same cycle of enshittification will continue.

      The question asked by decisionmakers is never "How can we leverage this to make people's lives better?"

      It is always "How can we leverage this to make more money?"

      And those two aims are almost always in direct conflict.

      Companies are not trying to eliminate physical media because it is a benefit to people. They are doing it to drive growth for the sake of growth.

      • chanakya 0 minutes ago
        "It is not from the benevolence of the butcher, the brewer, or the baker that we expect our dinner, but from their regard for their own interest.

        When companies make profit, it doesn't mean we get screwed, nor even that they have a high profit margin. The grocery chain stores are a prime example. They typically make 2-3% net margins on revenue, and provide fresh, clean, reasonably well-tested groceries close to our home for that margin. Imagine trying to get equivalent quality fresh groceries ourselves from the producers, and I'm guessing it would cost at least three times as much.

    • TacticalCoder 15 minutes ago
      I'll add some...

      Videos are shittier. Most songs are even shittier than the autotuned pieces of shit already were. Ads are worse (even though nobody thought that was even possible). Customer support is worse.

      > But hey...on the flip side...a lot of people are rapidly building software (that nobody is using)

      That nobody is using and that are worse than what we used to have. And because nobody's using them, the countless issues plaguing them aren't even reported by users anymore (there are no users anymore). Full of bugs, usability issues and gigantic security holes (but which doesn't really matter because nobody's using them).

      It's not even clear that it's making long-standing software (like say Linux or Emacs or Blender or QEMU) better at all.

      Now I'd say though: most of the issues you mentioned have nothing to do with LLMs / AI. But it's very clear that AI ain't solving much of the world's problems and that the promises we got years ago didn't materialize.

    • scotty79 21 minutes ago
      You are talking like we had this capabilities for a decades but it's been just a few months of truly capable AI. Virtually all of IT switched to agents that quick. Non IT people just started to make software for themselves. Vision and translation capabilities of LLMs help millions. People talk with AI's now verbally to get advice or to just hear a friendly voice. The math thing is just few weeks old. We are just getting started. How fast did the telegraph meaningfuly changed lives for the better for bulk of people? Radio? TV? Computers? Internet? Cellphones?

      Has any technology made an impact on physical world faster?

      • LaurensBER 14 minutes ago
        Exactly, we're all software engineers here and even for us it took weeks/months to get to a position where we can really leverage LLMs (think platform, security, observability, etc).

        Imagine being in charge of a physical factory, sure you might have a 200 USD Claude subscription but you also have a floor full of physical machines that have, if you're lucky, an undocumented debug interface and a paper manual. Not to mention safety inspections, insurance, etc.

        You can't vibe code a factory floor. Not yet, perhaps in 5 - 10 years but there are some very hard, non technical, problems that need to be fixed before we'll see meaningful improvement there.

      • oblio 14 minutes ago
        > to just hear a friendly voice

        This is far from being proven as a net gain and there are massive lawsuits incoming related to AI recommending suicide to kids and others.

        • LaurensBER 8 minutes ago
          In all fairness, depression and suicide have existed before LLMs.

          Safety systems are not optimal but improving everyday and although I doubt we'll ever get a perfect safety system, it seems that we're not very far from a "good" safety system. The days of Google recommending glue on pizza and generating diverse Nazis seem behind us.

          Not to mention all the untold good that LLMs have done. Personally I've seen that LLMs are incredibly good at pointing conspiracy nuts and extremists into the right (moderate) direction.

          In a polarised world, we definitely need that.

  • rico735 1 minute ago
    Don't think the 0.2% really are in agreement that "large, complex projects that used to take them weeks/months can now be completed by agents with a prompt".

    The divide seems to be between people who think you may need to understand the codebase for maintenance, potential future token costs, optimisation, code erosion & being wary of the assumptions/mistakes that they have seen LLMs make, and people who don't review code and are confident that the AI has made no mistakes and will be available for next-to-nothing forever.

    It's certainly allowing the latter to move fast... but will their clients be happy if/when they break things. We'll find out soon, I'm sure.

  • kooi 38 minutes ago
    The 1%ers are in the vertical, but the question is vertical to where?

    It needs to a potential field with practical, economical, "real life" attraction well. I.e, robotics, real economic efficiency gains, manufacturing novelties.

    The worry is that the 1% is attracted towards a non-practical money hole. I.e: Token burn for the lols, sophisticated software systems that dont provide actual value outside of giving NVIDIA cash.

    • freecodeio 29 minutes ago
      he's just trying to say we're in a cool kids club and you ain't in it, without specifying what the cool is about
  • comeonbro 53 minutes ago
    I would propose another mechanism: even the free-tier models have already completely saturated what most people are capable of appreciating.
    • oh_my_goodness 2 minutes ago
      Google's free AI responses are moving the bar on what I expect from search. But maybe not in the direction you think.
    • gammarator 49 minutes ago
      Or maybe needing.
    • DebtDeflation 34 minutes ago
      Honestly, outside of coding tasks, the AI Summary at the top of every Google search is adequate for 99% of what I need and I hardly even use ChatGPT any more.
      • throwuxiytayq 17 minutes ago
        God damn. I respect you for being honest, but… god damn. That’s a pretty low bar.
        • antonvs 5 minutes ago
          Just in the last few days I had an example where the Google Search version of Gemini had a much better and more comprehensive answer, by far, than ChatGPT. It had to do with Amazon’s use of Data Matrix barcodes on shopping bags. Gemini was able to fully describe how they’re used in Amazon’s logistics chain. ChatGPT essentially said it didn’t know. Your preconceptions may not be accurate.
  • ilovecake1984 39 minutes ago
    I’ll say this until I am blue one the face. Nerds (software dev, maths etc) see how good LLMs are at things they care about and assume they will be broadly applicable in future.

    There’s no reason to think this.

    • howunfortunate 27 minutes ago
      > There’s no reason to think this.

      There are lots of good reasons to think this.

      The first is that LLMs used to be bad at each of these nerd things, then toppled them like dominos. There's a pattern over time.

      Another reason to think this is g. Across every known measurement, human (and animal) intelligence is convergent. Being better at one thing correlates with being better at another thing. Although the reason is not perfectly clear, the pattern is well established, and seems to apply to LLMs too - GPT-6 is smarter than GPT-3 at everything, not just math. There's no reason to expect different for GPT-9.

      • throwawayy6767 9 minutes ago
        That's just cocktail party level of thinking. LLMs are getting better at math, code and logic (and marginally better at science and general knowledge) because these are domains that can be objectively verified and thus there is a potentially infinite supply of 'facts' to generate and train on. These are very powerful but ultimately very abstract domains. For everything else the messy real world and its physical bottlenecks gets in the way and there's little reason to expect progress to accelerate. It still takes months to get mice to reproduce and run experiments on, no matter how knowledgeable about biology the models have become.

        If anything, in some domains frontier models have become worse - claudisms and chatgpt idioms are making them notoriously bad at prose without a considerable amount of prompting and tweaking.

      • flecomet 16 minutes ago
        I don't agree that they get uniformly better at everything, especially subjective skills.

        For example, they have become more and more unintelligible when you ask for explanations or descriptive text. They assume you see the same context as them and shortcut explanations.

        I've had to craft a skill to get them to produce remotely understandable explanations of even mildly complex/non-mainstream subjects.

        This may be Curse of Knowledge https://en.wikipedia.org/wiki/Curse_of_knowledge on their part, and it also impacts human experts but still. Becoming better at one thing does not mean you become better at everything else, although I do agree with you that the better they become the more things there are they become good at, but their ability is still quite jagged and maybe increasingly so.

      • oblio 12 minutes ago
        > Across every known measurement, human (and animal) intelligence is convergent. Being better at one thing correlates with being better at another thing.

        LOL. Human (and animal!) intelligence is notoriously jagged. We're all basically idiots except for very narrow areas where we focus.

        • howunfortunate 7 minutes ago
          This is simply untrue. I tried to think of an animal example so I could give a concession, and there just...aren't any.

          Smarter just means smarter.

    • mindwok 3 minutes ago
      Just like the nerds thought about computers, or the internet, or video games, or smartphones, or crypto, or... whatever.

      The most passionate, intelligent people are usually a decent indicator of where culture is going to go.

    • aleph_minus_one 7 minutes ago
      > Nerds (software dev, maths etc) see how good LLMs are at things they care about and assume they will be broadly applicable in future.

      I don't think this explanation is sufficient:

      For example, many nerds care about 3D printing. On the other hand, my experiments (and the experiments of many nerds who I know) to let LLMs create files for 3D printing lead to horrendously lacking results.

      Or many nerds care about linguistics or puns. Whenever a new LLM comes out to which I have easy access, I do the test, and let it explain some specific German jokes to me that are based on convoluted puns in the German grammar. Until now, no LLM that I had access to could give satisfying explanations of the puns on which these jokes are based.

      Or even for coding (many nerds do care about elegant, sophisticated code): the LLMs that I could test were already overchallenged with the following task: I had written some code that is a very "artisanal", "clever" improvement of an algorithm over the version that one would find in a textbook. The task for the LLM was simply to write some code comments/documentation about the mathematical ideas upon which my improvement over the textbook version of the algorithm is based. It wasn't capable to do this. On the other hand, for a junior programmer, I would in such a situation expect that he goes through every single line and thinks through the algorithmic ideas so that he can learn from them.

      ---

      Thus: even for "nerdy" topics, it is very easy to find tasks where LLMs still suck. So, my hypothesis about the nerds that you mention in your post is that these nerds are rather people who want to believe (with religious fervor) that the current LLMs are exceptionally good instead of just looking into a slightly different direction than where the tech billionaires want the users of LLMs to look at.

    • majkinetor 36 minutes ago
      There is every reason to think this. Its about available quality data. The data that was easiest to fetch was already there, then we got some more data by asking experts to create datasets for post training. Once this is over, we will get to the outer world that didn't get to hoard it for bots to take it. This will certainly change. Put on a smart glasses and record what you do to fix a pipe. In 3-5 years, rinse and repeat.
      • RealityVoid 26 minutes ago
        Maybe, I'm at the point I don't know what to think anymore. But somehow, this feels still non-human. Humans don't need millions of hours of training and millions of samples to learn how do to something. We have proof that systems can deal with low-shot training. So why can't these? My point is when we have human level learning ability, then we're moving, if we expect expansion of capabilities to come from data alone, I'm prepared for disappointment.
    • bigstrat2003 8 minutes ago
      > Nerds (software dev, maths etc) see how good LLMs are at things they care about and assume they will be broadly applicable in future.

      It's worse than that. I am one of those nerds (a programmer), and I see that LLMs are shit at the things I care about. I have zero reason to believe that they will be good at other things either.

  • theturtletalks 14 minutes ago
    Claude Code accelerated this divide. Programmers using Claude Code realized that giving LLMs access to a terminal made them feel 100 times smarter. Access to Bash, CLIs, and the ability to visit websites without the many restrictions made them powerful.

    Meanwhile, the average user was still using ChatGPT, which probably felt like it had plateaued over the last few releases. The real power of these models became apparent when paired with a terminal. Agents like Muse and OpenClaw are now attempting to bring that same Claude Code-like power to everyday users through a simple chat UI.

  • prometheus1992 1 minute ago
    so what was karpathy expecting?? everyone will give up on their life and start feeding the monster in the screen? that's literally what the tiny sliver is doing; they are food for the AI model.
  • rbehrends 4 minutes ago
    > Somewhere around 20M people (0.2%) see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt.

    Really? That's not my experience. In fact, I've experimented quite a bit with spec-driven design where the LLM does not just get a prompt, but an (informal) spec, often with design considerations included. And yet, regardless of the power of the model, I've never seen them autonomously deliver what I'd consider a finished product.

    There are always edge cases that are not properly considered, architectural oopsies, performance issues, duplicated code, and other problems, that then need to be identified and fixed, unless it's throwaway software (e.g. a one-off script or some quick experiment).

    This is not to say that agents aren't extremely powerful (as I think they are, they do often blow my mind) but the prompt-and-forget approach is IME not a productive use of them for software that is meant to last. This may change in the future, but for now, fully or largely autonomous agentic work does not seem to be the road to quality.

  • AvAn12 44 minutes ago
    Fair assessment. Maybe the messaging should focus on “these are great accelerators for software developers” rather than “AI will change everything for everyone everywhere…” It is understandable that non-technical folks are kind of underwhelmed - not due to lack of understanding so much as lack of a tangible need. Not everyone needs an electron microscope or gas chromatograph…
  • variety8675 11 minutes ago
    A more cynical view is, trust us we've got really good stuff internally you're not allowed to see, but give us more money
  • m101 57 minutes ago
    My interpretation of this is something like: if LLMs are to be mega useful token counts need to increase by many orders of magnitude -> broad adoption (and spending) would require token costs to drop by many orders of magnitude -> before the common folk get mega useful tools existing GPUs will be worthless
  • anukin 48 minutes ago
    Tbh building an agent swarm and the coordination layer is not exactly frontier level. They don’t achieve any meaningful outcome rather than producing pr puff pieces. Hacking huggingface and Australian govt etc is very much possible with a team of humans and agents and does not need agent swarms. The cost is also lower.
  • wg0 18 minutes ago
    That's fine but these companies having billions in funding by now should have rewritten chalk and other terminal libraries in Rust while bundling the whole thing in Rust as a single static binary. I'm talking about horrible Claude Code and friends.

    And then using emgui they could expose full fidelity IDE+Agent workflow interface (like DeepSeek Harness and similar) that's also exposed as MCP to be operated by another agent and JSON RPC for over the network but no. Skill issue?

    PS: Even full fidelity Photoshop is possible with such high fidelity UI that has CLI, JSON API and MCP server all in the same single binary < 90 MB. For reference, see PhotoCraft or VectorCraft or WordCraft or PdfCraft.

  • tripleee 39 minutes ago
    He's intentionally forgoing all nuance in order to make this sound dramatic

    > Somewhere around 20M people (0.2%) see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt.

    No, they can't, at least not any semblance of quality. The cases we're seeing where this does kinda work is in ports and translation where all the rules are already documented in the best specification language possible with a way for the LLM to verify itself: code. We saw this close to a year ago now with Cloudflare and NextJS

    > The impact scales with ambition, problem size, and horizon. A question with a paragraph answer barely stresses the system. You need a reservoir of big, difficult problems that you really care about

    These are operating on different capabilities - AI's ability to answer informational queries as a chatbot frankly sucks and can't be trusted without verifying it. I run up against this every day. A problem with a verifiable answer on the other hand it's very good at solving. He knows this (his next paragraph) but he's putting them on the same scale of "stressing the system" to attempt to add proof to his introductory claim

    And then there's the completely unverifiable scare that there are internal frontier models way beyond anything we've seen "swarms of thousands of agents collaborating over weeks on software mega projects: minting zero days, running cyber attacks and defenses at machine speeds, discovering new science, advancing the frontier of mathematics"

    I dunno. I haven't been able to set up OpenAIs remote codex connection, their shit is buggy as hell and the web UI keeps crashing and making messages disappear. Is this what their internal superhuman "Things that would have taken top professionals in the industry years of work" looks like? Granted Claude has been really smooth, but still..

    • jfrbfbreudh 32 minutes ago
      Congrats, you’ve discovered that you are not part of this group.
      • bigstrat2003 5 minutes ago
        That group does not exist. It's all hype, zero substance.
  • skybrian 43 minutes ago
    > see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt

    Really? I start with a conversation for maybe 5 turns or so, where I ask it what would need to change, what the API might be, any database schema changes, URL schemes, and so on, and finally ask it to break it down into commits. Then I let it go, implementing 3-10 commits at a time via subagents. It usually gets the UI somewhat wrong, so there are followups to fix it. This is with Sol and Luna subagents.

    Is that what other people see?

    • sumedh 23 minutes ago
      Try using Astra, Fable, Opus.
      • derwiki 4 minutes ago
        Sol 6.1 is really good too though
  • skippyboxedhero 46 minutes ago
    Text generation is not the bottleneck. Does everyone work for Accenture and TCS?
  • socializer 34 minutes ago
    I think it's a weird take because it implies that the 6 billion people he's talking about actually have some interest in knowing about the capabilities of LLMs to solve frontier math? This is simply not something they care about or can evaluate. There are maybe several thousand people in the world who have some (abstract and barely-monetizable) use for this information, plus probably another 100,000 who don't understand any of the math, but like to cheer on.

    We have already reached "peak LLM" in terms of what normal people realistically need it to know or reason about. In fact, I'd say we reached that point about 1.5-2 years ago. There are two other barriers that remain unsolved:

    1. They're less dependable than humans and can't be meaningfully punished or forced to make up for mistakes, so you can't really replace humans with them, not without having a human babysit.

    2. Most people don't really have a special need for an LLM in their life. They may like that it answers questions or helps you polish a resume or, I guess the labs' favorite, helps you make restaurant reservations. But this sure isn't worth $200/mo for most people. Probably not even worth $5/mo.

    It'd be kinda funny if we create superhuman AGI and then no one has any real use for it, perhaps except for military murder-bots. There's always market for that.

    • HolyLampshade 26 minutes ago
      I know I’m bordering on repetitive and strongly negative, but the combination of the two barriers you mentioned has all but halted all of my interaction with agents.

      Even ones that cannot be easily dismissed (like the agent at the top of Google search) are often subtly wrong in a way that requires significantly more effort on my end to parse the loquacious output to determine where the inconsistencies are.

      Look, can they be useful to generate the html or jscript for humanity’s 9 billionth iteration of a web form? Sure. But man, I wish sanity had reigned and people had applied ML models in general to more substantive and beneficial projects.

      (Before anyone chimes in, I’m familiar with implementations of ML applied to esoteric domains; but by their very nature these don’t get all the news cycles, or hiring, or any of the other absolute insanity that the domain seems to contain)

    • iugtmkbdfil834 14 minutes ago
      This is by the weirdest possible opinion on llms as a whole. Not the AI booster club, not the AI doomer, not the undecided, but somehow "I can't think of anything to use it for." It is like watching the beginning of the internet in 90s and saying there is nothing to browse. The whole concept is just weird to me. At least the first 3, I can conceptually understand. It is admittedly hard for me to understand someone saying 'i tried this and best it can do is summarize emails'; I am not saying it is not true ( my boss said that exact thing ), but it is hard for me to reconcile with my view of the world.
  • andy99 45 minutes ago
    > Meanwhile, human review and comprehension are starting to fall behind.

    I think LLMs are valuable and spend most of my professional life working with them.

    I do wonder though whether there’s a Ponzi scheme aspect here where as long as the “frontier” can keep outrunning human review and comprehension, LLMs are always going to looked way more valuable than they are and the bubble will continue.

    This started with deep learning, expectations weren’t met and people started looking for value, then GPT came out and people got wooed again and forgot, then coding, then math, cyber, etc. As long as the dust doesn’t settle we never have to reflect on all the shortcomings and can just stare mesmerized at demos.

  • hn_throwaway_99 25 minutes ago
    The fundamental question I have that honestly I haven't been able to find any answers for: With all the talk of LLM-based AI capabilities "going vertical", what evidence is there (for or against) that LLM-based approaches won't eventually "hit a wall", that is find some aspect of intelligence where humans will still have primacy, and no amount of scaling will change that.

    E.g. some prominent folks definitely think LLM-based approaches will hit a wall, perhaps LeCun most notably. I also read a guest post on Terry Tao's site (which I really liked) that argues that, for all the very impressive recent AI results, they still operate within the "convex hull" of their training data: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t...

    I'm just curious if there is any actual data or evidence that takeoff (i.e. RSI, "the singularity", whatever you want to call it) is inevitable with current approaches.

  • ashleyn 1 hour ago
    >Meanwhile, human review and comprehension are starting to fall behind. For example, people are still involved in the "archeology" of the OpenAI-HF incident from many months ago. Mathematicians may be poring over the 722 manuscripts on frontier mathematics for a while.

    Amid all the discussion of sigmoid curves, and where the "LLM wall" will materialise, I think few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output.

    What I fear is that people simply eschew human review altogether, considering we're talking about the industry that came up with the "move fast and break things" credo. Human review of LLM-produced code where I work is already a farce, and we're not special enough to be one of Karpathy's 5,000. I do my best to manually review anything that's my responsibility, but I'm literally one of very few people left working on my team, so in practice what happens is I submit PRs that are at best glossed over by completely unrelated teams for security, malware/prompt injection, and other serious concerns. Quality insofar as vetting others' code has completely gone out the window and it shows in the number of bug reports that come back, often themselves written in Claudease. This is all on top of everyone cynically phoning it in in the first place, due to the omnipresent sword of Damocles that is additional AI-driven layoffs.

    Worse yet all the incentives point to this being the most economically viable thing individual companies can do. I think it goes without saying some type of regulation here is urgently needed, and that an unexpected cause of an AI bubble pop may end up being that humans simply aren't able to keep up with the pace of the output - leading either to precautionary plateauing of capability, or major liability risks related to a decline in quality.

    • michaelchisari 1 hour ago
      | few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output

      That was the dominant concern in the circles I’m in, so it’s worrisome it’s being treated as rare.

      • JBits 33 minutes ago
        I would question the narrative that humans lack the capacity to verify the output and would instead argue the people lack the incentive to verify the output.

        The response of many mathematicians to the recent dump is a good example: verifying these proofs amounts to unpaid labour for OpenAI and wastes time that could be spent doing publishable work which ultimately results in money or personal success. The slop factor also compounds the work required to verify the output considerably.

        For mathematicians, programmers or anyone, if the work required to deal with slop passes the limit, it is no longer in their own self interest to use LLMs. The expectation that people will use LLMs for the betterment of humanity against their own financial interest is baffling.

      • skydhash 43 minutes ago
        Humans are not immortal and cannot spend all their time into review (especially unpaid). Even today, there’s so much knowledge around that you have to be specialist of a narrow domain to get to the frontier. Even in computing which is just approaching a century of existence.
  • freecodeio 35 minutes ago
    I don't understand how "swarms of thousands of agents collaborating over weeks on software mega projects" works with the current context limits and at this point I'm too afraid to ask cause I'm afraid an AI bro is gonna punch me.
    • riffraff 18 minutes ago
      You can have agents make a plan and then other agents spec tickets and yet more agents do implementation and yet more do reviews and QA and whatever. It's possible.

      Is it effective? I'm not sure, cause if it was I don't see why OpenAI & co are not releasing such a project instead of demos.

    • dude250711 19 minutes ago
      Your product needs to be either the swarm itself or the PR about the swarm - not the outcome.
  • chevman 31 minutes ago
    I mean in late 2020/early 2021, Altman and others were saying the end of work was 6 months out.

    That clearly didn't happen :)

    • Legend2440 8 minutes ago
      Altman didn't say that. In March 2021, he did make some predictions, but they were much vaguer and farther out:

      > In the next decade, they will do assembly-line work and maybe even become companions. And in the decades after that, they will do almost everything, including making new scientific discoveries that will expand our concept of “everything.”

      https://moores.samaltman.com/

  • dude250711 17 minutes ago
    Am I the only one who thinks "yes, it is currently overhyped, but no, it will get there soon"?
  • notjes 22 minutes ago
    [flagged]
  • weinzierl 1 hour ago
    Most people see a clumsy chatbot, most professionals see modest gains, and a tiny group is watching the curve go vertical, all at once.

    The future is already here. It's just not very evenly distributed.

    • atmavatar 59 minutes ago
      > a tiny group is watching the curve go vertical

      Caveat: that same tiny group is employed by the AI vendors, meaning it's in their financial best interest to make it sound like the curve is going vertical.

      • pianopatrick 20 minutes ago
        Also, part of "seeing the curve go vertical" is that the projects that show those results are rather expensive. If you work at one of the AI companies or certain well funded customer companies then you can pay for the tokens. But, like it would make no sense for me personally to spin up a bunch of agents to try to write a browser or solve math problems.
      • Arkhaine_kupo 49 minutes ago
        Down is a perfectly valid direction for a vertical line when not given a ± in the vector.

        Considering the investment in AI, the lack of moat, and the increased inability of any of the big players to come even close to profitability (with OpenAI already breaking the "ads" emergency glass option)... perhaps he meant a tiny group is already seeing the line crater

      • majkinetor 33 minutes ago
        It can be both and it is.
      • njovin 51 minutes ago
        Another caveat: many of that same group seem to have a shared delusion that they’re birthing a super intelligence, and those are the same ones claiming the vertical curve.
        • sumedh 26 minutes ago
          > and those are the same ones claiming the vertical curve.

          Do you still write code by hand?

      • kmac_ 34 minutes ago
        I work for a typical software product company, and along with most of my colleagues, I clearly see that the curve is so steep that our software development process has already changed tremendously and will be different next year, and probably completely different in the following years. The revolution is real, undeniable, and the old days are gone. Some companies adapt to changes more slowly, some faster. LLMs, agents and harnesses are just a part of the bigger picture.
      • nullpoint420 37 minutes ago
        I hate to say it but this is cope. I believed this in the past but it's over.

        AI models can dismantle billion dollar industries. They can reverse engineer Adobe and Microsoft products that once were their moats and titans of their industry.

        Why do you think they'd need to lie?

        • riffraff 23 minutes ago
          Because if they can, why haven't they?

          LLMs are great and will get better but we have not seen an open source office suite built by LLMs come online yet, and even if there was one I'm pretty sure few would adopt it, cause that's not the only moat, LibreOffice has been around for decades.

    • msy 25 minutes ago
      The curve go vertical for what, precisely? For all the chest-puffing and ominous and cryptic comments from 'insiders' about these incredible capabilities every time they put something public it turns out to be a pale shadow of what was trumpeted. These are powerful useful tools but the quasi-cultish behaviour around them is getting old.
    • albatross79 47 minutes ago
      Congratulations, you've parroted something said by someone else.
    • lifeisloving 1 hour ago
      I use models all day everyday, have unlimited access to all models. The curve is not going "verticle". I have all the workflows and meta agentic tooling, im not holding it wrong. Its bad, not everything is a 20th percentile problem.

      There is in fact no indication of this, not evem the precious benchmaxxed benchmarks ya'll love to reference.

      There is however a exponential curve of slop, and an ever increasing number of people who's minds are completely captured by these things.

      • tkz1312 39 minutes ago
        As someone who has done software verification professionally for many years the last 6 months or so have looked extremely vertical. The robots are better proof authors than I probably ever could be even if I dedicated the rest of my days to the practice, and projects that once would have taken months now take a day or two.
        • lifeisloving 17 minutes ago
          Then you'd know well that 2hr of LLM code generation can easily be about 4-8hrs of review, and that review can be brutal.

          I'm not arguing that they cant write code, or write a proof. Its just not written or designed well and is absolutely brutal and soul crushing to work with. Look at these proofs they're producing also, they're millions of lines of Lean that are impossible to reason about.

          The way we're using the term 'verticle' to describe a curve means we're not being honest about this. This curve can actually be plotted, you can go look at the curve. It is not in fact 'verticle'. Each model release is climbing single digits on benchmarks it was overfit for.

          • tkz1312 6 minutes ago
            You don't need to review proof code.

            In the last 9 months or so llms have gone from just another useful proof tactic (like grind or sledgehammer or sat solvers) to being so good at writing proofs that I don't even bother to try myself anymore.

        • gr_norm 35 minutes ago
          I believe this, but it is also a unique case where the pitfalls of LLMs (producing weird errors that a human wouldn't) are zeroed out. Since you have a proof checker that tells you if the LLM did it right.
          • tkz1312 4 minutes ago
            I'm pretty convinced most serious software will have some kind of proof system inside within the next few years.
  • AdeptusAquinas 24 minutes ago
    "running cyber attacks and defenses at machine speeds"; worth noting that LLMs (even the vaunted frontier models) are exponentially slower at cyberattacks or defenses than your average WannaCry or Splunk automation from decades ago. Its this sort of delusion world the AI bros live in that is part of the reason there is a disconnect between what they think should be happening and what actually is.
  • mccoyb 1 hour ago
    If the software coming out of OpenAI and Anthropic is what we have to judge, I wonder about the 5000 ...

    Let's say, for the sake of argument, that the models are some multiplicative factor better on the inside.

    Doesn't that mean the demos should work?

    • j2kun 1 hour ago
      Unfortunately, marketing, hype, and venture capital overshadows any serious public discussion of capabilities.
      • AndrewKemendo 51 minutes ago
        Be the change you want to see in the world: Attend or host an AGI society event to have that conversation:

        agi-society.org

    • LastTrain 53 minutes ago
      It’s like the aliens paradox. If AI can build killer software already, where is it?
      • equinumerous 45 minutes ago
        Couldn't agree more. I find a new bug in the VSCode Codex extension every day... quantity != quality!
        • agentdev001 19 minutes ago
          I have a feeling that the team working on the VSCode Codex extension is one guy, who begrudgingly takes down a couple tickets every other week.
      • shermantanktop 32 minutes ago
        One solution to the aliens paradox is that they are so advanced they can hide from us. Maybe the killer software is kept inside the labs? I doubt it.
    • spiderice 1 hour ago
      I'm confused.. are you suggesting that Claude Code / Codex don't work? Because if you're still saying that in October 2026, it's a you problem. You're doing something wrong.
      • mccoyb 54 minutes ago
        No, I’m talking about the recent DevDay.

        Also, yes there are still bugs in Claude Code. I experience them nearly everyday.

        It is markedly better than early days, but still not the best harness.

        The best software written with agents seems to come from people outside of the labs (see pi, for instance — or all of cloudflare’s recent work)

        Which makes me question either the model, or the holders …

        • nullpoint420 38 minutes ago
          Cloudflare is where you lost me. I don't know anyone actually using them other than for their proxy, DNS servers, or DDOS protection.
          • mccoyb 34 minutes ago
            See, for instance: https://news.ycombinator.com/item?id=49182996 by Kenton Varda

            Also, not mentioned in my post:

            - Mitchell Hashimoto

            - Prime Intellect (and all their agent experiments)

            - Geoff Huntley (see Jiti, for instance)

            (many more)

            There's a ton of interesting software being developed with these models by people outside the in group, but I find most of the software from these big guys to be ... bland. Buggy copies of copies.

            • nullpoint420 25 minutes ago
              I don't see the Cloudflare stuff as groundbreaking, let alone people building companies around it. "Workers" are the wrong primitive, IMO. What if my code isn't in Javascript? They sandboxed the wrong part of the machine.

              I have a version of "Cloudflare OS" running in production, using Temporal as the orchestration and MicroVMs for the sandboxing.

              Sorry, I get your point but personally I don't get the hype around them.

      • lillesvin 14 minutes ago
        I don't use either myself but I'm in a position where I get to see what it produces for other people. Here's a recent (2-3 months old) example: A zsh script that immediately invoked a Python interpreter that immediately invoked a shell command and iterated over each line of output looking for one of a few strings to match... So, `<cmd> | grep "(stringA|stringB|stringC)"`, but a whole lot dumber.

        Thanks, but no thanks, I don't need that kind of code in anything I'm working with.

      • plorkyeran 24 minutes ago
        Claude code is an incredibly buggy mess. I have never used any other TUI that regularly has rendering errors or that is anywhere as laggy as Claude.
      • well_ackshually 30 minutes ago
        Claude Code still doesn't have a working scroll back buffer in their new renderer, it regularly flips its shit and mixes different pieces of history.

        Claude Code still can't get reasonable performance without writing a "game renderer" (that doesn't work)

        Claude Code is still written in JavaScript, eating hundreds of megabytes to make a shitty TUI whose literal sole role is to send API calls.

        Claude Code is software made by amateurs.

  • rolosa 42 minutes ago
    [flagged]
  • kydanet 1 hour ago
    [flagged]
  • 8484848484 1 hour ago
    [dead]
  • kittikitti 1 hour ago
    [flagged]