Abundant Intelligence, Scarce Judgment: a woman draws back a curtain to reveal a mechanical parrot and servers.

Just some notes. Let’s get the obvious out of the way. LLM tech is truly incredible. It’s a game changer much like the internet was. It’s unlocked a new level of capability and creativity for people everywhere. It’s changing the paradigm of labor including the structure and nature of knowledge work.

Ok with that said, there are some real frictions and constraints. It has no realistic model of the world. It is still probabilistic with strange failure modes and lack of reasoning. I can already hear the AI acolytes defending their new god, “but human work was bad too, humans generated lots of slop software and code”. Yes that’s true but we’re not playing the game about whataboutism. We’re evaluating the utility of the tech vs hype vs cost. Well cost is actually a part of the utility function but I’ll leave it separate for now.

The truth is if you use these tools within a domain you actually understand you recognize the flaws immediately whenever you take a closer look. Much like the Wizard of Oz, people are getting tricked by the stochastic parrot when they peer behind the curtain. There are real limitations to its ability to translate your intent into reality and by reality I mean a functioning, robust, maintainable codebase.

“SlopCodeBench was introduced in 2026 to measure how AI coding agents affect a codebase over many rounds of changes. Instead of testing one isolated task, it has agents repeatedly modify and extend their own previous work. Researchers found that structural erosion increased in 77% of trajectories, verbosity increased in 75.5%, and agent-generated code was about 2.3× more verbose than comparable open-source human code (Orlanski et al., 2026). Human-maintained repositories were comparatively stable, while agent-generated code tended to degrade over time. The key takeaway is that a codebase can keep passing tests while still becoming harder to maintain, understand, and extend.Reference: Orlanski, G., Roy, D., Yun, A., Shin, C., Gu, A., Ge, A., Adila, D., Roberts, N., Sala, F., & Albarghouthi, A. (2026). SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks. arXiv:2603.24755v2 (revised May 7, 2026). https://arxiv.org/abs/2603.24755v2.“This is the compound effect making software worse over time without proper intervention and management.

So the question is: why have these limitations been relatively easy to ignore? Part of the answer is that these weaknesses are being papered over by some of the most generous technological subsidies we have ever seen. When the solution to an unreliable model is often to use more tokens, use a bigger model, run more agents, or try again, cheap inference can hide architectural weaknesses that become much more obvious when compute is constrained.

However, we seem to be closing in on the limits of this, as great models are consistently degraded with increased time, usage and adoption. What was once difficult to discern, AGI or stochastic parrot, slides down the spectrum towards stochastic parrot whenever the frontier labs choose to do so.

The next point is nuanced and I need more data to be certain. Quoted token prices are decreasing overall. More “intelligence” is available to almost everyone. This access is being driven by Chinese open weight models. The labs there are running a classic strategy of dumping LLM access on the market cheaply. They gain market share while forcing competing labs in the private space to operate at a loss.

However at the frontier of intelligence, from the user perspective costs are increasing as the usage per dollar is declining. Subscriptions are buying less and less usage over time. The companies try to paper over it by giving free resets to disguise the corresponding service reductions. OpenAI as an example cut the $200 max 20x usage in half to 10x. Non-trivial work is becoming more expensive and users are having to do more and more cost optimization to maximize access to capable models.

If you have infinite tokens then the complaints of the peasants sound like foolishness or just rumors and thus the disconnect between the new releases and user feedback.

Regardless, this situation creates interesting opportunities for us humans. There will be increasing demand for people who can utilize these tools efficiently. To do so requires developing domain knowledge. LLMs can accelerate learning when used with intention. Otherwise constant delegation to the LLM will cause one’s skills to atrophy.

This is where I think the opportunity lies. The advantages accrue to the person who understands the application well enough to know what to delegate, what to verify, what requires judgment and where the model’s output is likely to fail. Real leverage comes from the critical thinkers using LLMs to compress the time required to develop competence and then leveraging the speed and scale of this technology. Prompt monkeys hacking at the keyboard will instead look at LLMs as a substitute for competence which creates a toxic cycle of dependence.

Some people will still be required to walk the hard path, to understand the domain in which they seek to impact. They will have to learn and develop intuition, skill, and judgment at some point even if the LLM can generate an answer instantly. If one does not do this, they will eventually fall prey to one of the LLM’s most dangerous failure modes: substituting a correct-looking answer, or correct-looking implementation, for what is truly correct.

The strange outcome of this AI revolution seems to be that as intelligence becomes abundant, judgment becomes more scarce.

Best,

Brian Christopher, CFA