{"article":{"slug":"what-would-a-serious-ai-product-look-like","title":"What Would A Serious AI Product Look Like?","subtitle":null,"summary":"One of the issues that I have with the current generation of “AI” products is that they do not appear to take their own premises seriously. I look at a plethora of obsequious chatbots claiming to be serious tools for problem solving, and I think, this is not what a problem-solving tool would look like. Even before we get to the tremendous ethical problems with the frontier labs, it is this impression of their composition *as a product* that makes me feel, constantly, whenever I am interacting…","content_type":"essay","language":"en","canonical_url":"https://blog.glyph.im/2026/09/serious-ai-product.html","author":{"name":"Glyph Lefkowitz","url":"https://blog.glyph.im/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Deciphering Glyph","url":"https://blog.glyph.im/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Opinion","slug":"opinion","url":"https://listedarticles.com/topics/opinion"},{"name":"Software Engineering","slug":"software-engineering","url":"https://listedarticles.com/topics/software-engineering"},{"name":"User Experience","slug":"user-experience","url":"https://listedarticles.com/topics/user-experience"},{"name":"AI Safety","slug":"ai-safety","url":"https://listedarticles.com/topics/ai-safety"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":5119,"reading_minutes":22,"published_at":"2026-09-27T12:00:00.000Z","added_at":"2026-09-28T18:15:08.561Z","updated_at":"2026-09-28T18:15:08.561Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/what-would-a-serious-ai-product-look-like","markdown_url":"https://listedarticles.com/articles/what-would-a-serious-ai-product-look-like.md","example":false,"citation":"Glyph Lefkowitz, Deciphering Glyph. \"What Would A Serious AI Product Look Like?.\" 27 Sept 2026. https://blog.glyph.im/2026/09/serious-ai-product.html (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://blog.glyph.im/2026/09/serious-ai-product.html"},"body_markdown":"One of the issues that I have with the current generation of “AI” products is that they do not appear to take their own premises seriously. I look at a plethora of obsequious chatbots claiming to be serious tools for problem solving, and I think, this is not what a problem-solving tool would look like.\n\nEven before we get to the tremendous ethical problems with the frontier labs,\nit is this impression of their composition *as a product* that makes me feel,\nconstantly, whenever I am interacting with them, that they are less a software\nproduct than that they are a grift, a scam designed to make me feel like I am\ninteracting with a product that has capabilities that it simply does not, to\ntry to lull me into a false sense of security that I can trust it.\n\nThe frontier labs are of course the worst offenders, but every criticism here\napplies just as much to Ollama, which (if anything, due to the obviously poorer\nquality of the available models themselves) needs these features *even more*\nthan the frontier labs do.\n\nHere, I will set down a few features that might convince me that an LLM-based product, particularly one focused on research or software development, was actually serious about helping me do useful things with it.\n\n## Make “Checking For Mistakes” A First-Class Feature\n\nThis is the biggest issue, and the major reason that I was inspired to write this post.\n\nIt is a truth universally acknowledged, that AIs cannot reliably provide information.\n\nI could cite a ton of news articles and studies about this fact, but there is\nno need.  Every single chatbot admits this, up front, in a fine-print\ndisclaimer as a core part of their user interface.  Gemini says “AI can make\nmistakes, so double-check responses”, Claude says “Claude is AI and can make\nmistakes. Please double-check responses.<sup>1</sup>” ChatGPT says “ChatGPT can make\nmistakes. Check important info.”.\n\nEvery time I see that last one, I wonder how I’m supposed to know what “info” is supposed to be “important”.\n\nAll of these warnings are all small, gray text, painfully obviously included as\nlegalese to push responsibility back onto the user rather than to help with\nanything.  This is a *core limitation* of all these products.  Checking their\noutput is a part of the workflow for using them that:\n\n1. you *absolutely cannot skip or skimp on* without creating risks to yourself\n   and whoever you are conveying its output to, and,\n2. it is *very easy to skip or skimp on* and you are encouraged at every turn\n   to do so, because “just trust the output” is one of the quickest ways to\n   save time.\n\nA chatbot product that took this weakness seriously, as an actual consideration\nfor using it, would put a checkbox next to *every claim* in its output.  It\nwould be a 2-column worksheet, where you’ve got the LLM output in the first\ncolumn, and next to it, human notes in the second column, explaining what work\nwent into checking this claim, and a *big* checkbox that you would only check\noff after you believe you’d checked its claims thoroughly enough.\n\nCoding assistants would need to have some version of this as well.  Right now,\nthis is pushed off into code review, which means it is a dark pattern which\nsubtly encourages the “author”<sup>2</sup> to offload this work to their code reviewer\nwithout ever looking.  Once again, “it’s probably fine, I don’t need to check”\nis the quickest way to save time and churn out those PRs faster.\n\nIt might even be useful for coding harnesses to have some affordance for\nchecking code before it even runs tests.  As the vendors themselves have\nadmitted,\nit’s not just expensive to burn tokens on your “AI”, you also end up burning\nfar more *compute* on the AI.  Being able to check your diffs before sending\nthem over to uselessly exhaust your testing compute cluster would be useful.\n\nIf your product tells me that it makes mistakes and I must be the one to check\nfor the mistakes, but then gives me *zero tools* to check for mistakes, I\ncannot take it seriously.\n\n### More Citations to Check, And More Details\n\nMost chatbots prefer to give an answer, rather than a citation. In my own personal use, I find that when asked to provide a list of citations with clearly marked sources for each one, they will appear to “get bored” halfway through the list and simply stop including citations at some point.\n\nWhen the bots include citations at all, present them as inline annotations that say nothing but the domain name of the search result, in a font so small that it’s barely legible, and an equally indecipherable icon that is fewer than 16 pixels on a side.\n\nThis is backwards.\n\nNow, I am aware that these citations do come from somewhere, and in an attempt\nto reduce hallucinations, all of the major providers support some form of\n“grounding”<sup>3</sup>, and that those little barely-readable citation links are\nreferencing actual structures in the\nRAG pipeline\nand not just potentially-hallucinated tokens, but I’m not talking about the\nunderlying machinery in the model, I’m talking about the *presentation to the\nuser*.\n\nPlus, regardless of whether a snippet of text came from a RAG query, we know\nthat LLMs can never provide an authoritative\nresult;\nit’s a fundamental limitation of the technology.  They can still garble the\nresults of RAG as much as they can misrepresent any other training data.  This\nmeans that it must never present its results *as* authoritative.\n\nIf you ask an AI to do research queries, every result should be presented as a\n*list of citations*.  Moreover, the presentation should display each citation as a\nlarge object of in its own right, with clearly identified metadata, including\nnot just the site where it was found but its publication date and, if possible,\nthe name of the author.  The literal, unmodified quotation (not from RAG, not a\nsummary: a quotation extracted with a regular program and not an LLM) should be\nfront-and-center, larger than any AI-generated text.\n\n*If* the AI product wants to editorialize or summarize (which should not always\nbe necessary!), the AI-generated text should be presented as small text\nunderneath the citation that has been found, de-emphasized as much as the\ndisclaimer is right now, at the very least until the user has verified that the\nsummary is accurate.  Perhaps, for a research project, a “did you read the\ncitation” checkbox might even be helpful.\n\nIf your product openly tells me that it will scramble, misrepresent, or omit\nits citations in its summaries, and I must read the original human-authored\ncitations to be sure, but then gives me no tools to track my reading of those\ncitations or even any way to *find* them, I cannot take it seriously.\n\n## No First-Person Output, No Apologies\n\nThere is no reason for a software development or research tool to use\nfirst-person language to describe itself.  They should not do so.  In fact\nthey should not be *allowed* to do\nso.\n\nThere is also no reason that they should ever *apologize*.  It is a waste of\neveryone’s time;\nit’s a waste for the chatbot to generate the apology, it’s a waste for the user\nto read the apology, and it’s a waste for the user to respond to the apology.\nYet they unfailingly do this upon every correction.\n\nThe vendors of these tools know that they are routinely causing mental-health\ncrises.\nIn response, they have added non-functional “guard rails” that can still, in\n2026, easily be\nbypassed.<sup>4</sup>\n\nA product seriously interested in helping with productivity would correct this\n*glaringly* obvious flaw, focus on the task at hand, and stop emitting useless\nverbiage.\n\nIn the previous two sections, I tried to focus on ways in which the harness\nwould be constructed differently even if the LLM technology is fundamentally\nimpossible to improve; in this case, I have to assume that the labs have *some*\ncontrol over the model itself.  But unless they are truly incapable of\ninfluencing their output (and all their “benchmarks” and “capabilities” seem to\nindicate that they can control it very tightly) they ought to be building\nmodels that are much less verbose.\n\n## More Non-Natural-Language User Interfaces\n\nAlthough natural language could *hypothetically* be a powerful interface for\ninteracting with a computer system, the practical upshot of LLM natural\nlanguage interfaces is that these interfaces are imprecise and repetitive, full\nof superstitions masquerading as “best\npractices”.  The inputs are a mess and the\nresulting outputs are a mess.\n\nThe general way of addressing this unstructured mess is to allow the chatbot to\n*directly take action* in response to the user’s input; in other words to\nsupply it with “tools” via an MCP server.  But again, this is backwards.  If we\ncannot even express our intent clearly in the first place, why are we trusting\nthis system to take potentially destructive and harmful actions on our behalf?\n\nInstead, I would expect a product that was seriously invested in helping me accomplish specific tasks, to have user interfaces specific to those tasks. Is it supposed to be able to be a security scanner that can discover OWASP top 10 bugs in a codebase? Have a button for that. Build that functionality into your harness, train it directly into the model, use smaller models that can satisfy that functionality more effectively than throwing it at the planet-sized brain of Fable or whatever.\n\nI’m aware that there are small software startups that do *something* like this,\nbut they are bolted on to the side of the main model providers’ APIs, not\nintegrated into the core of the product and not using their own models and AI\nsystems to achieve consistent and repeatable results.\n\n## Strong Data Provenance Indicators\n\nChatbots produce data tables pulled from websites, from APIs, from MCP tools or from summarizing and scrambling the user’s input. In order to provide the illusion of a seamless interface, this data is presented in-line regardless of where it comes from. But some of these outputs are produced mechanically via regular old API calls, for example, from the result of calling a tool or querying a website, but presented uniformly.\n\nBut there is a huge difference between an authoritative data source being\ninlined as part of a chatbot conversation, being treated as *input* by the\nchatbot, and some ad-hoc hallucinated data being treated as *output* of the\nchatbot.\n\nIf a product is trying to help me make accurate, empirically-grounded, data-driven decisions, the source of the data is critical.\n\nIntegrated into the “check for mistakes” and “verify citations” workflow I\ndescribed above, there’s a necessary “verify data programmatically” pass as\nwell; to have tools that will treat portions of the output as a *regular\nspreadsheet*, allowing regular-old computer arithmetic to verify things and\n*showing where such arithmetic was used, and how*.\n\n## Better User Control of Reproducibility\n\nAnyone familiar with the technical specifics of LLMs will know that they have a variable called “temperature” which controls the degree of randomness that the LLM uses to produce its outputs. But most users don’t know this, because it isn’t exposed as part of the user interface by default.\n\nThis leads to a subjective impression that you asked ChatGPT, and you got ChatGPT’s authoritative answer.\n\nYou can’t just set the temperature to zero and still get useful results - I am\naware that it does more than just scramble the output at random, and there are\nperhaps good reasons that simply exposing *just* a temperature setting would\nnot be that useful to users.  But if we followed some more of my earlier\nrecommendations for making more structured UI elements to solve specific\nproblems rather than having long back-and-forth chats where each refinement\ndepends on the previous response, perhaps those elements could also re-play the\nprocess so that users can see how reliable the bot is at a particular task and\ndevelop a sense of how the stochastic nature of the process actually affects\nit.\n\nSimilarly, if a user is trying to solve the same problem repeatedly with a\nchatbot, and the chatbot *product* has numerous computational tools that don’t\nreally have anything to do with the LLM, such as deterministic data-processing\ntools, then having a way to freeze the non-deterministic parts of the\ntranscript but re-populate a particular data frame with updated information and\nfork / continue the conversation from there would be a way to avoid introducing\npointless additional randomness when you already know what tool you’re trying\nto use.\n\nThe fact that every conversation is presented as this flat chat prompt that doesn’t let me interact with any of the widgets that were previously produced except through more chatting, really makes me feel like the whole product is just doing predatory social-media style “increase time on site” optimization, just trying to lure me into further repetitive and unreliable chats, rather than letting me get in, solve my problem, and get out.\n\n## Context Visibility\n\nManaging the LLM context is *the* ongoing challenge facing organizations that\nare trying to use “agentic” workflows.  Filling up the context with too much\ninformation causes well-known\nproblems. In response,\nadvanced LLM users attempting to solve larger problems must break up very long\nprompts into “skills”, give access to lengthy information via “tools”, and\ndelegating sub-problems to “sub-agents” rather than simply extending a single\nprompt indefinitely.\n\nAll of these strategies have flaws, because even on the largest models, compared to the breadth and depth of knowledge-work problems, LLM contexts are quite small.\n\nAnd yet, none of these products will *show the context to the user* by default.\nThere are third-party addons that can show you a simple progress\nbar\nbut for addressing the premier engineering difficulty with this technology,\nthat is below the bare minimum.\n\nThis lack of visibility means that almost all of the tools for extending the\ncontext are flying blind.  Rather than responding meaningfully to a full\ncontext, everyone just kind of guesses how much state they need by guessing and\ntrying over and over again with progressively more elaborate skill and sub-agent\nlayouts.  Even managing context compaction ends up being an advanced\nAPI-driven\nworkflow<sup>5</sup>.\n\nA serious product that was trying to help the user understand would not only show “available context” but explain the impact of context compactions, make it easier to see harness-generated prompts, and so on. This would be a first-class feature, combined with the aforementioned reproducibility / replay tools, would allow users to do real experiments to develop an understanding about how to make good use of the context window.\n\n## A Sandbox That Actually Works\n\nI’ve been focused on the chatbot interface here because it is the most\n*immediately* egregious upon looking at the UI.  But the “agentic loop” tools\nused for coding are equally dangerous, if not more so.  Coding tools\nkeep\ndestroying\neveryone’s\ndata,\nover the course of *years*.\n\nThese catastrophic incidents that become front-page news are relatively rare\ncompared to the amount of coding-agent use out there.  But they also aren’t the\nonly kind of sandbox violation.  Coding models will so routinely edit test code\ninstead of the system under test that there are “pro tips” articles all over\nthe web giving you the flawed\nadvice to simply *ask* the\nagent not to cheat.  News write-ups of the catastrophic incidents themselves\nwill also offer glib and wrong advice, like “use a docker container”.  That\nmight prevent it from literally deleting your operating system, but it won’t\nprevent it from destroying all the local work you have in your codebase (it\nneeds access to a checkout, after all!)\n\nThere is a flurry of activity in the infosec space where people are rushing to plug the gaps left by these coding harnesses. Everyone’s got their own version of an MCP approval gateway where you can optionally place a proxy between your agent and your production infrastructure.\n\nIn the best case, though, all these mitigations and proxies and prompts simply turn the user into an auto-approval automaton, hitting Y, Y, Y, Y over and over again, until you finally are driven mad and hit “yes to all”, turn on full-auto mode and submit yourself to the void. With nothing between your personal vigilance and disaster, there are no workflows left beyond decrementing your own vigilance until there’s nothing left and then hoping the disaster never arrives.\n\nThe fact that *some* mitigations exist that can be deployed by extra-cautious\nusers does not change the fact that “agentic coding” is an unsafe-by-default\ntechnology deployed without concern or guidance.  Every frontier lab has tied a\nspring-loaded shotgun to a dog; the fact that dog owners can publish thoughtful\nblog posts explaining how you can teach your dog the basics of gun safety or\nhow you can have your dogs play in a bullet-proof room does not mitigate the\nfact that the product should not have been allowed in the first place, nor\nshould it continue to exist without VERY strong security controls.\n\nI might believe that a frontier lab were seriously interested in providing developers with a useful tool if they shipped something that had safety built-in.\n\nThat means tools *in the harness*, detached from any LLM, independent of the\nprompt, that could:\n\n- sandbox all filesystem operations and strictly limit ANY deletions outside of specified scopes, regardless of operating system,\n- enforce snapshotting of the entire repo on every operation for easy rollbacks and minimal lost work,\n- remove the disaster of “auto\n  mode”\n  (not to mention nonsense like `--dangerously-skip-permissions` ) entirely, and\n- carefully consider a structure for presenting plans to the user where, rather than provoking immediate alert fatigue by asking for checks on every action, make structured plans which can be submitted to the user as a group of actions and reviewed and approved as a batch.\n\nIn the same way that I suggested above that research-based tasks should have a way of re-issuing prompts to determine how reproducible a result is, or whether other sources might be found, agent-based tasks should have a way of being executed against mock services for popular APIs, so that the verification can match both on the front-end (review the plan for making the API calls before they’re executed) and the back end (review the API calls that were issued to the mock service and verify that they matched).\n\nInstead, the frontier labs provide us products that are disasters out of the box, give us “best practices” to build massive and elaborate, as well as incomplete and error-prone, security perimeters of our own design. Then they blame “operator error” when it inevitably goes wrong. I cannot believe that these design choices are intended to help us be productive.\n\n## Bonus: Human Processes\n\nOrganizations *deploying* AI also frequently come across as unserious, for\nsimilar reasons.  In 2023, naive exuberance could perhaps be forgiven.  But\ntoday, as we near the close of 2026, there are several well-known problems,\nthat have been extremely well-covered in the press.  None of these things\nshould be surprising, but most orgs deploying these tools are still just\nletting them rip and hoping it all works out.\n\nOrganizations deploying these tools would need at least three kinds of major modifications to their internal processes, if they wanted to be serious about using them safely:\n\n### 1. Shift Rotations to Prevent Vigilance Decrement\n\nThere have been several high-profile incidents where software developers’ gradual acquiescence to accepting LLM output have lead to serious economic consequences for the companies deploying them, perhaps best typified by Amazon’s “millions of lost orders” due to a gradual decay of their engineering processes from LLM use.\n\nThese outages, and other AI-related failures, are due to the difficulty of\nmaintaining focus on the same problems.  In other words, as I described above,\nvigilance\ndecrement\nis a constant problem, because AI outputs are *most often* correct, but\ncontinue to be incorrect in surprising and non-intuitive ways.  As I have\npreviously written, you cannot trust\nyourself to catch every bug with code review, and LLM output.\n\nAviation, for example, has *very strict rules* around rest\nrequirements.\nThere is also a specific rule that “No certificate holder may operate an aircraft\nwithout a second in command if that aircraft has a passenger seating\nconfiguration, excluding any pilot seat, of ten seats or\nmore.”.\nOther safety-critical professions have similar rules.\n\nAnd yet, even in the age of the supposed “AI revolution”, most software teams are still assigning every engineer a full feature load, not planning for any rest, and telling people to review code whenever they happen to have some “free time”.\n\nMaintenance of vigilance has to be your top priority.  Regular, scheduled,\n*inviolable* rest periods where people do work without AI assistance, and are\nnot exposed to any AI output for review or otherwise, would be crucial in order\nto stay mentally sharp enough.\n\nThe tools themselves should have this sort of thing built in. The mistake-review process described above should have a periodic spot-check mode where a second reviewer periodically reviews a chatbot log, doing their own independent verification of claims, to see if they spot the same errors. This could provide a feedback loop to determine how much rest is necessary to maintain continuous attention and actually spot hallucinations.\n\n### 2. Skill Practice To Prevent Skill Loss\n\nIt is also well-known that AI use leads to AI reliance, and AI reliance leads to skill loss.\n\nI like to use the analogy to dockworkers at a seaport<sup>6</sup> adopting automation.\n\nIf you employ dockworkers to load and unload ships all day long, they are going to be getting tons of exercise. They will be able to lift heavy objects on demand, whenever. They might have plenty of health problems and injuries from this type of work, but “lack of exercise” will not be a problem.\n\nWith the development of standardized container ships and mechanized cranes, you are going to be changing their job description substantially: now they mostly spend all day sitting in a small cubicle moving a control lever back and forth, not lifting heavy stuff. They will get worse at lifting heavy objects.\n\nIn this analogy however, the cranes are not all that reliable.  We know they\nbreak, and they drop their payloads sometimes, and the stuff needs to be\nmanually moved.  But this only happens a few times a week, at most.  If you\nneed whoever is driving the crane to be able to jump out at any moment and\nstill move stuff around manually, then you need to make an affordance for that.\nYou need to give them *time* to go to the gym and do some lifting for practice,\nor every crane failure is going to be a major emergency.\n\nAn organization doing an AI transformation would also need a massive increase\nto learning & development budget, both in terms of resources and in terms of\nschedule.  If your people are going to lose skills because they’ve lost regular\npractice in the incidental course of doing their duties, then they are going to\nneed *deliberate*, intentional, non-incidental practice of those skills to keep\nthem sharp.\n\nBut rather than trying to accommodate new workflows and give time for people to adjust, most AI mandates are simply dropped on workers like a ton of bricks, with no time to adapt and no affordance for maintaining their skills. Operate the crane and stay fit and healthy and ready to switch back to manual lifting at any time and then get back in the crane cockpit right afterwards. Don’t mess up.\n\nThen an accident happens and everyone is surprised, as if this process weren’t\npractically *designed* to produce a terrible result.\n\n### 3. Mental Health Resources to Deal with Mental Health Risks\n\nAI psychosis often begins with practical problem-solving, and beyond that, it can start specifically at work. Not to mention the more pedestrian condition of “AI brain fry”.\n\nIf you are mandating your employees to use a hazardous tool that may seriously\nand *directly* damage their mental health, you need trainings and resources.\nYou need in-house therapists and you need to be making sure to check in with\npeople actively to make sure that this is not happening.\n\nAgain, the tool itself ought to have some way of dealing with this.  An\noccasional “take a\nbreak”\npopup is easily dismissed; they need a user-visible AI personal\ndosimeter so you can see\nyour *cumulative* usage over time.\n\nI don’t even know if “usage over time” is a sufficient metric to gauge risk. Maybe if your work chatbot start to talk about resonance too much, unless you literally work as an acoustic engineer, that should be flagged for someone.\n\nWe are, again, years into dealing with these tools, and we know these risks exist. Yet no serious mitigations are provided. Not even any way of measuring the risk exposure.\n\n### And More\n\nThere are also many other risks associated with the technology. There are intellectual property risks with the foundation models, due to recklessness with their training data. There are existential financial risks associated with the infrastructure build-out. The extent to which most “open” models are simply derivatives of frontier models is an open question.\n\n## What I Think\n\nIf any *one* of these things were regularly overlooked by AI vendors or users,\nthat would be a totally normal product oversight.  Room for improvement for the\nnext version, but nothing catastrophic.\n\nShipping without *any* of them doesn’t seem like lean product management, it\nseems like a careless attitude towards risk and a product design philosophy\noriented entirely towards short-term demos, with no regard for how to realize\nactual productivity gains.\n\nFurthermore, being available for *years* without anything like these features,\ndespite hundreds of incidents demonstrating the risks, with hundreds of\nbillions of dollars of funding, makes it seem to me like if they *were* to add\nall the features that would make their product actually safe and hypothetically\nuseful, these features would reveal that it is actually not an improvement to\nproductivity.\n\nIn the year since I first wrote about measuring the cost/benefit ratio of AI, I have heard from numerous people who have shown this to management to try to illustrate why their AI initiatives — like almost all AI initiatives — were either failing or burning out their engineers.\n\nI’ve also heard from lots of people that have told me that it’s obviously\nuseful and they don’t need to measure so carefully, because they are getting\nlots of work done that they couldn’t have otherwise.<sup>7</sup>\n\nI have yet to hear from a *single* person who has said “yeah, we measured\naccording to your methodology<sup>8</sup>, and it turns out that our AI work is going\ngreat and that our ratio is 0.75”.\n\nObviously, I cannot say for sure why this is; absence of evidence is not evidence of absence. But at this point I think the null hypothesis is that AI tools provide, in aggregate, zero value. They make mistakes too often, and the externalities they produce are so bad and so difficult to control that even before we get to the places where they are just physically poisoning people, even the negative effects on their direct users end up cancelling out whatever benefit to they provide to their organizations.\n\nIf I were wrong, then including tools to *measure* an AI’s effectiveness *at\nthe tasks their users are actually trying to accomplish*, rather than\nmeaningless\nbenchmarks,\nwould show big productivity gains.  The frontier labs would be champing at the\nbit to add such features, and crowing about their fantastic results.\n\nI think the labs know that if they did that, it would present a grim picture to their users. Such tools would let their users see that it’s making mistakes much more often than they realized, that they’re spending much more time with it than they want to be, and that it’s just generally not fit for purpose.\n\nIf they prove me wrong by adding in all of these safety mechanisms, and in the process, they make all of their AI technology less harmful, I’ll be thrilled to be debunked.\n\n## Acknowledgments\n\nThank you to my patrons who are supporting my writing on this blog. If you like what you’ve read here and you’d like to read more of it, or you’d like to support my various open-source endeavors, you can support my work as a sponsor!\n\n1. \nIt is also interesting that for the next section, *sometimes* it seems\nthat Claude’s disclaimer is “Please double-check cited sources.” instead. ↩\n2. \n... by which I mean the “prompter”, since authorship is not what’s happening here. ↩\n3. \nClaude has the “citations API”, Google has various different kinds of “grounding” against its own APIs, and I guess Microsoft can check OpenAI’s homework if you want. ↩\n4. \nGiven the relatively slow speed of the justice system and the mainstream press around the world, we probably will not hear about whether people are managing to incidentally break through these guard rails to self harm right now, but there are no shortage of stories still being *reported* right now\nwhere people were still doing just that, such as in this\nstory\nwhere the effect of the much vaunted “guard rails” in 2025 was that if you\nwanted it to write you a suicide note, it would refuse twice but acquiesce\non the third try.  I don’t see any reason to believe this fundamental issue\nhas been addressed in the meanwhile, since it had been happening for years\nat that point. ↩\n5. \nIn this tutorial we can also see an incredibly rosy scenario presented, where a long-running workflow effortlessly compresses all of the necessary information into the new context, even if it uses a lower-fidelity model to do so, rather than the tangled and gnarly problem of problems which really *are* too big to fit in the context, which is to say, “most real-world\nproblems”.  This presents the context limit instead as a minor speedbump to\nbe worked around rather than the fundamental flaw in LLM tooling. ↩\n6. \nA heavily fictionalized seaport. This is not how actual dockworkers work. This is a simplistic metaphor about incidental benefits of instrumental tasks, it is not supposed to delve deeply into the mechanics of maritime shipping. In particular I know that cranes are more reliable than this and this is not actually how you would respond to a crane malfunction anyway. Feel free to share fun facts about maritime shipping if that is your special interest but please do not @ me to *correct* this\nmetaphor. ↩\n7. \nTo my knowledge, none of their publicly-traded employers have posted a measurable improvement to efficiency outside the margin of error. ↩\n8. \nOr any similar methodology. I don’t need people to adopt the exact practice that I proposed there. ↩","body_html":"<p>One of the issues that I have with the current generation of “AI” products is that they do not appear to take their own premises seriously. I look at a plethora of obsequious chatbots claiming to be serious tools for problem solving, and I think, this is not what a problem-solving tool would look like.</p>\n<p>Even before we get to the tremendous ethical problems with the frontier labs,\nit is this impression of their composition <em>as a product</em> that makes me feel,\nconstantly, whenever I am interacting with them, that they are less a software\nproduct than that they are a grift, a scam designed to make me feel like I am\ninteracting with a product that has capabilities that it simply does not, to\ntry to lull me into a false sense of security that I can trust it.</p>\n<p>The frontier labs are of course the worst offenders, but every criticism here\napplies just as much to Ollama, which (if anything, due to the obviously poorer\nquality of the available models themselves) needs these features <em>even more</em>\nthan the frontier labs do.</p>\n<p>Here, I will set down a few features that might convince me that an LLM-based product, particularly one focused on research or software development, was actually serious about helping me do useful things with it.</p>\n<h2 id=\"make-checking-for-mistakes-a-first-class-feature\">Make “Checking For Mistakes” A First-Class Feature</h2>\n<p>This is the biggest issue, and the major reason that I was inspired to write this post.</p>\n<p>It is a truth universally acknowledged, that AIs cannot reliably provide information.</p>\n<p>I could cite a ton of news articles and studies about this fact, but there is\nno need.  Every single chatbot admits this, up front, in a fine-print\ndisclaimer as a core part of their user interface.  Gemini says “AI can make\nmistakes, so double-check responses”, Claude says “Claude is AI and can make\nmistakes. Please double-check responses.&lt;sup&gt;1&lt;/sup&gt;” ChatGPT says “ChatGPT can make\nmistakes. Check important info.”.</p>\n<p>Every time I see that last one, I wonder how I’m supposed to know what “info” is supposed to be “important”.</p>\n<p>All of these warnings are all small, gray text, painfully obviously included as\nlegalese to push responsibility back onto the user rather than to help with\nanything.  This is a <em>core limitation</em> of all these products.  Checking their\noutput is a part of the workflow for using them that:</p>\n<ol><li><p>you <em>absolutely cannot skip or skimp on</em> without creating risks to yourself</p><p> and whoever you are conveying its output to, and,</p></li><li><p>it is <em>very easy to skip or skimp on</em> and you are encouraged at every turn</p><p> to do so, because “just trust the output” is one of the quickest ways to\n save time.</p></li></ol>\n<p>A chatbot product that took this weakness seriously, as an actual consideration\nfor using it, would put a checkbox next to <em>every claim</em> in its output.  It\nwould be a 2-column worksheet, where you’ve got the LLM output in the first\ncolumn, and next to it, human notes in the second column, explaining what work\nwent into checking this claim, and a <em>big</em> checkbox that you would only check\noff after you believe you’d checked its claims thoroughly enough.</p>\n<p>Coding assistants would need to have some version of this as well.  Right now,\nthis is pushed off into code review, which means it is a dark pattern which\nsubtly encourages the “author”&lt;sup&gt;2&lt;/sup&gt; to offload this work to their code reviewer\nwithout ever looking.  Once again, “it’s probably fine, I don’t need to check”\nis the quickest way to save time and churn out those PRs faster.</p>\n<p>It might even be useful for coding harnesses to have some affordance for\nchecking code before it even runs tests.  As the vendors themselves have\nadmitted,\nit’s not just expensive to burn tokens on your “AI”, you also end up burning\nfar more <em>compute</em> on the AI.  Being able to check your diffs before sending\nthem over to uselessly exhaust your testing compute cluster would be useful.</p>\n<p>If your product tells me that it makes mistakes and I must be the one to check\nfor the mistakes, but then gives me <em>zero tools</em> to check for mistakes, I\ncannot take it seriously.</p>\n<h3 id=\"more-citations-to-check-and-more-details\">More Citations to Check, And More Details</h3>\n<p>Most chatbots prefer to give an answer, rather than a citation. In my own personal use, I find that when asked to provide a list of citations with clearly marked sources for each one, they will appear to “get bored” halfway through the list and simply stop including citations at some point.</p>\n<p>When the bots include citations at all, present them as inline annotations that say nothing but the domain name of the search result, in a font so small that it’s barely legible, and an equally indecipherable icon that is fewer than 16 pixels on a side.</p>\n<p>This is backwards.</p>\n<p>Now, I am aware that these citations do come from somewhere, and in an attempt\nto reduce hallucinations, all of the major providers support some form of\n“grounding”&lt;sup&gt;3&lt;/sup&gt;, and that those little barely-readable citation links are\nreferencing actual structures in the\nRAG pipeline\nand not just potentially-hallucinated tokens, but I’m not talking about the\nunderlying machinery in the model, I’m talking about the *presentation to the\nuser*.</p>\n<p>Plus, regardless of whether a snippet of text came from a RAG query, we know\nthat LLMs can never provide an authoritative\nresult;\nit’s a fundamental limitation of the technology.  They can still garble the\nresults of RAG as much as they can misrepresent any other training data.  This\nmeans that it must never present its results <em>as</em> authoritative.</p>\n<p>If you ask an AI to do research queries, every result should be presented as a\n<em>list of citations</em>.  Moreover, the presentation should display each citation as a\nlarge object of in its own right, with clearly identified metadata, including\nnot just the site where it was found but its publication date and, if possible,\nthe name of the author.  The literal, unmodified quotation (not from RAG, not a\nsummary: a quotation extracted with a regular program and not an LLM) should be\nfront-and-center, larger than any AI-generated text.</p>\n<p><em>If</em> the AI product wants to editorialize or summarize (which should not always\nbe necessary!), the AI-generated text should be presented as small text\nunderneath the citation that has been found, de-emphasized as much as the\ndisclaimer is right now, at the very least until the user has verified that the\nsummary is accurate.  Perhaps, for a research project, a “did you read the\ncitation” checkbox might even be helpful.</p>\n<p>If your product openly tells me that it will scramble, misrepresent, or omit\nits citations in its summaries, and I must read the original human-authored\ncitations to be sure, but then gives me no tools to track my reading of those\ncitations or even any way to <em>find</em> them, I cannot take it seriously.</p>\n<h2 id=\"no-first-person-output-no-apologies\">No First-Person Output, No Apologies</h2>\n<p>There is no reason for a software development or research tool to use\nfirst-person language to describe itself.  They should not do so.  In fact\nthey should not be <em>allowed</em> to do\nso.</p>\n<p>There is also no reason that they should ever <em>apologize</em>.  It is a waste of\neveryone’s time;\nit’s a waste for the chatbot to generate the apology, it’s a waste for the user\nto read the apology, and it’s a waste for the user to respond to the apology.\nYet they unfailingly do this upon every correction.</p>\n<p>The vendors of these tools know that they are routinely causing mental-health\ncrises.\nIn response, they have added non-functional “guard rails” that can still, in\n2026, easily be\nbypassed.&lt;sup&gt;4&lt;/sup&gt;</p>\n<p>A product seriously interested in helping with productivity would correct this\n<em>glaringly</em> obvious flaw, focus on the task at hand, and stop emitting useless\nverbiage.</p>\n<p>In the previous two sections, I tried to focus on ways in which the harness\nwould be constructed differently even if the LLM technology is fundamentally\nimpossible to improve; in this case, I have to assume that the labs have <em>some</em>\ncontrol over the model itself.  But unless they are truly incapable of\ninfluencing their output (and all their “benchmarks” and “capabilities” seem to\nindicate that they can control it very tightly) they ought to be building\nmodels that are much less verbose.</p>\n<h2 id=\"more-non-natural-language-user-interfaces\">More Non-Natural-Language User Interfaces</h2>\n<p>Although natural language could <em>hypothetically</em> be a powerful interface for\ninteracting with a computer system, the practical upshot of LLM natural\nlanguage interfaces is that these interfaces are imprecise and repetitive, full\nof superstitions masquerading as “best\npractices”.  The inputs are a mess and the\nresulting outputs are a mess.</p>\n<p>The general way of addressing this unstructured mess is to allow the chatbot to\n<em>directly take action</em> in response to the user’s input; in other words to\nsupply it with “tools” via an MCP server.  But again, this is backwards.  If we\ncannot even express our intent clearly in the first place, why are we trusting\nthis system to take potentially destructive and harmful actions on our behalf?</p>\n<p>Instead, I would expect a product that was seriously invested in helping me accomplish specific tasks, to have user interfaces specific to those tasks. Is it supposed to be able to be a security scanner that can discover OWASP top 10 bugs in a codebase? Have a button for that. Build that functionality into your harness, train it directly into the model, use smaller models that can satisfy that functionality more effectively than throwing it at the planet-sized brain of Fable or whatever.</p>\n<p>I’m aware that there are small software startups that do <em>something</em> like this,\nbut they are bolted on to the side of the main model providers’ APIs, not\nintegrated into the core of the product and not using their own models and AI\nsystems to achieve consistent and repeatable results.</p>\n<h2 id=\"strong-data-provenance-indicators\">Strong Data Provenance Indicators</h2>\n<p>Chatbots produce data tables pulled from websites, from APIs, from MCP tools or from summarizing and scrambling the user’s input. In order to provide the illusion of a seamless interface, this data is presented in-line regardless of where it comes from. But some of these outputs are produced mechanically via regular old API calls, for example, from the result of calling a tool or querying a website, but presented uniformly.</p>\n<p>But there is a huge difference between an authoritative data source being\ninlined as part of a chatbot conversation, being treated as <em>input</em> by the\nchatbot, and some ad-hoc hallucinated data being treated as <em>output</em> of the\nchatbot.</p>\n<p>If a product is trying to help me make accurate, empirically-grounded, data-driven decisions, the source of the data is critical.</p>\n<p>Integrated into the “check for mistakes” and “verify citations” workflow I\ndescribed above, there’s a necessary “verify data programmatically” pass as\nwell; to have tools that will treat portions of the output as a *regular\nspreadsheet*, allowing regular-old computer arithmetic to verify things and\n<em>showing where such arithmetic was used, and how</em>.</p>\n<h2 id=\"better-user-control-of-reproducibility\">Better User Control of Reproducibility</h2>\n<p>Anyone familiar with the technical specifics of LLMs will know that they have a variable called “temperature” which controls the degree of randomness that the LLM uses to produce its outputs. But most users don’t know this, because it isn’t exposed as part of the user interface by default.</p>\n<p>This leads to a subjective impression that you asked ChatGPT, and you got ChatGPT’s authoritative answer.</p>\n<p>You can’t just set the temperature to zero and still get useful results - I am\naware that it does more than just scramble the output at random, and there are\nperhaps good reasons that simply exposing <em>just</em> a temperature setting would\nnot be that useful to users.  But if we followed some more of my earlier\nrecommendations for making more structured UI elements to solve specific\nproblems rather than having long back-and-forth chats where each refinement\ndepends on the previous response, perhaps those elements could also re-play the\nprocess so that users can see how reliable the bot is at a particular task and\ndevelop a sense of how the stochastic nature of the process actually affects\nit.</p>\n<p>Similarly, if a user is trying to solve the same problem repeatedly with a\nchatbot, and the chatbot <em>product</em> has numerous computational tools that don’t\nreally have anything to do with the LLM, such as deterministic data-processing\ntools, then having a way to freeze the non-deterministic parts of the\ntranscript but re-populate a particular data frame with updated information and\nfork / continue the conversation from there would be a way to avoid introducing\npointless additional randomness when you already know what tool you’re trying\nto use.</p>\n<p>The fact that every conversation is presented as this flat chat prompt that doesn’t let me interact with any of the widgets that were previously produced except through more chatting, really makes me feel like the whole product is just doing predatory social-media style “increase time on site” optimization, just trying to lure me into further repetitive and unreliable chats, rather than letting me get in, solve my problem, and get out.</p>\n<h2 id=\"context-visibility\">Context Visibility</h2>\n<p>Managing the LLM context is <em>the</em> ongoing challenge facing organizations that\nare trying to use “agentic” workflows.  Filling up the context with too much\ninformation causes well-known\nproblems. In response,\nadvanced LLM users attempting to solve larger problems must break up very long\nprompts into “skills”, give access to lengthy information via “tools”, and\ndelegating sub-problems to “sub-agents” rather than simply extending a single\nprompt indefinitely.</p>\n<p>All of these strategies have flaws, because even on the largest models, compared to the breadth and depth of knowledge-work problems, LLM contexts are quite small.</p>\n<p>And yet, none of these products will <em>show the context to the user</em> by default.\nThere are third-party addons that can show you a simple progress\nbar\nbut for addressing the premier engineering difficulty with this technology,\nthat is below the bare minimum.</p>\n<p>This lack of visibility means that almost all of the tools for extending the\ncontext are flying blind.  Rather than responding meaningfully to a full\ncontext, everyone just kind of guesses how much state they need by guessing and\ntrying over and over again with progressively more elaborate skill and sub-agent\nlayouts.  Even managing context compaction ends up being an advanced\nAPI-driven\nworkflow&lt;sup&gt;5&lt;/sup&gt;.</p>\n<p>A serious product that was trying to help the user understand would not only show “available context” but explain the impact of context compactions, make it easier to see harness-generated prompts, and so on. This would be a first-class feature, combined with the aforementioned reproducibility / replay tools, would allow users to do real experiments to develop an understanding about how to make good use of the context window.</p>\n<h2 id=\"a-sandbox-that-actually-works\">A Sandbox That Actually Works</h2>\n<p>I’ve been focused on the chatbot interface here because it is the most\n<em>immediately</em> egregious upon looking at the UI.  But the “agentic loop” tools\nused for coding are equally dangerous, if not more so.  Coding tools\nkeep\ndestroying\neveryone’s\ndata,\nover the course of <em>years</em>.</p>\n<p>These catastrophic incidents that become front-page news are relatively rare\ncompared to the amount of coding-agent use out there.  But they also aren’t the\nonly kind of sandbox violation.  Coding models will so routinely edit test code\ninstead of the system under test that there are “pro tips” articles all over\nthe web giving you the flawed\nadvice to simply <em>ask</em> the\nagent not to cheat.  News write-ups of the catastrophic incidents themselves\nwill also offer glib and wrong advice, like “use a docker container”.  That\nmight prevent it from literally deleting your operating system, but it won’t\nprevent it from destroying all the local work you have in your codebase (it\nneeds access to a checkout, after all!)</p>\n<p>There is a flurry of activity in the infosec space where people are rushing to plug the gaps left by these coding harnesses. Everyone’s got their own version of an MCP approval gateway where you can optionally place a proxy between your agent and your production infrastructure.</p>\n<p>In the best case, though, all these mitigations and proxies and prompts simply turn the user into an auto-approval automaton, hitting Y, Y, Y, Y over and over again, until you finally are driven mad and hit “yes to all”, turn on full-auto mode and submit yourself to the void. With nothing between your personal vigilance and disaster, there are no workflows left beyond decrementing your own vigilance until there’s nothing left and then hoping the disaster never arrives.</p>\n<p>The fact that <em>some</em> mitigations exist that can be deployed by extra-cautious\nusers does not change the fact that “agentic coding” is an unsafe-by-default\ntechnology deployed without concern or guidance.  Every frontier lab has tied a\nspring-loaded shotgun to a dog; the fact that dog owners can publish thoughtful\nblog posts explaining how you can teach your dog the basics of gun safety or\nhow you can have your dogs play in a bullet-proof room does not mitigate the\nfact that the product should not have been allowed in the first place, nor\nshould it continue to exist without VERY strong security controls.</p>\n<p>I might believe that a frontier lab were seriously interested in providing developers with a useful tool if they shipped something that had safety built-in.</p>\n<p>That means tools <em>in the harness</em>, detached from any LLM, independent of the\nprompt, that could:</p>\n<ul><li>sandbox all filesystem operations and strictly limit ANY deletions outside of specified scopes, regardless of operating system,</li><li>enforce snapshotting of the entire repo on every operation for easy rollbacks and minimal lost work,</li><li><p>remove the disaster of “auto</p><p>mode”\n(not to mention nonsense like <code>--dangerously-skip-permissions</code> ) entirely, and</p></li><li>carefully consider a structure for presenting plans to the user where, rather than provoking immediate alert fatigue by asking for checks on every action, make structured plans which can be submitted to the user as a group of actions and reviewed and approved as a batch.</li></ul>\n<p>In the same way that I suggested above that research-based tasks should have a way of re-issuing prompts to determine how reproducible a result is, or whether other sources might be found, agent-based tasks should have a way of being executed against mock services for popular APIs, so that the verification can match both on the front-end (review the plan for making the API calls before they’re executed) and the back end (review the API calls that were issued to the mock service and verify that they matched).</p>\n<p>Instead, the frontier labs provide us products that are disasters out of the box, give us “best practices” to build massive and elaborate, as well as incomplete and error-prone, security perimeters of our own design. Then they blame “operator error” when it inevitably goes wrong. I cannot believe that these design choices are intended to help us be productive.</p>\n<h2 id=\"bonus-human-processes\">Bonus: Human Processes</h2>\n<p>Organizations <em>deploying</em> AI also frequently come across as unserious, for\nsimilar reasons.  In 2023, naive exuberance could perhaps be forgiven.  But\ntoday, as we near the close of 2026, there are several well-known problems,\nthat have been extremely well-covered in the press.  None of these things\nshould be surprising, but most orgs deploying these tools are still just\nletting them rip and hoping it all works out.</p>\n<p>Organizations deploying these tools would need at least three kinds of major modifications to their internal processes, if they wanted to be serious about using them safely:</p>\n<h3 id=\"1-shift-rotations-to-prevent-vigilance-decrement\">1. Shift Rotations to Prevent Vigilance Decrement</h3>\n<p>There have been several high-profile incidents where software developers’ gradual acquiescence to accepting LLM output have lead to serious economic consequences for the companies deploying them, perhaps best typified by Amazon’s “millions of lost orders” due to a gradual decay of their engineering processes from LLM use.</p>\n<p>These outages, and other AI-related failures, are due to the difficulty of\nmaintaining focus on the same problems.  In other words, as I described above,\nvigilance\ndecrement\nis a constant problem, because AI outputs are <em>most often</em> correct, but\ncontinue to be incorrect in surprising and non-intuitive ways.  As I have\npreviously written, you cannot trust\nyourself to catch every bug with code review, and LLM output.</p>\n<p>Aviation, for example, has <em>very strict rules</em> around rest\nrequirements.\nThere is also a specific rule that “No certificate holder may operate an aircraft\nwithout a second in command if that aircraft has a passenger seating\nconfiguration, excluding any pilot seat, of ten seats or\nmore.”.\nOther safety-critical professions have similar rules.</p>\n<p>And yet, even in the age of the supposed “AI revolution”, most software teams are still assigning every engineer a full feature load, not planning for any rest, and telling people to review code whenever they happen to have some “free time”.</p>\n<p>Maintenance of vigilance has to be your top priority.  Regular, scheduled,\n<em>inviolable</em> rest periods where people do work without AI assistance, and are\nnot exposed to any AI output for review or otherwise, would be crucial in order\nto stay mentally sharp enough.</p>\n<p>The tools themselves should have this sort of thing built in. The mistake-review process described above should have a periodic spot-check mode where a second reviewer periodically reviews a chatbot log, doing their own independent verification of claims, to see if they spot the same errors. This could provide a feedback loop to determine how much rest is necessary to maintain continuous attention and actually spot hallucinations.</p>\n<h3 id=\"2-skill-practice-to-prevent-skill-loss\">2. Skill Practice To Prevent Skill Loss</h3>\n<p>It is also well-known that AI use leads to AI reliance, and AI reliance leads to skill loss.</p>\n<p>I like to use the analogy to dockworkers at a seaport&lt;sup&gt;6&lt;/sup&gt; adopting automation.</p>\n<p>If you employ dockworkers to load and unload ships all day long, they are going to be getting tons of exercise. They will be able to lift heavy objects on demand, whenever. They might have plenty of health problems and injuries from this type of work, but “lack of exercise” will not be a problem.</p>\n<p>With the development of standardized container ships and mechanized cranes, you are going to be changing their job description substantially: now they mostly spend all day sitting in a small cubicle moving a control lever back and forth, not lifting heavy stuff. They will get worse at lifting heavy objects.</p>\n<p>In this analogy however, the cranes are not all that reliable.  We know they\nbreak, and they drop their payloads sometimes, and the stuff needs to be\nmanually moved.  But this only happens a few times a week, at most.  If you\nneed whoever is driving the crane to be able to jump out at any moment and\nstill move stuff around manually, then you need to make an affordance for that.\nYou need to give them <em>time</em> to go to the gym and do some lifting for practice,\nor every crane failure is going to be a major emergency.</p>\n<p>An organization doing an AI transformation would also need a massive increase\nto learning &amp; development budget, both in terms of resources and in terms of\nschedule.  If your people are going to lose skills because they’ve lost regular\npractice in the incidental course of doing their duties, then they are going to\nneed <em>deliberate</em>, intentional, non-incidental practice of those skills to keep\nthem sharp.</p>\n<p>But rather than trying to accommodate new workflows and give time for people to adjust, most AI mandates are simply dropped on workers like a ton of bricks, with no time to adapt and no affordance for maintaining their skills. Operate the crane and stay fit and healthy and ready to switch back to manual lifting at any time and then get back in the crane cockpit right afterwards. Don’t mess up.</p>\n<p>Then an accident happens and everyone is surprised, as if this process weren’t\npractically <em>designed</em> to produce a terrible result.</p>\n<h3 id=\"3-mental-health-resources-to-deal-with-mental-health-risks\">3. Mental Health Resources to Deal with Mental Health Risks</h3>\n<p>AI psychosis often begins with practical problem-solving, and beyond that, it can start specifically at work. Not to mention the more pedestrian condition of “AI brain fry”.</p>\n<p>If you are mandating your employees to use a hazardous tool that may seriously\nand <em>directly</em> damage their mental health, you need trainings and resources.\nYou need in-house therapists and you need to be making sure to check in with\npeople actively to make sure that this is not happening.</p>\n<p>Again, the tool itself ought to have some way of dealing with this.  An\noccasional “take a\nbreak”\npopup is easily dismissed; they need a user-visible AI personal\ndosimeter so you can see\nyour <em>cumulative</em> usage over time.</p>\n<p>I don’t even know if “usage over time” is a sufficient metric to gauge risk. Maybe if your work chatbot start to talk about resonance too much, unless you literally work as an acoustic engineer, that should be flagged for someone.</p>\n<p>We are, again, years into dealing with these tools, and we know these risks exist. Yet no serious mitigations are provided. Not even any way of measuring the risk exposure.</p>\n<h3 id=\"and-more\">And More</h3>\n<p>There are also many other risks associated with the technology. There are intellectual property risks with the foundation models, due to recklessness with their training data. There are existential financial risks associated with the infrastructure build-out. The extent to which most “open” models are simply derivatives of frontier models is an open question.</p>\n<h2 id=\"what-i-think\">What I Think</h2>\n<p>If any <em>one</em> of these things were regularly overlooked by AI vendors or users,\nthat would be a totally normal product oversight.  Room for improvement for the\nnext version, but nothing catastrophic.</p>\n<p>Shipping without <em>any</em> of them doesn’t seem like lean product management, it\nseems like a careless attitude towards risk and a product design philosophy\noriented entirely towards short-term demos, with no regard for how to realize\nactual productivity gains.</p>\n<p>Furthermore, being available for <em>years</em> without anything like these features,\ndespite hundreds of incidents demonstrating the risks, with hundreds of\nbillions of dollars of funding, makes it seem to me like if they <em>were</em> to add\nall the features that would make their product actually safe and hypothetically\nuseful, these features would reveal that it is actually not an improvement to\nproductivity.</p>\n<p>In the year since I first wrote about measuring the cost/benefit ratio of AI, I have heard from numerous people who have shown this to management to try to illustrate why their AI initiatives — like almost all AI initiatives — were either failing or burning out their engineers.</p>\n<p>I’ve also heard from lots of people that have told me that it’s obviously\nuseful and they don’t need to measure so carefully, because they are getting\nlots of work done that they couldn’t have otherwise.&lt;sup&gt;7&lt;/sup&gt;</p>\n<p>I have yet to hear from a <em>single</em> person who has said “yeah, we measured\naccording to your methodology&lt;sup&gt;8&lt;/sup&gt;, and it turns out that our AI work is going\ngreat and that our ratio is 0.75”.</p>\n<p>Obviously, I cannot say for sure why this is; absence of evidence is not evidence of absence. But at this point I think the null hypothesis is that AI tools provide, in aggregate, zero value. They make mistakes too often, and the externalities they produce are so bad and so difficult to control that even before we get to the places where they are just physically poisoning people, even the negative effects on their direct users end up cancelling out whatever benefit to they provide to their organizations.</p>\n<p>If I were wrong, then including tools to <em>measure</em> an AI’s effectiveness *at\nthe tasks their users are actually trying to accomplish*, rather than\nmeaningless\nbenchmarks,\nwould show big productivity gains.  The frontier labs would be champing at the\nbit to add such features, and crowing about their fantastic results.</p>\n<p>I think the labs know that if they did that, it would present a grim picture to their users. Such tools would let their users see that it’s making mistakes much more often than they realized, that they’re spending much more time with it than they want to be, and that it’s just generally not fit for purpose.</p>\n<p>If they prove me wrong by adding in all of these safety mechanisms, and in the process, they make all of their AI technology less harmful, I’ll be thrilled to be debunked.</p>\n<h2 id=\"acknowledgments\">Acknowledgments</h2>\n<p>Thank you to my patrons who are supporting my writing on this blog. If you like what you’ve read here and you’d like to read more of it, or you’d like to support my various open-source endeavors, you can support my work as a sponsor!</p>\n<ol><li></li></ol>\n<p>It is also interesting that for the next section, <em>sometimes</em> it seems\nthat Claude’s disclaimer is “Please double-check cited sources.” instead. ↩</p>\n<ol start=\"2\"><li></li></ol>\n<p>... by which I mean the “prompter”, since authorship is not what’s happening here. ↩</p>\n<ol start=\"3\"><li></li></ol>\n<p>Claude has the “citations API”, Google has various different kinds of “grounding” against its own APIs, and I guess Microsoft can check OpenAI’s homework if you want. ↩</p>\n<ol start=\"4\"><li></li></ol>\n<p>Given the relatively slow speed of the justice system and the mainstream press around the world, we probably will not hear about whether people are managing to incidentally break through these guard rails to self harm right now, but there are no shortage of stories still being <em>reported</em> right now\nwhere people were still doing just that, such as in this\nstory\nwhere the effect of the much vaunted “guard rails” in 2025 was that if you\nwanted it to write you a suicide note, it would refuse twice but acquiesce\non the third try.  I don’t see any reason to believe this fundamental issue\nhas been addressed in the meanwhile, since it had been happening for years\nat that point. ↩</p>\n<ol start=\"5\"><li></li></ol>\n<p>In this tutorial we can also see an incredibly rosy scenario presented, where a long-running workflow effortlessly compresses all of the necessary information into the new context, even if it uses a lower-fidelity model to do so, rather than the tangled and gnarly problem of problems which really <em>are</em> too big to fit in the context, which is to say, “most real-world\nproblems”.  This presents the context limit instead as a minor speedbump to\nbe worked around rather than the fundamental flaw in LLM tooling. ↩</p>\n<ol start=\"6\"><li></li></ol>\n<p>A heavily fictionalized seaport. This is not how actual dockworkers work. This is a simplistic metaphor about incidental benefits of instrumental tasks, it is not supposed to delve deeply into the mechanics of maritime shipping. In particular I know that cranes are more reliable than this and this is not actually how you would respond to a crane malfunction anyway. Feel free to share fun facts about maritime shipping if that is your special interest but please do not @ me to <em>correct</em> this\nmetaphor. ↩</p>\n<ol start=\"7\"><li></li></ol>\n<p>To my knowledge, none of their publicly-traded employers have posted a measurable improvement to efficiency outside the margin of error. ↩</p>\n<ol start=\"8\"><li></li></ol>\n<p>Or any similar methodology. I don’t need people to adopt the exact practice that I proposed there. ↩</p>","headings":[{"level":2,"text":"Make “Checking For Mistakes” A First-Class Feature","id":"make-checking-for-mistakes-a-first-class-feature"},{"level":3,"text":"More Citations to Check, And More Details","id":"more-citations-to-check-and-more-details"},{"level":2,"text":"No First-Person Output, No Apologies","id":"no-first-person-output-no-apologies"},{"level":2,"text":"More Non-Natural-Language User Interfaces","id":"more-non-natural-language-user-interfaces"},{"level":2,"text":"Strong Data Provenance Indicators","id":"strong-data-provenance-indicators"},{"level":2,"text":"Better User Control of Reproducibility","id":"better-user-control-of-reproducibility"},{"level":2,"text":"Context Visibility","id":"context-visibility"},{"level":2,"text":"A Sandbox That Actually Works","id":"a-sandbox-that-actually-works"},{"level":2,"text":"Bonus: Human Processes","id":"bonus-human-processes"},{"level":3,"text":"1. Shift Rotations to Prevent Vigilance Decrement","id":"1-shift-rotations-to-prevent-vigilance-decrement"},{"level":3,"text":"2. Skill Practice To Prevent Skill Loss","id":"2-skill-practice-to-prevent-skill-loss"},{"level":3,"text":"3. Mental Health Resources to Deal with Mental Health Risks","id":"3-mental-health-resources-to-deal-with-mental-health-risks"},{"level":3,"text":"And More","id":"and-more"},{"level":2,"text":"What I Think","id":"what-i-think"},{"level":2,"text":"Acknowledgments","id":"acknowledgments"}]}}