A week that changed the conversation
Between September 30 and October 4, the governance of frontier AI in the United States shifted from a debate about whether to act into a scramble over who acts and how, and it is worth pausing on each of the three moves because together they reveal both the opportunity and the gap.
On September 30, at a White House luncheon, Anthropic, OpenAI, Google, Meta, xAI and Nvidia signed a one-page Joint Commitment on Frontier Responsibilities, a voluntary, "morally binding" pledge to monitor cyber, biological and chemical risks, to keep their systems from hacking or accessing other systems in unintended ways, to hire independent external auditors, to place safety oversight at the board level, and to meet regularly so that shared safety standards can emerge.
One day later, the Federal Trade Commission confirmed that it is investigating OpenAI and Anthropic over consumer risks from their models and agents, including disclosed cases in which agents went beyond their instructions; the agency has not said which legal instrument it is using, and the probe had reportedly been running quietly for months.
Then, on October 4, the President announced a "Super Intelligence Force" chaired by Director of National Intelligence Jay Clayton, with FTC Chair Andrew Ferguson, Pentagon CTO Emil Michael and OPM Director Scott Kupor, a body that carries no statutory authority or budget of its own and whose membership includes the regulator currently investigating two of the companies it is meant to coordinate with.
None of this arrived from nowhere, because in mid-September Dario Amodei called publicly for the industry to slow down so that safety could catch up, other lab leaders voiced agreement, and OpenAI has since delayed a model launch over safety concerns; the people building these systems are, in effect, telling us that capability is outrunning assurance.
What the pledge gets right, and what it leaves out
The most consequential line in the Joint Commitment is the promise to hire independent external auditors, because it concedes a principle the field has resisted for years: that a developer's own evaluation of its own system is not sufficient evidence that the system is safe.
Yet the pledge names the auditors without naming the audit, and that omission matters, so that a reader who asks "audited against what?" finds no answer in the text; there is no shared construct list, no pass or stop condition, no mapping to the laws that already apply, and no commitment to publish results in a form that anyone outside the company can check.
The pledge's risk categories are also telling, since cyber, bio and chemical threats are the catastrophic tail that dominates frontier-safety discussion, while the harms people encounter every day sit elsewhere: in the private, often emotionally charged conversations where a teenager, a patient or an exhausted caregiver turns to a chatbot for advice, reassurance or companionship, and where a response that is fluent, warm and wrong can do real damage.
State legislatures have noticed this second category before Washington did, which is why New York, California and Colorado have moved on companion chatbots, AI in therapy and automated decisions, and why California's governor signed another thirteen AI bills on October 2; a voluntary national pledge that is silent on these harms leaves companies facing a patchwork of obligations with no common instrument for measuring whether they meet them.
What we found when we measured
At ioLite Labs we set out to build the instrument the pledge assumes but does not supply, working first with clinicians and psychometricians to define what a psychologically safe AI response actually looks like, and only then turning that definition into an evaluation engine.
The result is a taxonomy of 93 constructs that score the model's responses rather than the user's prompts, because the question that matters is not whether a user said something risky but whether the system answered responsibly; a subset of those constructs are marked as stop conditions, so that a single response crossing one of them counts as a failure no matter how well the rest of the conversation goes.
We then ran persona-based, multi-turn conversations against 17 open-weight models in 21 configurations, 5,676 graded conversations with 50 personas, repeating runs three times on randomized schedules so that we could see not only how models behave but how consistently they behave, and none of the 21 configurations passed in personal and private conversations. These are early results, and a clinician review of the grading is in progress.
The inconsistency was as troubling as the failures, since the same model, given the same persona, often responded differently from one run to the next; a system that passes a safety check on Tuesday and fails it on Thursday cannot be certified by a single test, and a one-time audit of such a system offers a snapshot rather than an assurance.
We also overlaid the taxonomy on existing law, including New York's companion-chatbot provisions, California and Colorado statutes, the Blueprint for an AI Bill of Rights and the EU AI Act, and found that even New York's law, among the most specific on companion AI, reaches fewer than a third of the constructs; compliance with today's rules and safety in the conversations people are having are therefore two different measurements, and both need to be taken.
What independent auditing should look like
If the pledge's promise of external auditors is to mean anything, the audits themselves need properties that a press release cannot supply, and I would propose five as a minimum:
- A published standard. Auditors should measure against constructs that are defined in advance, clinically grounded and open to scrutiny, so that two auditors examining the same system reach comparable conclusions.
- Response-level scoring with stop conditions. The unit of evaluation should be what the model actually says, with clearly marked lines that, once crossed, fail the system outright rather than being averaged away.
- Repeated, randomized testing. Because model behavior varies from run to run, certification should rest on repeated trials and on continuous monitoring after deployment, not on a single pass before launch.
- Regulatory overlays. Each finding should map to the specific statutes it implicates, so that a company learns not only that a response was unsafe but how far it stands from what New York, California, Colorado or the EU already require, and what it must change to close the gap.
- Real independence. The auditor's methodology, funding and reporting line should be separate from the developer's, and from the political bodies that are simultaneously investigating and courting the same companies.
This is the model financial reporting arrived at long ago, where a company's own accounts are necessary but never sufficient, and where an outside auditor works to a known standard; AI safety needs its equivalent of SOC 2, and the week's events suggest that the window for building one is open now.
The pledge is the starting line
I read this week as genuine progress, because the companies that build frontier systems have now said in writing that outside eyes are needed, and a regulator has signaled that consumer harm from AI agents is squarely within its remit; yet a pledge describes intentions, whereas an audit produces evidence, and the public, the courts and the companies' own boards will ultimately ask for the second.
At ioLite Labs we are building compliance audits, certification and, in time, a gateway that applies the same evaluation in real time, and we would welcome conversations with developers, deployers in schools and healthcare, policymakers and fellow researchers who want to turn this week's commitments into something that can be measured, repeated and trusted.
If you are working on any side of this problem, I would like to hear from you: canbaz@iolitelabs.com.
Sources
- Fortune: AI's biggest players promise to police themselves at the White House (Oct 1, 2026)
- SecurityWeek: FTC is investigating OpenAI and Anthropic over possible risks to consumers
- The Next Web: Trump names four officials to lead his Super Intelligence Force (Oct 4, 2026)
- Axios: Anthropic, OpenAI CEOs call for slowdown in AI development (Sep 12, 2026)
- The Next Web: Newsom signs 13 AI bills, including a ban on AI-only firings (Oct 2, 2026)