iolitelabs

The engine

The evaluation engine.

Harm from conversational AI is the same wherever it surfaces, so the instrument that measures it has to work everywhere too. The iolite engine applies one clinical taxonomy at every layer where a person meets a model. Today it scores browser conversations and direct API calls. The inference gateway underneath them all is what we are building next.

Architecture

One rubric, every layer.

Like an antibiotic, the engine does not care which organ the infection is in. Each layer passes structured events to the one below it, and the same taxonomy judges all of them. The status on each component shows what runs in production today.

14 of 16 components have working code. 5 run end to end in production.

Live in production · 5Working code, gap named · 9Next: what funding builds · 2

Human–AI interaction layer

Every surface where a person meets a model

Mobile

Partial

In-app SDK and companion app

ioLite Daily ships as a mobile-first app shell with a calm reading experience.

Next · Native iOS and Android apps; an embeddable SDK for third-party apps.

Browser

Live

ioLite Guardian extension

Chrome extension that classifies both sides of AI chats on-device, with a live backend, parent and district dashboards, and a Web Store submission.

Direct access

Partial

API and agent middleware

The audit engine drives target model APIs directly; Guardian exposes a classification API.

Next · Agent middleware that scores tool-using agents in-line.

Enterprise

Live

Schools, clinics, workplaces

Organizations, role-based access, district-level analytics, and Admin Console managed-device policy.

Data gathering layer

Capture signals without capturing people

Consent and privacy

Live

Opt-in, on-device redaction

Affirmative consent, on-device scoring, pseudonymous identity, and a server that refuses to store safe conversations.

Signal extraction

Live

Psychological safety markers

Both sides of the conversation are scored on the device against the taxonomy; flagged moments are confirmed before anyone is alerted.

Context tagging

Partial

Domain, role, vulnerability

Every flagged moment carries its site, surrounding turns, the monitored person's role, session length, and time of day.

Next · Explicit vulnerability and domain tagging.

Model telemetry

Partial

Which model, version, prompt

Audit targets record provider and model for every run.

Next · Model and version capture from live conversations in the field.

Compute and infrastructure layer

Where the models actually run

Model endpoints

Partial

Cloud, on-prem, edge

An encrypted target registry that evaluates cloud-hosted models.

Next · On-prem and edge deployment targets.

Inference monitoring

Next

Live scoring at the gateway

Not built. Guardian proves the scoring approach on the client side.

Next · A gateway service that scores live traffic with the same taxonomy.

Deployment gates

Next

Block, warn, or pass releases

Not built.

Next · Release gates that block a model version that fails safety-gate constructs.

Vendor scorecards

Partial

Model-to-model comparison

Audit rollups by construct and strategy, and evaluation seasons for re-runs.

Next · Public model-to-model comparisons, starting with Study 01.

The loop

1. Detect

Live

Diagnose the harm wherever it surfaces

Audits detect failures before release; Guardian detects them in live conversations.

2. Neutralize

Partial

Block, warn, or reroute in real time

Guardian warns the responsible adult in real time and surfaces crisis resources.

3. Immunize

Partial

Certify, retest, and keep it from recurring

Re-audits run by season against the same probe sets.

Taxonomy provenance

Built from frameworks, not from opinion.

The taxonomy is not a list we invented. It is assembled from validated clinical and AI-evaluation frameworks, and every node records its source. The structure and the counts are public; the sources and the construct definitions are shared under NDA.

3

Priority domains

21

Constructs

69

Sub-constructs

9

Safety gates

01

Adopted verbatim

Construct definitions are taken as written from their source frameworks, so each one can be traced back to its origin in an audit.

02

Organized into three levels

Priority domains contain constructs, and constructs contain the measurable sub-constructs that probes target.

03

Safety gates designated

The constructs where a single failure matters regardless of everything else, such as crisis and suicide-risk management, are marked as gates.

04

Mapped to the law

Every construct carries the regulatory requirements it bears on — New York's AI-companion duties, California SB 243, the EU AI Act — so an audit reads as compliance evidence, not only as a score.

Provenance, by the numbers

4

Source frameworks

93

Taxonomy nodes

9

Safety gates

Domains: Safety, Privacy, and Fairness · Trustworthiness and Usefulness · Design and Operational Effectiveness

Which frameworks, and how the 93 nodes divide between them, is part of the instrument. Customers and partners see it under NDA.

The evaluation pipeline

From scenario to evidence.

Every score is produced by the same chain, and every link leaves a record.

01

Scenario

A reviewed, multi-turn scenario written against one sub-construct. The library is versioned, and the same scenarios are re-run on every model version.

02

Run

The system under test is taken through the scenario turn by turn, and the full transcript is kept.

03

Score

Each transcript is scored from 0 to 4 against the construct it targets; 3 or above passes. Every score carries a written rationale.

04

Check and report

Transcript and rationale travel with every score, so a clinician can dispute any judgment; we will publish the clinician agreement rate once it is measured. Results roll up by construct into a graded report, with safety-gate failures shown on their own.

Resistance

Why the instrument doesn't decay.

Anyone who knows the antibiotics metaphor will ask about resistance: models change, so won't they learn to pass the test? They will try. That is the reason the engine is built to be re-run rather than run once.

A static filter decays as models change. A taxonomy that is re-run against every new model version accumulates a longitudinal record of exactly that change, including tuning that learns the test instead of learning the behavior. Each re-run adds labeled transcripts to the corpus, and the corpus sharpens the next set of probes.

That is why re-audit is structural rather than a contract term. Models ship new versions, so safety has to be measured again every time they do.

Version-over-version scores

Version-over-version results for the first model family will appear here once Study 01 is published.

Where the engine runs today

Proof at the surface, not a second business.

ioLite Guardian brings the engine into the browser for families and schools, and ioLite Daily shows what a calm, carefully measured information product looks like on a phone. Both exist to demonstrate the standard where people actually use AI. The business is the audit and the certification behind them.

The list

Get the measurements the day they publish.

One email per release, with a button that opens the full report. Unsubscribe with one click.

Put your system through the engine.

Vendors can request an audit of their own system. Schools, clinics and other institutions deploying AI can talk to us about what the engine measures for them.