Structural AI systems

An AI system with an org chart.

Most AI in business is a chatbot bolted to a workflow, or a pile of point automations nobody owns. Neither survives contact with a real operation. We build systems that are designed, organized and architectured to think and operate like a business — with hierarchy, decision rights, cross-functional teams and a security function — so the work has an owner even when the work changes.

The anatomy

Four layers, and a security function through all of them

Each layer answers a question the one below it cannot. Governance says who may decide. Teams say who does the work. Agents say what each one is allowed to touch. Models say which intelligence handles the task. Skip a layer and you get the thing most companies have now: capability with no accountability.

The four layers of a structural AI system — governance, teams, agents, models — with security and governance running through all four
Drawn from our own org record and security scorecard, redrawn on every deploy.

Hierarchy

The org chart is the product

A system that cannot say who owns a decision will make that decision badly, repeatedly, at machine speed. So the first deliverable of every engagement is the chart: the units, what each one decides without asking, what it must escalate, and where a person sits in the loop. The agents are hired into it afterwards — role card, exam, placement, certification — the same way a company staffs a department.

The SteelWorks org chart: command, control and assurance, delivery, revenue and capability tiers with agent counts per unit
Our own chart, from the live org record. Units moving to sister companies are shown moving.

Model selection

The strongest model for each task — decided by the record, not the brochure

Every turn the fleet takes is graded and written to a scoring database. A model earns a task by out-scoring the others on that task, and routing is frozen below an evidence floor: under thirty graded turns a lead is noise, and the lane keeps its default rather than chasing it. That restraint is the point. It is why the table below has a row that openly says no model has earned the lane yet.

Model selection by measured score: task types, models tried, the leader on measured score, and graded turn counts
From the live scoring database, last 30 days. A leader must itself clear the evidence floor.

What we do today

Build and operate proprietary systems that select, govern and measure the strongest available model for each task — local, open and frontier — and route on the result.

What is on the roadmap

Vertical model programs: corpora, evaluation sets and tuned models per market. We will say so when one exists and is measured. We do not claim a trained model we have not built.

Security & governance

Scored against a framework, on ourselves, every week

An agent that can act can be made to act wrongly — by a bad instruction hidden in a document it reads, by a permission nobody revoked, by a secret sitting somewhere it should not. We score our own operation against the NIST AI Risk Management Framework and publish the result, including the controls where we score zero. A finding with no evidence scores zero; nothing is graded on intent.

Twenty-five NIST AI RMF controls across Govern, Map, Measure and Manage, each coloured by our own measured score
Our live scorecard. The same instrument produces your report.

The AI Security & Risk Audit

The reference implementation

We are the first system we built

SteelWorks runs on the thing SteelWorks sells. The org chart above is ours. The controls are ours. The routing table is ours. When we say a structural system keeps working without a person in the loop, the evidence is that this company does.

61

agents, each with a written role and a tool scope

552

scheduled jobs, no hour of the day empty

5,878

graded turns behind the routing decisions

25

security controls measured weekly, published

Straight answers

What makes a system structural rather than just automation?

Structure means a shape that survives change: named units with decision rights, approval gates where a person should decide, agents with tool scopes they cannot exceed, and a report line for every one of them. Point automations break when the work changes, because nothing owns the work. A structural system reassigns it.

Do you train your own models?

Not today. We build proprietary systems that select the strongest available model per task on measured score. Vertical model programs are the roadmap, and we will say so when one exists and is measured rather than claiming it early.

After the build

Who runs it once it is built?

You do, on hardware you own, with the documentation and the org chart handed over. Ongoing care is a separate engagement, never a dependency you cannot leave.

How long before it does real work?

The assessment is first and it is measured in days. A first governed slice — three to five operations with approval gates and reporting — is a fixed-scope sprint. We do not sell a twelve-month transformation before anything runs.


Start here

Tell us what breaks, and we will tell you what it would take.

Leave your company and the operation that costs you the most time. You get a written reply with the shape of the system that would carry it — and an honest note if we think you do not need one.