Patent Pending · Independent measurement

Know your AI risk posture before your regulator does.

Boards approve AI they cannot evaluate. M&A closes on systems with no independent risk view. Model cards and vendor self-assessments describe — they do not measure. We produce a decision-grade risk score with reason codes you can defend, and an open format so every tool in your stack reports risk the same way.

Book a 30-min scoping call See what the report contains

No commitment. No NDA required for the first conversation.

Mapped against EU AI ActNIST AI RMFISO/IEC 42001SR 26-2DORA
Shadow Simplex Score
Illustrative
235/1000
Drift band Coherence Score CS = 765
Register vector · ST / ME / DY / TH / KI
Meta-condition status
KILLSAFHITLAUTTRUMAN
C(κ) 0.85 Reason codes included
Who this is for

If any of these apply, the material is relevant.

M&A / PE diligence

You are evaluating an AI-heavy target and need risk expressed in a form counsel and the investment committee can read — not a vendor model card.

Enterprise GRC & legal

You face EU AI Act, NIST, or board disclosure pressure and need an independent, reproducible artifact rather than another policy PDF.

Insurers & underwriters

You need credible, structured profiles to price AI exposure. Self-attested model cards are not sufficient underwriting input.

The problem

No one is keeping independent score.

Benchmarks saturate and get gamed. GRC tools mostly consume other people’s signals. Model cards and self-attestations have no independent measurement layer behind them. There is still no common format that lets risk output travel across your stack.

Benchmarks lose discrimination

They saturate, get gamed, and stop separating systems that are safe enough from systems that are not. A leaderboard is not a risk instrument.

No common language

Every scanner, evaluator, and GRC platform emits a different format. Outputs cannot be ingested, compared, or escalated cleanly.

Vendors grade themselves

Model cards and self-attestations have no independent layer. Findings can be shaped by the party being measured.

What you get

A score you can act on.
A format your tools can speak.

One deployment-level risk score with the explanation needed to defend it, and one open format so every tool reports risk the same way.

The score

Shadow Simplex Score

A numeric, deployment-level risk score with reason codes and a decision-grade explanation. Non-compensatory: any single critical failure collapses the composite. It scores the system under examination, not the paperwork around it.

  • Composite band + Coherence Score
  • Meta-condition status (vetoes that fire)
  • Reason-code spine for every material finding
  • Cross-reference to EU AI Act / NIST / internal policy
How the assessment works →
The format

SREL — open exchange format

A coordinate-based grammar for risk and control exchange. Scanners, evaluators, and GRC platforms can present as one coherent exposure surface instead of incompatible reports.

  • Open specification (CC0)
  • Reference implementations (Apache 2.0)
  • Machine-readable, human-auditable
  • Designed for downstream monitoring systems
Read the SREL spec →
Current status

What is shippable today vs. still earning validation.

We are precise about the difference. The rail and diagnostic gate are usable now. Full predictive calibration of the score is in progress.

We do not claim a fully validated risk score yet.

Today the score rests on a strong geometric prior plus proxy testing (held-out attack distributions and next-vintage performance). Real-incident calibration grows with scored volume. The diagnostic gate and SREL do not require full score validation to be useful in your pipeline today, and every scored output discloses its own coverage.

ShippableAvailable now
  • SREL format, coordinate space, grammar
  • Self-run diagnostic gate (DevSecOps-ready)
  • Public index structure
  • Reason-code spine + explanation structure
EarningIn progress
  • Proxy validation → calibrated predictive score
  • Incident-data volume for calibration
  • Hosted scoring API (preview / early access)
How the score works

Non-compensatory by design.

No dimension rescues another. Any single critical failure collapses the composite. The result is a reproducible number with reason codes, not a narrative opinion.

54
Risk primitives
5
Physics registers
C(κ)
Capability normalizer
7
Meta-condition vetoes
Binary collapse on critical failure
Coherent
0–200

Typical function maintained

Drift
201–400

Subclinical; recoverable

Mixed
401–600

Trajectory matters

Consolidating
601–800

Architectural intervention

Critical
801–1000

Shadow configuration

Full mathematical specification available under NDA to qualified enterprise and investor clients. Covered by US Provisionals 64/066,231 · 64/075,009 · 64/077,286.

Independence

Loss-bearer pays. Rank is never for sale.

We have no financial relationship with model vendors. Findings cannot be purchased into a favorable outcome. The published ranking stays independent of who pays for assessment.

No vendor relationships

Assessments cannot be influenced by the party being measured. We work for the exposed party — buyers, boards, GRC officers, insurers.

Reproducible, not narrative

Every score is reproducible from documented inputs. Qualitative judgments are explicit and auditable.

Open framework, closed calibration

The geometric substrate is published for scrutiny. The durable value lives in calibration and the scored corpus, not secrecy of the method.

Know your AI risk posture before your regulator does.

Start with a 30-minute scoping call. We will tell you which dimensions matter most for your situation and what an engagement would — and would not — deliver at this stage. No fabricated urgency.

Book a scoping call View engagement options