Celagenix AI Assurance · Free tool

The AI Confidence Assessment

AI is fluent by default and correct only by design. In ten minutes, score where your AI can be confidently wrong, and whether anyone would catch it before it reaches a customer, a regulator, or a board paper.

30 questions6 dimensionsNo email requiredScored live
0%Confidence
Your standing
Not yet scored
Answer the six dimensions below to score where your AI can be confidently wrong, and what to do about it. No email required, your answers stay in this browser.
How it works

Six dimensions of AI confidence

Each dimension follows our Assurance method, Expose, Ground, Gate, Evaluate, Monitor, Account. Answer honestly: score your organisation as it operates today, not as you intend it to. Progress saves automatically in this browser.

Yes, in place = 2
Partial = 1
No = 0
01
Expose
Knowing where AI can be wrong
0/10
You cannot control what you have not found. This dimension asks whether you know every place an AI or LLM shapes a decision, a document, or a customer interaction, including the tools staff adopted on their own, and whether you have ranked them by the cost of a confident error.
C-EXP-1.1
Do you have a current, maintained inventory of every place AI or an LLM produces output that informs a decision, a document, or a customer interaction, including tools staff have adopted on their own?
C-EXP-1.2
Has each AI use been risk-ranked by the consequence of a wrong-but-confident answer (regulatory, financial, safety, reputational), so effort concentrates where it matters most?
C-EXP-1.3
Do you know, for each AI use, whether a human meaningfully reviews the output before it is acted on, or whether it flows straight through?
C-EXP-1.4
Have you identified the specific ways each system can be confidently wrong (fabricated facts, outdated rules, wrong jurisdiction, misread document), rather than assuming errors would be obvious?
C-EXP-1.5
Is there a named owner accountable for the correctness of each AI-assisted output, rather than ownership diffused across IT, vendors, and the business?
02
Ground
Anchoring answers in your truth
0/10
A model is fluent by default and correct only by design. This dimension asks whether your AI answers from your own authoritative, current, jurisdiction-correct sources, whether it can cite them, and whether it declines rather than guesses when it has nothing to stand on.
C-GRO-2.1
Are your AI systems answering from your own authoritative, current source material (policies, contracts, data), rather than from the model's general training?
C-GRO-2.2
When an AI gives an answer, can it cite the specific source passage it relied on, so a reviewer can verify it in seconds?
C-GRO-2.3
Are the sources the AI draws on kept current, so it is not confidently quoting a superseded policy, price, or regulation?
C-GRO-2.4
For regulated or jurisdiction-specific work, is the AI grounded in the law that actually applies to you, rather than defaulting to the most common (often foreign) framework in its training?
C-GRO-2.5
When the AI has no good source for a question, is it built to say so, rather than to produce a fluent guess?
03
Gate
Catching it before it lands
0/10
Assurance is executable, not aspirational. This dimension asks whether real checks sit between the AI's output and its use, whether high-stakes work is genuinely held for sign-off, and whether those controls test the things that actually go wrong.
C-GAT-3.1
Are there automated checks between the AI's output and its use that can block or flag an answer before it reaches a person or a customer?
C-GAT-3.2
On high-consequence outputs, is human sign-off required and enforced by the system, not merely expected by policy?
C-GAT-3.3
Do your controls check the things that actually go wrong (numbers, dates, names, citations, prohibited content), rather than only formatting or tone?
C-GAT-3.4
When a check fails, is there a defined path (hold, escalate, correct), rather than the output simply proceeding anyway?
C-GAT-3.5
Are the controls themselves version-controlled and testable, so you can show which were in force at a given time?
04
Evaluate
Proving it, before it ships
0/10
You cannot manage a confident-error rate you have never measured. This dimension asks whether AI uses are tested against known-correct cases, whether hallucination is quantified, whether systems are stress-tested, and whether the evidence would survive an auditor.
C-EVA-4.1
Before an AI use goes live, is it tested against a set of known-correct cases to measure how often it is wrong?
C-EVA-4.2
Do you measure and track a confident-error (hallucination) rate for each system, rather than relying on general impressions of quality?
C-EVA-4.3
Has each high-stakes system been deliberately stress-tested with hard, adversarial, or edge-case inputs to find where it breaks?
C-EVA-4.4
When you change a model, prompt, or data source, do you re-test before shipping, rather than assuming an improvement did no harm?
C-EVA-4.5
Do you hold evidence of these tests (what was tested, the results, who signed off) that would stand up to an auditor or regulator?
05
Monitor
Watching it in production
0/10
A model that was right at launch can drift, and a vendor can change it under you without notice. This dimension asks whether live output is logged, watched for drift, alerted on, and re-verified when the ground moves.
C-MON-5.1
Are live AI outputs logged and traceable, so any given answer can be reconstructed and reviewed after the fact?
C-MON-5.2
Do you monitor production output for quality drift, rather than only testing once at launch?
C-MON-5.3
Are there alerts that fire when the AI behaves abnormally (error spikes, refusals, unusual inputs), rather than waiting for a complaint?
C-MON-5.4
When a vendor updates the underlying model, do you detect it and re-verify, rather than being silently exposed to a changed system?
C-MON-5.5
Is there a defined response when monitoring catches a problem (roll back, disable, notify), with someone on the hook to act on it?
06
Account
Standing behind the output
0/10
In the end someone has to answer for what the AI did. This dimension asks whether there is an audit trail, an incident process, honest disclosure where it matters, a mapping to the rules that bind you, and the ability to prove an answer was correct, not merely plausible.
C-ACC-6.1
Is there an audit trail showing, for a given AI-assisted decision, what the AI produced, what a human did with it, and on what basis?
C-ACC-6.2
Do you have an incident process for when a confidently-wrong AI output causes harm, on par with other operational incidents?
C-ACC-6.3
Where it matters, are people told when AI materially shaped a decision, document, or answer that affects them?
C-ACC-6.4
Is your AI use mapped against the obligations that apply to you (POPIA, GDPR, sector rules, the EU AI Act where relevant), rather than assumed compliant?
C-ACC-6.5
Could you demonstrate to a board, auditor, or regulator that a given AI output was correct and controlled, not merely plausible?
Your result

Where you stand, and what to do next

Your confidence band updates as you answer. The lower the score, the more likely a confident error reaches someone before a control catches it.

0–25% · 0–15 pts
Exposed
26–50% · 16–30 pts
Reactive
51–75% · 31–45 pts
Controlled
76–100% · 46–60 pts
Assured
Answer the dimensions above to reveal your confidence band and your recommended next step.

Book a Confidence Review

Your exposure is high and the controls are thin. A Confidence Review starts with the Expose stage of our Assurance method: we map where your AI can be confidently wrong, rank it by consequence, and hand you the two or three gates that remove the most risk fastest.

  • Where AI touches decisions, ranked by the cost of a wrong answer
  • The specific failure modes your systems are exposed to
  • A prioritised set of controls, starting with your highest-stakes output
Book a Confidence Review

Turn checking into control

You are catching some errors, but by hand and after the fact. A Confidence Review shows where grounding and executable gates would replace manual review, so the wrong answer is stopped before it reaches a person, consistently rather than occasionally.

  • A map of what is checked today versus what actually goes wrong
  • Where grounding in your own sources closes the biggest gaps
  • Executable gates for the output that carries the most consequence
Book a Confidence Review

Close the gaps and prove it

The foundation is in place. The remaining value is in coverage, measurement, and evidence: quantifying your confident-error rate, extending controls to the pipelines still uncovered, and building the trail that proves control rather than asserting it.

  • A measured confident-error rate for your key systems
  • Coverage extended to the AI uses still running unchecked
  • Evidence a board, auditor, or regulator would accept
Book a Confidence Review

Keep it assured as the ground moves

You are in strong shape. The risk from here is drift: a vendor changing a model, a new use going live unchecked, a regulation shifting. We help you hold the line with monitoring, re-verification, and a standard that every new AI use must clear before launch.

  • Monitoring and drift-detection across live systems
  • A launch gate every new AI use must pass
  • Re-verification when models, sources, or rules change
Book a Confidence Review
From assessment to assurance

See where your AI is confidently wrong

This assessment is the Expose stage of our Assurance method, run yourself. When you want the controls built, we ground, gate and monitor the pipelines that carry the most risk, proven on the products we run.

Book a Confidence Review