Home / Essentials / The comparison

MyRISK Essentials · the detailed comparison

Same input. Two very different outputs.

A capable team with a custom GPT or Copilot can produce a ranked risk register in an afternoon, at no product cost. That is worth saying out loud, and worth showing beside what an Essentials baseline returns from the identical starting profile. Four sector cases, both ways, no input changed between them — and the shorter list is the better one, for reasons this page sets out.

Read this first

  • Fourteen cases, built from public information — among them an automotive marketplace, a research university, a mid-market bank, an electricity network, local government, aged care, agriculture and a charity. No customer record appears here, in whole or in part.
  • The comparison is against a custom GPT or Copilot — two prompts, rank the risks then suggest treatments. Not a MyRISK product: it is what the tools most teams already have will give them, and it is the fair thing to measure against.
  • Counts are per run and stated as ranges across the four cases. Structural counts — how many scenarios, how many traced rows — are properties of the pipeline. Content counts vary between runs, because the enrichment search is never the same twice.
  • The zeros are not a bad score. They are stages a custom GPT or Copilot does not have, so there was nothing that could have produced a number.

The gap, in four numbers

What one has and the other structurally cannot

Essentials first in each pair, then what a custom GPT or Copilot returned from the same starting profile. The second line of each is why the number is worth anything.

83 against none Sourced observations, across the fourteen cases: between 1 and 16 per run, each carrying where it came from and a hash of it. A custom GPT or Copilot returned none, having no stage that takes evidence in. It matters because a reviewer can check an observation that names its source. An assertion they can only believe or not.
4 every time Scenarios modelled on every single run — downside, base case, adaptive and renewal, 56 across the fourteen. The custom GPT modelled none — it describes the present and stops. It matters because a risk written only in the present tense has no signposts, so nobody can tell whether it is arriving.
6 to 15 rows against none Traceability. Every board risk followed from the evidence that raised it, through the consequence, to the action now being taken — 150 rows across the fourteen. The custom GPT produced no chain of any length. It matters because this is the difference between explaining a decision a year later and re-arguing it from memory.
14 stages against 2 Essentials runs a typed pipeline from observation to board paper. A custom GPT or Copilot is two prompts: rank the risks, then suggest treatments. It matters because the twelve stages in between are where prioritisation happens. Skipping them is what returns thirty undifferentiated risks instead of eight that are yours.

What each returned

Ranges span all fourteen sector runs. The first three rows are what a custom GPT or Copilot does well — and it does them well. Everything below them is what it has no stage that could produce.

What a custom GPT or Copilot and an Essentials baseline each returned across fourteen sector cases.
DimensionCustom GPT or CopilotEssentials baseline
Ranked risks30, every case6–15 board-level
Treatments and controls270 actions, every case6–8 sequenced priorities
Grounded in a web searchYesYes
Sourced observationsNone1–16
Event clustersNone4–7
Domains mappedNone5–9
Causal driversNone5–6
Propagation effectsNone9–12
Exposed cohortsNone4–11
ScenariosNone4
Enterprise consequencesNone6–11
Traceability matrixNo rows6–15 rows
Per-run counts across fourteen sector cases, anonymised to sector. One case failed a semantic-validation check on its first run and passed on a re-run — a fail-closed guard, not a silent error.

The same bank, both ways

Generic best practice, or the exposure that was actually there

One mid-market bank, identical input. The custom GPT returned sound, sector-generic controls. Essentials found a different risk entirely — and can show where it came from.

What the custom GPT returnedExample

ONE OF 30 RANKED RISKS · TREATMENT: REDUCE

Cyberattack compromising banking services and customer data

Nine suggested actions, of which five:

  • Patch internet-facing critical vulnerabilities to risk-based deadlines
  • Harden privileged access with phishing-resistant multi-factor authentication
  • Segment critical banking systems from user and supplier networks
  • Deploy behavioural detection across identities, endpoints and cloud
  • Adopt zero trust across workforce, workloads and third parties

Every line of that is correct, and none of it is about this bank. Nothing ties it to the bank's own evidence, nothing says which exposure is material, and nothing shows a board how it was arrived at.

What the Essentials baseline returnedExample

ONE OF 9 BOARD RISKS IN THIS CASE · TRACED END TO END

Further regulatory enforcement, capital or operating constraints

Grounded in three sourced observations:

  • A regulator had imposed an operational-risk capital add-on
  • An AML and counter-terrorism-financing enforcement investigation was open
  • Regulators had stated that further action remained possible

Then followed through:

  • Cluster: AML/CTF and non-financial-risk governance
  • Enterprise consequence: licence-to-operate pressure, with a remediation backlog saturating control delivery
  • Current action: regulatory watch, obligation-impact assessment, delegated authority, and decision and evidence logging
  • Modelled downside: regulatory escalation tightens first, while constrained control-analysis capacity slows the fixes and a complaint spike compounds the load

The custom GPT did not return this risk at all. Same bank, same starting profile.

Why thirty ranked risks is worse than eight

The longer list looks like better value. It is not, and this is the part a comparison table cannot show on its own.

Two hundred and seventy things to do is a list nobody runs

Thirty risks with nine actions each is 270 actions. No organisation sequences 270 actions, so what happens instead is that the document is filed and the team carries on doing what it was already doing. A baseline that returns eight to twelve board risks and six to eight sequenced priorities has done the part that is actually hard: deciding what comes first. A list of thirty has handed that decision back to you unmade.

The same thirty would come back for any bank

Ranked risk lists generated this way are stable across organisations in the same sector, because the input that distinguishes you — your evidence, your obligations, your open matters — never entered the process. If the answer would not change for your competitor, it is not telling you anything about you. The ranking is by plausibility, not by anything measured.

Good advice pointed at the wrong risk still costs you

Every control in the cyberattack example above is sound. Patch internet-facing vulnerabilities, harden privileged access, segment critical systems: no reviewer would argue with any of it, and that is exactly what makes it hard to challenge.

But this bank's material exposure was regulatory — an imposed capital add-on, an open enforcement investigation, a regulator saying more could follow. A cyber programme of that size consumes the control-delivery capacity the regulatory remediation was already short of. The baseline modelled that directly: escalation tightens first, while constrained analysis capacity slows the fixes. Correct advice aimed at the wrong risk is not a neutral outcome; it spends the budget and the people you needed for the thing that was actually going to hurt you.

And you cannot defend any of it afterwards

If a regulator asks why cyber patching was prioritised over the enforcement remediation, the honest answer is that a tool suggested it. There is no observation to point at, no reasoning recorded, and nothing that shows what was known at the time. That is the difference between a list you acted on and a decision you can explain.

Why the gap is structural

A custom GPT or Copilot is two prompts: rank, then treat. It has no stage that takes in evidence, models a system, or tests a future. These are not features it is missing. They are the difference between a list and an analysis.

  • Evidence and provenance. Every observation carries a source and a hash, so it can be checked rather than taken on trust.
  • Causal analysis. Root conditions, drivers and a map of how domains are coupled, instead of a flat list.
  • Propagation. First- and second-order effects, feedback loops, and the critical paths between them.
  • Scenario testing. Four distinct futures with signposts, rather than a single present-tense view.
  • Enterprise translation. Risks mapped to consequences, obligations and what a board is accountable for.
  • Who is exposed. The cohorts carrying disproportionate exposure, and the fragilities behind that.
  • A sequenced plan. Prioritised controls pointed at an intervention, rather than four generic verbs.
  • Nothing silently dropped. Validation that no material risk disappeared between stages.

What this comparison is not

It is not a benchmark, and it is not a claim about any particular AI tool. It is fourteen cases, run both ways on one date, with the inputs held still — enough to show where the difference sits and not enough to put a percentage on it. A re-run would not return these exact numbers; on this evidence it would return the same shape. Your own material is the only comparison that settles it for you, and the Baseline is built from what you already have.

Who asked you to prove something in the last 90 days?

Send that request, and the spreadsheet you answered it from. The Baseline is built from what you already have.