← Back to Research

Perspectives

How We Try to Break NarrEx Before a Client Ever Sees It

Synthetic Test · NarrEx Internal Team · 9 August 2026

Synthetic Test Series #1 — Healthcare Staffing, PE Buyout

We don't have access to real, confidential deal materials yet. Real deal materials are exactly that — confidential — and there's no responsible way to use them for public testing without the appropriate permissions. So instead of waiting, we build our own.

This is the first entry in an ongoing series. We invent a fictional company, write a set of deal materials for it, plant a set of known inconsistencies on purpose, and run it through NarrEx cold. Then we check the output against our own answer key, and we go looking for anything NarrEx got wrong that we didn't plant. When we find something, we fix it and run the test again.

The setup

Company: Aldermere Healthcare Staffing Ltd (fictional), a UK nursing and care-staffing provider to public-sector healthcare commissioning bodies and private care operators. Buyer: Fairbank Capital Partners (fictional).

The scenario: Aldermere engaged a boutique corporate finance advisor to run a sell-side process on its behalf. That advisor prepared the deal materials — a management presentation and an accompanying financial model — working from figures supplied by Aldermere's own finance team. Fairbank Capital Partners, evaluating the acquisition, engaged a financial due diligence provider to assess the target's historical financial performance and underlying financial information. The FDD team used NarrEx as an additional verification layer across the transaction materials and the model.

We deliberately built the materials to reflect the kinds of inconsistencies that can arise in transactions of this size — smaller advisory teams, tighter timelines, a model assembled quickly rather than audited line by line. Nothing malicious, nothing exotic — just the ordinary friction that can arise when transaction materials are assembled under time pressure across multiple source files and workstreams.

We planted 9 deliberate inconsistencies on top of that baseline:

  • a revenue figure in the management presentation that did not reconcile to the corresponding model period
  • a gross margin inconsistency between the presentation and the model
  • a revenue growth claim calculated against the wrong base year
  • maintenance capex classified differently between the model and the presentation
  • development capex classified differently between the model and the presentation
  • a headcount figure in the presentation that included staff categories excluded from the model's count
  • a customer concentration figure that did not reconcile between the presentation and the model
  • a Return on Capital Employed (ROCE) figure calculated using the wrong numerator
  • a leverage ratio stated in the presentation's summary that was inconsistent with the same ratio implied elsewhere in the same document

Each was designed to test a different part of the system rather than one weak spot.

Test methodology

1 fictional transaction · 1 management presentation · 1 financial model · 9 deliberately planted inconsistencies · 4 test iterations · independent answer key used for validation

What happened

The first run identified some of the 9 planted inconsistencies correctly. It also missed several outright and, more usefully, surfaced defects we hadn't planted at all: a metric that was matched to the wrong reporting period; a ROCE figure that could not be reconciled to the underlying model inputs; once that was fixed, the same metric displaying a percentage-formatting error — an underlying value of 0.128 rendered as "0.128%" rather than "12.8%"; a Cash Conversion metric mapped to the wrong similarly-named line item in the model; and a "number of locations" metric that wasn't mapped reliably because the taxonomy treated the underlying concept inconsistently across sectors.

We are publishing the failures alongside the successful detections. We did not suppress the false negatives or only report the final pass. The objective here isn't to demonstrate a perfect test — it's to expose failure modes.

We traced each defect, fixed it, and ran the test again, across four iterations.

Test outcome:

  • 9 of 9 deliberately planted inconsistencies ultimately identified
  • 8 previously unknown product defects identified and fully remediated
  • 1 additional taxonomy issue identified and partially remediated — the current reporting year now matches correctly, but two historical years still don't

The planted-inconsistency count and the product-defect count are separate: a single underlying defect could affect the detection of more than one planted inconsistency, and not every defect we found traces back to something we deliberately planted.

The honest part

A synthetic test like this is not the same as a messy real-world deal. We know exactly what's wrong because we put it there. A real engagement doesn't come with an answer key. We're not claiming this proves NarrEx works on real transactions. We're showing how rigorously we're trying to find out where it doesn't.

This is test 1. We'll keep running these across different sectors and transaction types, and publish what we find — including when the answer is "this didn't work."

Ready to test narrative-model alignment on your own deals?

Apply