Flutterfrog

AI transformation

We automate one document flow. Properly.

Quotations, tenders, technical submittals, compliance filings — the work your team sends out over and over, where there is a right answer and somebody senior checks every one. We automate that flow, and check every document against your own standard before anyone sees it.

Why this usually fails

The output looks right, so nobody trusts it.

Most AI pilots produce drafts that look correct and might not be. So somebody senior reviews every one by hand — which is the work you were trying to remove. MIT’s Project NANDA studied enterprise adoption across structured interviews with 52 organisations, survey responses from 153 senior leaders, and a review of over 300 publicly disclosed AI initiatives, and found around 95% of organisations reported no measurable effect on profit and loss.

MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 — research conducted January to June 2025.

Our reading of it

Nobody wrote down what a good output was supposed to look like.

If it was never written down, nothing can be checked. If nothing can be checked, nobody puts their name to it. So it stays a demo.

Which is why we start with your standard, not with the AI.

What we build

Your standard, written down and built in.

Ask what makes one of your quotations correct and you get a specific list — twelve to twenty checks for most businesses. They already exist. They live in the head of whoever does the final read before it goes out.

We write them down with your people and build them in, so every document is checked before it reaches a desk. That is what makes the output predictable — not that the model is clever, but that you can see which check passed and which did not.

Reviewing this for security or procurement? Where your data sits, and what we do not claim.

A real standard, for a quotation

  • Every line priced from the current price list, not a superseded one
  • Tax treatment matching the customer's registration and delivery state
  • Payment terms as agreed with this customer, not the default
  • The clauses legal mandates, present and in the approved wording
  • Specification matching the enquiry, with deviations listed rather than buried

The written standard is yours, in your words. It is the part no general-purpose AI product can supply, because nobody outside your company knows what your quotation is supposed to say.

Is this for you

Five questions. Take one document and answer honestly.

This works on non-creative work — where output is judged on being correct and complete, not on being original.

1

Is the output a document?

Not a physical product, not a conversation.

2

Does the same document recur — more than fifty a month?

With the details changing every time.

3

Can you state what “correct” means?

A checklist, standard or regulation exists. If the honest answer is “you'd know it when you see it”, stop here.

4

Are the inputs scattered?

Somebody has to go and ask, dig through a drive, or remember.

5

Is getting it wrong expensive?

Rework, a lost tender, an audit finding, a penalty.

Yes to all five

Quotations and tenders · technical submittals · quality and customer documentation · regulatory and compliance filing · export documentation · contract and vendor administration.

Not this

Brand, design, campaigns — anything judged on originality. You cannot write a check for taste, and we will say so rather than sell you something that cannot be measured.

How we keep it predictable

A result you can check, not one that sounds right.

The reason most AI work fails in a business is not that the model is weak. It is that nobody decided in advance what a correct answer looks like — so there is nothing to check the output against, and a confident wrong answer sails straight through.

The shape of every flow we build

The model proposes. Something that cannot be talked round decides. A person confirms before anything leaves the building.

InYour existing recordsAs they are. No clean-up first.
AIProposes a resultWith its confidence, field by field.
GateDeterministic checkRules, not judgement. Passes or stops.
PersonConfirmsAgainst the evidence, not the claim.
OutResult you can auditEvery field traceable to its source.

Stops when: any required field is missing, malformed, or below the confidence bar

Measured
The gate is the difference between automation and a fast way to be wrong.

01

Correct is defined before anything is generated

The rubric is written first, from your requirements. A standard invented after the fact is not a standard.

02

The output is judged, not the thing that made it

We score the finished document as you would see it. Source that looks right and renders wrong is the common failure.

03

The generator never edits the rubric

The most reliable way to pass a test is to change it. If the same process writes the standard and meets it, the score means nothing.

04

The loop is bounded

A fixed number of attempts, then it stops and asks a person rather than burning budget converging on nothing.

05

Nothing mechanical is judged by a model

Totals, dates, references and formats are checked by rules. A model is only asked what a rule genuinely cannot settle.

The honest limit

Nothing is correct every time, and anyone who tells you otherwise is selling. What this gives you is different and more useful: when it is wrong it stops instead of shipping — and when something does slip through, you can trace exactly which check should have caught it and add that check.

How it works

One flow at a time.

We do not start with a platform rollout. We take one document flow, make it work, and measure it — then extend once it holds.

01

Pick the flow

One document, one team. The one where somebody senior checks every item before it goes out — that is the one worth automating first.

02

Write the standard

We sit with your people and write down what makes that document correct. Twelve to twenty checks, in your words. You keep this whether or not we build anything.

03

Build and check

Your knowledge loaded, the document produced, and every one checked against your standard before a person sees it. Anything that fails is corrected and re-checked.

04

Measure, then extend

We baseline turnaround and rework before go-live, so the change is measurable rather than asserted. Once that flow holds, the next one is much faster.

What we are still building

The checking engine is being extended now. The knowledge layer and the production toolchain are in daily use — we run our own company on them, and the checker runs deterministic checks today. What an engagement buys is your standard written into it, with your people. That work has to happen either way, and we would rather say so now than have you find out in month three.

Which document is eating your team’s week?

Tell us that, and we will tell you honestly whether it is a fit, what automating it would take, and what it would cost. If the answer is that you do not need us, we will say that too.

Looking for something built rather than a flow automated? That is our other service — Custom Project: the system your business runs on, built for you and owned by you.