Skip to main content
Success
[PRO SERVICES / SECURITY & GOVERNANCE]

AI Workflow
Safety Review

A live AI workflow can pass a demo while its permissions, failure handling and human oversight remain unclear. We review the system in use and give technology and risk owners a ranked set of fixes.

HOW IT WORKS
1

Map the workflow

2

Test how it can fail

3

Fix the weak points

You know what could go wrong and what to fix first.

10 risks

OWASP LLM TOP 10, 2025

Access

RETRIEVAL AND TOOL PERMISSIONS

Evidence

KEPT WITH EACH FINDING

[THE PROBLEM]

Live AI workflows often miss an independent review

A feature can produce useful answers while its data access, failure handling or support arrangements remain unclear. Those gaps matter once staff or customers rely on it.

We review the implemented workflow and the evidence behind its controls, including what happens when a model gives a wrong answer or a connected service fails.

The report records confirmed findings, recommended changes and what the testing did not cover.

WHAT YOU'VE GOT

  • A chatbot taking customer questions
  • Agents that send email, file invoices, edit your CRM
  • RAG over docs nobody's checked for PII
  • Copilot turned on across the company
  • Staff using ChatGPT on personal accounts
  • No eval suite, no logs, no kill switch

WHAT YOU NEED

  • A written threat model per workflow
  • Agents scoped to least-privilege tools
  • Retrieval sources classified and reviewed
  • Tenant boundaries enforced by the application
  • Identified tools with owners and permitted uses
  • Evals, traces, alerts and a way to switch it off
[WHAT WE LOOK FOR]

What we test first

The test plan uses relevant OWASP risks, NIST risk-management guidance and ICO expectations. Coverage depends on the workflow's data and authority.

01 / LLM01

Prompt injection

Instructions in messages, documents or tool results may redirect the model. We test the entry points and the application's response to unwanted instructions.

02 / LLM02

Data leakage

Prompts, retrieved records, logs and outputs can disclose data beyond its intended audience. Tests check both provider handling and access between users.

03 / LLM06

Excessive agency

An agent may be able to send messages, issue refunds or edit records. We check whether those permissions exceed the task and whether approval is enforced outside the model.

04 / LLM09

Confident wrong answers

We test answers against current sources and check how the workflow handles missing or conflicting evidence. A human escalation route may be needed for consequential advice.

05 / SHADOW AI

The AI you don't know about

Personal accounts, extensions and connected automations can sit outside the approved inventory. Discovery covers the agreed teams and technical evidence, with any gaps recorded.

[HOW WE WORK]

What the review gives you

One report, fixed price. We talk to the people running the workflows, read the prompts and the code, run the attacks ourselves, then write it up in plain English with a ranked list of fixes.

The report gives the operational owner and developers evidence, priority, recommended controls and a clear route to verification.

BOOK A REVIEW
01

Map the workflows

We record the workflows found within the agreed scope, with owners, models, data sources and connected tools. Interviews and technical evidence help identify unregistered use.

02

Read it, then attack it

System prompts, retrieval pipelines, tool definitions, model configs. Then we try to break them: direct and indirect prompt injection, tool abuse, tenant escape, output rendering tricks, data exfiltration via markdown images. The OWASP LLM Top 10, run for real against your stack.

03

Check the regulator angle

We assess data-protection requirements and any EU AI Act exposure against the system and your role. NIST and ISO/IEC 42001 can support the control review, while certification is a separate process.

04

Hand over a ranked fix list

One document. Findings ranked by impact and effort, each with the evidence, the OWASP or RMF reference, and the fix written so a developer can apply it. A page on shadow AI with the tools you ought to be giving people instead. Optional follow-up if you'd like us to do the fixes.

[RELEVANT VU WORK]

We apply these controls to software we operate

Our own AI systems use scoped access, human approval, logs, cost controls and recurring security review. A client review uses the same engineering questions against the workflow that is live in their business.

[A USEFUL FIRST CONVERSATION]

When this is worth discussing

We work best when there is a real operating problem, enough volume to measure and people from the affected teams who can make decisions.

Usually a good fit

  • An established UK business, usually with annual revenue above £10m
  • A repeated process with a known cost, delay, error rate or capacity problem
  • A senior sponsor and a day-to-day owner who understand the work
  • Access to the relevant staff, systems, sample records and security requirements

We may point you elsewhere

  • A standard product already covers the process well
  • The requirement is a one-off small build with no wider operating case
  • There is no owner or access to the people and data needed to test the result
  • The plan relies on AI making high-impact decisions with nobody responsible for review
[QUESTIONS]

Questions from IT, legal and compliance

Q.01

Isn't this just a pen test?

The scope overlaps with penetration testing. This review adds model inputs, retrieval, tool use, output handling and operational oversight to the application checks. We agree coverage with any existing testing programme.

Q.02

Does the EU AI Act apply to us?

If you put AI on the EU market, or its output is used in the EU, the Act may apply. It became generally applicable on 2 August 2026. Following Regulation (EU) 2026/1744, Annex III high-risk rules apply from 2 December 2027 and product-related high-risk rules from 2 August 2028. We map the system and your role before stating which duties apply.

Q.03

What about the UK? There's no AI Act here.

UK data protection, equality, consumer and sector rules can apply to AI systems. We identify technical evidence your privacy and legal owners need to assess the relevant duties.

Q.04

We're going for ISO 42001. Does this help?

It can supply findings and evidence for an AI management system. ISO/IEC 42001 certification covers a wider organisational scope and requires an independent certification process.

Q.05

What do you need from us?

We agree read-only access to the relevant code, prompts, retrieval sources, provider settings and logs, plus time with system owners. Test data and the permitted environment are settled before testing.

Q.06

How much does it cost?

Fixed price, scoped against the number of workflows. We tell you the number before we start. If the right answer is "you don't need us yet", we'll say so on the first call.

Q.07

Can you fix what you find?

Yes, but only if you want us to. The review stands on its own and the fixes are written so your developer can apply them. If you'd rather we did the work, we'll quote it separately once the report is in your hand.

Vu Agency AI safety review session

Book an AI workflow safety review

Send us the current workflow list, owners and connected systems. We will identify which one carries the highest operational or data risk and define the scope of an independent review.

[MORE PRO SERVICES]

More from Security & Governance

Every Pro Service page covers what it is, who it fits and how to start. The full list is in the footer below.

Message us on WhatsApp