Skip to main content

Test the model as an attack surface

Jailbreak and injection testing, content safety and trustworthy AI governance

Once an application calls a large model it gains a new attack surface: prompt injection, jailbreaks, data leakage and adversarial inputs. We test LLM security with offensive methods, assess AIGC content safety and privacy controls, and map the result to ISO/IEC 42001 and trustworthy AI requirements as a risk assessment and governance framework.

Our approach

One team scopes, executes and retests; conclusions are delivered as evidence, not checklists.

01

Tested with offensive methods, not only a compliance checklist

Assessors come from the penetration testing team and organise attack cases by OWASP LLM Top 10 and MITRE ATLAS, covering prompt injection, jailbreaks, data leakage, tool-call abuse and adversarial inputs, each finding with reproduction steps.

02

The application is assessed, not only the model

Most risk sits in the application layer: prompt templates, retrieval pipelines, tool calls and permission boundaries. Testing runs on your real application, so conclusions map directly to hardening you can implement.

03

Results map to ISO/IEC 42001 and domestic regulation

The report and governance framework map to ISO/IEC 42001 controls, the Interim Measures for Generative AI Services and algorithm filing requirements, for use in internal audit and external compliance.

Deliverables

01

Assessment report

Techniques, success rates, impact and hardening advice, organised by application layer rather than model layer.

02

Guardrail design

Input and output filtering, permission boundaries and monitoring metrics, ready for the development team to implement.

03

Governance framework

Risk register, audit points and compliance mapping aligned to ISO/IEC 42001.

How we deliver

Four stages, each with defined inputs, outputs and a client sign-off.

01Week 1

Scope and threat modelling

Map application architecture, data flows and tool permissions; define the attack surface and test cases.

02Week 2

Assessment

Run injection, jailbreak, leakage and abuse tests; record success rates and impact.

03Week 3

Report and hardening plan

Deliver the assessment report, guardrail design and compliance mapping; present to the team.

04After hardening

Retest

Verify hardening and issue a retest conclusion; retest per release thereafter.

Frequently asked questions

Notes on scope, execution and delivery standards. Contact us for anything not covered here.

Testing is organised by OWASP LLM Top 10 and MITRE ATLAS and covers direct and indirect prompt injection, jailbreaks, training data and system prompt leakage, insecure output handling, tool-call and plugin permission abuse, adversarial inputs and content safety. Each category is reported with attack success rate and impact.

Start from where you stand

Security, data and AI each start with a review of where you stand. The report and its findings are yours, whether or not the engagement continues.