Test the model as an attack surface
Jailbreak and injection testing, content safety and trustworthy AI governance
Once an application calls a large model it gains a new attack surface: prompt injection, jailbreaks, data leakage and adversarial inputs. We test LLM security with offensive methods, assess AIGC content safety and privacy controls, and map the result to ISO/IEC 42001 and trustworthy AI requirements as a risk assessment and governance framework.
Our approach
One team scopes, executes and retests; conclusions are delivered as evidence, not checklists.
Tested with offensive methods, not only a compliance checklist
Assessors come from the penetration testing team and organise attack cases by OWASP LLM Top 10 and MITRE ATLAS, covering prompt injection, jailbreaks, data leakage, tool-call abuse and adversarial inputs, each finding with reproduction steps.
The application is assessed, not only the model
Most risk sits in the application layer: prompt templates, retrieval pipelines, tool calls and permission boundaries. Testing runs on your real application, so conclusions map directly to hardening you can implement.
Results map to ISO/IEC 42001 and domestic regulation
The report and governance framework map to ISO/IEC 42001 controls, the Interim Measures for Generative AI Services and algorithm filing requirements, for use in internal audit and external compliance.
Deliverables
Assessment report
Techniques, success rates, impact and hardening advice, organised by application layer rather than model layer.
Guardrail design
Input and output filtering, permission boundaries and monitoring metrics, ready for the development team to implement.
Governance framework
Risk register, audit points and compliance mapping aligned to ISO/IEC 42001.
How we deliver
Four stages, each with defined inputs, outputs and a client sign-off.
Scope and threat modelling
Map application architecture, data flows and tool permissions; define the attack surface and test cases.
Assessment
Run injection, jailbreak, leakage and abuse tests; record success rates and impact.
Report and hardening plan
Deliver the assessment report, guardrail design and compliance mapping; present to the team.
Retest
Verify hardening and issue a retest conclusion; retest per release thereafter.
Frequently asked questions
Notes on scope, execution and delivery standards. Contact us for anything not covered here.
Testing is organised by OWASP LLM Top 10 and MITRE ATLAS and covers direct and indirect prompt injection, jailbreaks, training data and system prompt leakage, insecure output handling, tool-call and plugin permission abuse, adversarial inputs and content safety. Each category is reported with attack success rate and impact.
Start from where you stand
Security, data and AI each start with a review of where you stand. The report and its findings are yours, whether or not the engagement continues.