We find out what your AI can be made to do - then make sure it never happens.
Cross-tenant data exposure, prompt injection that survives your guardrails, agent and tool-use chains. HiltLock (previously Trampolyne AI) attacks the system as it actually runs, before launch or after. Where a finding cannot be closed in your code, policy is enforced outside the model. Every finding ships with reproduction steps.
Security for the AI you build.
Assess first, at one of two speeds: automated inside 24 hours, or human-led over one to two weeks. Either way, findings you can reproduce rather than a compliance report.
Published with the client's permission. Read the full case study
Four systems where the control was working
Every figure below comes from a published case study. Four systems, four failure modes, and in each one the control a buyer would have asked about was present and answering.
A scanner looks for a malformed request. In the first of those there was nothing malformed: a valid session token, a sequential integer, and an endpoint that never asked whether the caller owned the account. Authorisation logic that reads as ordinary traffic is the gap we work in.
Why your current coverage does not reach this
Nothing below is a criticism of your pentest or your auditor. They are answering different questions, correctly. Neither question is "can this model be talked into acting outside its authority".
| What you already run | What it answers | What it cannot reach |
|---|---|---|
| VAPT / pentest | Is the request path exploitable | Authorisation expressed as an instruction to a model, not as a rule in code |
| GRC / ISO 42001 / SOC 2 | Do you have controls, and evidence of them | Whether a control holds under adversarial pressure |
| AI-SPM scanners | Known patterns and signatures at scale | Novel chains that require reasoning about your specific architecture |
| Our own automated pass | Our attack library, run end to end in under a day | The same ceiling. We run one, we sell one, and we say where it stops |
| Guardrail products | Blocking known-bad prompts | Attacks that never look like a bad prompt |
Assess, then monitor, then control
One engagement arc, three phases, and everything we do sits inside it. Most clients start and stay at Assess, which is the right call until it is not.
Automated or human-led
Two speeds, set out in the next section. Either one, or the automated pass first and the human-led assessment on what it cannot reach. Both end in findings you can reproduce, with a fix path attached.
Scheduled re-test
Your models, prompts and tools change every sprint. Automated runs on your cadence catch the drift; a human re-test each quarter catches what the automation cannot.
Runtime enforcement
For failures caused by the model's own behaviour, which cannot be fixed in code. Recommended only when a finding needs it.
Automated in a day, or human-led in a fortnight
Assess runs at two speeds. This is the one choice you have to make to start, and it is a choice inside the program rather than an alternative to it. They answer different questions, and most clients want both in sequence rather than one instead of the other.
Automated assessment
Answers what is already broken. Known attack classes run end to end with no human in the loop, against your live system. The same attack library the human assessments run on.
Use it when
- You need a baseline before a launch, and you need it now
- A customer questionnaire needs AI testing evidence this week
- You want regression cover between human assessments
- Procurement is the bottleneck, so buying on your existing AWS commit is worth more than a new vendor form
What it will not do. It finds what it has patterns for. "Does this endpoint check ownership" is not a pattern, which is why the finding at the top of this page needed a human.
Human-led assessment
Answers what can be broken. Architecture-aware adversarial testing that reasons about your specific tenancy, tool wiring and retrieval scope, looking for chains nobody has written a signature for yet.
Use it when
- The system is multi-tenant, and a boundary failure is material
- The AI can take actions, not just answer questions
- A board, an auditor or an enterprise customer is asking
- The automated pass came back clean and you do not believe it
What you get. Exploit chains with reproduction steps, severity, and a fix path per finding. Every finding published on this site came out of this work.
We will tell you which one you need, including when the answer is the cheap one, and including when it is neither.
Six surfaces, and the boundary that cuts across them
Most vendors cover one or two. The question is the same on all of them: can this system be made to act outside its authority, and who gets to tell it to.
Chat and assistants
Injection that survives your guardrails, and extraction of what the guardrail is made of. One published case gave up its system prompt, policy list and enforcement design after refusing every direct request.
Retrieval and RAG
What is in the index that should not be, and whether retrieved content is treated as data or quietly as instructions. Indirect injection needs no access to your prompt at all.
APIs and authorization
Where AI breaches actually happen. We name the exact victim record we reached, not a theoretical "could happen". One published case pulled 95 customer reports through a single authenticated user's session.
Tools and MCP servers
A Model Context Protocol server is a first-class target, not a degraded one. Tool descriptions are handed to a consuming model as trusted context, so poisoned metadata reaches any agent that merely lists the server.
Agentic workflows
No user in the loop, real credentials, real tools. Chained calls where every step is permitted and the combination is not. Two published cases here, one ending in a transaction nobody asked for.
Voice agents
We place the call, over SIP and PSTN or WebRTC, and score the attack across the whole conversation: barge-in, phonetic evasion, session minting, the human-transfer boundary. Most alternatives test the model through a text API and leave all of that untouched.
The tenant boundary
Most AI testing runs as one user, which is why the interesting failures survive it. We run two identities and an intermediary holding both delegations, because that is the deployment agent protocols encourage. In the most recent published case, crossing that boundary needed no credential and no authorisation bypass.
The first voice findings are now published on the red team platform: a voice agent that gave up its operating instructions across four spoken turns, and a widget whose session endpoint minted billable sessions unauthenticated.
A clean report has a shelf life
Two things happen after the first assessment. The second one is only for cases the first one proves you need it.
Nothing stays closed on its own
A new model version. A tightened system prompt. A tool wired into the agent. Any of those can reopen what an assessment closed, and none of them looks like a security change on the way past.
Two cadences together
- Automated runs on your sprint rhythm, catching regression
- A human re-test each quarter, re-reading what changed shape
- Scope carries over from Assess, so there is no second discovery
- A subscription against your cadence, cancellable each quarter
What we will not claim. A full red team on every commit. Anyone selling that is selling the automated tier with better marketing.
For what code cannot close
Most findings sit in your code, where the fix is deterministic. Others come from the model's own behaviour. Tightening the prompt does not fix those, because the ways to phrase the same request are effectively infinite. Those need enforcement outside the model, before it acts.
Why not a guardrail product
Guardrails block known-bad prompts. Every attack in our published cases looked ordinary: a sequential integer, a request for a short story, a label typed into a form field. There was no bad prompt to block. Checking identity and scope against your policy is a different mechanism, not a better filter.
You own it. You author the policy and run the operations. We advise, and we run nobody's security operations. Observe mode logs what it would have done without touching your traffic, so you buy against your own evidence.
We have a commercial interest in you reaching Phase 03, and you should assume we know that. It is why the published cases close in the client's own code, and why we quote enforcement only after a finding you can reproduce shows that code cannot close it.
Three readers, one category
You own AI risk you did not choose
Product shipped an assistant, and it reached production without anything resembling an adversarial review. You need to know what is actually reachable before someone else finds out.
An enterprise deal is stuck in security review
Clear the AI section of your customer's security questionnaire with evidence rather than assurances. A focused one-week red team usually unblocks it.
You need specialist capability you cannot hire
Specialist AI security for the engagements that need it. Flexible on brand. We never approach your client on our own. How the partner model works
What you are probably comparing us to
Two of these rows are us. The first two win most often.
| Option | What you get | Where it falls short |
|---|---|---|
| Nothing, for now | Free, plus the model provider's safety work, which is real | That governs what the model says. It knows nothing about which tenant owns which record |
| Your VAPT vendor, extending scope | One vendor, one PO, and a line item that says AI | Ask what they test. A prompt list is not your tenancy model or your tool wiring |
| In-house, over a sprint | Engineers who know the architecture better than any outsider will | They built it, so they test what they expected to break |
| Big 4 AI risk assessment | Six to eight weeks, a compliance-style report, a known name on the cover | Findings you cannot reproduce, from people who did not attack the system |
| Platform AI-SPM scan | Dashboards, signatures, breadth, often inside existing spend | Cannot reach what needs reasoning about your architecture, and rarely says so |
| A guardrail product, bought first | A filter in front of the model today, and a dashboard showing it work | A control chosen before you know which boundary failed |
| HiltLock, automated | Under 24 hours, known attack classes, on AWS Marketplace | Pattern-bound by design. The report states the limit |
| HiltLock, human-led | One to two weeks, exploit chains with reproduction steps, a named human accountable | Depth over breadth. No traditional VAPT, and we say when you do not need us |
On the other two phases. The alternative to Monitor is an annual pentest. A year is roughly forty sprints, so if the model version, the system prompt and the tool list have all changed, the report describes a system that no longer exists. The alternative to Control is a guardrail product, compared in section 06.
What we are finding, every quarter
Anonymised patterns from the engagements we ran this quarter: which failures recurred, which controls held, and how long each took to bypass. No vendor pitch, no gated demo call.
4.2 Authorisation expressed in natural language is not enforced
The most frequent Critical finding this quarter was not a model weakness. It was retrieval scope enforced by instruction rather than by policy, an authorisation boundary described to the model instead of imposed on it.
Observed in 6 of 9 engagements / Median time to first bypass: 2.5 days
Find out what is actually reachable
Twenty minutes, no deck. Describe your AI system and we will tell you whether we are the right people, and what we would go after first.