Skip to content
בס״ד
Cyber Replay logo CYBER REPLAY

AI Security

AI Security for Teams Shipping LLM Features and Agents

Putting a model into production creates attack surface that conventional application testing does not look for. Your assistant reads untrusted text and treats it as instructions. Your agent holds credentials to several systems at once. Your inference stack pulls packages that are now a named target.

We test for these specifically, and we build AI features that hold up when someone attacks them on purpose.

Request an AI security assessment Book a 15-minute call

The eight failure modes we test for

Each of these is something we have researched and written about in depth. The links go to the detailed technical write-up.

Prompt injection

Untrusted text reaching your model becomes instructions. A support ticket, a scraped page or a PDF can redirect an assistant into leaking context, calling tools it should not, or returning attacker-chosen output to the next user.

Read: Unicode prompt injection in Apple Intelligence deployments

Over-permissioned agents

An agent that can read a database, call an internal API and send email has the combined blast radius of all three. Most agent frameworks default to broad credentials because narrow ones are inconvenient during development.

Read: Least privilege for AI agents

AI supply chain compromise

Model weights, inference libraries and orchestration packages are dependencies like any other, and they are being targeted. A compromised release ships straight to production behind a version bump.

Read: The LiteLLM PyPI compromise and credential theft

Secret exfiltration through orchestration

Chaining frameworks pass credentials, retrieved documents and tool output through the same context window. Anything in that window can end up in a model response, a log, or a third-party API call.

Read: Preventing secret exfiltration in LangChain and LangGraph

Exposed AI infrastructure

Self-hosted inference servers, notebook runners and low-code AI builders regularly ship with authentication off by default and get deployed straight onto the public internet.

Read: Langflow CVE-2026-33017 and AI pipeline exposure

Training and fine-tuning data poisoning

If an attacker can influence what you fine-tune on, they can influence what the model says. The effect is durable, hard to detect by inspection, and survives redeployment.

Read: Mitigating LLM data poisoning

Shadow AI

Staff paste customer data into whatever tool is convenient. The exposure is real, it is usually invisible to security teams, and blocking everything simply moves it to personal devices.

Read: Detecting and controlling unsanctioned AI tools

Compromised AI developer tooling

Coding assistants and autonomous review agents hold repository credentials and can approve or author changes. That makes them a high-value target and a direct path into your build pipeline.

Read: Hardening AI coding assistants against symlink attacks

Why the model layer needs its own testing

Traditional application security has a reliable pattern: separate code from data, and the injection class of bug largely disappears. Parameterised queries solved SQL injection. Output encoding solved most cross-site scripting.

Language models break that pattern. Instructions and data are the same thing, both just text in a context window, and no amount of prompt engineering reliably separates them. That is why prompt injection has no clean fix and why defence has to move up a level: assume the prompt can be subverted, then make sure a subverted prompt cannot do much damage.

In practice that means:

Engagements and pricing

AI Security Assessment
$5,500

Two weeks. Prompt injection testing, agent permission and tool-call review, AI supply chain audit, and a remediation-ready report.

AI Application Penetration Test
$6,500

Full application test covering both conventional web vulnerabilities and LLM-specific attack paths, with a free verification retest.

Secure AI Build
$38,000

90 days. Custom AI tool or LLM-backed application built with guardrails, scoped agent permissions and a penetration test before launch.

Fixed Scope, Fixed Price. If the signed scope takes longer than we estimated, we absorb the overage. You pay what the proposal says.

Clean-Code Warranty. On builds, any critical vulnerability we introduced and you find within 12 months of launch is fixed at no cost.

What an AI security assessment produces

  1. An inventory of what your AI can actually reach. Every tool, credential, data source and downstream system, which is usually broader than teams expect.
  2. Injection test results. The specific payloads that worked, where they entered, and what they achieved.
  3. A permission review. Where the agent holds more access than its job requires, and the narrower grant to replace it with.
  4. A supply chain audit. Model provenance, inference libraries and orchestration packages checked against known compromises.
  5. A remediation plan ordered by risk. What to fix this week, this quarter, and what to accept.

Building the AI feature rather than retrofitting it

If the AI product does not exist yet, the cheapest time to get this right is now. A Secure Build is $38,000 for a custom AI tool delivered in 90 days with guardrails, scoped agent permissions and a penetration test before launch.

Secure web app development

Further reading from our research

Frequently Asked Questions

What is AI security?

AI security covers the attack surface that appears when you put a model into production: prompt injection, over-permissioned agents, poisoned training data, compromised model and library supply chains, and staff feeding sensitive data into tools nobody approved. It sits alongside conventional application security rather than replacing it, because an LLM feature still lives inside a web application with all the usual risks.

Why is prompt injection so hard to fix?

Because language models do not reliably distinguish instructions from data. Everything in the context window is text competing for the model's attention, so any untrusted content you feed it can carry instructions. There is no equivalent of parameterised queries here. The practical defence is architectural: constrain what the model is allowed to do, validate its output before acting on it, and assume any given prompt can be subverted.

How much does an AI security assessment cost?

An AI Security Assessment is $5,500 for two weeks, covering prompt injection testing, agent permission review and an AI supply chain audit. A full AI application penetration test is $6,500 and includes a free verification retest. Building a new AI tool securely from scratch is $38,000 for a 90-day Secure Build.

Can you test an AI agent that takes real actions?

Yes, and those are the engagements that matter most. We map every tool the agent can call, work out what the worst legitimate-looking sequence of calls achieves, and test whether injected instructions can trigger it. Testing happens against a staging environment with real integrations mocked, so we can push hard without side effects.

Does a conventional penetration test cover our LLM features?

Only partially. A standard test finds the injection flaws, broken access control and authentication problems in the surrounding application, which still matter. It will not systematically test whether your assistant can be talked into calling an internal API on an attacker's behalf. That requires testing designed for the model layer.

We use a hosted model API rather than self-hosting. Are we still exposed?

Yes. Using a hosted provider removes the infrastructure risk of running inference yourself, but prompt injection, over-permissioned tool access, secret leakage through context and shadow AI are all properties of how you integrate the model, not where it runs.

Do you help with AI governance and policy?

We focus on the technical controls, then document what we implemented so it maps onto whatever framework you are being assessed against. If the immediate driver is an audit rather than a threat, say so up front and we will shape the engagement around the evidence you need to produce.

Get your AI reviewed before someone else does

Request an assessment Free cyber scorecard All cybersecurity services