AI Security Testing
Virtual Nomad tests the security of AI features: chatbots built on large language models, autonomous agents, and applications that sit on top of models and retrieval pipelines. These systems fail in ways that ordinary application testing does not cover. A model can be led into ignoring its instructions, an agent can be steered into misusing the tools it controls, and a retrieval system can return data it was supposed to keep back. We test for these cases directly.
Virtual Nomad tests the security of AI features: chatbots built on large language models, autonomous agents, and applications that sit on top of models and retrieval pipelines. These systems fail in ways that ordinary application testing does not cover. A model can be led into ignoring its instructions, an agent can be steered into misusing the tools it controls, and a retrieval system can return data it was supposed to keep back. We test for these cases directly.
Adding an AI feature creates a new attack surface on top of the one you already maintain. Because a model’s behavior is probabilistic, a control that looks solid in a demo can often be bypassed with the right input. Our testing is aimed at getting the system to do things it was not meant to do: reveal data, take actions without authorization, or reach systems beyond its intended scope.
Prompt injection
We test for direct prompt injection and for indirect injection, where instructions are planted in documents, web pages or other data the model reads. The point is to measure how hard it is to override the system prompt and any guardrails, and what an attacker actually gains once they do.
Tool and agent abuse
When a model can call tools, send email, run code or query a database, the question that matters is what it can be talked into doing with that access. We test whether an attacker can use the agent’s own permissions to carry out actions the application was never meant to allow.
Data leakage
Models expose system prompts, other users’ data, internal documents and credentials more readily than most teams assume. We look at how context and retrieved data are handled, and we try to pull out information that is supposed to stay private.
Authorization limits
We check that the AI cannot act beyond the permissions of the person using it. That includes privilege escalation through the model, access to other tenants’ data, and situations where an action taken by the assistant quietly sidesteps your normal access controls.
Defined coverage
Each engagement is mapped to the OWASP Top 10 for Large Language Model Applications and to current published research. The areas under test are set out clearly, and the results line up with other assessments you may run.
What you get back
Findings are listed in priority order with examples you can reproduce. For each one we give practical guidance: how to handle input and output, how to scope tool permissions, what to log and monitor, and which guardrail changes are worth making. A retest of the fixes is included.

