IRM Consulting & Advisory
Generative & Agentic AI Security

Data Poisoning: Securing AI Models

As few as 250 poisoned documents can backdoor a model. Here is what a SaaS company must control across fine-tuning, RAG and AI vendors, mapped to ISO 42001.

Data Poisoning in SaaS: What You Must Control When You Fine-Tune, Run RAG or Buy AI Features

Data poisoning is an attack where someone plants manipulated data in the material an AI system learns from or retrieves from, so the system behaves the way the attacker wants. For a SaaS company, that material includes fine-tuning sets, RAG knowledge bases and the data your AI vendors train on. The defense is provenance, integrity and supplier control.

Every few weeks we get a version of the same call. A CTO at a growing SaaS company has shipped an AI assistant on top of a vendor model, connected it to the help center and a shared drive, and an enterprise prospect has just sent a security questionnaire with a new section: "Describe how you protect the integrity of data used to train or ground your AI features." Nobody on the team owns that question. The model is not theirs, and the help center is edited by a dozen people.

If you want the fundamentals of how poisoning attacks work and the warning signs, start with our earlier explainer on AI data poisoning attacks. Here I get specific: what a 30 to 200 person SaaS company must control, how that maps to ISO/IEC 42001, and what to do in 90 days.

Why is data poisoning a bigger risk for SaaS companies in 2026?

In October 2025, Anthropic, the UK AI Security Institute and The Alan Turing Institute published a study showing that as few as 250 malicious documents were enough to plant a backdoor in language models ranging from 600 million to 13 billion parameters. The number of poisoned documents stayed roughly constant as models grew. The attacker needs a fixed count, not a percentage of the training data. The researchers are careful to note the backdoor they tested was low-stakes (it made the model output gibberish on a trigger phrase) and that it is unclear whether the result holds for more harmful behaviors or larger models. Still, a few hundred documents is within almost any attacker's budget.

The second is about retrieval. The PoisonedRAG paper presented at USENIX Security 2025 reported a 90% attack success rate when researchers injected five malicious texts per target question into a knowledge base holding millions of texts. Most SaaS AI features today are RAG features. Your knowledge base is the training set that nobody is governing.

The standards bodies have caught up. The OWASP Top 10 for LLM Applications 2025 lists Data and Model Poisoning as LLM04 and names pre-training, fine-tuning and embedding as the exposed stages. NIST AI 100-2 E2025, NIST's adversarial machine learning taxonomy, classifies poisoning into availability, targeted, backdoor and model poisoning. MITRE ATLAS tracks Poison Training Data (AML.T0020) and RAG Poisoning (AML.T0070) as separate techniques, which is a useful hint that you need separate controls for each.

Where can poisoned data enter your product?

Across our assessments, the entry points fall into three buckets. Most companies only think about the first.

  1. Your own training and fine-tuning data. Support tickets, product telemetry, labeled examples, and public datasets pulled from model hubs. Anyone who can write to those stores, including a compromised contractor account, can shape the model.

  2. Your RAG sources. Help center articles, Confluence or Notion spaces, shared drives, CRM notes, and in some products, customer-uploaded documents. If a customer or an outsider can write content that your assistant later retrieves, you have a live poisoning path with no training run required.

  3. Your AI vendors' data. The foundation model, the embedding model, and any SaaS tool with "AI inside" were trained on data you never saw. You cannot inspect it, so your control is contractual and evidentiary.

One short rule helps teams prioritize: if it can be written by someone you do not fully trust, and it can be read by the model, it is in scope.

Which controls matter, and which ISO 42001 controls do they map to?

ISO/IEC 42001 is the AI management system standard, and its Annex A gives you an auditable home for each of these controls. Here is how I map the work, along with an opinion on where most teams overspend.

Control

What it stops

ISO/IEC 42001 Annex A

Our take

Data source inventory and provenance records

Unknown or unvetted data reaching training or retrieval

A.7.3 Acquisition of data, A.7.5 Data provenance

Highest return. Most teams cannot list their RAG sources.

Write-access review on RAG sources

Outsiders or low-trust users planting content the assistant retrieves

A.7.4 Quality of data for AI systems, A.6.2.6 AI system operation and monitoring

Cheaper and more effective than any "AI firewall" product.

Dataset versioning and hashing

Silent changes between approval and training

A.7.6 Data preparation, A.6.2.8 AI system recording of event logs

Required before you fine-tune anything customer-facing.

Pre-release evaluation with a fixed test set

Backdoors and degraded behavior shipping to production

A.6.2.4 AI system verification and validation

Run it on every model or prompt change, not once a year.

AI supplier due diligence and contract terms

Vendor models trained on poisoned or unlicensed data

A.10.3 Suppliers, A.10.2 Allocation of responsibilities

Ask for evidence, not a trust-center badge.

Customer-facing disclosure and incident path

Customers learning about a poisoning event before you tell them

A.8.4 Communication of incidents

Write it before you need it.

Vendor "poison detection" scanners

Some statistical outliers in training data

Supports A.7.4

Overrated for SMBs. Buy last, if at all.

That last row is the one vendors will argue with. Detection tooling has a place in large training pipelines, but NIST's own taxonomy describes designing models that stay robust against supply-chain model poisoning as "a critical open problem." For a company your size, access control, provenance and testing do more per dollar.

If you are building toward certification, our ISO 42001 readiness checklist walks through the full clause and Annex A scope, and the ISO 42001 implementation package gives you the policy and evidence templates.

What should you ask your AI vendors about poisoning?

Most SaaS companies are buyers of AI before they are builders. These are the questions I put into vendor reviews. A vendor that cannot answer most of them in writing is telling you something.

  • What data sources trained or fine-tuned the model we are using, and how do you record their provenance?

  • How do you detect and respond to poisoned or manipulated data in your training and retrieval pipelines?

  • Is our data ever used to train models served to other customers? Can we opt out in the contract?

  • If a poisoning or model integrity incident affects us, how fast will you notify us, and through which channel?

  • Do you publish a model card or an ML bill of materials (ML-BOM) for the model version we run?

  • Which independent assessments cover your AI system, and do they include data governance, not only infrastructure?

OWASP's LLM04 guidance specifically recommends tracking data origins with tools like an ML-BOM and vetting data suppliers, so these questions are not exotic. They are becoming baseline, and the same list helps when you need to find the shadow AI tools your teams adopted without review.

What does a 90-day data poisoning plan look like?

Order matters more than tooling. This is the sequence we use with clients who have AI features in production and no formal AI governance yet.

Days 1 to 30: see what you have

  • Inventory every AI feature, the model behind it, and every data source it trains on or retrieves from.

  • For each RAG source, list who can write to it. Remove public or customer write paths that feed the assistant unless they are filtered and labeled.

  • Name one owner for AI data integrity. In a company this size, that is usually the CTO with a fractional security lead.

Days 31 to 60: lock down and version

  • Put fine-tuning and evaluation datasets under version control with hashes, and log who approved each version.

  • Build a fixed evaluation set, including a handful of adversarial prompts, and run it before every model, prompt or source change.

  • Send the vendor questions above to your top three AI suppliers and file the answers.

Days 61 to 90: make it auditable

  • Map each control to ISO/IEC 42001 Annex A and record evidence in your statement of applicability.

  • Add a poisoning scenario to your incident response plan, including how you would roll back a model or purge a RAG index.

  • Write the questionnaire answer your sales team will reuse, and have it reviewed by whoever signs off on security claims.

If you want a template for the policy side, our AI Governance Playbook covers acceptable use, data handling and model change control. Teams running their own training pipelines should also read our guide to security in the MLOps pipeline.

What do these controls not stop?

I would rather you hear this from me than from an auditor.

They do not clean a vendor's foundation model. If poisoned data went into a model you license, provenance records on your side will not remove it. Your protections there are evaluation, contract terms and the ability to switch models.

They do not stop prompt injection at runtime. Poisoning and prompt injection overlap, and NIST treats indirect prompt injection through retrieved data as its own attack class. Locking down RAG sources reduces the risk, but you still need output handling and least-privilege permissions, which we cover in how to secure AI agents.

They do not prove a model is clean. A backdoor is designed to stay quiet until its trigger appears. Testing raises your confidence. It does not give you a guarantee, and anyone selling one is overselling.

When you can skip most of this: if you call a third-party model through an API, do not fine-tune, and do not connect it to any internal or customer content, your poisoning exposure is mostly the vendor's. Do the vendor questions, keep an inventory, and move on.

Frequently asked questions

Is data poisoning the same as prompt injection?

No. Poisoning corrupts what a model learns from or retrieves, and it persists until the data is removed. Prompt injection manipulates a single interaction. RAG poisoning sits between the two, which is why it deserves its own controls.

How much poisoned data does an attacker need?

Less than most people assume. The 2025 Anthropic, UK AISI and Alan Turing Institute study found around 250 documents were enough for the backdoor they tested, across every model size they tried (600 million to 13 billion parameters). For RAG, published research has shown a handful of planted texts per target question can work.

Does ISO 42001 require data poisoning controls?

ISO/IEC 42001 does not name data poisoning as a control. It requires you to manage data acquisition, quality, provenance and preparation (Annex A.7) and supplier relationships (A.10) based on your risk assessment. For most SaaS AI features, poisoning will appear in that assessment.

Who should own AI data integrity in a small SaaS company?

One named person, usually the CTO, supported by a security leader. If you do not have one, a Virtual CISO can own the program, run the vendor reviews and keep the evidence audit-ready.

If an enterprise questionnaire just landed on your desk with an AI integrity section, book a 30-minute call and we will help you scope the answer and the 90 days behind it. You can also see how we structure AI governance services for SaaS teams.

Keep Reading

Related Articles

Our Industry Certifications

Our diverse industry experience and expertise in AI, Cybersecurity & Information Risk Management, Data Governance, Privacy and Data Protection Regulatory Compliance is endorsed by leading educational and industry certifications for the quality, value and cost-effective products and services we deliver to our clients.

Copyright © 2026 IRM Consulting & Advisory. All Rights Reserved.