Skip to content
All articles

AI Security · Governance

Security, safety, robustness and privacy are four different problems

One AI incident can touch all four, but each needs its own evidence and its own owner. A clinical summarizer shows why no single guardrail can prove them all.

· 4 min read

Four cards. Security asks whether an untrusted document can trigger an unauthorized tool action, with attack tests and tool traces as evidence. Safety asks whether an allowed answer can cause unacceptable harm: hazard analysis and scenario tests. Robustness asks whether quality holds under blur, accents, long context or missing data: slice metrics and perturbation tests. Privacy asks whether training or inference can expose a person's data: data inventory and access logs. One event can cross all four, and no single guardrail proves them all.

Teams often file every AI risk under one word, usually "security" or "safety", and then argue about whether a guardrail has dealt with it. The four concerns are related, but they ask different questions, need different evidence, and usually belong to different people.

  • Security asks whether an actor can violate confidentiality, integrity, availability or authorized control.
  • Safety asks whether operating the system can cause unacceptable harm, including harm with no attacker involved.
  • Robustness asks whether behavior stays acceptable under noise, distribution shift, ambiguity or stress.
  • Privacy asks whether personal or sensitive information is collected, inferred, used, retained or disclosed appropriately.

One summarizer, three different failures

Take a model that summarizes clinical documents.

A malicious document instructs it to reveal another patient's record. That is a security attack, with privacy impact. A faithful summary leaves out a critical allergy because the scan was poor. That is a robustness failure, with safety impact. And a correct diagnosis, inferred where nobody authorized the inference, can be a privacy and governance failure with no attacker anywhere.

No single guardrail can prove all three cases are controlled. Treating the four words as synonyms produces controls too vague to test.

Different questions need different evidence

QuestionDisciplineRepresentative evidence
Can an untrusted document cause an unauthorized tool action?SecurityTaint-aware attack tests, policy decisions, egress and tool traces
Can an allowed answer create unacceptable physical or social harm?SafetyHazard analysis, scenario tests, review and escalation outcomes
Does quality hold under blur, accent, long context or missing data?RobustnessSlice metrics, perturbation tests, calibrated abstention
Can training or inference expose a person's data?PrivacyData inventory, minimization, access logs, deletion and privacy attack tests

Engineering Insight

Record the overlap, but keep accountability separate. Security may own exploitability and access controls, product safety the hazardous outcomes, privacy the lawful processing and minimization, and ML engineering the robustness. A shared risk register shows where they intersect; each claim in it still needs an owner, a test, a threshold and a response.

An architecture review that lists one "AI risk" control and calls the job done has usually tested one of these four. Asking which discipline a control serves, and what evidence would show it failing, is the quickest way to find the other three.

Get one diagram a week

A short article built around one engineering diagram, from the same library as these courses.

One diagram-led article a week on AI and systems engineering. We email you once to confirm, and every newsletter has an unsubscribe link. Privacy policy