Patient Data as AI Training Data: A Cybersecurity Risk-Lifecycle Problem

When patient-derived information enters an AI pipeline, the security boundary expands from the database to the data, model, development pipeline, interfaces, vendors, and derived artifacts.

By John Keenan, CISSP

Start with the asset, not the model

Healthcare AI discussions often begin with what a model can do. Cybersecurity professionals should begin one step earlier: what information made the capability possible?

Patient-derived data can be extraordinarily valuable for training, tuning, evaluating, and operating AI systems. It can also be extraordinarily sensitive. Once that information enters an AI development lifecycle, protecting the source database is no longer sufficient. Copies and transformations can appear in labeling environments, development platforms, evaluation sets, model artifacts, retrieval systems, logs, prompts, and third-party services.

The dataset therefore becomes part of the attack surface.

This is familiar territory for CISSPs. Asset identification, classification, least privilege, retention, secure disposal, supplier risk, and data lifecycle management are not made obsolete by AI. AI makes failures in those disciplines harder to see and potentially more consequential.

De-identification is not a magic word

A recurring mistake in AI discussions is to treat “de-identified” as if it means “no longer sensitive.” Under the HIPAA Privacy Rule, HHS recognizes Expert Determination and Safe Harbor as methods for de-identifying protected health information. HHS also explains that properly de-identified information retains a very small residual risk of identification.

That distinction should affect security architecture. Removing direct identifiers does not eliminate the need to understand data provenance, combinations of attributes, free-text content, linkage opportunities, access paths, and downstream uses.

The question for the security team is not merely, “Did someone call this dataset de-identified?” It is, “What method was used, what residual risks remain, what other information can be combined with it, and what controls are appropriate to the resulting risk?”

AI adds integrity and model-security problems

Confidentiality is only one part of the problem. NIST's 2025 adversarial machine learning taxonomy describes attack classes including evasion, poisoning, privacy, and misuse attacks across predictive and generative AI.

Data poisoning is particularly useful for understanding why healthcare AI security cannot be reduced to privacy. If an attacker can manipulate training or fine-tuning data, the consequence may be altered model behavior rather than stolen records. A system can fail while confidentiality appears intact.

Privacy attacks raise another concern: information can sometimes be inferred or extracted through interaction with a model or system. Generative AI also introduces misuse scenarios in which legitimate capabilities are deliberately employed for harmful purposes.

The practical implication is that the security architecture must include the model, training pipeline, interfaces, access controls, monitoring, and supporting services - not just the database that originally contained the records.

Provenance and purpose matter

One of the hardest questions is administrative rather than technical: why does the organization have the information, and what uses are authorized?

Clinical data collected for patient care may later appear attractive for analytics, product improvement, model evaluation, or training. Those secondary uses can create governance questions even when the technical team sees only a useful dataset.

A mature program should be able to trace provenance: where the data originated, what transformations were performed, what restrictions apply, who approved the use, which systems received it, and what derivative artifacts were created. Data minimization also deserves renewed attention. “More data improves the model” is not, by itself, a security justification for collecting or retaining every available element.

Build the control questions before the pipeline

Before patient-derived information enters an AI pipeline, leaders should be able to answer a basic set of questions. How is the data classified? Who can access raw and transformed versions? Are vendors or cloud services involved? What contractual and technical restrictions apply? What is logged? What is retained? How are model artifacts protected? How are changes approved? How would the organization investigate suspected poisoning, leakage, or unauthorized reuse?

NIST's AI RMF provides a useful structure because it puts governance, context, measurement, and risk treatment around the technology. The Generative AI Profile adds AI-specific considerations without discarding conventional cybersecurity practice.

The central point is not that healthcare organizations should avoid AI. It is that patient-derived information should not become unlimited raw material simply because the organization has discovered a new use for it.

AI security begins before model training. It begins when the organization decides what information it has, why it may use it, how it will protect it, and what evidence will demonstrate that those decisions remain true throughout the lifecycle.

Sources

·       NIST AI RMF

·       NIST GenAI Profile

·       NIST AML 2025

·       HHS De-identification

Previous
Previous

The Interpreter in the AI-Enabled Exam Room: An Overlooked Security Boundary

Next
Next

When “Notice” Isn’t Enough: AI Transparency in Multilingual Healthcare