CCSP - Cloud Data Security - Section 2.9

Comprehend data protection of Artificial Intelligence (AI) and Machine Learning (ML) data.

NEW in the outline effective 1 August 2026. Protecting training data, model artefacts and inference data, including provenance, poisoning and leakage risks, and how classification and retention rules apply to AI datasets. This subdomain does not exist in pre-2026 prep material.

training data protectiondata poisoningmodel inversion and data leakageAI data provenanceAI dataset retention

Practice question for this objective

Free sampleCloud Data Securityhard

A public-sector agency runs a language model on a private cloud. Its training corpus contains classified citizen records, and policy states such data may be kept only as long as the specific training purpose requires. An audit finds copies of the raw corpus, intermediate feature files and checkpoint snapshots scattered across storage buckets with no expiry. What should the agency do next to align with its obligations?

  • ADefine and enforce retention schedules for the corpus, feature files and checkpoints, with secure disposal once the training purpose is met Correct
  • BMove all copies to a colder storage tier to reduce the cost of holding the classified data indefinitely
  • CTokenise the citizen identifiers within the raw corpus so the records can be kept for future retraining
  • DGrant the data science team broader read access so the scattered copies can be consolidated into one bucket
Apply purpose-bound retention and secure disposal to AI training datasets and derived artefacts, not just raw records. AI data retention obligations cover the whole training footprint, including derived feature files and model checkpoints, which can still embody the sensitive source data. Aligning with a purpose-limitation policy means defining retention schedules for each artefact and disposing of them securely once the training purpose is satisfied; tiering, tokenising to prolong holding, or consolidating access all leave the over-retention unresolved.

Why A is correct: The finding is uncontrolled retention of AI training artefacts, so setting purpose-bound retention periods and securely disposing of data when no longer needed is the corrective action.

Why B is wrong: Changing storage tier lowers cost but leaves the data retained past its purpose, which is the exact policy breach the audit identified.

Why C is wrong: Tokenisation reduces identifiability but is offered here to justify keeping the data, which still violates the purpose-limited retention requirement rather than satisfying it.

Why D is wrong: Consolidating and widening access addresses sprawl and convenience, not the core problem that classified training data is being retained beyond its permitted period.

See more CCSP practice questions, answers explained.

More in this domain

Back to all Cloud Data Security objectives, or the CCSP cert hub.

Examworthy is not affiliated with or endorsed by ISC2. Original, blueprint-aligned practice material only.