AzureCIO BriefingsRetrospectives

CIO Brief: AI Projects Create New Data Exposure Paths

By OnCloudSec Research Team · Published Oct 6, 2026 · 1 min read

Retrospective: this article looks back at events from September 2023, written in 2026 with the benefit of hindsight.

The short version: In 2023, Microsoft's own AI researchers accidentally exposed 38 terabytes of internal data — including passwords and private messages — by sharing a single overly broad access link on GitHub. AI projects move fast, involve huge datasets and often sidestep normal data controls.

Why AI projects create new exposure paths

AI work involves collecting, copying and sharing large datasets and models — often by researchers or data scientists who aren't security specialists. Data moves to new storage locations, gets shared with external collaborators and ends up in public code repositories.

The business impact

  • Data exposure at very large scale.
  • Tampered AI models if attackers gain write access.
  • Leaked credentials found in backups and datasets.

Questions to ask your team

  • Where do our AI and data science teams store training data and models?
  • How do they share data with partners or the public?
  • Are those storage locations covered by our normal security policies?
  • Who reviews data before it's used to train or ground AI?

What good looks like

AI data stored in governed locations, sharing through controlled, short-lived access, secret scanning on all repositories, and AI projects included in security reviews from the start.

The decision

Ask for an inventory of data used by AI initiatives and where it lives. If your AI team works outside normal data governance, bring it inside before the next project launches.

microsoft 38tb sas token leak impactMicrosoft AI SAS token 38TB2023

More on this story