Misc

Governing AI Across the Whole Portfolio

Part II of our series on InfoSec in the age of AI

In April, Christoph Klaassen and I wrote about the paradigm shift that AI brings to information security management and GRC. We argued that traditional governance, built for human-speed processes and annual cycles, can’t keep up with systems that reason, plan, and act. The blog post closed with four shifts every InfoSec professional should make.

Several readers came back with the obvious follow-up: “Fine. But what do I actually do on Monday morning?”

This post is our attempt at a first answer, given things in ($LARGE_ORGANIZATION) reality are often slightly more complicated ;-) It’s opinionated, occasionally provocative, and (we hope) a bit more practical.

First, Get the Picture Right

Much of the current debate is framed as an either-or: SaaS AI or self-hosted models, the hyperscaler’s copilot or the open-weight model in your own data center. That framing is neglecting operational reality in enterprises out there. It doesn’t describe how organizations actually use AI. In reality, four deployment modes coexist in almost every enterprise we see:

  1. Consumed SaaS AI: copilots, chat assistants, AI-native SaaS tools.
  2. Embedded AI: AI features that vendors add to products you’ve (hopefully) already approved, often switched on by a routine update. This is a very acute risk scenario, as we see this trend accelerating across all applications and services.
  3. API-built applications: your own apps and agents built on top of a commercial model API.
  4. Self-hosted models: open-weight models running on your own sovereign infrastructure, increasingly popular for sensitive data and cost reasons.

Risk does not disappear as one moves along this spectrum, it just changes. On the SaaS end, your control is mostly contractual, and you depend on the vendor’s transparency. On the self-hosted end, your control is technical, and the entire burden of proof is yours. Neither end is “safer.” They’re differently unsafe.

With that picture in mind, here are seven theses how the CISOs and their InfoSec departments might govern this portfolio of AI deployments.

Thesis 1: Your AI Policy Is a Static Document. Your Agents Don’t Read Those.

Most organizations responded to generative AI the way they respond to every technical advancement: they wrote a policy. It lists acceptable use, forbidden data categories, and approval requirements. It’s reviewed annually (rly?), and it’s read by pretty much nobody, least of all the agents that are taking over your business processes with an ever-increasing speed and scale.

In the agentic world, policy has to live where actions happen. What that means depends on the mode:

  • SaaS and embedded AI: policy is enforced through tenant configuration, such as which connectors are allowed, which data sources the assistant may index, and whether web grounding is on. Your admin console represents your policy to the agent. Ensure that they are “reading” it.
  • API-built and self-hosted agents: policy is enforced at the tool-call layer, through a gateway that checks each action against an allow-list before it executes.

The written policy doesn’t go away, but its role changes: it should describe where the enforcement happens and who owns it. If a policy statement has no organizational and technical enforcement points at the same time, it won’t control anything.

Thesis 2: An Agent Is an Employee You Hired Without a Background Check.

We’d never give a new hire admin rights to critical systems on day one without knowing who they are. However, that’s roughly what happens when an agent is connected to production systems.

Some things are constant across all modes. Every agent needs:

  • a named human owner,
  • scoped permissions,
  • a documented purpose,
  • and most importantly a way to revoke it quickly.
    Agents frequently act on behalf of a user and silently inherit everything that user can access, which is rarely what was intended.

What differs between different modes is what the background check examines:

  • SaaS and API: the vendor’s model governance, data retention, training use of your data, and sub-processors.
  • Self-hosted: the provenance and integrity of the model itself. Who trained it, on what, and could it contain hidden behavior? Research on “sleeper agent” models has shown that backdoored behavior can persist through standard safety training. The model files themselves are a supply chain artifact too; some serialization formats execute code on load. Approved sources, hash verification, and safe formats should be the minimum.
  • Embedded: whether anyone checked at all.

Thesis 3: Controllability Is What You Can Govern.

Regulators and boards ask for explainability. With large models, you won’t get it in any meaningful sense, and neither will your vendor. What you can govern is the blast radius. This leads us to the following practical approach to classify every action an agent can take:

  • Read-only: low risk, logged.
  • Reversible write: medium risk, logged and monitored.
  • Irreversible or external: payments, deletions, outbound communication. These require strong controls, including human approval where it’s meaningful (see Thesis 4).

A second, highly practical tool for architecture reviews is what Simon Willison calls the lethal trifecta: an agent that combines access to private data, exposure to untrusted content, and the ability to communicate externally can be made to exfiltrate data through prompt injection. This applies regardless of deployment mode. Your copilot reading external emails has the same structural problem as your self-built agent browsing the web. In case a design review finds all three properties, remove one, or document very explicitly why the residual risk is considered acceptable.

Thesis 4: Human-in-the-Loop Is Often a Compliance Fiction

Put a human in the loop, and the auditor is happy. That’s something you hear and read everywhere. However, it’s well known from work tasks comparable to assembly line work that repetitive tasks lead to lowered awareness and focus. We can reason that if a high number of agent actions per day is approved by a human, averaging three seconds each, with a 99.9% approval rate, you cannot expect to have a control. You then have a rubber stamp with a salary.

Meaningful human oversight, which the EU AI Act explicitly demands for high-risk systems, requires three things:

  • context: the human understands what they’re approving
  • time: they can actually evaluate it
  • authority: they can stop the system, and doing so has no career penalty

Visibility differs sharply across modes. With SaaS and embedded AI, you often only see what the vendor chooses to log, hence log export and audit trail access should belong in your mandatory procurement requirements. On-prem, full telemetry is possible, including prompts, tool calls, and outputs, but only if you build it. Either way, oversight should increasingly rely on observability and anomaly detection, with human approval reserved for the decisions that truly warrant it.

A practical test: measure your approval rates and times. If the numbers look like the ones above, redesign the control before an auditor – or an attacker – notices.

Thesis 5: Error Budgets, Not Zero Tolerance

In Part I, we argued for managing “ranges of acceptable behavior” rather than guaranteeing outputs. Here’s how to make that concrete: borrow from Site Reliability Engineering by defining an acceptable failure rate, an error budget, and continuously test against it.

The mechanism is a constant adversarial test suite covering prompt injection, data exfiltration attempts, and policy violations, relevant to each use case. Its pass rate becomes both your control and your audit evidence. That’s what “continuous assurance” looks like in practice, and this lowers audit preparation work at the same time ;-). A security win in every aspect.

What differs is who causes the change:

  • SaaS and API: the provider can change the model underneath you, sometimes even without notice. Continuous testing would be your drift detector in this case.
  • Self-hosted: changes can be a new model version, a different quantization, a fine-tune, an updated inference engine, a modified system prompt. Each one can shift behavior and that’s why testing becomes a release gate in your own pipeline.

Thesis 6: Your Most Critical Supplier May Be One You’ve Never Assessed

Third-party risk management (TPRM) was built for suppliers with contracts, SLAs, and questionnaires. AI stretches this model in three directions:

  • SaaS and API: classic TPRM still applies, but it needs AI-specific questions. Do you get notified of model changes? Can you pin versions? Is your data used for training? Where is inference running?
  • Embedded AI: a vendor you assessed two years ago silently adds AI features in a product update. A new AI feature must be a trigger for reassessment.
  • Self-hosted open weights: there may be no supplier to contract with at all, just a model publisher who doesn’t know you exist. TPRM is replaced by provenance checks and scrutiny of the inference stack (serving engines, orchestration frameworks, drivers), which is where classic vulnerability management reenters the picture. This is very comparable to the general discussion of supply chain security, where we see a strong increase of attacks throughout the last year.

Thesis 7: Shadow AI Is a Demand Signal, Not a Discipline Problem

When employees paste confidential data into unapproved AI tools, the reflex is to block and reprimand. Most of us know that this rarely works; usage just moves to personal devices, which you can’t observe, and you thereby lose control over your organization’s valuable information assets and potentially even intellectual property.

Shadow AI tells you that people have a real need your approved tools don’t meet. The answer is a paved road: an approved option that is easier to use than the forbidden one.

This is where the portfolio view pays off. The paved road doesn’t have to be one tool. Data classification can decide placement: a self-hosted model for sovereign, sensitive and regulated data, an approved SaaS assistant for everything else. Users get a clear rule, and your organization gets governance on both.

Who Owns What

The full problem space needs clear ownership. Our proposal:

  • GRC owns the inventory across all four modes, risk acceptance, and mapping of regulatory roles and obligations.
  • The CISO sets the strategy with the board and subsequently architecture principles and risk appetite, and reports the portfolio view back to the board.
  • SecOps owns telemetry, detection, adversarial testing, and the kill switch, whatever the deployment mode.
Control domain Owner SaaS Embedded API-built Self-hosted
Inventory GRC Procurement Update monitoring Registry Registry
Policy enforcement CISO Tenant config Tenant config Gateway Gateway
Supplier / provenance GRC TPRM Reassessment trigger TPRM Model intake
Change & testing SecOps Drift detection Drift detection Release gate Release gate
Telemetry & oversight SecOps Log export Log export Own logging Own logging
Kill switch SecOps Disable tenant feature Disable feature Revoke credentials Stop service

Table 1: Control Domains, Owners, and Deployment Modes

Five Metrics for the Board (Instead of a Heatmap)

We promised in Part I to stay away from heatmaps. Here’s what we’d report instead:

  1. Coverage: percentage of AI use cases in the inventory with a named owner, across all four modes.
  2. Least privilege: percentage of agents running with scoped, dedicated credentials rather than inherited user rights. You are already taking care of NHIs, right?
  3. Assurance: adversarial test pass rate against defined error budgets, trended over time.
  4. Response: mean time to revoke an agent or disable an AI feature.
  5. Oversight quality: ratio of irreversible agent actions to meaningful human reviews.

None of these are perfect. Still they all are more honest than a red-yellow-green square which is not providing any anchoring nor insight on what’s really happening in the machine room.

Closing

Part I ended with “Let’s build it together.” This post is the scaffolding and every thesis here is debatable. We’d like to hear where you disagree, especially if you’ve seen something work (or fail spectacularly) in practice.


These topics, and many more, will be discussed among practitioners at our upcoming AI Security Summit on November 5th.

These are our next trainings on AI Security: November, December, March.

Sources & Further Reading