All Posts
August 6, 2026 ·11 min read

The AI Governance For Board Members

A board-level checklist for governing AI systems and agents: inventory, delegated authority, security, regulatory readiness, and accountability.

Boards are now actively asking about AI. The issue is not awareness: it is the quality of the questions. “Are we using AI?” and “Do we have a policy?” are easy to answer, and the answers do not indicate whether an AI program is safe, compliant, or delivering value.

The governance surface has also widened. Organizations are no longer deploying only models that produce recommendations. They are deploying agents that hold credentials, call tools, update systems, and act across organizational boundaries. A model inventory is no longer enough; boards need visibility into the systems, data, tools, identities, and authority assembled around it.

The regulatory calendar has already moved. The EU AI Act’s Article 50 transparency duties took effect on August 2, 2026, while the July 2026 AI Omnibus moved most standalone high-risk obligations to December 2, 2027 and product-embedded high-risk systems to August 2, 2028. The distinction matters, since a deferred high-risk program does not defer the transparency duties that reach chatbots and synthetic content, and those are in force today. The scope of each duty, and the retroactivity condition buried inside it, are worked through separately.

In the United States, obligations continue to vary by jurisdiction, sector, and use case. NIST’s voluntary AI Risk Management Framework remains a useful operating reference, supplemented by its Generative AI Profile. As of August 2026 NIST states that AI RMF 1.0 is being revised and has published no completion date, so internal controls should track the revision instead of freezing to the 2023 text. The durable board posture is therefore not compliance with one static checklist. It is an operating model that can map a changing rule or threat to a current system inventory quickly.

What follows is a practical checklist for that discussion. The questions are phrased the way they should be asked in a governance setting.


1. Inventory and accountability

You cannot govern what you cannot see. The first failure mode in most organizations is that nobody has a current, accurate picture of where AI is actually being used, including shadow deployments, embedded vendor features, background agents, and tools connected after the original system was approved.

Questions to ask:

If ownership resolves to a committee, responsibility is diffused. Effective governance requires a named executive accountable for outcomes and a system record detailed enough to show what the organization has authorized.


2. Risk classification and proportionality

Not every AI system warrants the same level of oversight. A marketing copy assistant and a credit adjudication model should not be governed identically. Nor should a support assistant that drafts a response and an agent that can issue a refund. A useful classification considers both legal category and operational authority: what the system influences, what data it can reach, and what it can do without a person approving the specific action.

Questions to ask:

A practical test: review the last three AI systems that were blocked or materially modified through governance. If none exist, the process is likely performative.


3. Data lineage and training data provenance

Some of the most consequential AI incidents begin as data failures: training on data without appropriate rights, using customer information beyond its disclosed purpose, or allowing an agent’s shared memory to cross a tenant boundary. These become legal, contractual, and trust failures rather than ordinary engineering tickets.

Questions to ask:

This is often the point where clarity breaks down. Many organizations do not have a defensible answer, and the usual reason is that deletion was designed for the source system alone — every surface that can still return the content has to be on the list, including the indexes, caches, and derived summaries built from it.


4. System evaluation and ongoing monitoring

“We tested it before launch” is not evaluation. Production AI systems change as user behavior shifts, source data changes, vendors update models, and teams add tools or permissions. For an agent, the outcome depends on the model, prompt, context, memory, tool behavior, and authorization path. Evaluating only the final text misses most of the system.

Questions to ask:

Boards should request trend data on evaluation metrics for high-risk systems. If it does not exist, that absence is itself a finding. The machinery that produces those metrics belongs to a shared platform, since every product team otherwise rebuilds it at its own standard, which is the argument in the operating model behind production AI. The measurement set that applies to each system type is on the platform checklist.


5. Human oversight and escalation

Regulatory frameworks emphasize meaningful human oversight. In practice, this is often reduced to a confirmation button shown so frequently that approval becomes automatic. That is not oversight. A human gate is meaningful only when the reviewer has enough context, time, authority, and a safe way to disagree.

Questions to ask:

The measurable version of all this is the override rate, tracked per system and reviewed over time. A reviewer population that never disagrees is either unnecessary or unequipped, and a policy document cannot tell the board which one it is looking at.


6. Security and adversarial resilience

AI systems extend the traditional application-security surface. Prompt injection can arrive through a retrieved document, a tool result, or another agent, so the untrusted-input boundary sits wherever content enters the system, well upstream of the prompt box. Persistent memory can carry a poisoned instruction into later sessions, and a compromised tool or MCP server can turn a model error into an unauthorized action. OWASP’s current Agentic Security Initiative reflects this shift from securing a conversational model to securing an autonomous, connected system.

Questions to ask:

Ask when the last exercise ran and what it found. A programme whose findings are uniformly low-severity is usually testing the model in isolation, whereas the consequential failures live in the surrounding system of tools, memory, and permissions — which is where the testing methods that surface them are aimed.


7. Third-party and supply chain risk

Most enterprises consume AI through a chain of vendors: foundation model APIs, embedded SaaS features, agent platforms, and remote tool servers. A vendor review focused only on the model provider misses much of the operational dependency.

Questions to ask:

The practical test is whether a vendor can widen what its agent reaches without telling you. Where the contract requires no notice of capability changes, the inventory in section one goes stale on its own, and nobody inside the organization has done anything wrong.


8. Regulatory readiness and documentation

The Omnibus changed the schedule, not the need for a regulatory operating model, and two things follow from the dates above. Transparency duties are already live, so the board question is what shipped rather than what is planned. The deferred high-risk dates then create a funding question this year, because an evidence roadmap that starts in 2027 will not be ready for a December 2027 application date. The Omnibus also replaced the earlier company AI-literacy obligation with non-binding encouragement, though workforce competence remains a practical control, and U.S. requirements continue to vary enough by jurisdiction and sector that a single national checklist is not a defensible substitute for legal mapping.

Questions to ask:

Readiness here is mostly an evidence problem, and evidence is cheap to produce as the work happens and expensive to reconstruct a year later. A board can test the whole programme by asking for one artifact — the current documentation for a single high-risk system — and noting how long it takes to arrive.


9. Board-level reporting

Finally, a test of the governance program itself: what does the board see, and how often?

Questions to ask:

If AI appears only as a line item in a broader technology update, that is itself a governance gap.


The uncomfortable conclusion

Most AI governance programs are well-intentioned but structurally incomplete. Policies, committees, and presentations exist; operational mechanisms often do not. What is missing is execution: current system and authority inventories, active evaluation, enforceable permission boundaries, empowered reviewers, and evidence that controls work.

The point of a board-level checklist is not to become the expert. It is to ask questions precise enough that the absence of a real answer becomes visible. If executives cannot answer most of the questions above with specifics — names, numbers, artifacts — the governance program is not yet real, and the board now has a decision to make about what to do next.

The decision that follows is what matters. The rest is paperwork.

If your board is asking harder questions about AI and your team is working out how to answer them, I help executives build governance programs that are both compliant and operationally real. Get in touch.

Subscribe for more

Get posts on AI platforms, retrieval, agents, security, governance, and production engineering.

Subscribe on Substack