Board AI Oversight: What to Ask Every Quarter, Not Once a Year

Board AI Oversight: What to Ask Every Quarter, Not Once a Year

I've sat through the AI slide in enough board packs to recognise it on sight. Same five bullets, different quarter. "Programme on track." "Pilot expanding." "Exploring further use cases." Someone asks a question, gets a confident answer, and the item closes in under ten minutes. Nobody in the room could tell you whether the model's output quality moved since the last meeting, whether spend is tracking the business case, or whether a vendor dependency just became a problem. The slide looks like oversight. It isn't. That gap matters more in AI than it does almost anywhere else on a board agenda. A capital project drifts slowly enough that an annual check catches most of what matters. An AI system doesn't. Accuracy degrades, vendors consolidate, a team quietly works around a tool nobody trusts, and three months later all of that has compounded into something a board should have caught at the first sign. By the time it reaches the board as a problem, it usually didn't arrive as one quarter ago — it arrived as a trend that nobody was tracking because nobody had a standing question that would have surfaced it. This is about that standing question — or rather, the small set of them that belong on every quarterly board agenda, not just the annual strategy day. Not a maturity checklist. A recurring practice that keeps a board genuinely informed between the big strategic conversations, so those conversations start from an accurate picture instead of a recycled slide. Why an annual briefing isn't oversight Most boards already do the annual AI conversation reasonably well. There's a strategy day, a vision presentation, maybe an external speaker. The CIO or CTO walks through the roadmap, the board asks good questions, everyone leaves with a shared sense of direction. That's a real and useful exercise. It is also not oversight, and treating it as if it were is where boards get into trouble. Oversight is about exposure to risk and to reality in something closer to real time. A model's behaviour can shift within a single quarter — a vendor changes an underlying model version, a data source upstream degrades, usage patterns move outside what the system was validated against. None of that waits for the annual strategy day. Spend behaves the same way: AI infrastructure and API costs have a documented pattern of starting small and compounding faster than the budget anticipated, and a board that checks in once a year finds out about that compounding well after it's already locked in. The annual conversation answers "where are we going." The quarterly conversation has to answer "what changed since we last looked, and does it change the plan." Those are different questions, and a board that only asks the first one is flying with a twelve-month blind spot on a topic that moves on a much shorter clock. The five things that belong on every quarterly agenda What follows isn't a maturity framework. It's the minimum set of questions a board needs answered every quarter to stay ahead of the AI programme instead of catching up to it after something breaks. Value realisation. Of the AI initiatives this board approved, which ones are still tracking to the business case that justified them, and which have quietly missed it without anyone flagging it? Most board AI updates report activity — pilots launched, use cases identified — rather than whether anything launched two quarters ago is actually delivering the return it promised. Ask for the list of live initiatives against their original business case, every quarter, not just at renewal time. Risk and incidents. What AI-related incident, near-miss, or manual override happened this quarter that didn't make it to this table? Most AI failures don't look like outages. They look like a team quietly reverting to the manual process because they stopped trusting the output, or a customer-facing error that got caught and fixed before anyone escalated it. Those are exactly the things a quarterly incident review is supposed to surface — and exactly the things that get filtered out of a polished update slide. Spend against plan. What's the variance between this quarter's actual AI spend and what was budgeted, and what's driving it? Compute, model API costs, and retraining cycles are the line items that move fastest and least predictably. A board that only sees this annually discovers the overrun after several quarters have already compounded. Model and data integrity. Has output quality or accuracy moved since last quarter, and who would have caught it if it had? This is the question most boards never ask, because it requires management to have an actual monitoring answer rather than a reassurance. If the honest answer is "we don't have a way to know," that's the most useful answer the board will get all year — it tells you exactly where the next incident is going to come from. Decisions made without the board. What AI-related decision did management make this quarter that this board should have seen before it was finalised? Architecture choices, vendor selections, and data-handling decisions carry financial and risk implications that often get made at the technical layer and never surface until they're already locked in. Asking this explicitly, every quarter, is how a board stays ahead of decisions instead of inheriting them. None of these five require deep technical fluency to ask. They require the discipline to ask the same five things every quarter, in the same order, and to treat a vague answer as information rather than reassurance. What a real answer sounds like, and what a slide sounds like The difference between genuine oversight and a recycled update usually isn't in the question. Boards mostly do ask some version of "how's the AI programme going." It's in whether the answer has numbers attached and whether it changes from quarter to quarter. I worked with one organisation where the model accuracy on a customer-facing tool had drifted by roughly eleven points over two quarters before anyone outside the technical team noticed — not because the drift was hidden, but because nobody had ever asked the question that would have surfaced it. The board had been getting "the programme is performing well" for two consecutive quarters. Both statements were technically true and operationally useless. Accuracy was fine in aggregate. It had also been getting worse the entire time, in a segment that mattered, and the slide didn't have a line for that because nobody had asked for one. That's the tell to listen for. A real answer to "has output quality moved" comes with a number and a trend line, even if the trend is mildly bad. A slide answer comes with an adjective. "Strong." "On track." "Stable." Adjectives are what you get when the question hasn't been asked with enough specificity to force a number out of the person answering it. The fix isn't to distrust management — it's to ask the question in a form that can't be answered with an adjective. The same test applies to spend. "We're managing costs closely" is a sentence with no information in it. "Spend is 14% over plan, driven by retraining frequency on the fraud model" is a sentence a board can act on. If a quarterly update doesn't produce sentences like the second one, the board isn't getting oversight — it's getting reassurance dressed as a briefing, and the two are not the same thing even when they're delivered by the same confident person in the same conference room. Who should be answering, not just presenting A quarterly AI update usually gets delivered by whoever is best at presenting to a board — often someone several layers removed from the number being discussed. That's a reasonable choice for a strategy day. It's the wrong choice for oversight. The person answering "has accuracy moved" should be someone who would actually know, not someone summarising what they were told. The person answering the spend variance question should be able to explain the driver without checking a slide. This sounds like a small procedural point. It isn't. A polished intermediary will always round a complicated, slightly bad answer into something cleaner than the truth — not out of dishonesty, but because that's what presenting to a board selects for. The fix is structural: rotate who answers each of the five categories so the board is hearing from the person closest to the number, even if that person is less comfortable in the room than the usual presenter. This also changes what the CIO, CTO, and CFO need to bring to the table. Joint accountability for these five categories means none of them can treat the AI update as someone else's slide. The CFO owns the spend variance and increasingly the value-realisation number, since that's where the business case lives. The CTO or CIO owns risk, incidents, and model integrity, since that's where the operational reality sits. Decisions made without the board is a shared answer — neither function should be able to claim that one wasn't theirs to flag. Building the cadence so it doesn't slip None of this works as a one-time fix to the board pack. It works as a standing template that doesn't change shape from quarter to quarter, because the value of the exercise comes from the comparison across quarters, not from any single quarter's answer. Keep the same five categories in the same order every time. A board that lets management reorganise the update each quarter loses the ability to compare this quarter's spend variance to last quarter's, which is most of the point. Require numbers and trend lines, not status adjectives — if a category can't be answered with a number, that absence is itself the finding, and the board should treat "we don't measure that yet" as an action item, not a shrug. Send the actual figures before the meeting, not as a reveal in the room. A board reading numbers cold in real time asks shallower questions than a board that had ten minutes beforehand to notice that the spend variance looks odd or that the incident count is trending up. And separate this quarterly mechanism explicitly from the annual strategy conversation — the strategy day asks where the programme is going; the quarterly mechanism asks what changed since the board last checked. Boards that fold the two together end up doing neither well, because the questions pull in different directions and one set of slides can't serve both purposes. The cost-of-inaction conversation I've had with several CFOs usually centres on the spend and value side of this list — what waiting actually costs, measured properly. The board oversight conversation is the other half of the same problem: an organisation can be investing aggressively in AI and still have no functioning oversight of what that investment is actually producing quarter to quarter. Spend without oversight is not caution. It's exposure that hasn't been priced yet. What to take from thisSeparate the annual AI strategy conversation from quarterly oversight. They answer different questions and one cannot substitute for the other. Put five standing categories on every quarterly agenda, in the same order each time: value realisation, risk and incidents, spend against plan, model and data integrity, and decisions made without the board. Treat status adjectives as a non-answer. If a category can't be answered with a number or a trend line, that gap is the finding — log it as an action item, not a status update. Rotate who answers each category to the person closest to the number, not the most polished presenter. A summarised answer rounds off exactly the detail the board needs to hear. Send the figures ahead of the meeting. A board reading numbers cold in the room asks shallower questions than one that had time to notice what looks off beforehand. Review the cadence itself annually. A category that's been easy to answer for a year may need tightening once the programme's risk profile changes — oversight needs maintenance the same way the AI systems it's watching do.The boards that get this right aren't the ones with the most technically literate directors. They're the ones who stopped accepting an adjective where a number belongs, and who kept asking the same five questions long enough that the answers became comparable across quarters. That's a less glamorous practice than an annual AI strategy day with a guest speaker. It's also the one that actually catches the problem while it's still small enough to fix in a meeting instead of in a post-mortem.

Read full article
Delayed AI Adoption: The Financial Cost CFOs Aren't Modelling

Delayed AI Adoption: The Financial Cost CFOs Aren't Modelling

I've sat in enough quarterly reviews to recognise the posture. The AI agenda item comes up. The CFO gives the answer: "We're watching the space. We'll move when the ROI is clearer." The room nods. The item closes. And on the spreadsheet, there's a zero next to AI investment, which looks indistinguishable from prudence. The problem is that "watching" has a price that doesn't show up in that spreadsheet. It doesn't appear as a line item, a variance, or a write-down. It accumulates in places that only become visible later — in an attrition number that looks like a hiring problem, in a sales cycle that looks like a pricing problem, in a capability gap that looks like an execution problem. By the time it surfaces, the framing has already shifted from "we were careful" to "we fell behind." This isn't an argument for reckless AI spending. Most organisations that rushed in during 2023 and 2024 wasted money and generated reports nobody read. But the CFO who treats inaction as cost-free is making a financial modelling error. What follows is what that error actually looks like — and what a realistic cost-of-delay framework should account for. Why inaction looks free on the balance sheet CFO decision-making is built around visible costs. You can see what an AI pilot costs. You can see infrastructure spend, vendor contracts, and the headcount required for an AI team. What you cannot see — at least not immediately — is what a competitor's AI programme is costing you. This creates a structural asymmetry. Every pound spent on AI shows up as a cost. Every pound not spent shows up as nothing. The CFO's instinct, correctly shaped by years of capital allocation discipline, is to treat the nothing as neutral. It isn't. The parallel I find useful: data infrastructure in the early 2010s. Companies that invested in data warehousing and analytics capabilities between 2010 and 2015 weren't seeing obvious short-term returns. The value was in what they could do in 2017 and 2018 that competitors couldn't — run faster experiments, personalise at scale, catch fraud earlier. Organisations that waited for "clearer ROI" were, by the time they understood it, three years behind on the foundational work. AI has a shorter compounding cycle than data infrastructure did. An organisation that started building internal AI capability in 2023 has two years of institutional learning — how the tools fail in their specific context, which processes respond to automation and which don't, how their workforce adapts — that cannot be purchased. You can buy the tools. You cannot buy that understanding. Where the cost of delay actually lands "Competitive disadvantage" is too abstract to put in a financial model. Here is where delayed adoption creates measurable exposure. Talent is the most immediate and least visible cost. Engineers, analysts, and operations leads who want to work with AI tools are choosing employers who give them access to those tools. This is happening now. I've spoken to CTOs at organisations in the "we're monitoring" camp who are losing mid-level technical talent at a rate they're attributing to compensation. Some of it is compensation. A significant portion is that their competitors' engineers are doing more interesting work with better tools. The cost per lost hire — recruiting, onboarding, lost productivity during the gap — runs into six figures per person. That is quantifiable. It belongs in the model. Sales cycles are also affected in industries where AI is showing up in how deals are won. In financial services, insurance, logistics, and professional services, the competitor who can demonstrate AI-driven capability in a client meeting has a different conversation than the one who cannot. I've seen deals where the question is no longer "can you solve the problem" but "can you solve it at AI speed." The organisation that answers "not yet" is competing on a different basis. Process cost differential compounds every month. AI is reducing the unit cost of specific tasks — document processing, contract review, first-pass analysis, compliance checking — at rates between 40% and 80% depending on the task and the organisation. Every month a competitor runs those processes at that cost and you don't, the gap widens. For high-volume document-heavy operations, the monthly difference is not marginal. Data maturity debt is the cost that CFOs most consistently underestimate. AI systems improve with more data and better-structured data. An organisation that started an AI programme two years ago has two years of logged interactions, feedback loops, and model fine-tuning data. An organisation starting today has none of it. You cannot buy your way to that data — you can only accumulate it over time. The capability gap is not just in tools; it's in the training material that makes those tools work better for your specific context. Why the gap grows faster than expected The CFO who plans to "move when the ROI is clearer" is making an implicit assumption: that the cost of delay is roughly linear. Start now or start in twelve months, and you're twelve months behind. That's not how this works. AI advantages compound because data, models, and institutional knowledge reinforce each other. An organisation with a functioning AI programme is generating data from that programme. That data improves the models. Better models drive more adoption. More adoption generates more data. The gap between an early mover and a laggard doesn't grow at a steady rate — it accelerates. The organisation that is twelve months ahead today may be twenty-four effective months ahead in two years. Because the twelve months they had in advance were spent building capability that you have not yet started building. When you start, they are already iterating on systems you haven't built yet. In financial services, early adopters of AI-driven fraud detection are now running third and fourth-generation models trained on proprietary incident data. A bank entering the space now is not competing with those organisations' first-generation system — it's competing with what three years of production data and model iteration produces. That gap does not close in twelve months of effort. The point where an organisation can no longer close the gap through effort alone — where the advantage of the early mover becomes structural — varies by industry and use case. Across the sectors I work in, that inflection point is closer than most boards think. In some categories, it has already passed. What a cost-of-delay model actually looks like This is not a case for spending on AI without a plan. It's a case for modelling inaction honestly. A cost-of-delay model has three components that can be estimated with reasonable precision. The first is talent cost. Estimate the annual attrition rate in AI-adjacent roles and identify what fraction is tool-related rather than compensation-related — even 10% to 15% of technical attrition is significant. Multiply fully-loaded replacement cost per hire by that number. Add the cost of the skills gaps that accumulate while roles sit open. This gives a floor figure for what "watching and waiting" costs in people terms each year. The second is process cost differential. Map the three to five highest-volume internal processes where AI is demonstrably reducing costs elsewhere in your industry. Estimate current unit cost and monthly volume. Apply a conservative automation impact — 40% cost reduction is within what production deployments are showing for document-heavy processes. The monthly gap between your cost and a competitor's is the delay cost for that process cluster alone. The third, and often the largest, is capability catch-up cost. When the organisation eventually commits, it will need to hire specialists, build or acquire infrastructure, structure years of accumulated data, and run the pilots that should have been run earlier. Estimate this as 18 to 24 months of foundational work — the industry average from a standing start to meaningful production capability — multiplied by the fully-loaded cost of the required team. Add an opportunity cost factor for the efficiency or revenue the programme would have generated during that period. When I've helped CFOs run this model against a realistic "commit now" versus "commit in twelve months" scenario, the gap between the two paths is consistently larger than expected. The upfront investment in committing now is visible. The accumulated cost of waiting is spread invisibly across multiple budget lines over multiple years. What to take from thisModel inaction explicitly. The "watch and wait" position carries a real cost — it just doesn't appear in a standard budget. Build a cost-of-delay estimate and put it next to the investment case so the comparison is visible. Separate AI spending from AI capability building. Buying tools without building institutional knowledge is waste. Refusing to build capability because tools look expensive misses the point. The value accumulates in understanding and proprietary data over time, not in licence fees. Quantify talent exposure now. Pull your technical attrition data and have an honest conversation with your CTO about how much of it is tool-related. This is the fastest element of the cost to make concrete. Identify the two or three highest-volume internal processes where AI is reducing costs for competitors in your sector. These are your first comparison points for the process cost differential. Ask your CIO what starting in the next 90 days actually requires — not to scale, but to begin accumulating the data and institutional knowledge that makes later scaling possible. The 90-day question has a different, usually lower, barrier than the full programme question. Bring the compounding argument to your board. The question is not "when should we invest in AI" — it's "at what point does the gap become structural, and are we near that point in our sector." That is a strategic risk question that belongs in front of the board with the same treatment as any other structural competitive risk.The organisations that fell furthest behind on digital capabilities in the 2010s were not, in most cases, the ones that actively rejected the technology. They were the ones that treated the decision as something that could wait for better evidence. Better evidence kept arriving. So did the gap. I'm not suggesting evidence doesn't matter or that every organisation should move at the same speed. But the CFO's job is to model risk honestly, including the risk of standing still. In most sectors right now, that risk is underpriced.

Read full article
What the CFO Needs to Understand About AI Investment (That the Vendor Won't Tell Them)

What the CFO Needs to Understand About AI Investment (That the Vendor Won't Tell Them)

The deck looks great. There's a 3x ROI at month 12, a cost-per-decision metric your competitors would envy, and case studies from companies that look just like yours. The vendor has done this pitch a hundred times. They know what a CFO wants to see. The problem is the ROI model they're using was designed for software. And AI isn't software. That distinction sounds pedantic until you're twelve months in and wondering why the numbers don't match the deck. The gap between AI investment promises and P&L reality is probably the most expensive misalignment in enterprise technology right now. Not because AI doesn't deliver value — it does, for the right use cases, in the right organizations, under specific conditions. But because the financial model used to justify it was built for a different kind of purchase. Software procurement ROI runs on three assumptions: costs are predictable, value delivery is linear, and the failure mode is a delayed project. None of those hold for AI. Why the software ROI model breaks on AI Software has a cost structure that finance teams can work with. Licensing is known, implementation is estimated, ongoing support is a percentage of the license. The model is imperfect but manageable. AI cost structure doesn't map to any of those buckets cleanly. The largest cost variable in most enterprise AI programs isn't the AI itself. It's data. Before a model can be trained on anything useful, someone has to assess what data you actually have — which is usually different from what the business thinks it has — fix the quality problems, integrate sources that weren't built to talk to each other, and set up the governance to make sure the training data is legally usable. That work is slow, expensive, and almost never appears on a vendor proposal. It also doesn't end: data quality degrades, systems change, and each new use case adds new requirements. Compute is the second piece that gets undercounted. Training costs and inference costs are different things, and vendor estimates typically focus on training. Inference is what you pay for in production — every time the model scores a new input. For high-volume use cases like fraud detection, real-time pricing, or recommendation, inference costs at scale regularly exceed what the organization paid to train the model. Cloud pricing makes this easy to miss until the bills start arriving. Then there's talent. AI teams don't price like enterprise software teams. Data scientists, ML engineers, and MLOps specialists have their own market rates, and those rates aren't decreasing. More importantly, the team that builds a model is different from the team that runs it in production. Both need to be funded and sustained for as long as the model is in use. The last piece is governance and monitoring. Every production model needs drift detection, performance tracking, audit logging, and a scheduled retraining cadence. This is unglamorous, recurring spend that consistently goes missing from initial program budgets. A model without monitoring isn't a production model. It's a liability on a timeline. The time-to-value curve vendors don't show you Vendor decks show value beginning to accumulate somewhere around month six. The actual pattern is different enough to change how you fund the program. The first three months are almost entirely cost. Data assessment, infrastructure setup, hiring or contracting the team, use case definition. Nothing deployable. Months four through nine are where the model gets built and tested. Results exist but aren't trusted enough to act on. This is when programs are most at risk of being canceled — the spending is real, the returns aren't visible yet, and the business is getting impatient. Months ten through eighteen are shadow deployment and validation. The model scores live data. Outcomes get compared against what actually happened. Trust builds incrementally, or it doesn't build at all. Past eighteen months is typically where the value curve starts moving in a way that looks anything like the deck. And it does compound — more production data, a team that understands the operational patterns, a process that's been rebuilt around the AI output. The economics improve over time. But only if the program survives long enough to get there. If the board expects visible returns at month twelve and the program is in month nine with real costs and nothing to show yet, someone will pull funding. The time-to-value curve needs to be part of the approval conversation, not something the program team manages quietly while hoping performance picks up. The opex trap Most boards think of technology investment as capex: a project spend that produces an asset and then stops. AI programs don't work that way. The ongoing costs are material — compute that scales with usage, continuous data quality maintenance, model monitoring, and retraining when performance drifts. An organization that funds AI as a project will hit a wall when the project budget closes and someone realizes the model needs sustained investment to stay useful. This also changes the unit economics conversation. The question isn't just "what does it cost to build this?" It's "what does it cost to run this for three years?" Those are different numbers, and the second one is the one that matters for the actual investment decision. Red flags in vendor ROI models Two things in AI vendor decks deserve specific scrutiny. FTE displacement is the most commonly inflated line item. Many ROI models show cost savings by treating automated tasks as direct headcount reductions. In practice, organizations rarely convert FTE displacement into hard savings. People get redeployed to other work, absorbed into open roles, or kept on to manage the exception cases the model can't handle. The productivity gain is real — the cost reduction usually isn't, unless the organization explicitly plans a workforce reduction. A vendor model that treats FTE displacement as a direct cost saving is overstating the ROI. Efficiency gains disconnected from business outcomes are the other pattern. "Your team handles the same volume 30% faster" is a productivity improvement. It becomes a financial outcome only if the freed capacity generates revenue or the cost base actually decreases. Efficiency claims need to be connected to a specific result, not left as an assumption that value will follow. And case studies drawn from other companies at different scales in different industries are useful for direction only. The right ROI model uses your baseline, your data quality, your integration complexity, and your team's capability. A vendor can't assess any of those from a discovery call. What to actually track Total ROI — value divided by investment — tells you the aggregate return after the fact. It doesn't tell you whether a program in flight is working. The metrics that do: Model performance against baseline matters first. Is the model improving, and is that improvement translating to better decisions? The baseline needs to be set before the program starts — what was the business doing before the model existed? Without a documented baseline, there's nothing to measure against. Production adoption rate tells you whether the business integration is actually working. A model that produces output nobody acts on isn't delivering value regardless of how well it scores in testing. What percentage of model outputs are actually consumed by a decision-making process? Cost per decision at volume should decrease as throughput scales. If it isn't, the infrastructure design or use case economics have a problem worth investigating. Retraining cost trend should improve as models mature. If the cost and time to retrain keeps rising, the data architecture has a compounding problem that will only get worse. The success definition that usually goes missing AI programs get approved with vague success criteria because specificity feels like it creates accountability before the team has figured out what's achievable. That logic runs backward. Vague criteria are what allow programs to run for eighteen months without anyone agreeing on whether they're working. A complete success definition has four components: a specific metric, a documented baseline, a numeric target, and a date. "Improve fraud detection" is not a success definition. "Reduce the false negative rate from 4.2% to below 2.5% by Q3 of next fiscal year" is. The CFO's job is not to slow the program down by demanding this. It's to make the investment defensible when someone asks whether it's working. And in every organization I've seen do this at scale, someone eventually asks.

Read full article
The AI Talent Gap That Will Determine Whether Your Strategy Delivers

The AI Talent Gap That Will Determine Whether Your Strategy Delivers

The most common response to an AI talent gap is a senior hire. A Chief AI Officer, a VP of AI, a Head of Machine Learning — someone with the credentials to lead the function and signal organizational commitment. The hire is often necessary. It is not sufficient. The limitation is not the seniority of the hire. It is the assumption that AI capability is a function of a few experts at the top of a structure, when in practice AI delivery requires distributed capability across data engineering, software engineering, product management, and business functions. An excellent senior AI leader working with teams that lack data engineering depth, or with product managers who cannot translate business requirements into AI-ready specifications, will not solve the talent problem. What follows is a more granular account of where the talent gaps actually sit and what decisions the CTO and CHRO need to make to address them. The data engineering gap Of all the talent gaps in enterprise AI programs, the data engineering gap is the most consistently underestimated and the most consequential for delivery. AI models need clean, accessible, well-structured data. Producing that data at the quality and scale AI requires is data engineering work. It involves building and maintaining pipelines, managing schema consistency, handling data quality monitoring, implementing the access control infrastructure that AI systems need, and enabling the data freshness requirements that production AI applications demand. Most enterprise data engineering teams were built for business intelligence and analytics workflows: batch processing, monthly reports, data warehouse queries. These are different from what AI requires. AI applications often need lower latency, higher reliability, more granular access control, and better lineage documentation than analytics workflows do. The CTO who wants to deliver AI at scale needs a data engineering function capable of supporting it. That often means hiring, upskilling, and in some cases restructuring what the data function does — not just adding ML engineers to an existing team. The ML engineering versus data science distinction Organizations that are building their first production AI applications sometimes conflate data science — the role of developing models and validating their performance — with ML engineering, the role of deploying those models into production systems reliably and at scale. Both are necessary. They are not the same skill profile, and the market for each is different. Data scientists are relatively abundant at the senior level, because most organizations that have been investing in analytics have developed or hired them. ML engineers — people who can build model serving infrastructure, implement monitoring for production model performance, manage model versioning and rollback, and integrate AI components into existing software systems — are significantly scarcer. The consequence: organizations that have data science capability but limited ML engineering capability can develop models in research environments that never make it to production, or that make it to production but degrade gracefully without anyone noticing because the monitoring infrastructure does not exist. If the AI program's goal is production systems rather than research artifacts, the CTO needs to assess the ML engineering capacity specifically, not just the overall AI headcount. The product management capability gap AI products require a different kind of product management than traditional software products. The core difference: the behavior of an AI system is probabilistic, not deterministic. It does not do the same thing every time with the same inputs. It produces outputs that vary, that can be wrong in ways that are hard to predict, and that require different quality evaluation approaches than traditional software. Product managers who are excellent at defining functional requirements for traditional software often struggle with AI products because the tools for specifying and evaluating probabilistic behavior are different. Writing specifications for what an AI system should do, designing evaluation frameworks for outputs that are not right or wrong but better or worse, and building product intuition for what good AI performance looks like in a given context are skills that most PMs have not developed. The CHRO and CTO need to assess whether the organization's product management function has the capability to manage AI products effectively, and build a development plan for those who do not. This is a training and coaching question as much as a hiring question — the capability gap can often be closed more quickly through targeted development of existing PMs than through hiring. Business function AI capability The talent discussion in AI programs is usually focused on the AI team. The talent that is often more limiting in practice is the capability in the business functions that the AI program is serving. A demand forecasting AI system that produces excellent outputs is only valuable if the operations function can use those outputs to make better planning decisions. An AI-assisted underwriting tool only improves outcomes if underwriters can evaluate AI recommendations critically. An AI customer segmentation system only drives revenue if the marketing function knows how to act on the segments it produces. The capability gap in business functions shows up as underutilization: the AI system is deployed, adoption is technically measurable, but the organization is not capturing the value because the business users do not have the skills or the process changes required to convert AI outputs into better decisions. This is a CHRO problem more than a CTO problem. The CHRO needs to assess capability requirements in the business functions where AI is being deployed, build development programs that address specific skill gaps, and — where necessary — reassess role profiles to reflect the new capability expectations. The retention problem Building AI capability is hard. Retaining it is harder. The market for AI talent is competitive in ways that most enterprise organizations are not structured to compete with. The retention challenge is not purely about compensation, though that is a factor. It is about the work itself. AI engineers and data scientists who join an enterprise to build production AI systems stay when the work is technically interesting, when there is access to good data, when the organization moves fast enough to keep them engaged, and when there is a credible path to increasing impact. Enterprise AI programs that are slow to deploy, that are constrained by data access issues, or that cannot move past proof-of-concept into production create retention problems independent of compensation. The best AI talent leaves not because they were offered more money elsewhere, but because the organizational conditions did not support the work they wanted to do. The CTO's response to the retention problem is not primarily about retention packages — it is about building the organizational conditions that make the work worth staying for. That means clearing data access blockers, moving programs through proof-of-concept to production on a credible timeline, and giving AI teams genuine ownership over delivery. What to take from thisThe data engineering capability gap is more consequential for AI delivery than the data science or ML gap in most enterprises. Assess it specifically and address it before the program depends on it. ML engineering and data science are different roles with different skill profiles and different market availability. Both are required for production AI systems. Product management capability for probabilistic systems needs to be developed explicitly. Most PMs do not have it and it does not develop naturally through exposure alone. Business function capability to use AI outputs is a limiting factor that the CHRO needs to address. AI underutilization is usually a business function capability problem, not an AI system problem. Retention of AI talent depends more on organizational conditions than compensation. The CTO's retention strategy is about clearing blockers and maintaining momentum, not primarily about pay.

Read full article
The Organizational Change Nobody Plans for When AI Goes Into Production

The Organizational Change Nobody Plans for When AI Goes Into Production

Technical AI programs plan for model performance, infrastructure reliability, and user adoption. The change management plans in most AI programs cover communication, training, and rollout support. These are necessary. They are not sufficient. What does not make it into the program plan is the organizational change that AI deployment actually creates: changes in who makes decisions, where accountability sits, and how existing roles need to adapt. These changes are not side effects of the technical program — they are the substance of what AI deployment means for how the organization operates. And they tend to surface six to twelve months after go-live, in the form of confusion about accountability, conflict between roles that now overlap, and resistance from functions that feel their judgment has been displaced. Addressing these changes proactively requires treating AI deployment as an organizational design question, not just a technology one. The decision rights problem AI systems are good at making or informing decisions that humans previously made alone. When an AI system produces a recommendation — in credit assessment, in demand forecasting, in HR screening, in customer prioritization — the human who used to make that decision now has a different role. They are either ratifying the AI's recommendation, overriding it, or working alongside it in a way that requires a new kind of judgment. This changes the nature of the role without changing the job title or the org chart. The credit analyst who used to run the full assessment is now running exception management. The demand planner who used to construct forecasts is now reviewing and adjusting AI-generated ones. The recruiter who used to screen applications is now reviewing a pre-filtered shortlist. These are not simpler jobs. In some respects they are harder — they require a different kind of expertise, specifically the ability to evaluate AI outputs critically rather than produce analysis independently. The people in these roles may have been selected and developed for the original capability profile, not the new one. Organizations that do not address this explicitly produce two failure modes. People who struggle with the new role either resist the AI system — finding reasons not to use it, overriding it more than the data warrants — or over-defer to it, approving AI recommendations without the critical review that the accountability structure requires. The accountability gap When a decision goes wrong in a human-only process, the accountability is reasonably clear. When a decision informed by an AI recommendation goes wrong, the accountability is murkier, and organizations that have not thought about it in advance tend to discover this in a difficult situation. Was the decision wrong because the AI recommendation was wrong? Or because the human failed to apply appropriate judgment to a valid AI recommendation? Or because the system was deployed in a context it was not designed for? Or because the training data did not reflect the conditions that produced this specific case? None of these questions have clean answers in an organization that has not set up the accountability structure deliberately. The result is attribution conflict: the AI team points to the human decision-maker, the business function points to the AI system, and nobody has clear accountability for remediation. Defining accountability for AI-informed decisions before deployment is one of the most important organizational design questions in any AI program. It requires a clear statement of where human judgment is required, what escalation looks like when the AI recommendation is overridden, and what the process is for determining whether a decision-quality problem is an AI problem or a human judgment problem. The capability shift The capability an AI system replaces does not simply disappear from the organization's needs — it transforms. The expertise required to interpret and challenge AI outputs is often closely related to the expertise required to produce the underlying analysis manually. But they are not the same skill, and the transition period between the two is where the most significant organizational risk sits. In the immediate period after AI deployment, the organization typically has people who are capable of doing the work manually but are still learning to use the AI tool effectively. This is manageable. The medium-term risk is more significant: if the organization stops developing the underlying manual capability because AI handles it, and the AI system underperforms or becomes unavailable, the recovery is slower than anyone anticipated. This is not an argument against AI deployment. It is an argument for being deliberate about which capabilities the organization maintains independently of the AI system and which it allows to atrophy as AI performance becomes reliable. The role boundary conflicts When AI tools augment work across function boundaries — a customer-facing AI system that pulls from data owned by multiple departments, an AI planning tool used by both finance and operations — the organizational boundaries that existed for human work do not automatically translate. Who decides what data the AI system uses? Who is accountable if the AI produces an output that reflects poorly on a specific function? Who decides when the AI recommendation should be overridden? When the AI generates recommendations that one function disagrees with, what is the escalation path? These are organizational design questions dressed as AI governance questions. They arise because AI systems do not respect the lines between functions in the same way that human roles do. The AI pulls on data from wherever it can reach and produces outputs that may reflect or implicate multiple functions simultaneously. The organizations that handle this well have addressed it before go-live: clear data ownership, a governance structure for the AI system that is recognized by all the functions it touches, and escalation paths that do not depend on functional boundaries that the AI system has already made ambiguous. The human-in-the-loop question Most AI programs that include human review in their design treat it as a quality control mechanism: the human checks the AI output before it is acted on. This framing is correct but incomplete. The human in the loop is also, and more importantly, the locus of accountability for the decision. If the human review is a checkbox rather than a substantive check, the accountability protection is illusory — the organization has the appearance of human oversight without the substance. Regulators, clients, and courts are unlikely to accept "a human reviewed it" as a sufficient defense for a poor decision if the review process was not meaningful. What meaningful human oversight looks like — how long it should take, what the reviewer is expected to assess, what training they need, what recourse they have when they disagree with the AI — needs to be specified in the organizational design of the AI system, not left to individual judgment. What to take from thisMap how AI deployment changes existing decision rights before go-live. The people whose roles are affected need to understand the new expectation before the system is live, not after they have defaulted to the wrong behavior. Define accountability for AI-informed decisions explicitly. The accountability structure needs to specify where human judgment is required, what override looks like, and how decision-quality problems are attributed and remediated. Assess the capability implications of AI deployment over a three-to-five year horizon. Identify which capabilities the organization intends to maintain independently and which it is comfortable allowing AI to own. Address function boundary conflicts before deployment. The organizational design questions created by AI systems that cross functional lines need governance structures that all the affected functions recognize. Human-in-the-loop design needs to specify what meaningful review looks like, not just that review occurs. Checkbox oversight creates accountability risk, not accountability protection.

Read full article
AI Governance for Boards: What to Own and What to Delegate

AI Governance for Boards: What to Own and What to Delegate

There's a pattern I see in boardrooms that have added "AI strategy" to their agenda. An executive presents. The board listens. Someone asks a question that's technically about AI but actually about accountability. The executive answers with something about data governance and responsible use. The board nods. The item is closed. Nothing was actually governed. Boards are being asked to sign off on AI investments they can't fully interrogate, using governance frameworks that were designed for different kinds of risk. The result is a form of governance theater: the structures exist, the sign-offs happen, and the accountability is nowhere. This isn't a criticism of boards specifically. The frameworks genuinely don't fit. Audit committees are built around financial controls and statutory reporting. Risk committees are built around quantifiable risk exposures. AI introduces a risk profile that's different in kind — systems that make decisions at scale, that degrade silently over time, that can produce outcomes nobody explicitly designed, and that concentrate vendor dependencies in ways traditional procurement governance doesn't catch. Getting governance right doesn't require every board member to understand machine learning. It requires the board to own the right things and ask the right questions — and to know the difference between a real answer and a reassuring one. What the board needs to own Board-level AI governance has three genuine responsibilities. Everything else can and should sit with management. The first is the risk appetite. Not a list of approved use cases, but a real position on where the organization's tolerance for AI-driven decisions sits. What decisions can an AI make autonomously? What decisions require a human in the loop? What outcomes, if they occurred, would represent a failure of accountability at board level? These are governance questions, not technology questions. They need a board answer. The second is accountability structure. When an AI system produces a bad outcome — a biased recommendation, a pricing error at scale, a model that degrades and nobody notices for six months — who is accountable? The answer should never be "the model." It should be a named person in a named role with a documented process for how failures get escalated. The board should know what that structure is and should have satisfied itself that it's real, not just written down somewhere. The third is vendor concentration risk. Most enterprise AI programs now run on infrastructure from a small number of large providers. The board needs visibility into those dependencies — not at the technical level, but at the risk level. What happens to business continuity if a vendor relationship breaks? What proprietary data is in the hands of external providers, and under what terms? Everything else — model selection decisions, specific use cases, technical evaluation, operational monitoring — belongs with management and the relevant technical functions. The governance trap The trap boards fall into is trying to govern AI the way they govern everything else: by approving a strategy and reviewing a report. AI doesn't work that way. A strategy document approved eighteen months ago may bear no resemblance to what's actually in production today. Models evolve. Use cases expand beyond their original scope. The risk profile of a system that started as a recommendation tool changes when it starts making operational decisions at volume. Good AI governance requires a living understanding of what the organization is actually running, not just what it approved. That means the board needs reporting that tells it what AI systems are in production, what decisions those systems are making, and whether the performance monitoring is working — not just whether the program is "on track." Most board reporting on AI covers the program status, not the risk status. Those are different documents. 7 questions that matter These aren't technical questions. They're governance questions. A board member should be able to ask them in plain language and expect a plain-language answer. What decisions is AI making on behalf of this organization, and at what volume? Not what AI capabilities we have — what decisions it's actually making. If the answer requires a thirty-minute technical explanation, the governance reporting isn't working. Who is accountable when an AI system produces a wrong or harmful output? There should be a named person, not a process or a committee. What are we monitoring, and what triggers a review or a pause? Every production AI system should have defined performance thresholds. The board should know what those are and who owns the response when they're breached. What data are we using to train and run these systems, and do we have the rights to use it that way? Data licensing and privacy compliance create real legal exposure. This is a board-level question dressed as a technical one. Which external providers have access to proprietary or customer data, and under what terms? Vendor risk is real and underdisclosed in most AI reporting. How would we know if an AI system was producing discriminatory outcomes? The answer should describe a monitoring process, not a policy statement. What would we do if we had to take a system offline? Business continuity for AI systems is frequently underdeveloped. The board should be confident an answer exists. What good reporting looks like Board AI reporting that answers these questions would include: a register of AI systems in production and what decisions they're influencing, a summary of monitoring status and recent performance alerts, an update on data licensing and vendor contract status, and a brief note on any material changes to the risk profile since the last review. What boards typically receive: a slide on the AI strategy roadmap, a progress update against implementation milestones, and a chart showing the projected ROI. Those are different conversations. The strategy and roadmap conversation is important. So is the governance one. Both need time on the agenda, and conflating them is how organizations end up with AI programs that are well-funded and under-governed.

Read full article