Executive judgment

Most public-sector AI programmes begin in the wrong place. They begin with a technology demonstration, a general call for use cases or a procurement exercise. The institution then attempts to attach governance, data access, legal review and operational ownership to the resulting pilot. This sequence produces familiar outcomes: prototypes without production data, tools without accountable owners, approvals that take longer than development, and “successful” experiments that never alter a public service.

An AI-native institution starts from the function. It identifies a decision or service whose public purpose is clear, maps the governing law and actual workflow, separates routine judgment from exceptional judgment, and redesigns the end-to-end process around what machines and people can each do well. Technology selection follows institutional design. Assurance is built into the delivery system and scaled to consequence.

Zwarte Peper’s position: public institutions should pursue rapid, portfolio-wide AI adoption. Low-risk internal and assistive uses should move under standing rules rather than repeated permission. Systems that materially affect rights, safety or access to essential services require explicit legal authority, contextual testing, traceable decisions, effective human escalation and remedy. The relevant standard is not “human in the loop” by slogan; it is whether responsibility, control and correction are real.

This model is more ambitious than cautious pilot culture and more disciplined than indiscriminate automation. It treats delay as a public cost, legacy process as a risk source and organisational learning as a capability to be designed.

1. The unit of transformation is the institution, not the tool

Generative AI can draft, summarise, classify, retrieve, translate, code and support reasoning. Predictive systems can detect patterns, prioritise inspection and estimate demand. Yet an isolated model cannot determine which public objective has legal priority, which record is authoritative, which officer may depart from a rule or what remedy follows an error. Those are institutional questions.

The OECD’s 2026 Digital Government Outlook reports AI use in at least one area of government in 35 of 36 surveyed OECD countries, with the strongest adoption in internal processes and public services. It also finds that the foundations required for scale—interoperable data, digital public infrastructure, adaptive procurement, skills and organisational capacity—remain uneven.1 Adoption is therefore no longer primarily a problem of awareness. It is a problem of conversion: turning local technical capability into repeatable institutional performance.

Three shifts define that conversion:

  • From task assistance to service redesign. Saving ten minutes on drafting is useful. Removing avoidable hand-offs, repeated evidence and unowned queues from a licensing or benefits journey is transformational.
  • From project approval to standing governance. Requiring a new committee for every low-risk experiment slows learning. A reusable classification and control model lets ordinary uses proceed and directs scarce scrutiny to consequential ones.
  • From model accuracy to public outcomes. A technically accurate classifier can still worsen a service if it creates unmanageable false-positive queues, hides reasons or shifts work to citizens. Measures must include elapsed time, access, error, cost, remedy and distribution.

2. Build an automation portfolio around public value

A use-case inventory should rank opportunities by public value, feasibility and consequence. “AI potential” alone is not a priority. The first wave should combine meaningful scale with recoverable error and available data.

Use classExamplesDefault path
Internal productivityDrafting, summarisation, meeting records, coding assistance and knowledge retrieval.Standing approval for authorised tools, data rules, verification and professional accountability.
Operational triageDocument classification, routing, duplicate detection, inspection prioritisation and demand forecasting.Controlled production pilot with sampling, drift measures and accountable override.
Public interactionMultilingual guidance, service navigation, application support and status explanation.Clear disclosure, source-grounded answers, accessibility testing and rapid human escalation.
Decision supportEligibility analysis, risk scoring, resource allocation and professional recommendations.Legal and impact assessment, contextual validation, reason capture, monitoring and contestability.
Automated adverse decisionDenial, sanction, enforcement or withdrawal of an essential service without prior human judgment.Exceptional path; require clear authority, necessity, stringent evidence and effective remedy.

Portfolio governance prevents two errors. One is to begin only with trivial tools and declare success without affecting institutional performance. The other is to target the most politically visible high-risk decision before data, infrastructure and assurance exist. A balanced portfolio creates rapid value while building the evidence and capability required for more consequential uses.

4. Data and digital public infrastructure are the scaling layer

AI cannot compensate for missing identity, unreliable registers, inaccessible records or systems that cannot exchange authorised data. Retrieval over poor documents produces fluent confusion. Predictive models trained on operational artifacts reproduce the history of the process, including its exclusions. Before model work, teams need to establish provenance, authority, quality, permissible use, retention and correction.

The World Bank’s 2025 GovTech Maturity Index assesses 197 economies across core government systems, online service delivery, digital engagement and enabling institutions.3 Its architecture is instructive: AI maturity is inseparable from the transaction systems, shared platforms, standards and leadership on which digital government already depends.

Public institutions should build reusable components: identity and authorisation, consent or legal-basis records where appropriate, authoritative registers, event logs, notification, payment, case management, model gateways and evaluation services. Common infrastructure should remain modular and interoperable. A central platform that every agency must use can reduce duplication; it can also become a bottleneck or concentrated operational risk if agencies cannot substitute components or export their data.

5. A five-layer AI operating model

LayerInstitutional capabilityCore artefact
MandateDefine the public objective, authority, affected interests and accountable decision owner.Decision charter and legal map.
ServiceRedesign the user journey, workflow, exception path and remedy before selecting technology.Target operating model and service measures.
Data and technologyProvide governed data, secure model access, integration, evaluation and observability.System and data card with evaluation record.
AssuranceClassify impact, test controls, record approval and monitor incidents and drift.Living assurance case.
LearningCompare outcomes with the prior process, publish evidence and revise or stop.Benefits ledger and periodic review.

The UK Government’s AI Playbook provides ten principles covering lawful and responsible use, security, meaningful human control, lifecycle management, procurement, skills and assurance.4 Its practical strength is that these concerns are treated as delivery disciplines. Institutions should go one step further and make the disciplines executable: standard clauses, evidence templates, test harnesses, approved environments and risk-based delegated authority.

6. Procure outcomes and portability, not dependency

Traditional technology procurement assumes requirements can be specified in advance and acceptance occurs at delivery. AI systems change with model versions, data, prompts, retrieval sources and user behaviour. Contracts must govern a service over time.

A serious procurement should require benchmark performance in the institution’s context; documented limitations; security and incident cooperation; notice of material model changes; access to logs and evaluation evidence; data-use restrictions; export of records, prompts and configurations; transition support; and pricing that remains intelligible as volumes scale. Intellectual-property claims should not prevent the institution from understanding the decision process or moving supplier.

Competition requires architectural choices as well as clauses. A model gateway, open interfaces and separable retrieval, orchestration and user layers can allow components to be replaced. Framework agreements should preserve routes for smaller specialists and open-source solutions. “Sovereignty” should mean effective control over continuity, data, switching and essential capability—not necessarily public ownership of every server or model.

7. Replace compliance queues with continuous assurance

A central review board cannot manually approve AI at institutional scale. It should set the taxonomy, minimum evidence and escalation thresholds, then delegate ordinary decisions to trained service owners. High-impact uses receive independent challenge; low-risk uses proceed under standing controls and audit.

Controls should attach to failure modes. Hallucination requires grounding, constrained sources and verification. Discrimination requires subgroup evaluation and investigation of the underlying service. Privacy risk requires minimisation, access controls and retention. Automation bias requires workflow and interface tests. Cyber risk requires threat modelling, supply-chain controls and monitoring. A generic ethics checklist cannot demonstrate that any of these risks is controlled.

Assurance is continuous because the system and its environment change. Release gates should be followed by outcome monitoring, complaint analysis, incident reporting, red-team exercises and defined stop conditions. The objective is a shorter path from observed failure to correction—not a claim that deployment is failure-free.

8. Redesign work and preserve professional judgment

AI-native government should reduce administrative labour, not merely add output for officials to check. Every deployment should identify which tasks disappear, which are augmented, which become more important and which new responsibilities arise. Time released must be visible in workforce and service planning; otherwise productivity becomes additional activity rather than capacity.

Expertise remains essential. Experienced officials know where formal procedure diverges from reality, which facts are material and when an apparently ordinary case is exceptional. Their knowledge should shape rules, evaluations and escalation. At the same time, professional accountability cannot become a reason to preserve every manual step. Institutions should measure whether human review changes outcomes and place it where it does.

9. A 180-day route from pilots to institutional capability

Days 0–30: establish mandate and portfolio

Name an executive owner; approve the risk taxonomy and standing rules; select three service journeys and two internal functions; baseline cost, time, error and user experience; inventory lawful data and existing contracts.

Days 31–90: build reusable foundations

Deploy an authorised model gateway; create evaluation and logging services; adopt standard procurement clauses; train product, legal, security and operational leads together; document the target operating models and stop conditions.

Days 91–140: deploy into live work

Move beyond demonstrations. Introduce tools to bounded production cohorts, compare against the existing process, monitor exceptions and complaints, and publish an internal benefits ledger. Senior leadership should resolve policy and data blockers weekly.

Days 141–180: scale, revise or stop

Scale uses that improve final outcomes; redesign those that only accelerate one step; stop those whose control cost exceeds public value. Transfer evidence and components to the next service portfolio and publish an external transparency record for consequential deployments.

10. The executive scorecard

  • Public outcome: service time, access, quality, reliability and distribution compared with the prior process.
  • Productivity: total cost per resolved case and hours genuinely released, including assurance and correction.
  • Legality and remedy: valid authority, reason quality, challenge volume, reversals and time to correction.
  • Operational resilience: incidents, recovery, drift, supplier dependence and ability to continue or switch.
  • Adoption and capability: sustained use, professional competence, reuse of common components and time from idea to controlled production.

Conclusion: govern for deployment

Public institutions face a choice between governing AI as a sequence of exceptions and governing it as an ordinary institutional capability. The first path feels cautious but often preserves unmanaged experimentation, fragmented procurement and legacy failure. The second establishes clear authority, reusable controls and outcome evidence so that adoption can move faster.

AI-native government is not government by machine. It is government that understands which judgments belong in law, which tasks should be automated, which exceptions require expertise and how error becomes visible and correctable. The ambition should be substantial: fewer administrative queues, better decisions, more accessible services and public institutions able to learn at technological speed.

Authorities and selected research

  1. OECD, Digital Government Outlook 2026, “Adopting and governing AI in government”; OECD, Governing with Artificial Intelligence, 2025.
  2. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence, especially Articles 26, 27 and 49.
  3. World Bank, GovTech Maturity Index 2025: Tracking Public Sector Digital Transformation Worldwide.
  4. UK Government Digital Service, Artificial Intelligence Playbook for the UK Government, 2025; AI Insights series, updated 2026.
  5. US National Institute of Standards and Technology, AI Risk Management Framework.
  6. Government of India, India AI Governance Guidelines, November 2025.

Editorial note. This article is general policy research and does not constitute legal or procurement advice. Sources and regulatory status were last checked on 1 August 2026.