Sovereign AI Needs a Public Scorecard
National model programmes should be measured by capability, cost, language, control and continuity—not the number of launches.
States have begun to count artificial-intelligence models the way an earlier age counted factories. The number is visible, politically legible and easy to announce.
It does not establish sovereignty.
A national model may depend on foreign accelerators, cloud regions, software libraries and licences that can change. A domestically owned system may be too expensive or inaccurate for public use. An imported model may deliver excellent service while leaving a ministry unable to continue if the supplier withdraws.
Sovereign AI is not an origin label. It is the practical ability to choose, operate, inspect and replace the systems on which important institutions depend. Public investment should be judged through that definition.
The stack matters
Countries do not need to manufacture every chip, write every library and train every model alone. Autarky would waste resources and isolate research. Sovereignty means avoiding a dependency so concentrated that another actor can remove a critical capability without a credible substitute.
That requires visibility across compute, data rights, model access, evaluation, skills and deployment. Governments should know which components can be moved, which contracts can be terminated, how long substitution would take and what public service would fail in the meantime.
Partnership can increase sovereignty when it creates options. Domestic ownership can reduce it when performance is protected from scrutiny.
Measure public value
A sovereign-AI scorecard should begin with capability in the populations and tasks the programme claims to serve. Language results must be reported separately rather than hidden in a national average. Tests should cover dialects, code-switching, scripts and institutional vocabulary, with speakers from the relevant communities involved in evaluation.
Benchmarks should include real public problems. Can a model explain an official notice without altering its legal meaning? Does it retrieve the correct source? Can it recognise uncertainty in a health or benefits question? Does it preserve performance when documents are noisy and language is informal?
Next comes economics. Publish serving cost, latency, hardware needs and energy demand for representative tasks. A small model that runs affordably inside a regional public system may add more operational independence than a larger model that requires scarce infrastructure for every response.
Control must also be explicit. Are the weights available, and under what licence? Can an institution deploy the model in its own environment? Who can modify access terms? What is the continuity plan if the developer fails? An API can be convenient without being a sovereign asset.
Finally, report safety and failure. Public tests should cover data leakage, cyber misuse, discriminatory performance and adversarial instructions. Serious limitations should remain attached to the model as it changes versions.
The scorecard must resist gaming
Benchmarks can create their own illusion. Developers train against known tests, publish favourable averages and exclude invalid answers. The Stanford AI Index 2026 notes that invalid-response rates on some frontier evaluations reached 42 percent and that open- and closed-weight performance gaps vary substantially by task.
Independent evaluators should maintain hidden test sets and refresh them regularly. Results need sample sizes, uncertainty ranges and clear treatment of refusals and invalid responses. Models should be retested after significant updates and in the conditions where public agencies intend to use them.
The purpose is not a single national league table. Different systems may be appropriate for translation, research, administration or offline deployment. A scorecard helps buyers see the trade-offs and helps citizens see what public support purchased.
The strongest countercase
Young companies and research teams cannot disclose every training detail without surrendering commercial advantage. Premature rankings may discourage difficult work and reward developers that optimise for safe tests. A country also needs viable firms, not only transparent experiments.
Publication can be proportional. Proprietary recipes, sensitive datasets and security details may remain protected. But a project receiving public compute, data or grants owes the public reproducible performance, material limitations, licensing conditions and a clear account of support. Secrecy can protect an invention; it cannot prove the promised capability exists.
Failure should not automatically end support. Honest negative results teach a national programme where the constraints are. Milestones should reward evidence and correction rather than announcements that leave no room to admit a mistake.
Fund choices, not symbols
Governments should tie later funding to demonstrated language coverage, service-cost targets, independent testing and deployment beyond the developer's own laboratory. They should publish portfolio-level dependencies on accelerator vendors, clouds and foreign licences. That inventory will reveal which investment would create the most strategic choice next.
Sovereign programmes can build shared compute, lawful datasets and evaluation institutions that individual firms would not create. These are valuable public goods. Their success is visible when more institutions can choose reliable systems on sustainable terms—not when a minister reaches a larger announcement number.
Count models as projects. Measure sovereignty as capacity. The public scorecard is where the difference becomes impossible to hide.
The Global Federation defines sovereignty as durable democratic choice: the capacity to cooperate without becoming helpless when a partner, vendor or platform changes the terms.