What AI Can Find Is Not What AI Should Trust

Building a Public Interest Assurance Layer for Agentic Discovery

Petri Mikael Autio
Diagram showing the public-interest trust framework as the point of intervention in the AI discovery layer

TL;DR: As AI agents evolve from answering queries to executing real-world public workflows, how do they decide which software tools, datasets and models to trust? This is an emerging nexus of power that belongs to the global commons. We must build a public-interest trust and assurance layer into emerging Agentic Resource Discovery (ARD) standards now to ensure AI selects the best tool for the job instead of just the loudest one.

From a Prompt to an Operational System

Suppose a climate-change induced flood hits a rural district hard in the Himalayas. A local government officer wants a response plan based on current satellite data and the best available rapid flood-modeling tools. How can they get to a good result quickly?

They have learned to use AI in their daily routines, so imagine next that they create a good prompt. An agent could clarify the problem, fetch a locally runnable open-source flood model, spin it up and run it and overlay the result with water and sanitation information from mWater. It would still need to surface uncertainties, check local suitability and leave consequential judgments with responsible people. But the result could be a reviewable plan with prioritized targets, produced in record time.

Or take district-wide chlorine monitoring in schools. A district officer asks: “How do I set this up?” An agent could then discover the Solstice MCP server, ingest regional data, design surveys and workflows from successful precedents, and propose accounts and permissions for approval. The district could have an operational system in hours or days while retaining responsibility for institutional decisions, data protection and adoption.

These scenarios are becoming feasible. Solstice Institute has led the way in aid with an MCP server and a growing set of skills that help agents work with operational data and workflows. But the scenarios depend on AIs knowing how to use existing platforms, specialized open-source tools and inert data repositories when the user makes their prompt and not settling for something subpar randomly found, or confidently coding a less-than-ideal solution on the spot.

As people become accustomed to having agents do their work, where does power sit?

Making Good Work Legible to AI

In the public forum we no longer just write for humans. If you have built a great tool, today’s imperative is to write its documentation so AI can discover and understand it. In my experience, if you maintain a platform, it is now more than ever better to overcommunicate than undercommunicate its capabilities. AIs can parse far more information than most prospective users will read.

An easy first step toward this kind of legibility is an llms.txt file pointing tools that support the emerging convention toward authoritative material on your site.1

The next step is to make tools themselves discoverable. Agentic Resource Discovery, or ARD, is an open specification for publishing, indexing and searching resources such as tools, skills, APIs, MCP servers and specialized agents.2 The current specification uses /.well-known/ard.json as the listing backbone. A profile there describes what a resource does, how it can be invoked and, through representative queries, the problems it is suited to solve. ARD is an open, Apache 2.0 specification, still at proposal stage, that grew out of the Linux Foundation AI Catalog work. Hugging Face’s Discover Tool and GitHub’s Agent Finder are working implementations; Google has said native ARD support is coming to its Agent Registry. More important than the manifest path are the entries themselves: two to five representative queries help registries decide when a resource should be returned.

I see many good tools announced at conferences, used during a funded project and then forgotten. Documentation, stable metadata and a machine-discoverable account of what a tool does and where it fails give useful work a better chance of being rediscovered. These are now the minimum level needed. But how do we collectively ensure that AI surfaces the right tools and the best available knowledge for the task at hand?

Discovery Is a Governance Question

Tech giants determine much of our lives now, and may shape this space as well. This is a reason to act in the public interest now. Semantic ranking will favor resources with abundant English-language documentation, familiar terminology and a large online footprint. Large vendors and GitHub-star-heavy projects can become easier to retrieve than quieter tools with stronger evidence, better local fit or years of field use.

ARD already provides a slot for a stronger answer. Its optional trust manifest can carry identity, provenance, signatures and attestations. But while ARD communicates the claim, it does not decide whether it checks out or how much weight it deserves. That is delegated to the declared trust framework, while registries and clients make the actual trust decision. That opens the space for us to act and define the public-interest trust framework that ARD already expects someone to supply. It should extend beyond corporate compliance to record field evidence, maintenance, licenses, data sovereignty, operating contexts, offline capability, known failure modes and security assessments. It should distinguish publisher claims from independent assurance and explain why a resource was recommended.

Security must be designed into this layer. A catalog that steers agents is immediately also an attack surface because tool descriptions can carry prompt injection, skills can be poisoned, and publisher identity alone says little about whether a resource is safe. Claims must be verified, provenance preserved and consequential selections logged.3

The institutional foundation may already be in place. The Solstice Institute maintains expanding libraries designed to turn shared knowledge into operational systems. Its Global Indicator Library combines definitions, references, calculations and tested question sets that users can bring directly into surveys. The platform’s environmental datasets, expert workflows, open-source integrations, and MCP skills also already support global development at scale.

The Solstice Institute could act as an early publisher of rich agentic profiles, rather than as the body that assesses its own claims. A useful first push could add ARD profiles to well-used resources from its libraries, including carefully written representative queries, data requirements, operating contexts and known limitations. An independent registrar or assurer should own the verification criteria and judgments.

The Digital Public Goods Alliance is relevant here for more than the reach of its registry. It already performs something close to an independent assurance role, with an eligibility check, technical review against nine published indicators and annual reassessment. Extending that machinery to agentic trust would require additional criteria, including evidence of field use, local suitability and operational performance. But it makes DPGA an obvious institution with which to explore the assurance side of this work.4

Solstice could occupy a complementary role. Its libraries translate indicators, standards and expert knowledge into surveys, data structures and operational workflows that governments can and do put directly to use. Publishing these resources through rich ARD profiles would test whether agents can discover not only individual tools, but proven building blocks for stronger public systems. Solstice could also curate fit-for-purpose collections of skills for governments, with transparent evidence about where each resource works, what it requires and where its limitations lie.

What Organizations Can Do Now

A few steps for giving good solutions a fair chance of being found when AI agents reach for tools:

The capability frontier of large language models will continue to advance. Whether they reliably find the best of our accumulated tools and knowledge for a given context is a separate question. The answer will depend, in no small part, on what we make visible to them now, and on who governs the systems through which they find it.

A serious first implementation would define the trust framework, publish and assess profiles for a manageable number of established resources, connect them to a working registry and test whether assurance changes what agents select. This is a bounded program of work, not a new institution.

The next step as I see it is to define a public-interest trust framework for ARD, apply it to established resources and test whether it changes what agents select. The discovery layer is emerging, and moving early can matter disproportionately. Whose evidence and values will shape this layer?

Petri Mikael Autio
Head of Product, Solstice Institute and mWater Foundation


  1. llms.txt community proposal.↩︎

  2. Agentic Resource Discovery specification.↩︎

  3. ARD identity, trust and verification.↩︎

  4. Digital Public Goods Alliance, DPG Standard, DPG Registry, and DPG Standard review process.↩︎

← All postsAI Leadership →