Risk and compliance teams are not asking whether the agent is good. They are asking whether the decision to ship it is defensible. That calls for different artifacts than an engineering demo: a written statement of what the system must and must not do, reproducible results mapped to the framework they work in, evidence traceable to specific runs, an explicit account of what was not tested, and a clear line between what you measured and what remains their governance responsibility.
A reviewer working to the NIST AI Risk Management Framework has to file your evidence against GOVERN, MAP, MEASURE, and MANAGE. If your report doesn't speak that language, they translate it themselves and lose confidence in the process. The same is true for the CSA Agentic AI profile and for internal model-risk standards at banks and insurers.
Handing over a per-run appendix already mapped to the framework converts a long meeting into a review. It is also the moment to be honest about coverage: supported, partial, roadmap, out of scope — stated as such.
There is no "NIST compliance" certification for an AI system, and governance isn't something a tool can discharge on a client's behalf. A vendor who blurs that line hands the client a problem they'll find during their next audit.
The honest posture is a division of labour: we generate the measurement evidence, you govern. Saying so directly tends to increase trust with the exact people whose sign-off you need — they've usually heard the other version and know what it's worth.
For regulated buyers, "who holds our data" often gets asked before "how accurate is it." An evaluation that runs inside the client's environment, with no vendor system holding their content, removes an entire category of questions — vendor risk assessment, data residency, and for federal work the scope of what a prime has to assess.
Present that as architecture rather than policy. It's checkable, which a promise isn't.
The short answer above is ours, and we've written at length about how we got to it: