Short checklists, model cards for AI compliance
Security teams send 300-question AI vendor surveys that produce prose, not artifacts. A columnist proposes short, artifact-backed checklists and a standardized model card.
A security columnist warned that long, prose-driven AI vendor questionnaires produce essays instead of verifiable evidence and do not scale with risk. The columnist proposed shorter checklists tied to artifacts and a standardized model card as a way to make assessments more usable.
The columnist cited a 2009 study led by surgeon Atul Gawande, backed by the World Health Organization, in which a 19-item surgical checklist reduced complications and deaths across eight hospitals. The piece noted that enforcement for the EU AI Act’s general-purpose AI provisions begins in August, high-risk obligations are phasing in, and ISO/IEC 42001 is appearing in third-party risk questionnaires.
The columnist described the current landscape of frameworks and standards, noting that NIST’s AI Risk Management Framework is widely used in North America and that organizations also reference the OECD Principles, HITRUST work, sector regulators such as the FDA, and multiple state laws. Published crosswalks show overlap between ISO 42001, NIST AI RMF, and the EU AI Act, and the columnist wrote that a single, well-designed program can address multiple frameworks.
The columnist identified three recurring problems with vendor questionnaires. First, many questions request free-text descriptions that cannot be validated against artifacts. Second, questions often treat model behavior as a fixed property even though large language models change with versions, prompts, or configuration, which makes point-in-time attestations quickly out of date. Third, questionnaires frequently fail to scale with risk, applying the same long addendum to low-risk chatbots and to systems used for clinical decision support.
To address those issues, the columnist outlined five tests that each question should pass before inclusion: the answer must be supportable by an artifact such as a log, configuration file, evaluation report, data flow diagram, or architecture document; the question should be tied to a defined risk tier; it should be measurable or binary where possible; it should be decision-relevant so a bad answer would change procurement; and controls should be mapped once and reused across regulatory frameworks.
The columnist sketched a condensed checklist that could fit on a single card. Core items include where the model is deployed and who provides upstream models; what data flows in and out and where those flows are logged; where copies of data are stored, retention periods, and access controls; which inference parameters are logged per request and whether incidents can be replayed; the evaluation suite run against production and its cadence; human oversight points and their scope; incident response processes; and procedures for changing models, including testing gates and customer notifications.
The piece proposed a standardized model card using a fixed schema covering model version and lineage, categories of training data provenance, data retention and storage details, a core set of evaluation benchmarks with published scores, documented safety mitigations, default inference parameters, and a change log linked to model updates. The columnist added that a common model-card schema could answer many repetitive assessment questions and reduce duplicate documentation burdens.
The columnist summarized the approach with two statements in the piece: “a framework is only as good as its worst question” and “compliance is an evidence problem, and for AI, evidence is an observability problem.”








