Skip to main content

AI tooling evaluation & selection

Turns "which AI tools?" from a debate into a repeatable evaluation. A weighted scorecard plus a time-boxed trial protocol — so a selection is evidence-backed and can be re-run when the market moves (which, for AI tooling, is monthly).

The scorecard

Score each candidate against weighted criteria (template: toolkit/templates/tooling-scorecard.md). Suggested weights — tune per decision:

CriterionWhat to checkWeight
Model/agent quality on our codebasesmeasured via the golden-task harness, not vendor demos25%
Enterprise controlsSSO, audit logs, data residency (UK/India), RBAC20%
Security & data termsno-training on our data, retention, SOC 2 / ISO 2700120%
ExtensibilityMCP, hooks, headless/CI, scriptability15%
Costper-seat / per-token, at our expected volume10%
Licensing & supportUK + India availability, support SLAs10%

The trial protocol

  1. Shortlist from the radar assess ring and market scan.
  2. Time-boxed trial — a fixed window, the same golden tasks for every tool, the same reviewers.
  3. Score on the harness results + the scorecard criteria.
  4. Decide and record — the winner enters the radar trial/adopt ring; the decision is an ADR.

Rules

  • No data without terms. A tool that trains on our data or won't sign no-training terms is disqualified regardless of quality — insurance client data is in scope.
  • Measured, not demoed. Vendor benchmarks don't count; the harness score on our tasks does.
  • Re-runnable. When a provider ships a major update, re-run the harness and re-score — this playbook is also the machinery behind model change-management.

Standards referenced: NIST AI RMF, SOC 2, ISO/IEC 27001.