AI tooling evaluation & selection
Turns "which AI tools?" from a debate into a repeatable evaluation. A weighted scorecard plus a time-boxed trial protocol — so a selection is evidence-backed and can be re-run when the market moves (which, for AI tooling, is monthly).
The scorecard
Score each candidate against weighted criteria (template: toolkit/templates/tooling-scorecard.md). Suggested weights — tune per decision:
| Criterion | What to check | Weight |
|---|---|---|
| Model/agent quality on our codebases | measured via the golden-task harness, not vendor demos | 25% |
| Enterprise controls | SSO, audit logs, data residency (UK/India), RBAC | 20% |
| Security & data terms | no-training on our data, retention, SOC 2 / ISO 27001 | 20% |
| Extensibility | MCP, hooks, headless/CI, scriptability | 15% |
| Cost | per-seat / per-token, at our expected volume | 10% |
| Licensing & support | UK + India availability, support SLAs | 10% |
The trial protocol
- Shortlist from the radar assess ring and market scan.
- Time-boxed trial — a fixed window, the same golden tasks for every tool, the same reviewers.
- Score on the harness results + the scorecard criteria.
- Decide and record — the winner enters the radar trial/adopt ring; the decision is an ADR.
Rules
- No data without terms. A tool that trains on our data or won't sign no-training terms is disqualified regardless of quality — insurance client data is in scope.
- Measured, not demoed. Vendor benchmarks don't count; the harness score on our tasks does.
- Re-runnable. When a provider ships a major update, re-run the harness and re-score — this playbook is also the machinery behind model change-management.
Standards referenced: NIST AI RMF, SOC 2, ISO/IEC 27001.