Reference — generated from the toolkit
Source of truth: toolkit/templates/tooling-scorecard.md. Edit it there; this page is regenerated on build.
Tooling evaluation scorecard — {tool name}
- Candidate: … Category: IDE / agent / platform / model
- Evaluator(s): … Trial window: … Date: {YYYY-MM-DD}
| Criterion | Weight | Score (1–5) | Weighted | Evidence |
|---|---|---|---|---|
| Model/agent quality on our codebases (golden-task harness) | 0.25 | link to harness run | ||
| Enterprise controls (SSO, audit, residency, RBAC) | 0.20 | |||
| Security & data terms (no-training, retention, certs) | 0.20 | |||
| Extensibility (MCP, hooks, headless/CI) | 0.15 | |||
| Cost at expected volume | 0.10 | |||
| Licensing & support (UK + India, SLAs) | 0.10 | |||
| Weighted total | 1.00 | /5 |
Disqualifiers (any one = out, regardless of score)
- Trains on our data / won't sign no-training terms
- No data residency option for UK client data
- No audit logging
Recommendation
Adopt / Trial / Assess / Reject — and the one-line reason. Record the decision as an ADR.