How to evaluate enterprise AI
RFP questions, TCO, build vs buy, copilot vs work OS, and how to score an agent harness without buying a demo.
- Evaluation
How to Choose Between a Coding Harness and an Enterprise Harness
A coding harness runs a repository — Claude Code, Cursor, Codex. An enterprise harness runs company jobs with connectors and signers. Most organisations need both; they are not substitutes.
Read article - Evaluation
How to Evaluate an Agent Harness
Evaluating an agent harness means checking whether it can stop a write, replay who signed, swap the model without rewriting tools, and fail a real sensor — not whether the demo answered a question.
Read article - Evaluation
How to Evaluate Collaborative AI
An RFP sheet for shared jobs — not copilots. Questions on roster, stored rejection, second department, files in one place, and what survives after the session. Plain language for operators.
Read article - Evaluation
How to Evaluate Loop Engineering
Evaluating loop engineering means asking whether standing orders can skip quietly, record outcomes on a run page, parameterise a write without a model, pause, notify the roster, and reuse a recipe without copying the old Slack channel.
Read article - Evaluation
What are auditors asking for around AI?
Who decided, did the model write unchecked, and which rulebook applies. How to prepare a first evidence pack this quarter without a huge project.
Read article - Evaluation
What to Look for in Model Routing
Model routing is a policy that uses a cheaper model for simple steps and a stronger model only when the task needs it — not a dropdown labelled “best.”
Read article
See what governed AI looks like on your stack.
Connect your tools, run a workstream, and keep every decision on your ledger. Start on Free.