How do we know it is safe and working?
The evidence a procurement review asks for, produced by the system as it runs.
Evaluations, observability, security review and AI governance: what was asked, retrieved, answered, spent and approved — recorded as it happens rather than reconstructed later.
What it is
We make an AI system inspectable. Evaluation sets run on every change, so a prompt edit cannot quietly make answers worse; traces record what was asked, what was retrieved, what was answered and what it cost; tenant isolation is proved by a policy test per table. The result is an evidence trail that auditors and procurement teams accept, produced by the system as it runs instead of assembled the week before the review.
Who it is for
If none of these sound like your situation, this is probably the wrong arc — and we would rather tell you that than sell you a discovery.
- Your largest prospect sent a 90-question security review and it is sitting in someone's inbox.
- An agent has been live for a month and nobody can say what it told a customer last Tuesday.
- A regulator can ask you to show the basis of a recommendation, and the basis is a chat log.
- Someone changed a prompt and nobody knows whether the answers got better or worse.
How the work is shaped
Repeatable shapes, not bespoke proposals. Each one has been run before and has a duration we hold to.
- 01Eval Harness · 3 weeksA golden set per agent, scored on every pull request, with the failure cases written by the people who own the process.
- 02Observability & Guardrails · 5 weeksTracing, cost ceilings, approval queues and alerting on the three failures that matter: a wrong answer, no answer, and runaway spend.
- 03Governance Programme · 8–10 weeksIsolation tests per table, an access review, a subprocessor register, a DPA pack, and an AI usage policy your team can actually follow.
What it costs
Published, not gated. The assumptions column is the part that matters — a price without them is a guess you discover was wrong in week three.
| Band | From | What it buys | Typical duration | Assumes |
|---|---|---|---|---|
| Eval Harness | ₹3,50,000 | Golden question sets per agent, scored in CI, with the pass threshold published. | 3 weeks |
|
| Observability & Guardrails | ₹9,50,000 | Traces, per-agent cost reporting, approval queues, and alerts on wrong answer, no answer and runaway spend. | 5 weeks |
|
| Governance Programme | ₹19,00,000 | Isolation tests per table, access review, subprocessor register, DPA pack and AI usage policy. | 8–10 weeks |
|
What sits inside it
Evals
A golden set per agent, scored on every change, with the threshold published rather than negotiated after a failure.
Observability
What was asked, retrieved, answered and spent — per request, per agent, per tenant.
Security
Tenant isolation proved by a policy test per table, plus access review and secret handling.
Compliance
DPA, subprocessor register and retention rules kept current, so a questionnaire becomes a lookup.
AI governance
Where a person must approve, what an agent may never commit, and the log that shows both held.
What it has produced
Every number here comes from work on this page. Follow it to the story and check it.
- to answer a 90-question security review, from three weeks
- 4 hoursto answer a 90-question security review, from three weeksSee the story
- of recommendations carrying a cited suitability basis
- 100%of recommendations carrying a cited suitability basisSee the story
- prices, scopes or dates committed by an agent without human approval
- 0prices, scopes or dates committed by an agent without human approval
- isolation test per table, run on every merge
- 1isolation test per table, run on every merge
The work behind it
A security review answered in an afternoon
An accountancy practice stopped losing three weeks per client to evidence-gathering by making the trail a by-product of the work.
4 hours target time to answer a 90-question security review
55 staffEvery recommendation carrying its reasoning
An advice firm made the suitability basis a by-product of giving advice rather than a document written afterwards.
100% of recommendations carrying a cited suitability basis
Is Assure what you need?
Tell us what is slow and what it is costing. If the answer is a different arc, or no arc at all, we will say so before anyone writes a proposal.
Review your AI governance