How do we know it is safe and working?

The evidence a procurement review asks for, produced by the system as it runs.

Evaluations, observability, security review and AI governance: what was asked, retrieved, answered, spent and approved — recorded as it happens rather than reconstructed later.

What it is

We make an AI system inspectable. Evaluation sets run on every change, so a prompt edit cannot quietly make answers worse; traces record what was asked, what was retrieved, what was answered and what it cost; tenant isolation is proved by a policy test per table. The result is an evidence trail that auditors and procurement teams accept, produced by the system as it runs instead of assembled the week before the review.

Who it is for

If none of these sound like your situation, this is probably the wrong arc — and we would rather tell you that than sell you a discovery.

  • Your largest prospect sent a 90-question security review and it is sitting in someone's inbox.
  • An agent has been live for a month and nobody can say what it told a customer last Tuesday.
  • A regulator can ask you to show the basis of a recommendation, and the basis is a chat log.
  • Someone changed a prompt and nobody knows whether the answers got better or worse.

How the work is shaped

Repeatable shapes, not bespoke proposals. Each one has been run before and has a duration we hold to.

  1. 01Eval Harness · 3 weeksA golden set per agent, scored on every pull request, with the failure cases written by the people who own the process.
  2. 02Observability & Guardrails · 5 weeksTracing, cost ceilings, approval queues and alerting on the three failures that matter: a wrong answer, no answer, and runaway spend.
  3. 03Governance Programme · 8–10 weeksIsolation tests per table, an access review, a subprocessor register, a DPA pack, and an AI usage policy your team can actually follow.

What it costs

Published, not gated. The assumptions column is the part that matters — a price without them is a guess you discover was wrong in week three.

Published price bands and the assumptions behind them
BandFromWhat it buysTypical durationAssumes
Eval Harness₹3,50,000Golden question sets per agent, scored in CI, with the pass threshold published.3 weeks
  • Your team supplies 30–50 real questions with agreed correct answers
  • Agents already run somewhere we can call them
  • One scoring rubric per agent, revised once
Observability & Guardrails₹9,50,000Traces, per-agent cost reporting, approval queues, and alerts on wrong answer, no answer and runaway spend.5 weeks
  • Systems already emit logs, or can be instrumented by us
  • One alerting destination — email, Slack or WhatsApp
  • Log retention and storage are billed to your account
Governance Programme₹19,00,000Isolation tests per table, access review, subprocessor register, DPA pack and AI usage policy.8–10 weeks
  • One product and one production environment in scope
  • Legal review of the DPA is by your counsel, not ours
  • Certification audit fees, if you pursue one, are yours

What sits inside it

Evals

A golden set per agent, scored on every change, with the threshold published rather than negotiated after a failure.

Observability

What was asked, retrieved, answered and spent — per request, per agent, per tenant.

Security

Tenant isolation proved by a policy test per table, plus access review and secret handling.

Compliance

DPA, subprocessor register and retention rules kept current, so a questionnaire becomes a lookup.

AI governance

Where a person must approve, what an agent may never commit, and the log that shows both held.

What it has produced

Every number here comes from work on this page. Follow it to the story and check it.

to answer a 90-question security review, from three weeks
4 hoursto answer a 90-question security review, from three weeksSee the story
of recommendations carrying a cited suitability basis
100%of recommendations carrying a cited suitability basisSee the story
prices, scopes or dates committed by an agent without human approval
0prices, scopes or dates committed by an agent without human approval
isolation test per table, run on every merge
1isolation test per table, run on every merge

Is Assure what you need?

Tell us what is slow and what it is costing. If the answer is a different arc, or no arc at all, we will say so before anyone writes a proposal.

Review your AI governance