Template · after a paying Base exists

Private-eval template

After you pay for Base, run these named steps on your own workflows. We will not publish your scores or invent anyone else’s.

No public lift %No logosPaying Base first

What this page is

  • A public template that names the eval steps
  • For use after a paying Base exists
  • Private measurement on your facts under department tenancy
  • Scores stay in your tenant; this page is the contract, not a leaderboard.

What this page is not

  • A place that publishes lift percentages
  • A customer logo wall or named-reference roster
  • Invented case-study ROI
  • A staffed eval before Base billing is real

No public lift percentages. No customer logos. No invented names or case-study ROI. Measurement stays on your facts under department tenancy.

How the loop runs

Paying Base → Needle → Live facts → Named signals → Private score → Scale

  1. Paying Base
  2. Needle
  3. Live facts
  4. Named signals
  5. Private score
  6. Scale

Pay for Base first. Pick a workflow needle on your data. Land signed events. Name what you will count. Score in your tenant. Scale inference only if the needle moved.

Words on this page

Paying Base
A live workspace you pay for, with a connector and signed events on department-scoped products — not empty streams.
Needle
One workflow outcome you already run (for example incident evidence recall). Yours — not a vendor leaderboard.
dept.* products
Named department streams on the mesh (ops heartbeats under your tenancy). Agents cite those events.
Compound beat
Grow usage only after the needle moves on your private runs. If it does not, fix context and signals before spend.

Eval steps

A team can run this sequence on their own data. The names are the contract. The scores are not published here.

  1. Step 1 of 7

    Confirm Base

    Paying workspace, connector, and signed events before you measure.

    You already have a paying Base: a workspace, a connector, and signed events on department-scoped products. If that is not true yet, create a workspace first. Do not run this template on empty streams.

  2. Step 2 of 7

    Pick a needle

    One workflow outcome on your data, not a vendor bench.

    Choose one workflow outcome you already run — incident evidence recall, account-health narrative, governed metric freshness. The needle is yours. Do not import a vendor leaderboard.

  3. Step 3 of 7

    Land facts

    Signed dept.* heartbeats from tools you already run.

    Point the tools you already run at the mesh. Heartbeats publish as dept.* products under your tenancy. Agents cite those events, not a chat window.

  4. Step 4 of 7

    Name signals

    Leading and lagging blanks you will count — no public lift %.

    Write down what you will count. Leading: time to linked evidence, context attach. Lagging: repeat-incident rediscovery, renewal-judgment quality. Names of signals — not a public lift figure.

  5. Step 5 of 7

    Run the eval

    Private score on your facts in your tenant.

    Score the needle on your facts, in your tenant. Compare recall or workflow completion against the signals you named. Keep the run private.

  6. Step 6 of 7

    Keep score private

    No published lift %, logos, or invented names.

    Do not publish lift percentages, customer names, or logos from this run. Residual-honest measurement stays on your data.

  7. Step 7 of 7

    Scale after proof

    Grow meters only if the needle moved; else fix context first.

    Grow usage meters when the needle moves on your facts. If it does not, fix context and signals before spend. Proof before inference scale is the Compound beat.

Before you scale inference

After the run, pick one private outcome. Scores stay in your tenant.

Scale inference

Needle moved on your facts in private runs. Context and signals held. Grow usage meters.

Fix context first

Needle flat or noisy. Repair signals, connectors, or dept.* coverage before spend.

Do not scale

Proof missing — empty streams, no paying Base, or scores you would not trust in your tenant. Create a workspace or finish Base first.

No public lift percentage. Scores stay in your tenant.

Private-eval checklist (your tenant)

Fill the blanks on your facts. No lift percentages. No logos. No invented customers.

Needle name
________
Leading signal
________ (example shape only: time-to-X)
Lagging signal
________ (example shape only: repeat-Y)
Open checklist.md

# Private-eval checklist (your tenant)

Fill the blanks on your facts. No lift percentages. No logos. No invented customers.

## Fields

- Needle name: ________
- Leading signal: ________ (example shape only: time-to-X)
- Lagging signal: ________ (example shape only: repeat-Y)

## Steps

- [ ] 1. Confirm paying Base — workspace + connector + signed events on dept.* products.
- [ ] 2. Pick a needle — one workflow outcome you already run.
- [ ] 3. Land live ops facts — tools you already run → signed events under your tenancy.
- [ ] 4. Name leading and lagging signals — fill the blanks above; names only, not a public lift figure.
- [ ] 5. Run the private eval — score the needle on your facts in your tenant.
- [ ] 6. Keep the score private — do not publish lift %, customer names, or logos from this run.
- [ ] 7. Scale inference only after proof — if the needle did not move, fix context and signals first.

Scores stay in your tenant. This file is the contract, not a leaderboard.

Scores stay in your tenant. This file is the contract, not a leaderboard.

Needle shapes (your data, not our logos)

Workflow types only. Swap the facts for yours. These are not customer stories.

Incident evidence recall

Time to linked PR or change evidence on your incidents.

Account-health narrative

Merged support and CRM context your team already operates.

Governed metric freshness

Whether a briefing cites a freshness-governed product, not a spreadsheet dump.

FAQ

Is this Braintrust / an LLM-judge platform?

No. This page is a private workflow-eval contract on your mesh facts after a paying Base — not a hosted judge marketplace.

Where do scores live?

In your tenant. This page does not publish lift percentages, logos, or customer names.

Which connectors for the first needle?

Prefer PagerDuty then GitHub when the needle is ops. Slack waits (parked). Catalog listing is not Connected.

Do I need Base first?

Yes. Paying Base — workspace, connector, signed events — before staffed or private eval on this template.

Run it on your tenant

Create a workspace if you do not have a paying Base yet. Open console to continue on your tenant after a paying Base exists.

Create a workspace

Start here if you do not have a paying Base yet.

Open console

Continue on your tenant after a paying Base exists.

Paying Base is the floor. Template steps run on your tenant after that.