How it works
Legal judgment is the scarcest training data there is.
It isn’t on the internet. It lives in redlines, in what a partner refuses to sign, in which point gets traded away at 2am. We didn’t add AI to a traditional law firm. We built the firm as the place that data comes from.
The position
We are not building a tool for lawyers.
Legal AI is sold to law firms as software. The lawyer stays the operator, so the ceiling is their own throughput.
We built the opposite. The system runs the production work. The lawyer owns the matter and spends their hours on judgment: what position to take, which risk is worth carrying, what to concede.
We are the firm on the engagement letter, so we see how each matter ends. That is the label the training depends on.
The loop
Real matter
client work
Model drafts
first pass
Lawyer decides
edit · reject · escalate
Outcome
closed · negotiated
Training data · evals · RL environments
decisions and outcomes return as signal
What gets captured
Every decision a lawyer makes is a label.
Ordinary acts of doing the work. Each already has the shape a training signal needs.
| The lawyer | Becomes | Improves |
|---|---|---|
| Edits a generated draft | Preference pair: rejected vs. accepted text | House style, risk posture, what survives review |
| Rejects an approach outright | Negative example with the reason attached | Position selection, and what never to propose |
| Escalates to a partner | Routing label on a hard case | When to defer to a human instead of answering |
| Holds or concedes a point | Outcome-labelled trajectory | Which positions are worth the fight |
| Closes the matter | Full trajectory plus result, an RL environment | End-to-end judgment, not just the next token |
The training setup
Where each signal goes.
Ordinary post-training. The source is not: this data only exists at the moment a lawyer decides something on a live deal.
The version that went out
Supervised fine-tuning
The only unambiguous label of correct that legal work produces.
Draft, and the partner's edit of it
Preference pairs
Ranking the edit above the draft teaches taste, which does not generalise from public text.
What survived the negotiation
Outcome-weighted reward
A clause still standing at signing scores above one traded away in week three.
Deals the model has not seen
Held-out matters
Scoring a model on what it was trained on measures recall, not skill.
Why it compounds
Judgment normally is not stored.
A decision made on a deal usually survives in an email thread, a redline, or one lawyer’s memory. We record it against the matter it came from, so it stays after they leave.
The second model
A fixed fee is a forecast.
Hourly billing puts the variance of a matter on the client. A fixed fee moves it to us, which requires predicting what the matter will cost before it starts.
It reads
- Data room size and state
- Deal structure
- Who is across the table
- Scope we are asked to own
Estimator
- fine-tuned on closed matters
It predicts
- Documents to turn
- Model compute
- Partner hours
- Cost if it drags
One number, agreed before we start. We carry the difference.
Every closed matter is another labelled example of what that shape of deal really took. Same loop, pointed at cost instead of quality.
How we score it
What we measure.
- Survival to close
- Whether a position was still standing at signing.
- Estimate error
- How far the quoted fee sat from what the matter cost to produce.
- Partner delta
- How much changed between what the system produced and what went out.
- Round trips
- How many passes before the client stopped asking for changes.
- Recurrence
- Whether the same client brought us the next deal.
Live matters, so they stay under privilege. Numbers when enough deals have closed to mean anything.
