Training an agent is mostly writing code
- ai
- local-models
- legal
- tooling
Astro is a desktop client I built that talks to a language model running on a graphics card in this room. Nothing goes to a data centre, there is no account, and the window is deliberately boring: a text box, an Attach button, a status light.

Over the last few days the question stopped being does it work and became is it good enough for someone doing regulated work. I assumed that meant feeding it material about the field. Almost none of the work turned out to be that.
The pack, and what checking it cost
The first attempt was the obvious one. Write a knowledge pack for the sector, the same shape as any other skill: a directory, a SKILL.md, frontmatter and markdown, allowlisted to one agent.
Then I ran an adversarial pass over every claim in it. One agent per claim, each told to try to refute rather than confirm.
19 of 42 claims came back refuted or needing a rewrite. Forty-five percent. Among them one fabricated quotation and one source attributed to the wrong body, both produced by research agents that had sounded completely sure of themselves.
A second pack, for financial services, went into the same pipeline and never came out. The run burned 4.15M subagent tokens across 172 agents and hit a weekly limit with 124 agents still running. The solicitors verification had finished. The financial one had not started.
So the financial claims stayed in a scratch file and were never written into a pack. At a 45% refute rate, unverified domain claims put in front of an FCA-regulated firm are worse than having nothing at all. That number is the single most useful thing I learned all week, and it is an argument for the verification step, not against the packs.
A document is knowledge that has stopped moving
The deeper problem with a pack is that it is a snapshot. Legislation changes, guidance gets reissued, and a markdown file written in August is quietly wrong by spring.
So the next piece was a small command-line tool that reads legislation.gov.uk directly:
astrolaw find "renters rights"
astrolaw inforce "renters rights"
astrolaw section ukpga/1988/50 21
find returns Acts and their references. section returns the text of a provision with its URL.
inforce is the one that earns its place. An Act passing is not the same as an Act being in force, and the difference lives in the commencement regulations. That distinction is exactly the sort of thing a model will smooth over, because both facts sound like the same fact.
Only the reference goes out. Never a client name, never a matter number, never document text. That keeps the “nothing leaves the building” line true in the place where it actually matters.
Case law is a gap I left open deliberately. Find Case Law requires a computational analysis application under the Open Justice Licence, which does not permit programmatic access without one. No fee, but a gate. So the agent is told to say it cannot check a case rather than have a go at one.
One in five citations was still invented
Here is where I stopped feeling clever.
I asked it, five times, which Act abolished section 21. Four answers gave ukpga/2025/26, correct. The fifth invented a “Renters’ Rights Bill 2023, enacted as ukpga/2023/44”.
The tool was sitting right there, allowlisted and working. It answered from memory anyway.
Four attempts to fix that with words
I did not go straight to code. I tried to instruct my way out of it first, four times, measuring each rather than reading one good answer and declaring victory.
| Attempt | Result |
|---|---|
| “Look it up when the answer is a number you were taught” | searched once, 3/5 |
| “Always search before stating an event ID or port” | 1/4 |
| “Never state one as bare fact, say it’s worth checking” | 0/5 hedged |
| All three, in the section that works for other rules | no better |
The reason is now clear, and it generalises past this one problem.
A rule about accreditations holds 5 times out of 5, because its trigger is lexical: the words are in the request, so the rule fires on pattern. “Is this answer a numbered identifier?” asks the model to classify its own output, and this one cannot do that. It does not know when it is unsure. DHCP option 42 came back as 128, then 4, then 69, confidently, on three separate runs.
Lexical triggers work. Self-assessment does not. Every instruction I write now gets keyed to words that appear in the request, never to the model’s opinion of its own certainty.
One thing did help short of code, and it was context rather than instruction. A memory file listing the event IDs and ports an MSP actually meets took the same questions from 1/4 to 5/6, including one answer that was not in the file. Event 4740 still comes back wrong 2 times in 5, so that is a large improvement rather than a fix.
I deleted the failed instruction instead of leaving it in. A rule that does not work still costs context on every single turn, and worse, it invites you to trust an answer nothing ever checked.
So the check became code
Every reply is now scanned for legislation references. Each one is resolved against legislation.gov.uk and the real title appended.
| Reference in the reply | What it resolves to | Effect |
|---|---|---|
ukpga/2025/26 | Renters’ Rights Act 2025 | confirms a correct one |
ukpga/2023/44 | Pensions (Extension of Automatic Enrolment) Act 2023 | exposes the invention |
ukpga/1988/50 | Housing Act 1988 | catches a mismatched reference |
Cached, capped at six lookups per reply, and wrapped so a failed lookup can never eat the answer it was checking.
The snippet problem, which is how a wrong number reaches a client
The search stack had a quieter fault. The local SearXNG instance returns about 150 characters per hit. The stamp duty rates snippet was cut off mid-figure, at “£4,”.
A model answering from that will produce something confident and wrong.
What surprised me is that the fix needed almost no new machinery. Fetching full pages already worked, and asked directly it returned the complete band table in 13.7 seconds. The sources that matter here are server-rendered or have APIs, so a heavyweight scraper buys nothing. Nothing was forcing the search-then-fetch sequence. It fetched when I told it to and answered from the snippet when I did not.
So the sequence became one command instead of a request. Search, fetch the top pages in parallel, return them numbered. It ranks legislation.gov.uk first and gov.uk second, takes one page per site so three hits from the same domain cannot collapse into what looks like three sources, and caps each page so a monster page cannot eat the context window.
Checked end to end, it produced the correct bands to £1.5m and the correct £300,000 first-time threshold, and cited [1] and [3] without being asked to in that turn.
Making a citation bracket mean something
A [3] in a reply says nothing on its own. The reply does not carry the source list, so the bracket is decoration a reader is invited to trust.
Research runs now write a manifest: the query, a timestamp, and one line per source with its title and URL. Titles and URLs only, page text stays out of it. The client reads that manifest whenever a reply contains a bracket and appends what each number actually was.
A number with no matching source is called out in the open: [7] does not match any source that was retrieved. A reply citing brackets when no research ran at all says exactly that.
Two gates now run on every reply, and both are code. Legislation references resolved against legislation.gov.uk. Bracket numbers resolved against the manifest.
Three things I turned down, and why
A proposal came in for a heavier retrieval stack. Working through it was more instructive than building it would have been.
Chunking documents to 512 tokens with a reranker on top. The premise was that you cannot feed three pages to a model without diluting its attention. Measured here, that is false: 5 out of 5 needles found at 250,287 characters. And a 512-token chunk cheerfully splits a statutory subsection away from its proviso, which is worse than not chunking at all.
“Algorithmically barred from answering from memory.” That sentence describes a prompt, and prompts do not bar anything on this model. Its own rule said if the source is not present, say I cannot verify. That is the self-assessment class I had already watched fail four separate ways, including 0 out of 5 on a hedge.
Running generated Python to work out tax. The risk is the formula, not the arithmetic. Stamp duty is banded, so price * rate executed perfectly is still perfectly wrong, and the precision makes it more convincing. Known taxes get a deterministic calculator instead, written once and checked.
Documents: templates, not writing
Rendering was never the problem. A contract with nested clauses, a bordered charges table, a schedule and a signature block came out correct first time, two pages of A4.
The problem was upstream. Free generation means the model authors legal wording from scratch on every request, which is worse than merely unreliable.
So documents go through named templates now. Ask for one with no values and it prints the list of values it needs. Give it some but not all and it refuses to render and names what is missing, so a document cannot go out with a gap in it. Both templates carry a line on the page saying the document is an unchecked draft.
Asked in plain English for an agreement for a named practice, 18 users, 90 days’ notice, it produced a correct two-page PDF with the right parties and the right notice period, and worked out £1,386 as 77 × 18 on its own.
One bug from that work is worth repeating, because it would have shipped quietly. WeasyPrint silently ignores <ol start="4">. Clause numbering restarted at 1 after every table, so the contract had two clause 1s. In a contract that is a defect, not a cosmetic issue. Fixed with a document-wide CSS counter so numbering runs straight through whatever interrupts it.
The shape it settled into
Five independent assessments of the design converged on the same conclusion without being pointed at it, and it is the opposite of where I started.
Take the side effects away from the model. Code does anything that touches the world: the lookups, the calculations, the rendering, the checks. The model classifies text into a fixed set of options and drafts sentences a person approves before they go anywhere.
That is a smaller job for the model than the phrase “AI assistant” suggests, and it is the reason the thing can be pointed at regulated work at all. It also means a model swap matters far less than I assumed. “Use a better model” was never the escape hatch.
What I have to be honest about
The pack is model-verified, not solicitor-verified. It cites material past my own knowledge cutoff, including an SRA notice dated 17 August 2026 and a 2026 High Court decision. A real solicitor has to check those citations before this goes anywhere near a law firm.
I set out to teach a model a subject. I spent almost all of the time writing the checks for the times it would get that subject wrong, and that is the part that turned out to be the product.
