The challenge
An Australian government agency maintains a practice framework running to over a thousand pages, with more than five hundred cross-references between its documents. Practitioners in a number of separate organisations are expected to apply it consistently.
They reach for it in the field, between appointments, and on the phone. Until this work the options were a paper copy, keyword search across a folder of PDFs, the website, or asking a colleague.
The solution
- The guide set analysed as a whole first, every guide read against the others and against the legislation behind them.
- An assistant built over the same corpus. Tenant isolated, hosted in Australia, access limited to the trial cohort.
- Every answer cites the paragraphs it was built from, and the system withholds an answer it cannot cite.
- Four rounds of expert scoring against a fixed question set before any practitioner saw it, with retrieval and generation scored separately.
- A two-week trial across several organisations, with a cohort trained on the assistant and use left optional.
The impact
In twenty five of those places the two copies say different things, so two practitioners could each follow the guidance correctly and arrive at a different answer. Finding them previously meant reading the whole set.
The agency holds all 177 as a consolidation roadmap, each one showing the passage it was found in and the passage it conflicts with.
Practitioners rated about half the answers they received, so the 83 per cent covers the answers they chose to judge rather than all of them.
The agency set its stopping condition before the four rounds of expert scoring began: an answer that was wrong or misleading. None occurred in the final three rounds.
Delivered as a consolidation roadmap, a working assistant, the round-by-round evaluation, a practitioner trial findings pack and a forward plan.