Skip to content
KAZES.studio — home
All projects

AI systems

Concept study

A retrieval layer that clinical reviewers actually trusted

An evidence-grounded assistant for trial protocol review, built so that every answer could be traced to source and every change to a prompt could be measured before release.

Engagement
Dedicated engineering team
Duration
Five months, discovery through deployment
Year
2025
Client
Not applicable
TRUST BOUNDARY
Service topology and trust boundaries.. This is a generated schematic, not a product screenshot.

The challenge

What made this difficult.

Reviewers were spending hours cross-checking protocol language against prior amendments and precedent documents. A first-generation assistant had been abandoned because it produced fluent answers that could not be verified, and reviewers had no way to tell a reasonable inference from an invented one. The technical problem was not summarisation — it was provenance. Without traceability, no amount of answer quality would make the system usable in a regulated workflow.

Constraints

Non-negotiables we designed around.

01

Every claim must be attributable

Answers are only useful if a reviewer can jump to the exact clause that supports them, and confirm nothing was paraphrased in transit.

02

The corpus is versioned and amended

Protocols accrue amendments. A retrieval index that ignores document lineage will happily return superseded guidance.

03

Regulated domain, low tolerance for drift

Behaviour changes between model versions are a compliance concern, not just a quality concern.

04

Existing document infrastructure

Content already lived in a permissions-aware repository that had to remain the source of truth.

Approach

How we would build it.

Made provenance a hard requirement, not a feature

The retrieval layer returns structured spans — document, version, clause range — rather than free text. The interface is designed so a claim without a resolvable span cannot be displayed, which removes the failure mode structurally instead of asking the model to behave.

Modelled document lineage as first-class metadata

Amendment chains are represented explicitly in the index. Superseded clauses are retained but demoted, and the citation surfaces the amendment that introduced the language.

Built the evaluation harness before the assistant

A fixed set of reviewer-authored question-and-answer pairs, with unanswerable questions deliberately included. Every prompt, retrieval, and model change runs against this set as a gate. Quality became a number that could be compared across releases.

Isolated model providers behind a boundary

Inference sits behind an internal interface with prompt and retrieval versioning, so a provider upgrade is a controlled operation with a rollback rather than a surprise.

Designed the autonomy boundary explicitly

The system retrieves, cites, and drafts. A qualified reviewer approves anything that reaches the record. That boundary is enforced in code and stated in the interface, not left to convention.

Architecture

How the pieces fit together.

Architecture

Retrieval and evaluation architecture. Document lineage is represented in the index rather than inferred at query time.

  1. Sources

    • Protocol repository
    • Amendment history
    • Precedent set
    • Reviewer annotations
  2. Ingestion

    • Structure-aware parsing
    • Lineage graph
    • Clause segmentation
    • Change detection
  3. Retrieval

    • Hybrid keyword + dense
    • Version-aware filtering
    • Span extraction
    • Ranking with recency prior
  4. Generation

    • Provider abstraction
    • Prompt versioning
    • Output schema validation
    • Span enforcement
  5. Assurance

    • Evaluation set
    • Regression gate in CI
    • Trace store
    • Reviewer checkpoint

Stack

What it would run on.

Interface

  • TypeScript
  • React
  • Python
  • PostgreSQL

Retrieval

  • Hybrid lexical + dense search
  • Structured clause index
  • Document lineage graph

Infrastructure

  • Containerised services
  • Object storage for raw documents
  • Tracing and cost telemetry
  • CI with evaluation gate

Expected outcomes

What success would look like.

Answer attribution

100%

Design target: every displayed claim resolves to a citable source span. Structurally enforced, not model-dependent.

Unanswerable handling

Refusal tested

Deliberately included unanswerable cases in the evaluation set to measure refusal behaviour rather than answer volume.

Change safety

Gated

Prompt, retrieval, and model changes must pass the fixed evaluation set before reaching reviewers.

What we learned

The conclusions we would carry forward.

In a regulated workflow, traceability is a design constraint on the interface, not a property you ask the model for.

Representing document lineage at index time was cheaper and more reliable than reconstructing version context at query time.

Building the evaluation harness first changed the conversation with the client about what 'good' meant — it stopped being a matter of taste.