Healthcare · Clinical AI

Consultation Summarizer

A clinician talks to a patient. The record writes itself, correctly, in the structure the system expects. The transcription was the easy part. Everything that stops it inventing a symptom is the actual product.

Sector
Healthcare
Shape
Speech plus LLM
Handling
HIPAA and GDPR
Hard part
Not the model
Status
Running and supported
The problem

What it replaced.

Clinical notes were written after the fact, from memory, at the end of a shift. The record that resulted was thin, late and occasionally wrong in a way nobody could audit, because there was no artefact to check it against. Every clinician knew this. None of them had time to fix it.

The fourth node is why this took months rather than a weekend. A summary that adds a symptom is worse than no summary.
What we built

The pieces, and why each one exists.

What we built

Transcription that survives real rooms

Two people, overlapping speech, background noise, and a vocabulary where "hypertension" and "hypotension" differ by one phoneme and mean opposite things. General transcription is excellent at conversational English and noticeably worse here, which is where the tuning went.

What we built

Structuring into SOAP rather than prose

The output has to land in Subjective, Objective, Assessment and Plan as the record system expects them, not as a paragraph somebody then has to cut up. The model is constrained to that structure rather than asked politely for it.

What we built

A verification pass against the transcript

Every clinical claim in the summary is checked back against the words actually spoken. Anything the transcript does not support is dropped rather than softened, and the clinician sees what was dropped.

What we built

Handling built in from the first commit

HIPAA and GDPR requirements shaped where audio lives, how long it lives, who can reach it and what is logged. Retro-fitting that after a working prototype is how projects in this sector die.

The hard part

Where the time actually went.

The failure mode that matters is not a missed word, it is a confident sentence describing something that never happened. A language model asked to summarise will smooth a gap rather than leave one, which is exactly the wrong instinct in a medical record. Most of the engineering is the machinery that makes the system prefer an incomplete note to an invented one, and makes the omission visible to the person signing it.

What we took from it

Carried into later work

  • The clinical safety work is the product, not a phase at the end of it.
  • A model that refuses to guess is more useful here than a model that scores higher.
  • Compliance decided the architecture, so it had to be understood in week one.
The stack

What it runs on.

Speech
Speech to textSpeaker separationClinical vocabulary tuning
Language
LLM pipelineConstrained outputSOAP structuring
Verification
Claim checkingTranscript groundingClinician review step
Compliance
HIPAAGDPRRetention policyAccess logging
What is not on this page

No client name, and no numbers we were not given.

Most of our contracts ask us not to name the client, and we would rather keep the work than the logo. We have also not put a percentage improvement on this page, because we would be estimating it and you would be right to discount it.

What we will do is arrange a reference conversation with the client, with their agreement, for work that resembles yours. Ask on the first call.

Next step

Bring us the problem, not the spec.

Thirty minutes. If your problem rhymes with this one we will say what we would do differently the second time, which is usually the more useful half.