Held-out evals · clinical documentation
ChartForge
Eval results front and center: classic · flat · modern · BASE on locked fixtures. SOAP, transcripts, and notes sit beside the quality boards — FT create stays paused.
What the evals show
- Held-out compare: quality first across classic · flat · modern · BASE
- Prior SOAP six-metric board (structure · grounding · fluency · § judge)
- Trustworthy clinical docs clinicians will actually sign
- Medical necessity language that survives payer audit
At a glance
18 transcripts · held-out compare + six-metric board on /metrics (/eval alias)
91% section accuracy · +15% vs baseline LLM
Context Aggregation of Transcripts — evaluation on ACI-Bench / MTS-Dialog
Together for inference + FT · never OpenAI — classic CAT local; flat = direct-prompt baseline; FT launch paused.
System design →- “CAP-lite + span provenance: atomic claims with utterance IDs, not just clusters — SpecialtyScribe modularity.”
- “Omit > fabricate: safety checklist (SI, meds, follow-up, measures) gates the draft beside grounding coverage.”
- “Four parallel section generators, non-reasoning models — When Reasoning Hurts for fidelity-sensitive SOAP.”
- “Held-out eval: classic CAT · flat direct-prompt · modern hierarchical · Together BASE — FT paused.”