Held-out evals · clinical documentation

ChartForge

Eval results front and center: classic · flat · modern · BASE on locked fixtures. SOAP, transcripts, and notes sit beside the quality boards — FT create stays paused.

What the evals show

  • Held-out compare: quality first across classic · flat · modern · BASE
  • Prior SOAP six-metric board (structure · grounding · fluency · § judge)
  • Trustworthy clinical docs clinicians will actually sign
  • Medical necessity language that survives payer audit

At a glance

18 transcripts · held-out compare + six-metric board on /metrics (/eval alias)

91% section accuracy · +15% vs baseline LLM

Context Aggregation of Transcripts — evaluation on ACI-Bench / MTS-Dialog

Together for inference + FT · never OpenAI — classic CAT local; flat = direct-prompt baseline; FT launch paused.

System design →