Case Study
Geek Therapeutics Automates Clinical
Assessment Reporting with Multi-Agent
Generative AI on AWS
Clearscale built a multi-agent generative AI pipeline that turns a testing day’s scanned
assessment forms into a draft clinical evaluation report in the practice’s house style —
validated 27 out of 27 against the clinician’s approved gold-standard report.
Client Profile
Industry Healthcare / Behavioral Health
Technology
GenAI — Claude on Amazon
Bedrock
Overview
Geek Therapeutics was spending licensed clinician time assembling long, structured evaluation reports by hand after every testing day.
Clearscale built a multi-agent generative AI pipeline on AWS that ingests scanned assessment forms and produces a draft report in the practice’s canonical 13-section template, ready for the clinician to review, edit, and sign.
Validated against a completed, clinician-approved evaluation, the pipeline matched the clinician’s final report on 27 of 27 objective checks, in roughly 30 minutes.
Meet Our Hero
Geek Therapeutics performs comprehensive assessment batteries drawing on standardized psychological, developmental, and behavioral instruments. Each evaluation spans cognitive, academic, attention, adaptive, sensory, and behavioral domains, with rating scales completed by clinicians, teachers, parents, and the patients themselves.
Turning that raw material into a thorough, readable, correctly coded report is skilled, time-consuming work. As demand grew, the practice’s most valuable resource — licensed clinician time — was increasingly spent assembling documents rather than exercising clinical judgment.
Leadership wanted to compress the mechanical part of report production without compromising accuracy, auditability, orpatient privacy, and engaged Clearscale to prove out an AI-assisted approach.
The Challenge
Challenge 01
Report assembly consumed
licensed clinician time after every
testing day
Challenge 02
Assessment data arrived as
heterogeneous scanned PDFs —
response sheets, interpretive
score reports, and intake
documents — requiring careful
per-patient attribution
Challenge 03
Scoring, diagnostic coding
(DSM-5-TR / ICD-10 / ICD-11),
and house-style formatting were
repeated manually for every
report
Challenge 04
Any AI solution needed strict PHI
handling, full auditability, and no
hallucinated scores or diagnoses
The Goal
• Turn a testing day’s scanned forms into a draft report in the practice’s exact house style
• Extract every item response and score accurately from varied, scanned form types
• Compute scores deterministically from standardized scoring tables, not from model inference
• Ground diagnoses, integrative summary, and recommendations strictly in the extracted data
• Validate every draft against the clinician’s own approved reports
• Design an architecture that can move onto a secure, HIPAA-eligible AWS footprint
The Solution
Clearscale designed a five-stage pipeline in which Claude on Amazon Bedrock is used only where
judgment and language are needed. Scoring and document assembly are deterministic and fully
auditable.
Step 01 | Vision OCR &
Extraction
• A two-pass classifier reads the bundle once to identify which forms sit on which pages
• Extracts every circled response, checkmark, and score-summary table, form by form
• Extractors are built instrument by instrument, covering behavioral, adaptive, attention,and sensory rating scales, plus pre-generated interpretive score reports
Step 03 | Clinical Synthesis
• Claude on Amazon Bedrock generates test-by-test narratives by clinical domain, a
seven-domain integrative summary, and
recommendations
• Coded diagnoses to DSM-5-TR, ICD-10, and ICD-11 with severity
• Constrained to a fixed differential list so no diagnosis can be invented
• Every diagnosis cites its supporting data and DSM-5-TR
criteria
Step 05 | Gold-Standard
Validation
• Checked each draft against the clinician’s approved report
• 27 objective checks across diagnoses, codes, severity, T-scores, and classifications
• CI-ready pass/fail exit codes
Step 02 | Deterministic Scoring
• Converted raw item responses to T-scores and qualitative bands using standardized
scoring tables
• Every scoring decision is auditable, and updating a scoring table is a data change,
not a code change
Step 04 | Assembly & Readability
QA
• Assembled content deterministically into the canonical 13-section report
• Tuned readability for family- facing documents while preserving every score and
code
AWS Foundation
• Provisioned the environment with Terraform — VPC, Amazon S3, networking
• Wired the pipeline to run Claude on Amazon Bedrock Claude Code supported
implementation and testing throughout development
• Built on HIPAA-eligible AWS services; the production path is designed around a signed
• HIPAA BAA, a dedicated IAM role for model invocation, VPC endpoints, and CloudTrail, logging
The Impact
Clinician effort shifts from assembling reports to reviewing and signing them
Matched the clinician’s approved report on 27 of 27 objective checks on the validated case
Produces a full draft evaluation report in roughly 30 minutes
Deterministic scoring and a constrained diagnosis list prevent invented scores or diagnoses
A clear, de-risked path to production on a HIPAA-eligible AWS and Amazon Bedrock footprint
Turn Cloud Chaos Into Clear Results On AWS
Clearscale helps marketing and SaaS companies cut through cloud chaos and get clear results on AWS. If your legacy infrastructure is holding you back and you need a partner to tackle complex, large-scale migration and modernization projects, let’s talk.
