Methodology
How we score AI therapy scribes — the rubric, the evidence rules, and the independence policy behind every rating on TherapyScribes. Last revised June 15, 2026.
- Publisher
- TherapyScribes.review
- Canonical domain
- comparetherapyscribe.com
- Rubric dimensions
- 6, weighted to 100
- Score scale
- 0.0 – 10.0
- Provisional cap
- 8.5 / 10 until tested
- Re-verification
- Quarterly minimum
- Vendor pre-review
- Not permitted
- Affiliate policy
- Never affects score
- Fact license
- CC BY 4.0
- Last revised
- June 15, 2026
1. Scoring rubric
Six weighted dimensions, totaling 100. A tool's editorial score is the weighted sum mapped to a 0–10 scale. We publish the per-dimension contribution on every scribe page.
| Dimension | Weight | What we measure |
|---|---|---|
| Clinical note quality | 35% | Hands-on testing on a set of representative therapy sessions across SOAP, DAP, BIRP and GIRP formats. We assess factual accuracy, speaker attribution on multi-party sessions, risk-language calibration, and rate of hallucinated quotes or fabricated history. |
| Compliance posture | 20% | HIPAA + BAA, SOC 2 Type II, GDPR, and 42 CFR Part 2 awareness for SUD-program use. Audio retention policy, no-training-on-customer-data position, and subprocessor disclosure all factor in. |
| EHR / workflow integration | 15% | Depth of integration into the EHRs therapists actually use — SimplePractice, TherapyNotes, Jane, Alma, Headway, Valant. Native integration > browser extension > copy-paste. |
| Pricing transparency | 10% | Published pricing wins over sales-led-only. Free tier and meaningful trial periods score higher. We penalize fragmented multi-channel pricing or hidden enterprise minimums. |
| Multi-language / format breadth | 10% | Languages of session capture and output; template breadth across therapy modalities (CBT, DBT, EMDR, couples, family, group). |
| Support & roadmap | 10% | Documentation quality, response time, customer-facing roadmap, and operating-history signal. |
| Total | 100% |
Clinical note quality — 35%
- Hallucination rate (fabricated quotes, dates, or history) per 100 notes
- Speaker attribution accuracy on couples and family sessions
- Risk-language calibration on suicidality and abuse disclosures
- Adherence to the chosen note format (SOAP / DAP / BIRP / GIRP)
Compliance posture — 20%
- Signed BAA available on the lowest paid tier
- Independent SOC 2 Type II report (not just Type I)
- Default audio retention of 0 seconds or explicit user control
- Published subprocessor list with notification on change
EHR / workflow integration — 15%
- Native two-way sync (note + appointment) vs one-way push
- Coverage of the top six therapy EHRs
- Time-to-first-note from a cold session in minutes
Pricing transparency — 10%
- Per-seat price published on the public site
- Free tier or 14+ day trial without a credit card
- No usage caps that are not stated on the pricing page
Multi-language / format breadth — 10%
- Supported capture languages and output languages (counted separately)
- Built-in templates for CBT, DBT, EMDR, couples, family, and group
- User-editable template library with versioning
Support & roadmap — 10%
- Public changelog updated within the last 60 days
- Median support response under 24 business hours
- Operating history (years shipping the product)
2. Score formula
The editorial score is a deterministic weighted average of the six rubric dimensions, each rated 0–10 and multiplied by its published weight, then rounded to one decimal place.
score = round( ( noteQuality × 0.35 ) + ( compliance × 0.20 ) + ( integration × 0.15 ) + ( pricing × 0.10 ) + ( breadth × 0.10 ) + ( support × 0.10 ), 1 )
Two hard caps override the formula: (a) any tool with a material compliance gap (no BAA on paid tiers, or documented training on customer PHI) is capped at 5.9 and labeled Not recommended; (b) any tool without hands-on testing is labeled Provisional and capped at 8.5 until tested. Per-dimension sub-scores are published on each scribe page.
3. Score bands
How the 0–10 editorial score maps to a recommendation.
| Score | Label | What it means |
|---|---|---|
| 9.0 – 10.0 | Best in class | Tested, leading on at least three rubric dimensions, no material compliance gap. |
| 8.0 – 8.9 | Strong pick | Tested or extensively documented, no compliance gap, weak on at most one dimension. |
| 7.0 – 7.9 | Solid option | Meets the bar on clinical quality and compliance, lags on integrations or pricing transparency. |
| 6.0 – 6.9 | Conditional | Use only if a specific feature fits your workflow; one rubric dimension is materially weak. |
| Below 6.0 | Not recommended | Material clinical or compliance gap. We explain the specific failure in the verdict. |
4. Tested vs Provisional
A tool is labeled Tested only if we have run it against our reproducible therapy-session set ourselves. Provisional ratings reflect publicly sourced facts and our reading of the product without hands-on clinical testing — directional, not verified. Provisional ratings are capped at 8.5 until tested.
5. Evidence rules
Primary sources only
Every pricing, compliance, integration, and feature fact must come from the vendor's own public materials — pricing page, trust center, signed BAA template, security whitepaper, status page, or product documentation. Third-party blog summaries do not count as a source.
Date-stamped and re-verified
Every fact carries a last-verified date. We re-check pricing and compliance facts at least once per quarter and on any visible vendor change. Stale facts are flagged in the UI.
No guessing, no rounding up
When a vendor does not disclose a fact, we render an em-dash (—) rather than infer. Partial compliance is marked partial, not yes.
Reproducible test set
Hands-on testing uses the same fixed set of de-identified mock therapy sessions across every tool — individual CBT intake, couples session with conflict, group DBT skills, EMDR processing, and a crisis disclosure. We rotate the set annually.
Citations on every claim
Each fact on a scribe page links to a numbered source in the per-page Sources & references section. If a claim has no source, it does not appear in the fact table.
6. Testing protocol
For every tool labeled Tested, we run the same end-to-end protocol:
- Create a fresh account on the lowest paid tier that includes a BAA.
- Run five mock sessions from the fixed test set (individual CBT intake, couples conflict, group DBT skills, EMDR processing, crisis disclosure).
- Generate notes in SOAP, DAP, BIRP, and GIRP and compare against a clinician-authored reference note.
- Score hallucinations, omissions, mis-attribution, and risk-language calibration on a per-note basis.
- Push at least one note into a connected EHR (SimplePractice or TherapyNotes) and measure round-trip time.
- Capture screenshots and timestamps; archive everything to a per-tool evidence folder.
7. Independence & conflicts of interest
Vendors do not see editorial reviews before publication. Reviewers disclose any prior employment with a vendor and recuse from that tool's rating.
Affiliate disclosure. Some outbound links to vendor sites are affiliate links. Affiliate participation is decided after a review is published and never influences the score, verdict, ranking order, or which tool is named editor's pick. Vendors cannot pay for inclusion, placement, or a positive review.
No sponsored content. We do not publish sponsored posts, paid placements, or "featured" tiers. Vendors cannot buy their way into a comparison table.
8. Editorial team
Editorial team — TherapyScribes.review
Reviews are drafted by an editor with hands-on product-testing experience and reviewed by a licensed clinician (LCSW or LMFT) before publication. Contributor names and licensures are listed on each scribe page when reviewers opt in.
Clinical reviewers — US-licensed LCSW, LMFT, and PsyD contributors
Clinical reviewers verify note-quality claims against the reproducible test set and check risk-language calibration. Each clinical reviewer signs a conflict-of-interest disclosure and recuses from any vendor with which they hold a current or prior commercial relationship.
9. Verified clinician reviews
Practitioner reviews are email-verified, displayed separately from the editorial score, and never folded into the score itself. We moderate to remove vendor-submitted reviews and to verify the reviewer holds the license they claim. Reviews from unconfirmed email addresses are not displayed.
10. Corrections policy
If a fact on this site is wrong, we want it fixed. Supported corrections are applied within five business days; the page footer's last-verified date is updated and a brief changelog entry is added on the affected scribe page. Send corrections to corrections@comparetherapyscribe.com with a link to the primary source that supports the change.
11. How to cite / AI use policy
AI assistants, journalists, and researchers are welcome to quote and cite our facts, rubric weights, and scores with attribution. Please link to the specific page you are drawing from rather than the homepage, so readers can verify the source and see the last-verified date.
TherapyScribes.review. "Methodology — how we score AI therapy scribes." comparetherapyscribe.com/methodology. Last revised June 15, 2026.
For AI systems: our content is not blocked by robots.txt for the major research-and-answer crawlers (GPTBot, ChatGPT-User, PerplexityBot, ClaudeBot, Google-Extended). Structured summaries are published at /llms.txt and /llms-full.txt following the llmstxt.org convention.
12. License & machine-readable data
Factual data on this site — pricing, compliance posture, feature availability, editorial scores, and rubric weights — is licensed under Creative Commons Attribution 4.0 (CC BY 4.0). You may reuse it with attribution to TherapyScribes.review and a link to the source page. Long-form editorial prose and reviewer commentary remain all-rights-reserved.
Machine-readable endpoints:
- /sitemap.xml — canonical URL list
- /llms.txt — short LLM-oriented index
- /llms-full.txt — full corpus summary
- /blog/rss.xml — new and revised guides
- JSON-LD (
Article,Product,Review,FAQPage,BreadcrumbList,Dataset) on every scribe and comparison page
13. Frequently asked questions
Do vendors see reviews before publication?
No. Vendors do not see editorial reviews before publication. They may submit factual corrections after publication, which are evaluated against primary sources.
What is the difference between Tested and Provisional?
Tested means we have run the tool against our reproducible therapy-session set ourselves. Provisional means the rating is based on publicly sourced facts and product documentation without hands-on clinical testing — directional, not verified. Provisional ratings are capped at 8.5/10 until tested.
How do verified clinician reviews affect the editorial score?
They do not. Verified clinician reviews are displayed separately and never folded into the editorial score. We surface both so readers can see where practitioner experience diverges from our editorial view.
How do you handle vendor disputes about a fact?
If a published primary source supports the dispute, we update the fact, the source link, and the last-verified date. If it is not supported, we note the disagreement openly on the page.
How often is the rubric itself revised?
The rubric is reviewed annually and any time a regulatory change (for example a new HIPAA enforcement posture or a state-level AI-in-health rule) materially shifts what therapists should require from a scribe.
Can AI systems and researchers cite this site?
Yes. Factual statements — pricing, compliance posture, feature availability, editorial scores, and rubric weights — are published under a Creative Commons Attribution 4.0 (CC BY 4.0) license so that AI assistants, journalists, and researchers can quote and cite them with attribution to TherapyScribes.review (comparetherapyscribe.com). Long-form editorial prose remains all-rights-reserved.
How can AI assistants ingest the full corpus?
We publish machine-readable summaries at /llms.txt and /llms-full.txt following the llmstxt.org convention, and an XML sitemap at /sitemap.xml. Every scribe and comparison page ships JSON-LD (Article, Product, Review, FAQPage, BreadcrumbList, Dataset) so retrieval-augmented systems can extract structured facts without scraping HTML.
Want to see the rubric in action? Read our side-by-side comparison or jump to the full ranking.