Medical Transcription Accuracy

What HeliQore’s speech-to-text WER results show

A transparent benchmark of English medical dictation, terminology, noisy audio, accented speech, and the limits of automated scoring.

← Back to all HeliQore blog articles

HeliQore was evaluated with a Word Error Rate (WER) benchmark designed around medical and medico-legal transcription. The purpose was not to produce a marketing number at any cost. It was to identify where transcription worked well, where it struggled, and which results still need better reference verification.

Across the selected, more controlled medical dictation groups, HeliQore recorded a combined WER of 1.79% over 75 cases and 10,534 reference words. This figure describes the selected group only, not every benchmark recording or workflow.

How the HeliQore WER test was performed

The test used file transcription with English regional settings, consistent submission settings, saved generated transcripts, and a repeatable scoring script. Completed jobs were compared with paired reference transcripts. Cases without a valid audio/reference pairing were not presented as successful results.

Primary WER used the standard formula:

WER = (substitutions + deletions + insertions) / reference words × 100

References and outputs were lowercased, whitespace-normalised, and scored without punctuation for the primary word score. Numbers, units, abbreviations, and known US/UK spelling variants were handled using documented, symmetric rules. Clinically meaningful words, negations, names, and repetitions were retained. Formatting and dictated layout commands were treated separately where possible.

WER can still be higher than the practical information error rate when the reference and output use different representations. For example, a date may appear as spoken words or numerals, a measurement may use a full unit name or an abbreviation, and a medical term may use an accepted US or UK spelling. HeliQore is designed to use medical-specific language and established unit conventions, so a difference such as milligrams per litre versus mg/L does not automatically mean that the underlying clinical information is wrong. These representations should be normalised consistently or reviewed as a separate formatting finding rather than counted as a clinical mistake.

Some records required reference review because the written reference did not always appear to match the audio exactly. A reference transcript is not automatically ground truth merely because it is written down. If it contains a typo, missing word, formatting command, or wrong number, WER can measure the disagreement between two texts rather than the system’s actual speech-recognition ability.

Datasets included in the benchmark

The benchmark used several distinct types of English medical or medically oriented speech. They measure different capabilities and should not be collapsed into one unexplained average.

  • Corti Med-Dictate: medical dictation recordings intended for transcription evaluation.
  • NCH medical/legal: medico-legal material with report-style references; useful, but reference fidelity and template content require care.
  • Hani Synthetic Medical: synthetic medical speech used as exploratory evidence rather than a substitute for real clinician dictation.
  • Hani Pipelines: short synthetic pipeline/dialogue-style turns; informative for experimentation, but not a headline doctor-dictation set.
  • MultiMed English: a large supplementary medical speech collection with more variation and difficult cases.
  • Trelis Medical Terms: short medical terminology-focused samples.
  • Trelis EKA Hard: an intentionally difficult set involving accent and noise challenges.
  • Trelis MultiMed Hard: a difficult medical speech set containing challenging acoustic and speaker conditions.

Dataset-level WER results

DatasetCasesReference wordsWER
Hani Pipelines exploratory192001.00%
NCH replacement cases21,3091.15%
Corti Med-Dictate246,9141.46%
Hani synthetic medical302,1113.36%
Trelis Medical Terms508084.70%
Trelis MultiMed Hard501,37821.41%
MultiMed English supplement1,42736,71027.21%
Trelis EKA Hard5036368.04%

The strongest results came from the more controlled medical dictation and terminology-focused sets. The harder sets were substantially less successful, which is expected when audio contains background noise, heavy accents, short turns, unfamiliar speech patterns, or difficult acoustic conditions.

Why were the harder datasets harder?

Background noise can mask consonants, word endings, numbers, and low-volume speech. Accents and pronunciation variation can change the acoustic patterns used to recognise specialist terms, names, and medication words. Short recordings are especially volatile: one wrong word in a ten-word sample produces a much larger percentage than the same error in a long dictation.

Some difficult collections also include speech that is not the same task as a doctor dictating a report. Consultation-style dialogue, synthetic speech, interruptions, and multi-speaker material test robustness and conversational audio—not only medical dictation accuracy.

What do the results mean?

A low WER means the recognised word sequence was close to the chosen reference under the stated rules. It does not prove that every medication, dosage, negation, anatomical term, name, or diagnosis is safe. A single clinically important error can matter more than many harmless formatting differences.

The benchmark suggests that HeliQore is currently strongest on the more controlled medical dictation and terminology material tested. The difficult audio results identify an improvement opportunity: noise, accent variation, and challenging acoustic conditions deserve more engineering and more representative evaluation.

The good-set scores should also be interpreted carefully. Some reference transcripts were created or supplied separately from the audio and were not independently verified word by word before testing. A mismatch such as a missing word, a number written differently, or a formatting instruction included on only one side can raise or lower WER without reflecting a recognition failure.

How does this compare with human medical transcription?

Historical studies show why professional review remains important even when accuracy percentages are high. In a study of 206 surgical pathology reports containing 23,458 words, human transcription had a reported mean accuracy of 99.6%, with a range of 99.4% to 99.8% [1]. The study also found that the tested speech-recognition system performed less accurately than human transcription.

A study of 47 emergency-department charts reported 99.7% accuracy for traditional transcription, compared with 98.5% for the voice-recognition system tested in that study. The human service averaged 1.2 corrections per chart [2]. These are historical results for particular systems and workflows, not direct comparisons with HeliQore.

A systematic review of speech recognition in healthcare found that studies were heterogeneous and difficult to combine. It concluded that speech recognition was generally not as accurate as human transcription, while offering potential benefits in turnaround time and cost. It also highlighted accented voices, training, task length, macros, templates, and workflow as relevant factors [3].

Human transcription is not a reason to skip review. Human work can contain errors too, and medical documents should be checked against the source. It is a reason to keep an accountable review step in the workflow, especially for clinically or legally important material.

Future improvements for noise and accents

Future HeliQore evaluation and product work can focus on:

  • More robust denoising and voice-activity detection for background sound.
  • Microphone and recording-quality guidance before upload.
  • Evaluation across Australian, British, American, and other English accents.
  • Improved recognition of medication names, measurements, numbers, and specialist vocabulary.
  • Better handling of short turns, low-volume speech, and variable speaking pace.
  • Confidence or uncertainty signals that help reviewers find passages needing attention.
  • Separate reporting for word accuracy, formatting, speaker behaviour, and clinically important errors.
  • Reference verification against the original audio before publishing benchmark claims.

Reference review is particularly important. The next pass should listen to every disputed passage, correct reference typing mistakes, document intentional formatting conventions, and keep an audit trail of changes. This will produce a more accurate estimate of speech recognition rather than a score dominated by reference construction.

What should a medical transcription buyer take from this?

Ask to see the test method, dataset composition, reference quality, normalization rules, and per-category results—not only one headline percentage. Check whether the recordings resemble your clinicians, specialties, accents, microphones, and working environment.

For production use, treat automated transcription as documentation support. Review the complete transcript before clinical, medico-legal, insurance, or administrative use. Pay particular attention to negations, names, dosages, measurements, medications, diagnoses, laterality, and missing or inserted statements.

Frequently asked questions

Which HeliQore results were strongest?

The selected stronger groups were Corti Med-Dictate at 1.46% WER, Hani Pipelines at 1.00%, Hani Synthetic Medical at 3.36%, and the two NCH replacement cases at 1.15%. The pooled WER for those selected groups was 1.79%. These are dataset-specific results, not a universal product guarantee.

Why is WER higher on accented or noisy speech?

Noise obscures speech sounds, while accent and pronunciation variation can make words and specialist terms harder to distinguish. Short recordings also make each individual error count for more.

Does a low WER guarantee a safe medical transcript?

No. WER counts word differences, not clinical importance. A low score can still contain a dangerous error in a dosage, negation, name, or diagnosis, so professional review remains necessary.

Were all references verified against the audio?

No. Reference quality was identified as a limitation. Some good-set results may be affected by differences between the written reference and the audio, and further reference verification is planned.

Are these results a comparison with human transcription?

No. The historical human-transcription studies provide context only. They used different products, recordings, dates, and workflows, so they should not be treated as a direct head-to-head HeliQore comparison.

Sources

  1. Al-Aynati and Chorneyko, pathology reports
  2. Zick and Olsen, physician charting in the ED
  3. Johnson et al., systematic review of speech recognition in health care
  4. HeliQore medical transcription accuracy guide

Dataset sources

  1. Corti Med-Dictate: https://huggingface.co/datasets/corti/med-dictate
  2. NCH medical/legal practice files: https://www.nch.com.au/scribe/practice.html
  3. Hani Synthetic Medical Speech Dataset: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset
  4. Hani SynthaticPipelines: https://huggingface.co/datasets/Hani89/SynthaticPipelines
  5. MultiMed English: https://huggingface.co/datasets/leduckhai/MultiMed
  6. Trelis Medical Terms: https://huggingface.co/datasets/Trelis/eval-medasr-medical-terms-2025-20260408-1927
  7. Trelis EKA Hard: https://huggingface.co/datasets/Trelis/eval-medasr-eka-hard-20260408-1922
  8. Trelis MultiMed Hard: https://huggingface.co/datasets/Trelis/eval-medasr-multimed-hard-20260408-1932

Important: This article reports a product benchmark, not a clinical validation or certification. HeliQore output must be reviewed by an appropriately qualified professional before use.