All articles

Automatic session transcription: how it works, where it goes wrong and how to set it up properly

A transcription tool turns what is said in a session into text, and often into a session report. Here is how it works, what lowers its quality, the errors to look for when you review it and the simple settings that change the result.

Countries covered : France

Raw transcript and session report, two different texts

The raw transcript is the word-for-word text of the session. It is long and full of hesitations and unfinished sentences. For 45 minutes, it runs to several pages. The session report is a short summary written from it. The report is what goes into the file, once reviewed.

  • Content
    • Raw transcript. Everything that was said, in order
    • Session report. Themes, progress, points to come back to
  • Possible errors
    • Raw transcript. Words misheard, left out or made up
    • Session report. Those of the transcript, plus shortcuts in the summary
  • Use
    • Raw transcript. Checking exactly what was said
    • Session report. Following the patient from one session to the next

A transcription error can slip into the report with nothing to flag it. For an overview of AI note-taking, read our article on AI note-taking in therapy sessions. Here, we focus on the transcription itself.

How the machine turns sound into text

Speech recognition

The model cuts the audio signal into small pieces and works out, for each passage, the most probable sequence of words. Whisper, a widely used model published in 2022, learned from about 680,000 hours of audio found on the internet. The key word is "probable". The model does not understand the session. When the sound is poor, it does not leave a gap. It offers a plausible word, sometimes a wrong one, and nothing in the text shows that it hesitated.

Telling the voices apart

To know who said what, a separate step groups the passages that share the same voice features and assigns them to a speaker. It works well when people take turns and have distinct voices. It goes wrong more often when two voices sound alike, when people interrupt each other, or when a third person is in the room.

From transcript to session report

A language model then reads the transcript and writes the report. It can make its own mistakes, for example by summarising a nuanced passage too quickly. Errors from the two steps add up. The HAS (Haute Autorité de Santé, the French national health authority) sums up the expected approach with the acronym A.V.E.C. (learn, verify, assess, communicate) in its first guidance on using generative AI in healthcare (October 2025). The draft HAS and CNIL (the French data protection authority) guide on AI in care settings (working document of 16 February 2026) asks professionals to keep human oversight and to compare the tool's output with their own clinical reasoning.

What lowers quality in a session

  • Distance from the microphone
    • What happens. A distant voice blends into noise and echo. A quiet voice, such as a patient who is crying, is lost quickly.
    • What to do. Place the device between you, closer to whoever speaks softly.
  • Background noise
    • What happens. Ventilation, street noise, white noise machine, table vibrations.
    • What to do. Move the device away from noise sources and turn off notifications.
  • Accents
    • What happens. More errors for ways of speaking that are less present in the training data.
    • What to do. Review these sessions more carefully.
  • Children's voices
    • What happens. Recognised much less well than adults' voices.
    • What to do. Check the passages where the child speaks first.
  • Long silences
    • What happens. They encourage made-up sentences.
    • What to do. Read closely whatever follows a long pause.
  • Overlapping voices
    • What happens. Only one voice is kept, or both are mixed up.
    • What to do. In couple or family sessions, expect speaker labels to correct.
  • Specialised vocabulary
    • What happens. Medications, tests, acronyms and unusual first names are replaced by more common words.
    • What to do. Check each of these words.

The available studies mostly cover English. They show trends, not how a given tool performs in French. In 2020, a study published in PNAS tested five consumer systems (Amazon, Apple, Google, IBM, Microsoft). The average word error rate was 35% for African American speakers, against 19% for white speakers. In 2024, researchers showed (AIES conference) that Whisper's strong performance with adults does not carry over to children. Even after extra training on English-speaking children's voices, about 9% of words were still wrong on their test set.

For speech and language therapists and neuropsychologists, one more caution. These models are built to produce correct text. You can therefore expect them to smooth over some of what matters to you clinically, such as a paraphasia, a stutter or a pronunciation error. When the exact form of what the patient says is part of the assessment, keep your own notes of it.

Typical errors and how to spot them

Missing words

A study published in 2018 in JAMA Network Open analysed 217 documents dictated by 144 American physicians with the speech recognition software of the time. The machine's text contained 7.4 errors per 100 words. The most common error was a missing word (34.7% of errors), ahead of an added word (27%). After human review, the rate fell to 0.3% in signed documents. Tools have improved since, but the lesson holds. A missing word cannot be seen, and review makes all the difference.

The negation that disappears

In spoken French, people often drop the "ne". A patient says « j'ai pas envie de recommencer » (I don't want to start again) without the first half of the negation. The whole negation then rests on one short word, said quickly. If that « pas » is lost, the sentence says the opposite. The 2018 study also mentions a "no" added by mistake, which reversed a clinical finding. Check first the sentences about suicidal thoughts, substance use, treatment and violence.

Made-up sentences

In 2024, Allison Koenecke and colleagues presented a study on Whisper, OpenAI's speech recognition model, at the ACM FAccT conference. They had 13,140 English audio segments transcribed from AphasiaBank, a database that includes people with aphasia and people without language disorders.

  • 1.4% of transcriptions contained whole sentences that were not in the audio.
  • 38% of these inventions included harmful content, such as violence, false associations or invented authority (a "thank you for watching", a link to an official website).
  • Inventions were more frequent for people with aphasia (1.7% against 1.2%) and linked to a longer share of silence.

A new test in December 2023 showed an improvement (12 of the 187 segments concerned still produced an invention), but the authors conclude that the model still makes things up regularly and reproducibly. A therapy session contains a lot of silence. Any sentence in the report that you do not remember should be checked.

Names, numbers and swapped speakers

A misheard medication name can turn into a common word or into another drug with a similar name. In the 2018 study, medications accounted for 2.3% of errors. First names of relatives, doses and dates raise the same problem. Another trap is one of your questions being attributed to the patient, which can credit them with an idea they never expressed.

A report with no transcript behind it

If the microphone was not active or was badly placed, the transcript can be almost empty. Depending on the tool, a report may still be produced from a few fragments, and it will look correct. A 45-minute session does not fit in ten lines of transcript.

The three-minute review

  1. Check that the transcript covers the whole session.
  2. Reread every negation on a high-risk topic.
  3. Check medications, doses, first names and dates.
  4. Find in the transcript any sentence you do not remember.
  5. Check that none of your sentences is attributed to the patient.
  6. Add what the machine cannot know, such as non-verbal cues and your analysis.

Do it the same day. Later, you will no longer be able to tell a made-up sentence from a forgotten detail.

Setting up your practice for good transcription

  • Position. Between you and the patient, uncovered, on a stable surface. Not in a pocket, not behind a screen. On a phone, the main microphone is often at the bottom.
  • Phone or computer. A phone is easy to place between the two chairs (Do Not Disturb mode, battery charged). A computer, often on the desk, is further from the patient and its fan can be heard. An external microphone helps in a large room or in family sessions.
  • Video consultations. With headphones, the patient's voice is not in the room. Depending on the tool, only your voice may be captured. Check during a trial that both voices appear.
  • One-minute test. Sit in the patient's chair and speak for a minute at normal volume, then quietly. Do the same from your own chair, then read the transcript. If the quiet voice is missing, move the device closer. Repeat the test whenever you change room or device.
  • Say what matters out loud. The draft HAS and CNIL guide notes that these tools may require small changes in practice, for example saying part of the examination out loud. In a session, a rephrasing (« si je résume, vous avez repris le traitement lundi », so if I sum up, you started the treatment again on Monday) helps the patient and gives the transcript a clear version of the key point.

Live transcription or dictation after the session?

  • What is captured
    • Live transcription. What the patient and the practitioner say
    • Dictation after the session. Your summary, in your words
  • Sound quality
    • Live transcription. Depends on the room and the voices
    • Dictation after the session. One voice close to the microphone, usually more reliable
  • Suitable situations
    • Live transcription. Individual sessions where detail matters
    • Dictation after the session. Patient who refuses, young children, groups

Many practitioners combine the two, and add a short dictation after a transcribed session to include their clinical reading.

What the law says in France

Informing the patient. The GDPR (Article 13) requires people to be informed when their data is collected. For an AI tool used in routine care, the draft HAS and CNIL guide considers that fair information, together with a right to object, is in principle enough. It adds that consent may be needed under the « code civil » (French Civil Code) when a person's voice is captured. The practical steps are in our article on patient consent to AI note-taking.

Health data and confidentiality. A transcript contains health data, a special category of data under the GDPR (Article 9), covered by professional confidentiality. See our article on professional confidentiality and AI.

Hosting. Health data entrusted to a service provider must go to an HDS-certified host (article L1111-8 of the « code de la santé publique », the French Public Health Code). Ask the vendor where the audio is turned into text, how long the audio and the transcript are kept, and whether the data is used to train models. The draft HAS and CNIL guide recommends having subcontractors and their location listed in the contract. More detail in our article on HDS hosting.

Keep the raw transcript, or only the session report?

No text settles this question. Three general rules help you decide.

  • Keep as little as possible. The GDPR requires data to be "adequate, relevant and limited to what is necessary" (Article 5). A word-for-word transcript holds far more than follow-up needs, including information about third parties.
  • Anything you keep can be requested. The patient can obtain a copy of the data you hold about them (Article 15 of the GDPR), transcript included, with its errors. See our article on patient access to their file.
  • The transcript is there to check. Until the report has been reviewed, it is the only way to know what was said.

A simple rule. The approved report goes into the file and follows the retention periods for patient files. The raw transcript is used for review, and you keep it only for a clear reason. Check what your tool keeps by default and whether you can delete a transcript.

How Delta works

During the session, Delta transcribes what is said, then prepares a session report that you review and correct. You can also dictate your observations right after the session. The audio is used only to produce the transcript and is not kept. The report takes your specialty and your therapeutic approach into account. It is added to the patient's file, where you can see the progress specific to your approach over the sessions, for example in addiction care how the patient's substance use changes over several months. Delta also drafts your assessment reports, letters and certificates.

Data is hosted in France, with a host certified for health data (HDS). AI processing, transcription included, takes place on servers located in France. The patient's name and identifying information are pseudonymised before they go through the AI. No audio file is kept, data is never used to train models, and it is encrypted in transit and at rest.

Frequently asked questions

Is automatic session transcription reliable?

In good conditions (well-placed microphone, quiet room, adult voices), it captures most of what is said. It makes more mistakes with quiet voices, children, accents, overlapping voices and long silences. Always review negations, names and numbers.

Do I need a special microphone to transcribe a session?

Usually not. A phone placed between you and the patient is enough in a consulting room of normal size. An external microphone helps in a large room or in family sessions.

Can AI make up sentences that were never said?

Yes. A study presented in 2024 at the ACM FAccT conference found made-up sentences in 1.4% of transcriptions produced by Whisper, more often when the audio contained long silences.

Does transcription work in video consultations?

Yes, if the tool also captures the patient's voice. With headphones, this is not always the case. Run a trial first.

Should I keep the raw transcript in the file?

No text requires it. The GDPR asks you to keep only what is necessary, and the patient can obtain a copy of everything you hold. The approved session report is usually enough for the file.

What if the patient refuses transcription?

You do not transcribe the session, and the refusal changes nothing in their care. Take your notes as usual, or dictate your observations once they have left.

Sources

Cookie policy