---
updatedAt: 2026-09-29T20:37:00.000Z
agentTools:
  projectIndex: https://docs.recall.ai/llms.txt
---

# Diarization

Choose the right diarization mode so your transcripts use accurate speaker labels and names across remote, shared-mic, and hybrid meetings.

<Callout icon="📘" theme="info">
  For an introductory walkthrough on diarization, see this <Anchor target="_blank" href="https://www.recall.ai/blog/speaker-diarization">speaker diarzation</Anchor> guide.
</Callout>

Diarization is the process of assigning each part of a transcript to a speaker, so the transcript can show who said what.&#x20;

<br />

# Speaker labels

To diarize a transcript, the transcript assigns a **speaker label** to each part of the conversation. A speaker label is the identifier shown in the transcript for a speaker.

There are two types of speaker labels:

* **Participant speaker labels**: labels tied to a specific participant in the meeting platform, so you can identify which meeting participant the label refers to.
* **Generic speaker labels**: labels that distinguish speakers but are not tied to a participant from the meeting platform, such as numbers (`1`, `2`, `3`) or letters (`A`, `B`, `C`).

<br />

# Diarization methods

A **diarization method** defines how transcript text is assigned to speakers. It determines how speakers are identified in the transcript and whether the transcript uses **participant speaker labels** or **generic speaker labels**.

There are four diarization methods you can use:

<Accordion title="Perfect diarization">
  "Perfect diarization" transcribes each audio stream separately. Because each stream is transcribed independently, Recall can keep speech from each stream separate and return **participant speaker labels**.
</Accordion>

<Accordion title="Hybrid diarization">
  "Hybrid diarization" is the most accurate diarization method when some audio streams may contain more than one speaker. It uses perfect diarization to transcribe each audio stream separately, then applies machine diarization within each stream to distinguish between multiple speakers sharing that same stream.
</Accordion>

<Accordion title="Speaker-timeline diarization">
  "Speaker-timeline diarization" uses active speaker events from the meeting platform to map transcript text to the participant identified as speaking during each time range. The transcript is returned with **participant speaker labels**.
</Accordion>

<Accordion title="Machine diarization">
  "Machine diarization" uses a [supported third-party speech-to-text transcription provider](https://docs.recall.ai/docs/transcription#transcription-implementation-guides) to distinguish speakers based on voice characteristics. The transcript is returned with **generic speaker labels** (e.g., `A`, `B`, `C` or `0`, `1`, `2`), rather than participant speaker labels.
</Accordion>

<Callout icon="📘" theme="info">
  For more information on which diarization method is right for your use case, check out our guide on <Anchor target="_blank" href="https://www.recall.ai/blog/speaker-diarization">speaker diarization</Anchor>.
</Callout>

## Comparing diarization methods

The table below compares the available diarization methods and their tradeoffs.

| Method                       | Supported for                                                     | Transcribes               | Speaker labels                                        | Useful when                                                                                                                                                              | Caveats                                                                                                                                                                                                    |
| :--------------------------- | :---------------------------------------------------------------- | :------------------------ | :---------------------------------------------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Perfect diarization          | Bot async transcription and bot real-time transcription           | Separate audio streams    | Participant speaker labels                            | Each participant is joining from their own device, so each person has their own audio stream                                                                             | Does not distinguish between multiple people speaking in the same audio stream, and does not include raw provider data                                                                                     |
| Hybrid diarization           | Bot async transcription                                           | Separate audio streams    | Participant speaker labels and generic speaker labels | Most participants join from their own device, but some calls may include multiple people sharing the same device (e.g., people joining together from a conference room). | Requires additional code to map speaker labels to participants, still uses generic speaker labels for speakers sharing a stream, and does not include raw data from the third-party transcription provider |
| Speaker-timeline diarization | Bot/DSDK async transcription and bot/DSDK real-time transcription | Single mixed audio stream | Participant speaker labels                            | You want participant speaker labels and you need the raw data from the third-party transcription provider                                                                | Depends on meeting-platform active speaker events, which can sometimes be inaccurate or incomplete (e.g. noise, missed speaker changes, or overlapping speech).                                            |
| Machine diarization          | Bot/DSDK async transcription and bot/DSDK real-time transcription | Single mixed audio stream | Generic speaker labels                                | Multiple people may be speaking on the same mixed audio, such as when several participants join together from a conference room or shared device                         | Can be less accurate when different speakers have similar-sounding voices                                                                                                                                  |

<Callout icon="📘" theme="info">
  ### Why separate audio streams improve diarization accuracy

  Diarization accuracy depends in part on how meeting audio is provided for transcription. Meeting audio is typically provided in one of two ways:

  - **Single mixed audio stream**: all participant audio is combined into one stream before transcription. This makes speaker attribution harder, especially when multiple people are speaking at the same time or when multiple people are sharing the same device or microphone.
  - **Separate audio streams**: audio is provided in separate streams instead of one mixed stream. This makes speaker attribution more accurate because each stream can be transcribed independently.

  In general, diarization is more accurate when using **separate audio streams**.

  #### Separate streams platform support

  | Platform        | Real-time separate audio streams | Async separate audio streams |
  | :-------------- | :------------------------------- | :--------------------------- |
  | Zoom            | ✅ Supported                      | ✅ Supported                  |
  | Microsoft Teams | ✅ Supported                      | ✅ Supported                  |
  | Google Meet     | ✅ Supported                      | ✅ Supported                  |
  | Webex           | ❌ Not supported                  | ❌ Not supported              |
</Callout>

## Hybrid diarization

Hybrid diarization uses a combination of participant information from the meeting platform and machine diarization to handle the meeting's join pattern:

* If participants are joining from their own devices, the anonymous speaker label is replaced with the real participant name from the participants list.
* If Recall detects that multiple participants are effectively coming from the same device/microphone, the anonymous speaker labels are kept (so voices can still be separated).

### Requirements to use hybrid diarization

To use hybrid diarization:

* The meeting platform must support separate audio streams.
* You must use bots to record meetings (not supported for DSDK).
* You must use async transcription (not supported for real-time transcription).
* Your selected <Anchor target="_blank" href="doc:ai-transcription">third-party transcription provider</Anchor> must support **machine diarization**.

<Callout icon="🚧" theme="warn">
  ### Caveats when using hybrid diarization

  Hybrid Diarization can provide the most accurate results, but it comes with a few tradeoffs:

  - Your application will need additional logic/code to <Anchor target="_blank" href="https://github.com/recallai/sample-apps/blob/main/bot_async_transcription_hybrid_diarization/src/convert_to_hybrid_diarized_transcript_parts.ts">map the generic speaker labels to the participants</Anchor>.
  - You will still get generic speaker labels for participants sharing a stream.
  - When enabled, you will not be able to get raw data from the third-party transcription provider for either real-time or async transcription. For debugging, you can get provider ids for a specific transcription by [asking the MCP.](https://docs.recall.ai/docs/mcp)

  There are also some cost considerations when using hybrid diarization:

  - For **async transcription**, we typically see **0.6x to 1.2x** the transcription credit usage, with the average cost difference being around **1x**. This is because we optimize the audio output by trimming out silence.
</Callout>

### Enabling hybrid diarization

To use it, you will need to implement the logic/algorithm by doing the following:

* Enable <Anchor target="_blank" href="doc:perfect-diarization">Perfect Diarization</Anchor> by setting `diarization.use_separate_streams_when_available: true` as seen in the <Anchor target="_blank" href="ref:recording_create_transcript_create">Create Async Transcript</Anchor>.
* Enable machine diarization through one of the options listed in the machine diarization section of this guide.

You will then receive transcript parts where the `participant.id` is `null` and the `participant.name` contains a temporary composite key instead of a real participant name. For example, you will receive something like:

```json
[
  {
    "participant": {
      "id": null,
      "name": "200-0",
      "is_host": false,
      "platform": "mobile_app",
      "extra_data": {...},
      "email": null
    },
    "words": [...]
  },
  // ... other transcript utterances
]
```

This temporary participant name key follows the format `{participant_id}-{anonymous_label}`.

With this, you can then:

* Fetch the list of participants via the `recording.media_shortcuts.participant_events.data.participants_download_url` field
* Build a mapping of each `participant_id` to its set of anonymous labels across all transcript parts that looks like this:

```
{
  // participant_id: anonymous_labels[]
  100: [0],
  200: [0, 1]
  // other participants
}
```

* Get the list of participant ids with only one anonymous label
* Iterate over the transcript and:
  * For participants with exactly one anonymous label, replace the anonymous speaker with the real participant name and metadata
  * For participants with multiple anonymous labels (multiple people sharing a device), leave the anonymous labels unchanged

<Callout icon="📘" theme="info">
  You can use <Anchor target="_blank" href="https://github.com/recallai/sample-apps/tree/main/bot_async_transcription_hybrid_diarization">this sample app</Anchor> to see how to get the transcript using hybrid diarization.
</Callout>

## Perfect diarization

Perfect diarization is our default and recommended diarization method. It transcribes each participant's audio stream separately instead of using a single mixed audio stream for the entire meeting. This improves speaker attribution accuracy in the transcript, especially when multiple people are speaking at the same time.

When perfect diarization is enabled, the transcript is returned with **participant speaker labels** out of the box. Perfect diarization is also supported for both **real-time** and **async** transcription.

### Requirements to use perfect diarization

To use perfect diarization:

* The meeting platform must support separate audio streams.
* You must use bots to record meetings (not supported for DSDK).

### Caveats when using perfect diarization

<Callout icon="🚧" theme="warn">
  ### Caveats when using perfect diarization

  There are a few tradeoffs to be aware of when using perfect diarization:

  - It does not distinguish between multiple people speaking from the same audio stream, such as multiple participants joining together from a conference room or shared device.
  - It does not support raw provider data for real-time or async transcription. You will not receive `transcript.provider_data` real-time transcription events and the async transcription raw provider download data in `provider_data_download_url` will always be `[]`.
    - For debugging, you can get provider ids for a specific transcription by [asking the MCP.](https://docs.recall.ai/docs/mcp)

  There are also some cost considerations when using perfect diarization:

  - For **real-time transcription&#x20;**&#x70;ricing, we typically see around **1.8x** the transcription credit usage compared to standard transcription. This is because overlapping speech across separate streams must be transcribed independently.
  - For **async transcription&#x20;**&#x70;ricing, we typically see **0.6x to 1.2x** the transcription credit usage, with the average cost difference being around **1x**. This is because we optimize the audio output by trimming out silence.
</Callout>

### Enabling perfect diarization

To enable perfect diarization, set `diarization.use_separate_streams_when_available` to `true` in the transcription configs as seen below.

#### Enabling perfect diarization for real-time transcription

To configure perfect diarization in a <Anchor target="_blank" href="ref:bot_create">Create Bot</Anchor> request, set `recording_config.transcript.diarization.use_separate_streams_when_available` to `true`:

```json
{
  // other create bot request configs
  "recording_config": {
    // other recording_config configs
    "transcript": {
      // other transcript configs
      "diarization": {
        "use_separate_streams_when_available": true
      }
    }
  }
}
```

For details on how to implement/access the transcript using real-time transcription, see:

* [Meeting Bot Real-time Transcription](https://docs.recall.ai/docs/bot-real-time-transcription)
* [Desktop Recording SDK Real-time Transcription](https://docs.recall.ai/docs/dsdk-realtime-transcription)

#### Enabling perfect diarization for async transcription

To configure perfect diarization in a <Anchor target="_blank" href="ref:recording_create_transcript_create">Create Async Transcript</Anchor> request, set `diarization.use_separate_streams_when_available` to `true`:

```json
{
  // other create async transcript request configs
  "diarization": {
    "use_separate_streams_when_available": true
  }
}
```

For details on how to implement/access the transcript using async transcription, see [Async Transcription](https://docs.recall.ai/docs/asynchronous-transcription).

## Speaker-timeline diarization

Speaker-timeline diarization uses active speaker events from the meeting platform to assign transcript text to participants. Instead of transcribing separate audio streams, it relies on the meeting platform to indicate who the active speaker is over time.

When speaker-timeline diarization is used, the transcript is returned with **participant speaker labels**. This works best when each participant is joining from their own device, because the meeting platform can more reliably associate active speaker events with a specific participant.

### Requirements to use speaker-timeline diarization

To use speaker-timeline diarization:

* The meeting platform must support **active speaker events** (Google Meet, Microsoft Teams, Zoom support it out of the box; Webex meeting hosts must have a paid account and have closed captions enabled).

<Callout icon="🚧" theme="warn">
  ### Caveats when using speaker-timeline diarization

  There are a few tradeoffs to be aware of when using speaker-timeline diarization:

  - It depends on meeting-platform active speaker events, which can sometimes be inaccurate or incomplete (e.g. noise, missed speaker changes, or overlapping speech).
  - It is not a good fit when multiple people are speaking from the same device or microphone, such as a conference room setup.

  Speaker-timeline diarization can result in less accurate diarization given its dependency on the meeting provider and the limitations listed above. You can enable Perfect Diarization if available to fix these issues.

  It is typically most useful when you need both participant speaker labels and raw data from the third-party transcription provider.
</Callout>

### Enabling speaker-timeline diarization

Speaker-timeline diarization is used whenever machine diarization is not enabled and `diarization.use_separate_streams_when_available` is not set to `true`.

#### Enabling speaker-timeline diarization for real-time transcription

To configure speaker-timeline diarization in a <Anchor target="_blank" href="ref:bot_create">Create Bot</Anchor> request, ensure that machine diarization isn't enabled and set `recording_config.transcript.diarization.use_separate_streams_when_available` to `false`:

```json
{
  // other create bot request configs
  "recording_config": {
    // other recording_config configs
    "transcript": {
      // other transcript configs
      "diarization": {
        "use_separate_streams_when_available": false
      }
    }
  }
}
```

For details on how to implement/access the transcript using real-time transcription, see:

* [Meeting Bot Real-time Transcription](https://docs.recall.ai/docs/bot-real-time-transcription)
* [Desktop Recording SDK Real-time Transcription](https://docs.recall.ai/docs/dsdk-realtime-transcription)

#### Enabling speaker-timeline diarization for async transcription

To configure speaker-timeline diarization in a <Anchor target="_blank" href="ref:recording_create_transcript_create">Create Async Transcript</Anchor> request, ensure that machine diarization isn't enabled and set `diarization.use_separate_streams_when_available` to `false`:

```json
{
  // other create async transcript request configs
  "diarization": {
    "use_separate_streams_when_available": false
  }
}
```

For details on how to implement/access the transcript using async transcription, see [Async Transcription](https://docs.recall.ai/docs/asynchronous-transcription).

## Machine diarization

Machine diarization is produced by your **third-party transcription provider**, not Recall. Instead of using meeting-platform participant information, the transcription provider separates speakers based on voice characteristics and returns a diarized transcript to Recall.

When machine diarization is enabled, the transcription provider separates speakers using generic speaker labels such as `A`, `B`, `C` or `0`, `1`, `2`, rather than meeting participant identities (i.e. "Jake", "Conner").

### Machine diarization async

For async transcription, speaker names in the transcript are replaced by generic speaker labels.  That means the transcription will look like this:

* `A` said "X"
* `B` said "Y"

### Machine diarization real-time

For real-time transcription, the generic speaker labels from the transcription provider are available in `transcript.provider_data`.

The `transcript.data` event returns Recall’s normalized transcript output with speaker names.  These will contain the speaker names pulled from the meeting platform.  To retrieve the generic speaker labels, listen for `transcript.provider_data` webhook event.

### Requirements to use machine diarization

To use machine diarization:

* Your selected <Anchor target="_blank" href="doc:ai-transcription">third-party transcription provider</Anchor> must support **machine diarization**.
* Your selected <Anchor target="_blank" href="doc:ai-transcription">third-party transcription provider</Anchor> must support your chosen transcription mode, whether **real-time** or **async**.

<Callout icon="🚧" theme="warn">
  ### Caveats when using machine diarization

  There are a few tradeoffs to be aware of when using machine diarization:

  - The transcript uses **generic speaker labels**, not **participant speaker labels**, because the transcription provider does not know which meeting participant each voice belongs to.
  - Accuracy can vary by provider.
  - It can be less accurate when different speakers have similar-sounding voices.
</Callout>

### Enabling machine diarization

To enable machine diarization, set the provider-specific diarization field in your transcription provider configuration.

| Provider     | Real-time                                        | Async                                    |
| :----------- | :----------------------------------------------- | :--------------------------------------- |
| Deepgram     | `deepgram_streaming.diarize: true`               | `deepgram_async.diarize: true`           |
| ElevenLabs   | -                                                | `elevenlabs_async.diarize`               |
| Assembly     | `assembly_ai_async_chunked.speaker_labels: true` | `assembly_ai_async.speaker_labels: true` |
| Rev          | `rev_streaming.enable_speaker_switch: true`      | -                                        |
| Speechmatics | `speechmatics_streaming.diarization: "speaker"`  | -                                        |

#### Enabling machine diarization for real-time transcription

To configure machine diarization in a <Anchor target="_blank" href="ref:bot_create">Create Bot</Anchor> request, set the provider configs as seen above and `recording_config.transcript.diarization.use_separate_streams_when_available` to `false`. An example with Deepgram would look like:

```json
{
  // other create bot request configs
  "recording_config": {
    // other recording_config configs
    "transcript": {
      // other transcript configs
      "diarization": {
        "use_separate_streams_when_available": false
      },
      "provider": {
        "deepgram_streaming": {
          "diarize": true
        }
      }
    }
  } 
}
```

For details on how to implement/access the transcript using real-time transcription, see:

* [Meeting Bot Real-time Transcription](https://docs.recall.ai/docs/bot-real-time-transcription)
* [Desktop Recording SDK Real-time Transcription](https://docs.recall.ai/docs/dsdk-realtime-transcription)

#### Enabling machine diarization for async transcription

To configure machine diarization in a <Anchor target="_blank" href="ref:recording_create_transcript_create">Create Async Transcript</Anchor> request, set the provider configs as seen above and `diarization.use_separate_streams_when_available` to `false`. An example with Deepgram would look like:

```json
{
  // other create async transcript request configs
  "diarization": {
    "use_separate_streams_when_available": false
  },
  "provider": {
    "deepgram_async": {
      "diarize": true
    }
  }
}
```

For details on how to implement/access the transcript using async transcription, see [Async Transcription](https://docs.recall.ai/docs/asynchronous-transcription).

<br />

# FAQ

## Why am I seeing `Speaker A`, `Speaker B`, or `0`, `1`, `2` instead of names?

For async transcription, this indicates **Machine Diarization** was used via a third-party transcription provider. Machine diarization can separate voices, but the provider's labels are generic and are not tied to meeting participants.

To get participant names, remove provider diarization flags such as:

* `assembly_ai_async.speaker_labels`
* `deepgram_async.diarize`

## Why do multiple speakers calling from the same device appear as the same participant in the transcript?

This usually happens when multiple people are sharing one device or microphone and **speaker-timeline diarization** is being used (or machine diarization is not enabled or not available).

For conference rooms or other shared-device setups, use **machine diarization** when you only have a mixed audio stream, or **hybrid diarization** when separate audio streams are available and some streams may contain multiple speakers.

## Microsoft Teams: why are speaker names missing or diarization looks wrong?

Teams has a setting that affects whether speakers can be identified in captions/transcripts: **Transcription Caption Identification**. If this setting is turned off, transcripts will not get diarized properly with multiple speakers.

Where to find this setting in Teams: `Accessibility` -> `Captions and Transcripts` -> `Transcription` -> `Automatically identify me in meeting captions and transcripts`

Be aware that org-wide Teams policies can override individual user settings.

<Embed title="" typeOfEmbed="iframe" url="https://www.loom.com/embed/ffaf35d666164cc59d96704234043f7a?sid=701cb0f2-6f3f-403b-b04b-404d66918f84" height="300px" width="100%" href="https://www.loom.com/embed/ffaf35d666164cc59d96704234043f7a?sid=701cb0f2-6f3f-403b-b04b-404d66918f84" html="false" />

## Is there any additional costs for diarization?

There is no separate diarization feature fee. However, some diarization methods, such as perfect diarization and hybrid diarization, can increase transcription credit usage depending on how audio is processed. See those sections for more details.