> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pyannote.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Models

> Choose the right speaker diarization model for your audio processing needs

<CardGroup cols={2}>
  <Card title="Precision-3" icon="star">
    <Badge color="green" shape="pill" icon="sparkles">Latest</Badge>

    Our most accurate diarization model, with tunable detection and frame-level probability scores.
  </Card>

  <Card title="Precision-2" icon="star-half-stroke">
    Previous-generation high-accuracy diarization, and the current default.
  </Card>

  <Card title="Live-1" icon="signal-stream">
    Streaming diarization over WebSocket for live audio.
  </Card>

  <Card title="Community-1" icon="users">
    Open-source diarization for research, prototyping, and self-hosting.
  </Card>
</CardGroup>

## Precision-3 migration

Precision-3 is available as an opt-in model. Set `model: "precision-3"` in your request to use it.

<Warning>
  If your integration reads the `confidence` output, it needs updating. This is covered in the checklist below.
</Warning>

### Model migration

If you explicitly set `model: "precision-2"` in your requests, update this value to `"precision-3"`.

### Confidence output deprecation

The `confidence` output is deprecated. It is not renamed to a single equivalent: it is replaced by three separate scores, each answering a narrower question.

| Deprecated | Replacement | Answers |
| - | - | - |
| `confidence` | `speakerProbability` | Is this speaker active on this frame? |
| `confidence` | `speechProbability` | Is speech present at all? |
| `confidence` | `crosstalkProbability` | Is there overlap? |

If your integration reads `confidence`, update it to read whichever of the three matches what you actually need. `turnLevelConfidence` has not changed.

### Migration checklist

<Steps>
  <Step title="Replace or remove explicit model references">
    Replace any explicit `model: "precision-2"` references with `"precision-3"`.
  </Step>

  <Step title="Replace confidence reads">
    Replace reads of `confidence` with `speakerProbability`, `speechProbability`, or `crosstalkProbability`.
  </Step>

  <Step title="Confirm turn-level confidence">
    Confirm `turnLevelConfidence` usage still works as expected. No change is needed here.
  </Step>

  <Step title="Test a sample of your traffic">
    Test a sample of your traffic against Precision-3, particularly if your audio is closer to clean read speech, telephone, or broadcast material.
  </Step>
</Steps>

***

## Choosing the right model

### Precision-3

**Best for:** Teams who want the highest diarization accuracy available, plus control over detection behaviour and access to frame-level probability scores.

Precision-3 is 34.2% more accurate, on average, than Community-1.

<Tip>
  Self-hosted options for Precision-3 are available on Enterprise plans. Self-hosted deployments also choose an operating mode per deployment — accuracy (14.3 average DER), balance (14.9) or speed (15.5). Operating modes are not available on the API.
</Tip>

**Typical use cases:** phone call analytics, meeting transcription with speaker attribution, video dubbing and timestamp-critical workflows, building training data for voice assistants, and more.

Precision-3 provides optimization controls for workflows that need to tune detection to unusual audio conditions, or that consume probability curves directly: custom reconciliation with STT, and training data preparation pipelines.

**Advanced features:**

* **Speaker identification with voiceprints**: identify known speakers in your audio using pre-enrolled voiceprints
* **Exclusive diarization mode:** returns diarization where only one speaker is active at a time, making STT reconciliation easier
* **Flexible speaker count control:** set `minSpeakers`, `maxSpeakers` and `numSpeakers` for any number of speakers
* **Input controls:** `vadSensitivity` and `crosstalkSensitivity` trade precision against recall for voice activity and overlapping speech
* **Frame-level probability scores:** `speechProbability`, `crosstalkProbability` and `speakerProbability` sent every 20ms

[Learn more about Precision-3](https://www.pyannote.ai/blog/precision-3)<Icon icon="arrow-up-right" iconType="solid" />

***

### Precision-2

**Best for:** Existing integrations that have not yet moved to Precision-3.

Precision-2 is 28% more accurate, on average, than Community-1. It remains the default model.

<Tip>
  Self-hosted options for Precision-2 are available on Enterprise plans.
</Tip>

**Typical use cases:** phone call analytics, meeting transcription with speaker attribution, video dubbing and timestamp-critical workflows, building training data for voice assistants, and more.

**Advanced features:**

* **Speaker identification with voiceprints**: Identify known speakers in your audio using pre-enrolled voiceprints
* **Exclusive diarization mode:** Returns speaker diarization where only one single speaker (the most likely to be transcribed) is active at a time, making STT reconciliation easier
* **Flexible speaker count control:** Set `minSpeakers`, `maxSpeakers` and `numSpeakers` parameters for any number of speakers
* **Human-in-the-loop correction:** Use confidence scores to help streamline manual correction processes

[Learn more about Precision-2 ](https://www.pyannote.ai/blog/precision-2)<Icon icon="arrow-up-right" iconType="solid" />

***

### Live-1

**Best for:** Teams building live voice products that need speaker labels before a recording finishes: contact centers, meeting tools, broadcast workflows, and real-time voice agents.

**Typical use cases:** live meeting assistants, contact center agent assist, live captioning with speaker attribution, broadcast speaker attribution, and multi-party voice agents that need to track who is speaking.

**Advanced features:**

* **Sub-300ms latency:** Speaker labels arrive fast enough for live captioning and real-time agent assist, tested against noisy, real-world audio rather than clean studio recordings.
* **Native streaming architecture:** Processes audio in 100ms chunks over WebSocket, with a speaker tracking layer that holds consistency across the stream without needing the full conversation.
* **Built for real conditions:** Trained and validated on overlapping speech, background noise, and conversations with more than two participants.

**Technical specs:**

* Up to 8 speakers, up to 5 hours duration per stream
* Input: 16 kHz mono audio, 100ms chunks over WebSocket
* Output: `diarization_speaker_start` and `diarization_speaker_end` events, each with a start or end timestamp and speaker label

[Learn more about Live-1](https://www.pyannote.ai/blog/how-we-built-streaming-diarization)<Icon icon="arrow-up-right" iconType="solid" />

***

### Community-1 (hosted)

**Best for:** Teams who want the open-source model without managing infrastructure

**Typical use cases:** Prototyping, low-volume production workloads, testing and validation

**Key Benefits:**

* **Cost efficiency:** hosted at cost, ideal for experimentation and low-volume workloads
* **No infrastructure management:** Focus on your application while we handle the deployment
* **Easy migration:** Start with hosted Community-1 and upgrade to Precision-3 when needed
* **Same powerful model:** Access the same Community-1 model through our API without setup complexity

[Learn more about Community-1 ](https://www.pyannote.ai/blog/community-1)<Icon icon="arrow-up-right" iconType="solid" />

***

### Community-1 (self-hosted with pyannote.audio 4.0)

**Best for:** Researchers, developers, and personal hobby projects who want full control over their diarization models and workflows.

**Typical use cases:** Academic work, product-iteration, prototyping, and custom diarization deployment (e.g., dataset-specific fine-tuning or custom reconciliation with STT).

**Key Benefits:**

* **Best open-source speaker diarization model available** - outperforms pyannote.audio 3.1 across all key metrics
* **Open-source flexibility:** Full transparency into model weights and code allowing local and offline training and inference.

**Trade-offs:**

* Lower accuracy compared to Precision-2 and Precision-3
* No support for advanced features like speaker identification and voiceprints
* Requires deploying the model on your own infrastructure

[Learn more about pyannote.audio 4.0 ](https://github.com/pyannote/pyannote-audio)<Icon icon="arrow-up-right" iconType="solid" />

***

## How to specify a model in diarization requests

When making a diarization request, you can specify which model to use using the `model` parameter:

```bash theme={null}
curl -X POST "https://api.pyannote.ai/v1/diarize" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://files.pyannote.ai/marklex1min.wav",
    "model": "precision-3"
  }'
```

By default, if you do not specify a model, the API will use the Precision-2 model. Set `model` explicitly to use Precision-3.

<Note>
  Precision-3 is opt-in. See [Precision-3 migration](#precision-3-migration) for details.
</Note>

### Switch between models

You can easily switch between models by changing the `model` parameter:

* `"model": "precision-3"` for Precision-3
* `"model": "precision-2"` for Precision-2 (default)
* `"model": "community-1"` for Community-1

<Note>
  **Note:** Speaker identification and voiceprint features are not available for Community-1 models. These advanced features are exclusive to Precision-2 and Precision-3.
</Note>

### Compare results between models

To compare performance between models on your specific data:

1. Process the same audio file with both models
2. Compare the diarization results
3. Evaluate which model provides better accuracy for your use case

## Pricing

For detailed pricing information, visit our [pricing page](https://www.pyannote.ai/pricing).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.