Nemotron 3 Diarization β Core AI
Core AI (.aimodel) conversion of NVIDIA's
Nemotron 3 Diarization
streaming speaker-diarization model (up to eight speakers, 10 ms output frames)
for iOS 27+ / macOS 27+. Converted from checkpoint revision
a435e9867d79e789e90053f9b6d6834053af564a
(Nemotron-3-Diarization.nemo, SHA-256
867c53f552998f772e5b5e5c082962ae85ee7ca5669c2bc17d7f615133d4e96d) with
coreai-torch 0.4.1 (standard graphs) and 0.4.2 (Neural Engine encoder),
coreai-core 1.0.0b2. No weights were retrained, quantized or pruned.
The model is split into two graphs that a host streaming loop calls once per chunk: a preencoder (Mel features β 80 ms frame embeddings) and an encoder/head (packed speaker cache + FIFO + chunk embeddings β speaker probabilities). The Arrival-Order Speaker Cache, FIFO and chunking logic of Streaming Sortformer run on the host; see "Streaming loop" below.
Contents
| Path | Description |
|---|---|
low-latency/pipeline.json |
Live: FP32 preencoder + FP16 Neural Engine encoder/head (encoderComputePreference: neuralEngine) |
low-latency/models/preencoder_fp32.aimodel |
Live preencoder |
low-latency/models/encoder_head_fp16_ane.aimodel |
Encoder/head, Neural Engine layout (channel-first BC1S, per-head attention, float valid mask) |
low-latency/frontend/ |
Mel frontend parameters, Hann window and Mel filterbank |
low-latency/learned_silence_embedding.f32le |
Learned silence embedding (512 Γ float32) used to pad the speaker cache |
offline/pipeline.json |
Published offline profile: FP32 preencoder + FP16 Neural Engine encoder/head |
offline/models/ |
Offline preencoder_fp32.aimodel and encoder_head_fp16_ane.aimodel |
offline/frontend/, offline/learned_silence_embedding.f32le |
Offline frontend and cache-padding assets |
SHA256SUMS |
SHA-256 of every file |
.f32le files are raw little-endian float32 in C order. Load the Neural Engine
encoder with SpecializationOptions(preferredComputeUnitKind: .neuralEngine);
its first specialization takes about 35 s and is then cached by Core AI.
Ahead-of-time compilation (xcrun coreai-build compile --preferred-compute neural-engine) is optional and device-architecture specific, so no compiled
bundle is included.
Streaming profile
Graphs under low-latency/ are fixed-shape for the model card's "Low latency" (1.04 s)
configuration, applied synchronously (cache update inside each step):
| Parameter | Value |
|---|---|
| Chunk length | 9 encoder frames (0.72 s) |
| Left / right context | 1 / 4 encoder frames (0.08 / 0.32 s) |
| Speaker cache / FIFO | 264 / 264 frames |
| Speaker-cache update period | 222 frames |
| Encoder frame | 8 Mel frames (80 ms) |
| Output resolution | 10 ms (8Γ upsampled), 8 speakers |
For file analysis, load offline/pipeline.json. Its separate graphs use the
published offline profile: central 340, left 1, right 40, cache 264, FIFO 40,
update period 300, and packed capacity 685. Preencoder input is
[1,3048,128], output [1,381,512]; encoder input is [1,685,512], validity
mask [1,685], native output [1,5480,8], and cache output [1,685,8].
Names and types match the live interface below. Resolve paths relative to
the selected manifest; never use the live shapes with offline graphs.
Only these two production pipelines are included. Alternate precision
manifests, unused model variants and validation reports are omitted. Keep
each selected model directory intact; SHA256SUMS covers both profiles.
Frontend
16 kHz mono float32 audio β pre-emphasis 0.97 β centered STFT (FFT 512,
Hann window 400, hop 160, constant zero padding) β power spectrum β 128
Slaney Mel bins β log(x + 2^-24). No dither, no feature normalization.
floor(samples / 160) valid frames. The window and filterbank are the
checkpoint's own buffers, which NVIDIA stores in bfloat16; they differ
slightly from freshly computed float32 values. Project each frame's power
spectrum onto the filterbank separately (a batched matrix product changes the
float rounding and makes long streams diverge).
Graph interfaces
Live preencoder
| Name | Shape | Type |
|---|---|---|
in mel_features |
[1, 112, 128] time-major (left 8 + central 72 + right 32 frames, zero-padded) | float32 |
in mel_length |
[1] valid frames | int32 |
out chunk_embeddings |
[1, 14, 512] | float32 |
out chunk_embedding_length |
[1] = ceil(mel_length / 8) | int32 |
Live encoder/head, Neural Engine layout (encoder_head_fp16_ane.aimodel)
| Name | Shape | Type |
|---|---|---|
in packed_embeddings |
[1, 542, 512]: speaker cache, FIFO, chunk embeddings, zero-padded | float16 |
in valid_mask |
[1, 542]: 1 for the first packed_length rows, 0 after |
float16 |
out native_probabilities |
[1, 4336, 8] | float16 |
out cache_probabilities |
[1, 542, 8] | float16 |
Probabilities are sigmoid outputs; the model card's default decision threshold is 0.5.
Streaming loop
Per chunk: slice left/central/right Mel frames, run the preencoder, pack
[speaker cache | FIFO | chunk embeddings] up to 542 rows, run the encoder/head,
emit the central chunk's rows of native_probabilities, then update the FIFO
and speaker cache from cache_probabilities following NVIDIA NeMo's Streaming
Sortformer (Arrival-Order Speaker Cache compression every 222 frames, learned
silence embedding for padding). On the final call, flush the remaining chunks
with zero right context and emit only rows up to the true end of the audio.
Validation (live profile; includes reference variants not distributed here)
Against the original NeMo FP32 model with the same streaming profile:
| Graph | PSNR vs FP32 reference |
|---|---|
| Preencoder FP32 / FP16 | exact / 81.4 dB |
| Encoder/head FP32 (native / cache) | 164.6 / 171.5 dB |
| Encoder/head FP16 (native / cache) | 99.4 / 103.4 dB |
| Neural Engine encoder FP16 (native / cache) | 100.8 / 104.7 dB |
- FP32 closed-loop 30 s stream: strict agreement (absolute 2e-6 + relative 2e-5), identical segments, DER 4.52% (zero collar, overlap included).
- FP16 closed loop: DER identical to FP32; a few segment boundaries move by 10 ms.
- Recommended pipeline: on iPhone 17 Pro (iOS 27.2) the 30 s sample is analyzed in 1.36 s (about 22Γ real time); on a Mac, a 98.6 s six-speaker recording matched NeMo FP32 with mean absolute probability difference 0.0004.
These are checks against the reference model on a few recordings, not a new accuracy benchmark; see NVIDIA's model card for accuracy.
Offline validation
On iPhone 17 Pro, the offline pipeline completed 30 s and 98.64 s recordings with 3,000 and 9,864 output frames. Speaker activity differed from the released FP32 offline reference on one and two 10 ms frames, respectively. The 30 s sample DER was 2.3819% (zero collar, overlap included). Mixed precision does not pass the strict FP32 probability tolerance. These sample checks do not establish corpus-wide accuracy. First model loading took about 53 s; subsequent loading took less than a second.
License
Use is governed by the OpenMDW License Agreement, version 1.1
(see LICENSE), the license of the original model. See NOTICE for origin.
Model tree for smdesai/Nemotron-3-Diarization-CoreAI
Base model
nvidia/Nemotron-3-Diarization