Nemotron 3 Diarization β€” Core AI

Core AI (.aimodel) conversion of NVIDIA's Nemotron 3 Diarization streaming speaker-diarization model (up to eight speakers, 10 ms output frames) for iOS 27+ / macOS 27+. Converted from checkpoint revision a435e9867d79e789e90053f9b6d6834053af564a (Nemotron-3-Diarization.nemo, SHA-256 867c53f552998f772e5b5e5c082962ae85ee7ca5669c2bc17d7f615133d4e96d) with coreai-torch 0.4.1 (standard graphs) and 0.4.2 (Neural Engine encoder), coreai-core 1.0.0b2. No weights were retrained, quantized or pruned.

The model is split into two graphs that a host streaming loop calls once per chunk: a preencoder (Mel features β†’ 80 ms frame embeddings) and an encoder/head (packed speaker cache + FIFO + chunk embeddings β†’ speaker probabilities). The Arrival-Order Speaker Cache, FIFO and chunking logic of Streaming Sortformer run on the host; see "Streaming loop" below.

Contents

Path Description
low-latency/pipeline.json Live: FP32 preencoder + FP16 Neural Engine encoder/head (encoderComputePreference: neuralEngine)
low-latency/models/preencoder_fp32.aimodel Live preencoder
low-latency/models/encoder_head_fp16_ane.aimodel Encoder/head, Neural Engine layout (channel-first BC1S, per-head attention, float valid mask)
low-latency/frontend/ Mel frontend parameters, Hann window and Mel filterbank
low-latency/learned_silence_embedding.f32le Learned silence embedding (512 Γ— float32) used to pad the speaker cache
offline/pipeline.json Published offline profile: FP32 preencoder + FP16 Neural Engine encoder/head
offline/models/ Offline preencoder_fp32.aimodel and encoder_head_fp16_ane.aimodel
offline/frontend/, offline/learned_silence_embedding.f32le Offline frontend and cache-padding assets
SHA256SUMS SHA-256 of every file

.f32le files are raw little-endian float32 in C order. Load the Neural Engine encoder with SpecializationOptions(preferredComputeUnitKind: .neuralEngine); its first specialization takes about 35 s and is then cached by Core AI. Ahead-of-time compilation (xcrun coreai-build compile --preferred-compute neural-engine) is optional and device-architecture specific, so no compiled bundle is included.

Streaming profile

Graphs under low-latency/ are fixed-shape for the model card's "Low latency" (1.04 s) configuration, applied synchronously (cache update inside each step):

Parameter Value
Chunk length 9 encoder frames (0.72 s)
Left / right context 1 / 4 encoder frames (0.08 / 0.32 s)
Speaker cache / FIFO 264 / 264 frames
Speaker-cache update period 222 frames
Encoder frame 8 Mel frames (80 ms)
Output resolution 10 ms (8Γ— upsampled), 8 speakers

For file analysis, load offline/pipeline.json. Its separate graphs use the published offline profile: central 340, left 1, right 40, cache 264, FIFO 40, update period 300, and packed capacity 685. Preencoder input is [1,3048,128], output [1,381,512]; encoder input is [1,685,512], validity mask [1,685], native output [1,5480,8], and cache output [1,685,8]. Names and types match the live interface below. Resolve paths relative to the selected manifest; never use the live shapes with offline graphs.

Only these two production pipelines are included. Alternate precision manifests, unused model variants and validation reports are omitted. Keep each selected model directory intact; SHA256SUMS covers both profiles.

Frontend

16 kHz mono float32 audio β†’ pre-emphasis 0.97 β†’ centered STFT (FFT 512, Hann window 400, hop 160, constant zero padding) β†’ power spectrum β†’ 128 Slaney Mel bins β†’ log(x + 2^-24). No dither, no feature normalization. floor(samples / 160) valid frames. The window and filterbank are the checkpoint's own buffers, which NVIDIA stores in bfloat16; they differ slightly from freshly computed float32 values. Project each frame's power spectrum onto the filterbank separately (a batched matrix product changes the float rounding and makes long streams diverge).

Graph interfaces

Live preencoder

Name Shape Type
in mel_features [1, 112, 128] time-major (left 8 + central 72 + right 32 frames, zero-padded) float32
in mel_length [1] valid frames int32
out chunk_embeddings [1, 14, 512] float32
out chunk_embedding_length [1] = ceil(mel_length / 8) int32

Live encoder/head, Neural Engine layout (encoder_head_fp16_ane.aimodel)

Name Shape Type
in packed_embeddings [1, 542, 512]: speaker cache, FIFO, chunk embeddings, zero-padded float16
in valid_mask [1, 542]: 1 for the first packed_length rows, 0 after float16
out native_probabilities [1, 4336, 8] float16
out cache_probabilities [1, 542, 8] float16

Probabilities are sigmoid outputs; the model card's default decision threshold is 0.5.

Streaming loop

Per chunk: slice left/central/right Mel frames, run the preencoder, pack [speaker cache | FIFO | chunk embeddings] up to 542 rows, run the encoder/head, emit the central chunk's rows of native_probabilities, then update the FIFO and speaker cache from cache_probabilities following NVIDIA NeMo's Streaming Sortformer (Arrival-Order Speaker Cache compression every 222 frames, learned silence embedding for padding). On the final call, flush the remaining chunks with zero right context and emit only rows up to the true end of the audio.

Validation (live profile; includes reference variants not distributed here)

Against the original NeMo FP32 model with the same streaming profile:

Graph PSNR vs FP32 reference
Preencoder FP32 / FP16 exact / 81.4 dB
Encoder/head FP32 (native / cache) 164.6 / 171.5 dB
Encoder/head FP16 (native / cache) 99.4 / 103.4 dB
Neural Engine encoder FP16 (native / cache) 100.8 / 104.7 dB
  • FP32 closed-loop 30 s stream: strict agreement (absolute 2e-6 + relative 2e-5), identical segments, DER 4.52% (zero collar, overlap included).
  • FP16 closed loop: DER identical to FP32; a few segment boundaries move by 10 ms.
  • Recommended pipeline: on iPhone 17 Pro (iOS 27.2) the 30 s sample is analyzed in 1.36 s (about 22Γ— real time); on a Mac, a 98.6 s six-speaker recording matched NeMo FP32 with mean absolute probability difference 0.0004.

These are checks against the reference model on a few recordings, not a new accuracy benchmark; see NVIDIA's model card for accuracy.

Offline validation

On iPhone 17 Pro, the offline pipeline completed 30 s and 98.64 s recordings with 3,000 and 9,864 output frames. Speaker activity differed from the released FP32 offline reference on one and two 10 ms frames, respectively. The 30 s sample DER was 2.3819% (zero collar, overlap included). Mixed precision does not pass the strict FP32 probability tolerance. These sample checks do not establish corpus-wide accuracy. First model loading took about 53 s; subsequent loading took less than a second.

License

Use is governed by the OpenMDW License Agreement, version 1.1 (see LICENSE), the license of the original model. See NOTICE for origin.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for smdesai/Nemotron-3-Diarization-CoreAI

Finetuned
(9)
this model