fix: keep non-streaming TTS conditioning within the current turn

#22
by Movix - opened

Summary

Fix the non-streaming TTS edge cases described in OpenBMB/MiniCPM-V#1141.

  • Select the first TTS EOS after the last TTS BOS, so a completed turn in the history cannot close a truncated current reply.
  • Slice token IDs and hidden states to the same available end before merging their embeddings. The last sampled token in a truncated reply has not been forwarded through the LLM, so it is excluded from speech conditioning until a corresponding hidden row exists.

Completed turns keep their full pre-EOS span. The teacher-forcing terminal-EOS exclusion, no-BOS fallback, stream path, and existing first-sample batch behavior are preserved. No extra LLM forward pass or generation-cache mutation is introduced.

Validation

python3 -m pytest tests/test_tts_boundaries.py -q

The CPU regressions execute the complete chat and _generate_speech_non_streaming method bodies extracted from the checked-in source, with model generation and waveform backends stubbed. They cover history/current-turn boundaries, multiple EOS markers, completed and truncated turns, teacher forcing, missing BOS, empty consumed spans, projected-hidden normalization, and the chat-to-TTS handoff.

The same 25 cases fail on the original source (11 failed, 14 passed) and pass on the modified source. Removing either fix independently makes the suite fail again. Real pretrained-model inference and acoustic-quality validation have not been run.

Related work: open HF PR #20 changes nearby batch-chat code but leaves this EOS selector unchanged. This change is scoped to non-streaming TTS boundaries and consumed token/hidden-state alignment.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment