Streaming microphone transcription emits Unicode replacement characters for Chinese text

#49
by zhengsihua - opened

Environment

  • macOS 27.0 on Apple M4 Pro
  • Python 3.13.5
  • voxmlx==0.0.2
  • mlx==0.32.0
  • Model: mlx-community/Voxtral-Mini-4B-Realtime-6bit

Description

Chinese microphone transcription sometimes prints Unicode replacement characters (�). For example:

你好竹子��

Transcribing an audio file does not show the same problem.

Steps to reproduce

voxmlx

Speak Mandarin into the microphone and watch the incremental output.

Expected behavior

Complete Chinese characters should be printed without replacement characters.

Actual behavior

One or more � characters occasionally appear in the live transcript.

Suspected cause

stream.py decodes and prints each token separately:

text = sp.decode(
    [token_id], special_token_policy=SpecialTokenPolicy.IGNORE
)
print(text, end="", flush=True)

Some Chinese characters span multiple byte-level tokens. Decoding an incomplete token sequence produces U+FFFD. The file transcription path collects all output token IDs and decodes them together, which explains why it does not show the issue.
Buffering token IDs until they form a valid decoded sequence should prevent the replacement characters while preserving streaming output.

Your diagnosis looks right. Decoding one token at a time cannot work with byte-level pieces, because a single CJK character is spread across several tokens and each partial decode produces a replacement character. That also explains why file transcription is clean: the whole sequence gets decoded in one call.

The usual fix is to keep a pending buffer rather than printing per token. Accumulate the ids, decode the accumulated sequence on each step, and emit only the delta against what you have already printed. If you would rather stay at byte level, hold back any trailing bytes that form an incomplete UTF-8 sequence and flush them once the next token completes the character.

Sign up or log in to comment