Open-source research project · Speech tokenization

BiMTokenizer: Preserving Semantic-Acoustic Balance in Low-Bitrate Speech Tokenization via Bidirectional State-Space Modeling

Xin Zhang · Lin Li · Chuanbo Liu · Jianquan Liu · Kong Aik Lee


01Architecture

A single-tower acoustic codec trained end-to-end with discriminators, supervised additionally by a frozen semantic encoder (Whisper / SenseVoice) via reconstruction alignment and cosine-similarity objectives.

BiMTokenizer architecture diagram
Figure 1. Overview of BiMTokenizer. The acoustic codec encodes a mel-spectrogram via stacked Mamba blocks into latent zc, quantizes it through Residual Spherical Leech Quantization, and decodes through a Vocos head. During training, a frozen semantic encoder enforces alignment between the original and reconstructed speech in the semantic space (dashed paths are training-only).

02Reconstruction · LibriSpeech test-clean

Each column is the same utterance reconstructed by a different codec. Bitrates are reported next to each method. Our variants are highlighted in red.

Method Sample 1Sample 2Sample 3Sample 4
Ground Truth
EnCodec 1500 bps
DAC 1500 bps
SpeechTokenizer 1000 bps
BigCodec 1040 bps
Mimi 1100 bps
XCodec 1000 bps
XCodec2.0 800 bps
DualCodec 1075 bps
XY-Tokenizer 1000 bps
SimWhisper-Codec 1100 bps
BiMTokenizer-Whisper 1100 bps
BiMTokenizer-SenseVoice 1100 bps

All baselines reproduced using official pre-trained checkpoints at their respective bitrates.

03Reconstruction · LibriSpeech test-other

The noisier test-other set, drawn from speakers and recording conditions held out from training.

Method Sample 1Sample 2Sample 3Sample 4
Ground Truth
EnCodec 1500 bps
DAC 1500 bps
SpeechTokenizer 1000 bps
BigCodec 1040 bps
Mimi 1100 bps
DualCodec 1075 bps
XY-Tokenizer 1000 bps
BiMTokenizer-Whisper 1100 bps
BiMTokenizer-SenseVoice 1100 bps