Index-Echo-S2TT-2B-FP8

Official FP8 (W8A8) quantization of IndexTeam/Index-Echo-S2TT-2B, part of the Index-Echo speech-to-text translation (S2TT) model family by bilibili.

This repository mirrors the original checkpoint layout; only the text LLM backbone (llm/) is quantized to FP8 - the audio tower, connector, and all speech-synthesis components remain in BF16. Use it exactly like the original repo (same infer.py / configs).

Quantization

  • Scheme: FP8_DYNAMIC (FP8 E4M3 weights, per-token dynamic FP8 activations), produced with llm-compressor (quantization_scheme recorded in recipe.yaml).
  • All Linear layers of the language model backbone (llm/) are quantized; the audio tower, connector, lm_head, embeddings and all other pipeline components are kept in BF16.
  • Format: compressed-tensors safetensors - load directly with vLLM (quantization="compressed-tensors") or transformers.

Consistency validation

Measured on an NVIDIA A100 against the original BF16 checkpoint (greedy decoding, official translation prompt):

Metric BF16 FP8 Delta
Perplexity (fixed corpus) 4.8757 4.9627 +1.79%
zh->en generation identical - - yes
en->zh generation identical - - no (semantically equivalent)

Usage

Identical to the original checkpoint - clone this repo and follow the README / infer.py of the base model (IndexTeam/Index-Echo-S2TT-2B). The quantized LLM backbone loads via compressed-tensors; make sure compressed-tensors (or vLLM / a recent transformers) is installed.

Quantized and published by the Index team, 2026-10-03.

Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IndexTeam/Index-Echo-S2TT-2B-FP8

Finetuned
(2)
this model

Collection including IndexTeam/Index-Echo-S2TT-2B-FP8