Instructions to use IndexTeam/Index-Echo-S2TT-2B-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IndexTeam/Index-Echo-S2TT-2B-FP8 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("translation", model="IndexTeam/Index-Echo-S2TT-2B-FP8")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("IndexTeam/Index-Echo-S2TT-2B-FP8", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Index-Echo-S2TT-2B-FP8
Official FP8 (W8A8) quantization of IndexTeam/Index-Echo-S2TT-2B, part of the Index-Echo speech-to-text translation (S2TT) model family by bilibili.
This repository mirrors the original checkpoint layout; only the text LLM backbone (llm/) is quantized to FP8 - the audio tower, connector, and all speech-synthesis components remain in BF16. Use it exactly like the original repo (same infer.py / configs).
Quantization
- Scheme:
FP8_DYNAMIC(FP8 E4M3 weights, per-token dynamic FP8 activations), produced with llm-compressor (quantization_schemerecorded inrecipe.yaml). - All
Linearlayers of the language model backbone (llm/) are quantized; the audio tower, connector,lm_head, embeddings and all other pipeline components are kept in BF16. - Format: compressed-tensors safetensors - load directly with vLLM (
quantization="compressed-tensors") or transformers.
Consistency validation
Measured on an NVIDIA A100 against the original BF16 checkpoint (greedy decoding, official translation prompt):
| Metric | BF16 | FP8 | Delta |
|---|---|---|---|
| Perplexity (fixed corpus) | 4.8757 | 4.9627 | +1.79% |
| zh->en generation identical | - | - | yes |
| en->zh generation identical | - | - | no (semantically equivalent) |
Usage
Identical to the original checkpoint - clone this repo and follow the README / infer.py of the base model (IndexTeam/Index-Echo-S2TT-2B). The quantized LLM backbone loads via compressed-tensors; make sure compressed-tensors (or vLLM / a recent transformers) is installed.
Quantized and published by the Index team, 2026-10-03.
- Downloads last month
- 17
Model tree for IndexTeam/Index-Echo-S2TT-2B-FP8
Base model
IndexTeam/Index-Echo-S2TT-2B