Instructions to use tokimoa/uratori-ja-2b-mlx-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use tokimoa/uratori-ja-2b-mlx-8bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download tokimoa/uratori-ja-2b-mlx-8bit --local-dir uratori-ja-2b-mlx-8bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
uratori-ja-2b-mlx-8bit
MLX version of tokimoa/uratori-ja-2b v1.1 for Apple silicon, quantised to 8 bits (group size 64) with mlx-lm. uratori-ja-2b is a Japanese decision model: given a text state and typed questions (yes/no, multiple choice, ordinal score) it returns a probability distribution over the allowed answers in one forward pass, without generating text. It is trained for grounding checks against supplied evidence, document comparison and RAG judgments. See the parent model card for details, training and limitations.
The backbone (the text decoder of Qwen/Qwen3.5-2B with the LoRA merged) is stored in mlx-lm format; the scoring head is in head.safetensors and is applied by predict_mlx.py to the hidden state at the option markers.
Accuracy
Measured on uratori-ja-eval with this file, one question per call, on an Apple M4 Max.
| test (802) | challenge (300) | Answers changed vs PyTorch (challenge) | Time per question | |
|---|---|---|---|---|
| uratori-ja-2b-mlx-8bit | 0.783 | 0.797 | 0 of 300 | about 250 ms |
| uratori-ja-2b v1.1 (PyTorch, bf16) | 0.781 | 0.797 |
Usage
pip install mlx-lm huggingface_hub
python -c "from huggingface_hub import hf_hub_download; print(hf_hub_download('tokimoa/uratori-ja-2b-mlx-8bit', 'predict_mlx.py'))"
python <that path> # downloads the model (2.0 GB) and runs the example
Or from Python:
from pathlib import Path
from huggingface_hub import snapshot_download
import importlib.util
folder = Path(snapshot_download("tokimoa/uratori-ja-2b-mlx-8bit"))
spec = importlib.util.spec_from_file_location("predict_mlx", folder / "predict_mlx.py")
mod = importlib.util.module_from_spec(spec); spec.loader.exec_module(mod)
model = mod.MlxUratori(folder)
state = {"根拠": "返品は商品到着後14日以内に限り受け付けます。開封済みの商品は返品できません。",
"主張": "開封済みの商品でも、到着から7日以内なら返品できる。"}
questions = {"support": {"type": "choice", "instructions": "`主張` は `根拠` から支持されるか",
"criteria": {"支持": "根拠だけから主張の全体が成り立つ。", "矛盾": "根拠と両立しない部分がある。",
"情報不足": "矛盾はないが、根拠からは成否を決められない部分がある。"}}}
print(model.predict(state, questions)["support"])
state is a string or a dict of named texts; questions follows TypeSafe's /v1/systemone question format (noul, choice, score). Probabilities are temperature-scaled with the parent model's calibration. Inputs over 4,096 tokens raise an error. Memory use is about 3 GB.
Files
| File | Content |
|---|---|
model.safetensors, config.json |
Quantised backbone in mlx-lm format |
head.safetensors |
Scoring head (float32) |
uratori.json |
Marker token, input limit, calibration temperatures |
predict_mlx.py |
Runner (mlx-lm plus the head) |
License
Apache License 2.0, like the parent model. The base model Qwen/Qwen3.5-2B is Apache 2.0.
日本語
tokimoa/uratori-ja-2b の Apple silicon 向け MLX 版です(8bit 量子化、2.0 GB)。日本語の文章と質問を渡すと、文章を生成せずに答えの確率分布を返します。精度は test 0.783、challenge 0.797(PyTorch 版は 0.781、0.797)、M4 Max で 1 問 250 ms 前後です。使い方は上のとおりで、mlx-lm と huggingface_hub だけで動きます。詳しい説明と制限は元のモデルのページにあります。
- Downloads last month
- 33
8-bit