GravityOCR-KO (internal, private)

Korean continued fine-tuning of trillionlabs/GravityOCR, with replay from the original English/CJK training pools so the table and formula tasks are retained.

Like the base model, this is a standard GlmOcrForConditionalGeneration checkpoint plus block_diffusion.json. ar_loss_weight: 1.0 is preserved, so the checkpoint still has a trained AR path and can self-speculate (diffusion drafts, AR verifies).

Training

base trillionlabs/GravityOCR (post-GRPO release, not the GLM-OCR base)
objective joint AR + block diffusion, ar_loss_weight=1.0, bd_size=32, block-phase jitter on
caps max_length=16384, max_response_length=6144, max_vision_tokens=8192

Usage

import torch
from transformers import AutoProcessor, GlmOcrForConditionalGeneration

repo = "trillionlabs/GravityOCR-KO"
model = GlmOcrForConditionalGeneration.from_pretrained(repo, dtype=torch.bfloat16, device_map="cuda")
processor = AutoProcessor.from_pretrained(repo)

Stock transformers>=5.8 loads it for AR decoding. Self-speculative decoding needs the nanoVLM-DiffusionOCR repository (in process) or the SGLang patch (serving).

Downloads last month
57
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for trillionlabs/GravityOCR-KO

Base model

zai-org/GLM-OCR
Finetuned
(1)
this model