TinyBalls-110M-V1
A 113M-parameter LLaMA-style instruct model, trained from scratch on 2B tokens and then fine-tuned for chat and tool use.
The base (non-instruct) model is a separate repo:
igidn/TinyBalls-v1-Base.
Architecture
| params | 113.3M (151.0M counting the tied head twice) |
| layers | 12 |
| hidden size | 768 |
| attention | 12 query heads / 4 KV heads (GQA), head dim 64 |
| FFN | SwiGLU, intermediate 2048, pre-activations clamped at 10.0 |
| vocab | 49,154 (SmolLM2 byte-level BPE + ChatML and tool specials) |
| context | 32,768 (8,192 used in practice, see below) |
| tied embeddings | yes |
| RoPE theta | 500,000 |
Weights are float32 and load as a standard
transformers LlamaForCausalLM.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("igidn/TinyBalls-110M-V1")
model = AutoModelForCausalLM.from_pretrained(
"igidn/TinyBalls-110M-V1", torch_dtype="float32")
messages = [{"role": "user", "content": "What is the capital of France?"}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True,
return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=128,
do_sample=True, temperature=0.8, top_p=0.95)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:]))
Generation stops at <|im_end|> (id 2), which is the configured EOS. The
sampling settings in generation_config.json (temperature=0.8, top_p=0.95)
are the ones to use — greedy decoding is prone to repetition loops on this
checkpoint. Every number below was measured under sampling.
Chat format
<|im_start|>system
{system}<|im_end|>
<|im_start|>user
{user}<|im_end|>
<|im_start|>assistant
{content}<|im_end|>
<|im_start|>tool
<|tool_result|>{result}<|im_end|>
An assistant turn is optional <think>...</think> reasoning, then text, then
zero or more <|tool_call|>{...} segments, then <|im_end|>. There is no
newline after the closing <|im_end|> — the next turn's <|im_start|> follows
immediately.
Tool calling
Tool schemas go in the system message as full JSON, and the model replies
with <|tool_call|>{"name": "...", "arguments": {...}} and nothing after it:
<|im_start|>system
You are a helpful assistant with access to tools. Use the provided tools when
they can help answer the user's request. If a tool call is needed, put any
reasoning in <think>...</think> first.
Available tools:
[{"type": "function", "function": {"name": "get_weather", ...}}]
To call a tool, reply with <|tool_call|>{"name": "...", "arguments": {...}} and nothing after it.<|im_end|>
Use the chat_template in tokenizer_config.json as the format of record.
Training
Three stages, each from igidn/tinyballs-v1:
- Pretraining, 2.01B tokens (1.5B broad + 0.5B quality anneal) on an English-only mix of web, books, Wikipedia, code, math, papers and forums. Sequence length 2048, RoPE theta 10,000.
- SFT on 90,804 conversations / 161M tokens of distilled chat, tool-calling and agentic data, in three phases. RoPE theta is raised 10k → 50k → 500k alongside context 2048 → 8192 → 32768.
- Tool-call repair, 630 steps / 41.3M tokens on a targeted mix for schema binding, abstention, error recovery and grounded answers. This is the released checkpoint.
Files
| file | |
|---|---|
model.safetensors |
fp32 LlamaForCausalLM weights |
config.json |
architecture, rope_theta 500000, 32k context |
tokenizer.json |
the exact tokenizer these weights were trained with |
tokenizer_config.json |
specials + ChatML chat_template |
generation_config.json |
sampling defaults, EOS = `< |
Modelfile |
Ollama serving template |
License
MIT.
-igidn
- Downloads last month
- 317