Model Card for gpjt/1xrtx3090-moe-1

This model is gpjt/1xrtx3090-moe-1, a trained-from-scratch base model that adds Mixture-of-Experts support to the GPT-2-style architecture from Sebastian Raschka's book "Build a Large Language Model (from Scratch)".

Model Details

Model Description

  • Developed by: Giles Thomas, based on code by Sebastian Raschka
  • Model type: GPT-2 style transformers-based causal LLM.
  • License: Apache 2
  • Parameters: 446,410,752 (219,697,152 active per token)
  • Context length: 1,024
  • Embedding dimensions: 768
  • MHA heads: 12
  • Layers: 12
  • Experts per layer: 6 (2 active)
  • QKV bias: False
  • Weight tying: False

Don't have high expectations for the model! It has only 446M parameters (the GPT-2 "small" size plus some extra FFNs for the MoE support) and was trained on 8,928,264,192 tokens, which means that it doesn't know many facts and is not terribly smart. If you want to do serious work, use a serious model (I like Qwen's). But if you want to build on this and see what you can do with a 2021-vintage LLM, please do feel free to play with it!

Model Sources

How to Get Started with the Model

You can download and run the model for inference directly:

from transformers import pipeline
pipe = pipeline("text-generation", model="gpjt/1xrtx3090-moe-1", trust_remote_code=True)
out = pipe(
    "Every effort moves you",
    max_new_tokens=20,
    do_sample=True,
    temperature=1.4,
    top_k=25,
)
print(out[0]["generated_text"])

Note that because it uses custom code, you'll need to set trust_remote_code to True.

It supports AutoTokenizer, AutoModel and AutoModelForCausalLM:

>>> from transformers import AutoTokenizer, AutoModel, AutoModelForCausalLM
>>> tokenizer = AutoTokenizer.from_pretrained("gpjt/1xrtx3090-moe-1")
>>> model = AutoModel.from_pretrained("gpjt/1xrtx3090-moe-1", trust_remote_code=True)
>>> llm_model = AutoModelForCausalLM.from_pretrained("gpjt/1xrtx3090-moe-1", trust_remote_code=True)

You can also fine-tune it; this notebook has an example.

Again, don't expect too much from this model! It's a 446M-parameter GPT-2 one, trained on a limited number of tokens. It's both dumb and ignorant ;-)

Training Details

  • Machine type: Local machine with an RTX 3090
  • Tokens: 8,928,264,192 tokens rounded up to the nearest batch.
  • Dataset: gpjt/fineweb-gpt2-tokens
  • Micro-batch size: 3
  • Global batch size: 96
  • Dropout: 0.0
  • Gradient clipping: 3.5
  • Learning rate: 0.0014
  • Schedule learning rate: True
  • Weight decay: 0.01
Downloads last month
366
Safetensors
Model size
0.5B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train gpjt/1xrtx3090-moe-1