Mimans M1 (514M) - 5.01B Tokens Milestone

Mimans M1 is a 514M parameter decoder-only transformer pretrained from scratch on code, math, and technical text using TileLang custom Blackwell GPU kernels and Muon + AdamW hybrid optimization.

Model Summary

  • Parameters: 514,345,526 (~514M)
  • Architecture: GQA (10:2), SwiGLU FFN, AttnRes Skip Gating, RMSNorm
  • Context Length: 4,096 tokens (max 8,192, RoPE $\theta = 500,000$)
  • Vocabulary: 49,152 (Byte-level BPE with FIM support)
  • Training Tokens: 5.01B Tokens (Global Step 9,551)
  • Latest Loss: 1.9756
  • Hardware: NVIDIA GeForce RTX 5090 (Blackwell sm_120)

Usage & Generation

Loading with Hugging Face Transformers:

from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "ankushthakurr09/MimansM1_v1"
model = AutoModelForCausalLM.from_pretrained(repo_id, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)

prompt = "def fibonacci(n):"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Running Standalone Generation Script:

python generate.py --prompt "def hello_world():" --max-tokens 50
Downloads last month
1,283
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support