Upload folder using huggingface_hub

Browse files

Files changed (8) hide show

README.md +99 -0
config.json +31 -0
generation_config.json +6 -0
model.safetensors +3 -0
special_tokens_map.json +1 -0
tokenizer.json +0 -0
tokenizer_config.json +19 -0
training_metadata.json +15 -0

README.md ADDED Viewed

	@@ -0,0 +1,99 @@

+---
+language:
+- en
+library_name: transformers
+license: apache-2.0
+tags:
+- sparknet
+- causal-lm
+- text-generation
+- gpt
+- pytorch
+- 70m
+pipeline_tag: text-generation
+model-index:
+- name: SparkNet-70M-v5
+  results: []
+---
+# SparkNet 70M v5
+SparkNet 70M v5 is the final 70M-parameter checkpoint from the SparkNet research run by **DienerTech**. It is a compact GPT-2–style decoder (12 layers, 512 hidden size, 8 attention heads, 1024-token context) that was trained for ~1B tokens on a custom mixture of high-quality web and document corpora. The release ships with the SparkNet v5 tokenizer and weights stored in `model.safetensors`, ready for direct use via 🤗 Transformers.
+## Model Details
+- **Developer**: DienerTech
+- **Architecture**: GPT-2–style causal decoder (approx. 70M parameters), dropout 0.1, cosine LR schedule, AdamW (fused).
+- **Context length**: 1,024 tokens.
+- **Tokenizer**: SparkNet v5 byte-level BPE (vocab size 50,257, EOS = `` and `<|pad|>` padding).
+- **Framework**: PyTorch / 🤗 Transformers 4.46+.
+- **Checkpoint**: Converted to `model.safetensors` for safe loading; no `pytorch_model.bin` left in the repo.
+## Intended Use
+- Lightweight text generation experiments, story/note drafting, or as a base for instruction-tuning / domain adaptation (LoRA, QLoRA, etc.).
+- Research on small-model scaling laws or tokenizer experimentation.
+## Limitations & Risks
+- No RLHF / instruction tuning; outputs will be generic next-token predictions and may require prompting tricks.
+- Training data is predominantly public web/document text, so bias, toxicity, or outdated information may surface.
+- Not evaluated for safety-critical deployments—perform your own alignment and red-teaming before production use.
+## Training Data
+- 1B tokens packed into 1,024-token blocks (`datasets/sparknet-v5-1b`).
+- Sources sampled uniformly across: `codelion/finepdfs-1B`, `codelion/dclm-baseline-1B`, `codelion/fineweb-edu-1B`, plus curated DienerTech blog data.
+- Validation set: `wikitext-2-raw-v1` (standard Hugging Face split).
+## Training Procedure
+- **Optimizer**: AdamW (fused) with β₁=0.9, β₂=0.95, weight decay 0.1, gradient clipping at 1.0.
+- **Learning rate**: 1e-4 peak with 3% warmup then cosine decay.
+- **Batching**: per-device batch size 32, gradient accumulation 2 → 65,536 tokens/step.
+- **Budget**: 1,000,000,000 effective tokens (≈15,259 steps).
+- **Hardware**: Single 24GB+ NVIDIA GPU with TF32 + Flash Attention enabled.
+- **Best checkpoint**: step 14,000 with eval loss 4.99 on WikiText-2 (logged via `trainer_state.json`).
+## Evaluation
+Formal downstream evaluation has not been run yet. Inside `trainer_state.json`, the best validation (WikiText-2) cross-entropy reached **4.9869** at step 14k. If you benchmark the model (e.g., with lm-eval-harness), please consider contributing results back to the card via a PR.
+## Usage
+```python
+from transformers import AutoTokenizer, AutoModelForCausalLM
+import torch
+model_id = "DienerTech/sparknet-70m-v5"
+tokenizer = AutoTokenizer.from_pretrained(model_id)
+model = AutoModelForCausalLM.from_pretrained(
+    model_id,
+    torch_dtype=torch.bfloat16,  # or torch.float16 on older GPUs
+    device_map="auto",
+)
+prompt = "In a distant research lab, a tiny transformer model awakened and"
+inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
+output = model.generate(
+    **inputs,
+    max_new_tokens=120,
+    temperature=0.9,
+    top_p=0.9,
+    do_sample=True,
+)
+print(tokenizer.decode(output[0], skip_special_tokens=True))
+```
+## Citation
+```
+@software{sparknet70mv5,
+  author = {DienerTech},
+  title = {SparkNet 70M v5},
+  year = {2025},
+  url = {https://huggingface.co/DienerTech/sparknet-70m-v5}
+}
+```
+Please open an issue or PR on the DienerTech Hugging Face repo if you have feedback, evaluations, or fine-tuned variants to share.

config.json ADDED Viewed

	@@ -0,0 +1,31 @@

+{
+  "activation_function": "gelu_new",
+  "architectures": [
+    "GPT2LMHeadModel"
+  ],
+  "attn_pdrop": 0.1,
+  "bos_token_id": 50256,
+  "dtype": "float32",
+  "embd_pdrop": 0.1,
+  "eos_token_id": 50256,
+  "initializer_range": 0.02,
+  "layer_norm_epsilon": 1e-05,
+  "model_type": "gpt2",
+  "n_embd": 512,
+  "n_head": 8,
+  "n_inner": null,
+  "n_layer": 12,
+  "n_positions": 1024,
+  "reorder_and_upcast_attn": false,
+  "resid_pdrop": 0.1,
+  "scale_attn_by_inverse_layer_idx": false,
+  "scale_attn_weights": true,
+  "summary_activation": null,
+  "summary_first_dropout": 0.1,
+  "summary_proj_to_labels": true,
+  "summary_type": "cls_index",
+  "summary_use_proj": true,
+  "transformers_version": "4.57.1",
+  "use_cache": false,
+  "vocab_size": 50257
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,6 @@

+{
+  "_from_model_config": true,
+  "bos_token_id": 50256,
+  "eos_token_id": 50256,
+  "transformers_version": "4.57.1"
+}

model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:223dace1d2c6be60c9f8793863e1795b36a24e53ff188b83d0949bd6af0c49e6
+size 256356888

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1 @@


1	+ {}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,19 @@

+{
+  "added_tokens_decoder": {
+    "1": {
+      "content": "<|pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "",
+  "extra_special_tokens": {},
+  "model_max_length": 1024,
+  "pad_token": "",
+  "padding_side": "right",
+  "tokenizer_class": "PreTrainedTokenizerFast"
+}

training_metadata.json ADDED Viewed

	@@ -0,0 +1,15 @@

+{
+  "run_name": "sparknet-70m-v5",
+  "timestamp": "2025-11-15T07:24:08.029637",
+  "params": {
+    "n_embd": 512,
+    "n_layer": 12,
+    "n_head": 8,
+    "context_length": 1024,
+    "token_budget": 1000000000
+  },
+  "datasets": [
+    "sparknet-v5-1b"
+  ],
+  "notes": "V5 | Custom tokenizer, dropout, cosine LR, static 1B token dataset."
+}