End of training

Browse files

Files changed (14) hide show

README.md +130 -0
benchmarks.shelve.bak +0 -0
benchmarks.shelve.dat +0 -0
benchmarks.shelve.dir +0 -0
config.json +39 -0
generation_config.json +6 -0
logs/attn_layer_mapper=all, attn_loss_fn=raw_mse, attn_projector=mlp/events.out.tfevents.1724600288.e3f806ea38c9 +3 -0
merges.txt +0 -0
model.safetensors +3 -0
special_tokens_map.json +6 -0
tokenizer.json +0 -0
tokenizer_config.json +20 -0
training_args.bin +3 -0
vocab.json +0 -0

README.md ADDED Viewed

	@@ -0,0 +1,130 @@

+---
+base_model: gpt2
+datasets:
+- wikimedia/wikipedia
+library_name: Distily
+license: mit
+tags:
+- bitnet
+- 1.58b
+- generated_from_trainer
+model-index:
+- name: distily_test_attn_mlp
+  results: []
+---
+# Summary
+Distilled with [Distily](https://github.com/lapp0/distily) library
+using teacher model [gpt2](https://huggingface.co/gpt2)
+on dataset [wikimedia/wikipedia](https://huggingface.co/datasets/wikimedia/wikipedia).
+<!-- This model card has been generated automatically according to the information the Trainer had access to. You
+should probably proofread and complete it, then remove this comment.
+# Model description
+More information needed
+# Intended uses & limitations
+More information needed
+-->
+# Model Architecture:
+- **Architecture**: `GPT2LMHeadModel`
+- **Total Parameters**: 124,439,808
+- **Data Type (dtype)**: torch.bfloat16
+- **Model Size**: 0.24 GB
+# Benchmark Metrics Comparison
+| Metric |  |
+| :--- |
+# Resource Usage Comparison
+- VRAM Use: 7.7845 GB
+# Distillation (Teacher -> Student) Architecture Difference:
+- **Architecture**: `GPT2LMHeadModel` -> `GPT2LMHeadModel`
+- **Total Parameters**: 124,439,808 -> 124,439,808
+- **Data Type (dtype)**: torch.bfloat16 -> torch.bfloat16
+- **Model Size**: 0.24 GB -> 0.24 GB
+<details>
+<summary>Module Diff Details</summary>
+```diff
+```
+</details>
+<br/>
+# Train Dataset
+Trained on 145,738,108 tokens from the [wikimedia/wikipedia](https://huggingface.co/datasets/wikimedia/wikipedia) dataset.
+- Num Samples: `247,500`
+- Subset: `20231101.en`
+- Split: `train`
+# Training Objective
+```
+DistillationObjective(logits_loss_component=LossComponent(label=logits, weight=1, loss_fn=kl), attn_loss_component=LossComponent(label=attn, weight=25.0, loss_fn=raw_mse, layer_mapper=all, projector=mlp))
+```
+# Hyperparameters
+The following hyperparameters were used during training:
+<details>
+<summary>Expand</summary>
+- learning_rate: `0.0001`
+- train_batch_size: `4`
+- eval_batch_size: `8`
+- seed: `42`
+- optimizer: `Adam with betas=(0.9,0.999) and epsilon=1e-08`
+- lr_scheduler_type: `cosine_with_min_lr`
+- lr_scheduler_warmup_ratio: `0.5`
+- num_epochs: `1.0`
+- distillation_objective: `DistillationObjective(logits_loss_component=LossComponent(label=logits, weight=1, loss_fn=kl), attn_loss_component=LossComponent(label=attn, weight=25.0, loss_fn=raw_mse, layer_mapper=all, projector=mlp))`
+- train_embeddings: `True`
+- lr_scheduler: `<torch.optim.lr_scheduler.LambdaLR object at 0x7fde37bf8df0>`
+- student_model_name_or_path: `None`
+- student_config_name_or_path: `None`
+- student_model_config: `None`
+- reinitialize_weights: `None`
+- copy_teacher_modules: `[('lm_head', False)]`
+- student_model_as_bitnet: `True`
+- dropout: `None`
+- teacher_model_name_or_path: `gpt2`
+- teacher_load_in_8bit: `False`
+- teacher_load_in_4bit: `False`
+- dataset_uri: `wikimedia/wikipedia`
+- dataset_subset: `20231101.en`
+- dataset_split: `train`
+- dataset_column_name: `text`
+- dataset_sample_size: `250000`
+- dataset_test_size: `0.01`
+- gradient_accumulation_steps: `1`
+- weight_decay: `0.0`
+- max_grad_norm: `1.0`
+- warmup_ratio: `0.5`
+- warmup_steps: `0`
+- gradient_checkpointing: `True`
+</details>
+<br/>
+# Framework Versions
+- Distily 0.3.0
+- Transformers 4.44.1
+- Pytorch 2.4.0+cu121
+- Datasets 2.21.0

benchmarks.shelve.bak ADDED Viewed

File without changes

benchmarks.shelve.dat ADDED Viewed

File without changes

benchmarks.shelve.dir ADDED Viewed

File without changes

config.json ADDED Viewed

	@@ -0,0 +1,39 @@

+{
+  "_name_or_path": "gpt2",
+  "activation_function": "gelu_new",
+  "architectures": [
+    "GPT2LMHeadModel"
+  ],
+  "attn_pdrop": 0.1,
+  "bos_token_id": 50256,
+  "embd_pdrop": 0.1,
+  "eos_token_id": 50256,
+  "initializer_range": 0.02,
+  "layer_norm_epsilon": 1e-05,
+  "model_type": "gpt2",
+  "n_ctx": 1024,
+  "n_embd": 768,
+  "n_head": 12,
+  "n_inner": null,
+  "n_layer": 12,
+  "n_positions": 1024,
+  "reorder_and_upcast_attn": false,
+  "resid_pdrop": 0.1,
+  "scale_attn_by_inverse_layer_idx": false,
+  "scale_attn_weights": true,
+  "summary_activation": null,
+  "summary_first_dropout": 0.1,
+  "summary_proj_to_labels": true,
+  "summary_type": "cls_index",
+  "summary_use_proj": true,
+  "task_specific_params": {
+    "text-generation": {
+      "do_sample": true,
+      "max_length": 50
+    }
+  },
+  "torch_dtype": "bfloat16",
+  "transformers_version": "4.44.1",
+  "use_cache": true,
+  "vocab_size": 50257
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,6 @@

+{
+  "_from_model_config": true,
+  "bos_token_id": 50256,
+  "eos_token_id": 50256,
+  "transformers_version": "4.44.1"
+}

logs/attn_layer_mapper=all, attn_loss_fn=raw_mse, attn_projector=mlp/events.out.tfevents.1724600288.e3f806ea38c9 ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:859c4bb2bcf29bc0983cdb184f1a994342ab7503e6e3bafc3ce538af92dfeba4
+size 29625273

merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:8e9faae978a05408ab2bfd92ac4c4c80c628da9b54adbbb2d5b7be826c17ebf9
+size 248894656

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,6 @@

+{
+  "bos_token": "<|endoftext|>",
+  "eos_token": "<|endoftext|>",
+  "pad_token": "<|endoftext|>",
+  "unk_token": "<|endoftext|>"
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,20 @@

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "50256": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "bos_token": "<|endoftext|>",
+  "clean_up_tokenization_spaces": true,
+  "eos_token": "<|endoftext|>",
+  "model_max_length": 1024,
+  "pad_token": "<|endoftext|>",
+  "tokenizer_class": "GPT2Tokenizer",
+  "unk_token": "<|endoftext|>"
+}

training_args.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:2cb326353387e09fd1fb6f2a6c9d3c5718471b4d4551bccf34af78705979a2e7
+size 5432

vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff