davisrbr
/

Meta-Llama-3-8B-Instruct-AQLM-2Bit-1x16-r8_bs4

Generated from Trainer

Model card Files Files and versions Community

davisrbr commited on Aug 19, 2024

Commit

2c828d3

·

verified ·

1 Parent(s): 771ce62

Model save

Files changed (2) hide show

README.md +4 -4
adapter_model.safetensors +1 -1

README.md CHANGED Viewed

@@ -6,14 +6,14 @@ library_name: peft
 tags:
 - generated_from_trainer
 model-index:
-- name: Meta-Llama-3-8B-Instruct-AQLM-2Bit-1x16-r16
   results: []
 ---
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
 should probably proofread and complete it, then remove this comment. -->
-# Meta-Llama-3-8B-Instruct-AQLM-2Bit-1x16-r16
 This model is a fine-tuned version of [ISTA-DASLab/Meta-Llama-3-8B-Instruct-AQLM-2Bit-1x16](https://huggingface.co/ISTA-DASLab/Meta-Llama-3-8B-Instruct-AQLM-2Bit-1x16) on the red_pajama-data-1_t-sample dataset.
@@ -35,11 +35,11 @@ More information needed
 The following hyperparameters were used during training:
 - learning_rate: 0.0002
-- train_batch_size: 1
 - eval_batch_size: 8
 - seed: 42
 - gradient_accumulation_steps: 16
-- total_train_batch_size: 16
 - optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
 - lr_scheduler_type: linear
 - lr_scheduler_warmup_steps: 200

 tags:
 - generated_from_trainer
 model-index:
+- name: Meta-Llama-3-8B-Instruct-AQLM-2Bit-1x16-r8_bs4
   results: []
 ---
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
 should probably proofread and complete it, then remove this comment. -->
+# Meta-Llama-3-8B-Instruct-AQLM-2Bit-1x16-r8_bs4
 This model is a fine-tuned version of [ISTA-DASLab/Meta-Llama-3-8B-Instruct-AQLM-2Bit-1x16](https://huggingface.co/ISTA-DASLab/Meta-Llama-3-8B-Instruct-AQLM-2Bit-1x16) on the red_pajama-data-1_t-sample dataset.
 The following hyperparameters were used during training:
 - learning_rate: 0.0002
+- train_batch_size: 4
 - eval_batch_size: 8
 - seed: 42
 - gradient_accumulation_steps: 16
+- total_train_batch_size: 64
 - optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
 - lr_scheduler_type: linear
 - lr_scheduler_warmup_steps: 200

adapter_model.safetensors CHANGED Viewed

@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:e08b94cd718e3b61f74ad50c6d14db95a6941cd76dde4c345c715255720d6198
 size 83945296

 version https://git-lfs.github.com/spec/v1
+oid sha256:4105c8151e3a0b2030988d43b65baf62cb45b8a787696db58f22defb59546108
 size 83945296