cgifbribcgfbi
/

Meta-Llama-3.1-8B-Instruct-abliterated-chem-claude-5-comp3-sort-pat

@@ -35,7 +35,7 @@ strict: false
 datasets:
   - path: dset_comp3.0_sortpatent_count_pat400_in5_5000.jsonl
     type: chat_template
-    split: train
 dataset_prepared_path: last_run_prepared
 val_set_size: 0.04
@@ -60,7 +60,7 @@ wandb_log_model:
 gradient_accumulation_steps: 1
 micro_batch_size: 4  # This will be automatically adjusted based on available GPU memory
-num_epochs: 1
 optimizer: adamw_torch_fused
 lr_scheduler: cosine
 learning_rate: 0.00002
@@ -103,7 +103,7 @@ special_tokens:
 This model is a fine-tuned version of [mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated](https://huggingface.co/mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated) on the dset_comp3.0_sortpatent_count_pat400_in5_5000.jsonl dataset.
 It achieves the following results on the evaluation set:
-- Loss: 0.5731
 ## Model description
@@ -133,15 +133,24 @@ The following hyperparameters were used during training:
 - optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
 - lr_scheduler_type: cosine
 - lr_scheduler_warmup_steps: 10
-- num_epochs: 1.0
 ### Training results
 | Training Loss | Epoch  | Step | Validation Loss |
 |:-------------:|:------:|:----:|:---------------:|
 | 0.7           | 0.0061 | 1    | 0.8766          |
-| 0.6465        | 0.3354 | 55   | 0.6349          |
-| 0.5865        | 0.6707 | 110  | 0.5731          |
 ### Framework versions

 datasets:
   - path: dset_comp3.0_sortpatent_count_pat400_in5_5000.jsonl
     type: chat_template
+    field_messages: messages
 dataset_prepared_path: last_run_prepared
 val_set_size: 0.04
 gradient_accumulation_steps: 1
 micro_batch_size: 4  # This will be automatically adjusted based on available GPU memory
+num_epochs: 4
 optimizer: adamw_torch_fused
 lr_scheduler: cosine
 learning_rate: 0.00002
 This model is a fine-tuned version of [mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated](https://huggingface.co/mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated) on the dset_comp3.0_sortpatent_count_pat400_in5_5000.jsonl dataset.
 It achieves the following results on the evaluation set:
+- Loss: 0.4583
 ## Model description
 - optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
 - lr_scheduler_type: cosine
 - lr_scheduler_warmup_steps: 10
+- num_epochs: 4.0
 ### Training results
 | Training Loss | Epoch  | Step | Validation Loss |
 |:-------------:|:------:|:----:|:---------------:|
 | 0.7           | 0.0061 | 1    | 0.8766          |
+| 0.6414        | 0.3354 | 55   | 0.6293          |
+| 0.5608        | 0.6707 | 110  | 0.5473          |
+| 0.4733        | 1.0061 | 165  | 0.5161          |
+| 0.5142        | 1.3415 | 220  | 0.4954          |
+| 0.4771        | 1.6768 | 275  | 0.4824          |
+| 0.423         | 2.0122 | 330  | 0.4750          |
+| 0.4375        | 2.3476 | 385  | 0.4676          |
+| 0.4311        | 2.6829 | 440  | 0.4630          |
+| 0.4019        | 3.0183 | 495  | 0.4620          |
+| 0.4726        | 3.3537 | 550  | 0.4589          |
+| 0.4677        | 3.6890 | 605  | 0.4583          |
 ### Framework versions