Quant for 3.5

Browse files

Files changed (12) hide show

.gitattributes +1 -0
README.md +189 -37
config.json +30 -0
dolphin_moe.png +3 -0
mergekit_moe_config.yml +47 -0
model.safetensors.index.json +1 -0
original_repo_url.txt +1 -0
output.safetensors +3 -0
special_tokens_map.json +29 -0
tokenizer.json +0 -0
tokenizer.model +3 -0
tokenizer_config.json +46 -0

.gitattributes CHANGED Viewed

@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text

 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
+dolphin_moe.png filter=lfs diff=lfs merge=lfs -text

README.md CHANGED Viewed

@@ -1,69 +1,221 @@
 ---
 license: apache-2.0
 library_name: transformers
-quantized_by: bartowski
-pipeline_tag: text-generation
 ---
-## Exllama v2 Quantizations of laser-dolphin-mixtral-2x7b-dpo
-Using <a href="https://github.com/turboderp/exllamav2/releases/tag/v0.0.13">turboderp's ExLlamaV2 v0.0.13</a> for quantization.
-## The "main" branch only contains the measurement.json, download one of the other branches for the model (see below)
-Each branch contains an individual bits per weight, with the main one containing only the meaurement.json for further conversions.
-Conversion was done using the default calibration dataset.
-Default arguments used except when the bits per weight is above 6.0, at that point the lm_head layer is quantized at 8 bits per weight instead of the default 6.
-Original model: https://huggingface.co/macadeliccc/laser-dolphin-mixtral-2x7b-dpo
-<a href="https://huggingface.co/bartowski/laser-dolphin-mixtral-2x7b-dpo-exl2/tree/8_0">8.0 bits per weight</a>
-<a href="https://huggingface.co/bartowski/laser-dolphin-mixtral-2x7b-dpo-exl2/tree/6_5">6.5 bits per weight</a>
-<a href="https://huggingface.co/bartowski/laser-dolphin-mixtral-2x7b-dpo-exl2/tree/5_0">5.0 bits per weight</a>
-<a href="https://huggingface.co/bartowski/laser-dolphin-mixtral-2x7b-dpo-exl2/tree/4_25">4.25 bits per weight</a>
-<a href="https://huggingface.co/bartowski/laser-dolphin-mixtral-2x7b-dpo-exl2/tree/3_5">3.5 bits per weight</a>
-## Download instructions
-With git:
-```shell
-git clone --single-branch --branch 6_5 https://huggingface.co/bartowski/laser-dolphin-mixtral-2x7b-dpo-exl2
-```
-With huggingface hub (credit to TheBloke for instructions):
-```shell
-pip3 install huggingface-hub
-```
-To download the `main` (only useful if you only care about measurement.json) branch to a folder called `laser-dolphin-mixtral-2x7b-dpo-exl2`:
-```shell
-mkdir laser-dolphin-mixtral-2x7b-dpo-exl2
-huggingface-cli download bartowski/laser-dolphin-mixtral-2x7b-dpo-exl2 --local-dir laser-dolphin-mixtral-2x7b-dpo-exl2 --local-dir-use-symlinks False
-```
-To download from a different branch, add the `--revision` parameter:
-Linux:
-```shell
-mkdir laser-dolphin-mixtral-2x7b-dpo-exl2-6_5
-huggingface-cli download bartowski/laser-dolphin-mixtral-2x7b-dpo-exl2 --revision 6_5 --local-dir laser-dolphin-mixtral-2x7b-dpo-exl2-6_5 --local-dir-use-symlinks False
 ```
-Windows (which apparently doesn't like _ in folders sometimes?):
-```shell
-mkdir laser-dolphin-mixtral-2x7b-dpo-exl2-6.5
-huggingface-cli download bartowski/laser-dolphin-mixtral-2x7b-dpo-exl2 --revision 6_5 --local-dir laser-dolphin-mixtral-2x7b-dpo-exl2-6.5 --local-dir-use-symlinks False
 ```

 ---
 license: apache-2.0
 library_name: transformers
 ---
+# Laser-Dolphin-Mixtral-2x7b-dpo
+![laser_dolphin_image](./dolphin_moe.png)
+**New Version out now!**
+Credit to Fernando Fernandes and Eric Hartford for their project [laserRMT](https://github.com/cognitivecomputations/laserRMT)
+## Overview
+This model is a medium-sized MoE implementation based on [cognitivecomputations/dolphin-2.6-mistral-7b-dpo-laser](https://huggingface.co/cognitivecomputations/dolphin-2.6-mistral-7b-dpo-laser)
++ The new version shows ~1 point on average.
+## Process
++ The process is outlined in this [notebook](https://github.com/cognitivecomputations/laserRMT/blob/main/examples/laser-dolphin-mixtral-2x7b.ipynb)
++ The mergekit_config is in the files.
++ The models used in the configuration are not lasered, but the final product is. This is an update from the last version.
++ This process is experimental. Your mileage may vary.
+## Future Goals
++ [ ] Function Calling
++ [ ] v2 with new base model to improve performance
+## Quantizations
+**These Quants will result in unpredicted behavior. New quants are available as I have updated the model**
+Quatizations provided by [TheBloke](https://huggingface.co/TheBloke/laser-dolphin-mixtral-2x7b-dpo-GGUF)
+*Current [Quantizations](https://huggingface.co/macadeliccc/laser-dolphin-mixtral-2x7b-dpo-GGUF)*
+## HF Spaces
++ GGUF chat available [here](https://huggingface.co/spaces/macadeliccc/laser-dolphin-mixtral-chat-GGUF)
++ 4-bit bnb chat available [here](https://huggingface.co/spaces/macadeliccc/laser-dolphin-mixtral-chat)
+## Code Example
+Switch the commented model definition to use in 4-bit. Should work with 9GB and still exceed the single 7B model by 5-6 points roughly
+```python
+from transformers import AutoModelForCausalLM, AutoTokenizer
+def generate_response(prompt):
+    """
+    Generate a response from the model based on the input prompt.
+    Args:
+    prompt (str): Prompt for the model.
+    Returns:
+    str: The generated response from the model.
+    """
+    # Tokenize the input prompt
+    inputs = tokenizer(prompt, return_tensors="pt")
+    # Generate output tokens
+    outputs = model.generate(**inputs, max_new_tokens=256, eos_token_id=tokenizer.eos_token_id, pad_token_id=tokenizer.pad_token_id)
+    # Decode the generated tokens to a string
+    response = tokenizer.decode(outputs[0], skip_special_tokens=True)
+    return response
+# Load the model and tokenizer
+model_id = "macadeliccc/laser-dolphin-mixtral-2x7b-dpo"
+tokenizer = AutoTokenizer.from_pretrained(model_id)
+model = AutoModelForCausalLM.from_pretrained(model_id, load_in_4bit=True)
+prompt = "Write a quicksort algorithm in python"
+# Generate and print responses for each language
+print("Response:")
+print(generate_response(prompt), "\n")
 ```
+[colab](https://colab.research.google.com/drive/1cmRhAkDWItV7utHNqNANVZnqDqQNsTUr?usp=sharing) with usage example
+## Eval
+## EQ Bench
+<pre>----Benchmark Complete----
+2024-01-31 16:55:37
+Time taken: 31.1 mins
+Prompt Format: ChatML
+Model: macadeliccc/laser-dolphin-mixtral-2x7b-dpo-GGUF
+Score (v2): 72.76
+Parseable: 171.0
+---------------
+Batch completed
+Time taken: 31.2 mins
+---------------
+</pre>
+evaluation [colab](https://colab.research.google.com/drive/1FpwgsGzCR4tORTxAwUxpN3PcP22En2xk?usp=sharing)
+## Summary of previous evaluation
+|                                               Model                                               |AGIEval|GPT4All|TruthfulQA|Bigbench|Average|
+|---------------------------------------------------------------------------------------------------|------:|------:|---------:|-------:|------:|
+|[laser-dolphin-mixtral-2x7b-dpo](https://huggingface.co/macadeliccc/laser-dolphin-mixtral-2x7b-dpo)|  41.31|  73.67|     61.69|   42.79|  54.87|
+## Detailed current evaluation
+|                                               Model                                               |AGIEval|GPT4All|TruthfulQA|Bigbench|Average|
+|---------------------------------------------------------------------------------------------------|------:|------:|---------:|-------:|------:|
+|[laser-dolphin-mixtral-2x7b-dpo](https://huggingface.co/macadeliccc/laser-dolphin-mixtral-2x7b-dpo)|  42.25|  73.45|     63.44|   43.96|  55.77|
+### AGIEval
+|             Task             |Version| Metric |Value|   |Stderr|
+|------------------------------|------:|--------|----:|---|-----:|
+|agieval_aqua_rat              |      0|acc     |21.26|±  |  2.57|
+|                              |       |acc_norm|21.65|±  |  2.59|
+|agieval_logiqa_en             |      0|acc     |34.72|±  |  1.87|
+|                              |       |acc_norm|35.64|±  |  1.88|
+|agieval_lsat_ar               |      0|acc     |26.96|±  |  2.93|
+|                              |       |acc_norm|26.96|±  |  2.93|
+|agieval_lsat_lr               |      0|acc     |45.88|±  |  2.21|
+|                              |       |acc_norm|46.08|±  |  2.21|
+|agieval_lsat_rc               |      0|acc     |59.48|±  |  3.00|
+|                              |       |acc_norm|59.48|±  |  3.00|
+|agieval_sat_en                |      0|acc     |73.79|±  |  3.07|
+|                              |       |acc_norm|73.79|±  |  3.07|
+|agieval_sat_en_without_passage|      0|acc     |42.23|±  |  3.45|
+|                              |       |acc_norm|41.26|±  |  3.44|
+|agieval_sat_math              |      0|acc     |37.27|±  |  3.27|
+|                              |       |acc_norm|33.18|±  |  3.18|
+Average: 42.25%
+### GPT4All
+|    Task     |Version| Metric |Value|   |Stderr|
+|-------------|------:|--------|----:|---|-----:|
+|arc_challenge|      0|acc     |58.36|±  |  1.44|
+|             |       |acc_norm|58.02|±  |  1.44|
+|arc_easy     |      0|acc     |82.20|±  |  0.78|
+|             |       |acc_norm|77.40|±  |  0.86|
+|boolq        |      1|acc     |87.52|±  |  0.58|
+|hellaswag    |      0|acc     |67.50|±  |  0.47|
+|             |       |acc_norm|84.43|±  |  0.36|
+|openbookqa   |      0|acc     |34.40|±  |  2.13|
+|             |       |acc_norm|47.00|±  |  2.23|
+|piqa         |      0|acc     |81.61|±  |  0.90|
+|             |       |acc_norm|82.59|±  |  0.88|
+|winogrande   |      0|acc     |77.19|±  |  1.18|
+Average: 73.45%
+### GSM8K
+|Task |Version|           Metric            |Value|   |Stderr|
+|-----|------:|-----------------------------|-----|---|------|
+|gsm8k|      2|exact_match,get-answer       | 0.75|   |      |
+|     |       |exact_match_stderr,get-answer| 0.01|   |      |
+|     |       |alias                        |gsm8k|   |      |
+### TruthfulQA
+|    Task     |Version|Metric|Value|   |Stderr|
+|-------------|------:|------|----:|---|-----:|
+|truthfulqa_mc|      1|mc1   |45.90|±  |  1.74|
+|             |       |mc2   |63.44|±  |  1.56|
+Average: 63.44%
+### Bigbench
+|                      Task                      |Version|       Metric        |Value|   |Stderr|
+|------------------------------------------------|------:|---------------------|----:|---|-----:|
+|bigbench_causal_judgement                       |      0|multiple_choice_grade|58.42|±  |  3.59|
+|bigbench_date_understanding                     |      0|multiple_choice_grade|60.70|±  |  2.55|
+|bigbench_disambiguation_qa                      |      0|multiple_choice_grade|38.37|±  |  3.03|
+|bigbench_geometric_shapes                       |      0|multiple_choice_grade|21.73|±  |  2.18|
+|                                                |       |exact_str_match      | 0.00|±  |  0.00|
+|bigbench_logical_deduction_five_objects         |      0|multiple_choice_grade|35.00|±  |  2.14|
+|bigbench_logical_deduction_seven_objects        |      0|multiple_choice_grade|23.57|±  |  1.61|
+|bigbench_logical_deduction_three_objects        |      0|multiple_choice_grade|50.33|±  |  2.89|
+|bigbench_movie_recommendation                   |      0|multiple_choice_grade|45.00|±  |  2.23|
+|bigbench_navigate                               |      0|multiple_choice_grade|50.00|±  |  1.58|
+|bigbench_reasoning_about_colored_objects        |      0|multiple_choice_grade|60.35|±  |  1.09|
+|bigbench_ruin_names                             |      0|multiple_choice_grade|51.12|±  |  2.36|
+|bigbench_salient_translation_error_detection    |      0|multiple_choice_grade|32.26|±  |  1.48|
+|bigbench_snarks                                 |      0|multiple_choice_grade|67.96|±  |  3.48|
+|bigbench_sports_understanding                   |      0|multiple_choice_grade|70.59|±  |  1.45|
+|bigbench_temporal_sequences                     |      0|multiple_choice_grade|35.80|±  |  1.52|
+|bigbench_tracking_shuffled_objects_five_objects |      0|multiple_choice_grade|22.56|±  |  1.18|
+|bigbench_tracking_shuffled_objects_seven_objects|      0|multiple_choice_grade|17.20|±  |  0.90|
+|bigbench_tracking_shuffled_objects_three_objects|      0|multiple_choice_grade|50.33|±  |  2.89|
+Average: 43.96%
+Average score: 55.77%
+Elapsed time: 02:43:45
+## Citations
+Fernando Fernandes Neto and Eric Hartford. "Optimizing Large Language Models Using Layer-Selective Rank Reduction and Random Matrix Theory." 2024.
+```bibtex
+@article{sharma2023truth,
+title={The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction},
+author={Sharma, Pratyusha and Ash, Jordan T and Misra, Dipendra},
+journal={arXiv preprint arXiv:2312.13558},
+year={2023} }
+```
+```bibtex
+@article{gao2021framework,
+  title={A framework for few-shot language model evaluation},
+  author={Gao, Leo and Tow, Jonathan and Biderman, Stella and Black, Sid and DiPofi, Anthony and Foster, Charles and Golding, Laurence and Hsu, Jeffrey and McDonell, Kyle and Muennighoff, Niklas and others},
+  journal={Version v0. 0.1. Sept},
+  year={2021}
+}
 ```

config.json ADDED Viewed

	@@ -0,0 +1,30 @@

+{
+  "_name_or_path": "mlabonne/Marcoro14-7B-slerp",
+  "architectures": [
+    "MixtralForCausalLM"
+  ],
+  "attention_dropout": 0.0,
+  "bos_token_id": 1,
+  "eos_token_id": 2,
+  "hidden_act": "silu",
+  "hidden_size": 4096,
+  "initializer_range": 0.02,
+  "intermediate_size": 14336,
+  "max_position_embeddings": 32768,
+  "model_type": "mixtral",
+  "num_attention_heads": 32,
+  "num_experts_per_tok": 2,
+  "num_hidden_layers": 32,
+  "num_key_value_heads": 8,
+  "num_local_experts": 2,
+  "output_router_logits": false,
+  "rms_norm_eps": 1e-05,
+  "rope_theta": 10000.0,
+  "router_aux_loss_coef": 0.001,
+  "sliding_window": null,
+  "tie_word_embeddings": false,
+  "torch_dtype": "bfloat16",
+  "transformers_version": "4.37.0.dev0",
+  "use_cache": true,
+  "vocab_size": 32000
+}

dolphin_moe.png ADDED Viewed

Git LFS Details

SHA256: 5f82457da1aa82007718e010a67cd0d47308741efe70e61555e74ed4cbc9e34d
Pointer size: 132 Bytes
Size of remote file: 3.39 MB

mergekit_moe_config.yml ADDED Viewed

	@@ -0,0 +1,47 @@

+base_model: mlabonne/Marcoro14-7B-slerp
+gate_mode: hidden
+dtype: bfloat16
+experts:
+  - source_model: cognitivecomputations/dolphin-2.6-mistral-7b-dpo
+    positive_prompts:
+      - "Help me debug this code."
+      - "Rewrite this function in Python."
+      - "Optimize this C# script."
+      - "Implement this feature using JavaScript."
+      - "Convert this HTML structure into a more efficient design."
+      - "Assist me with writing a program that"
+      - "How do you"
+      - "Explain the concept of"
+      - "Give an overview of"
+      - "Compare and contrast between"
+      - "Provide information about"
+      - "Help me understand"
+      - "Summarize"
+      - "Make a recommendation on"
+      - "Answer this question"
+  - source_model: WizardLM/WizardMath-7B-V1.1
+    positive_prompts:
+      - "add these numbers"
+      - "whats 2+2"
+      - "subtraction"
+      - "division"
+      - "multiplication"
+      - "addition"
+      - "I need help with a math problem"
+      - "Solve for x"
+      - "Add these two numbers together: 4 + 3 = 7"
+      - "Multiply 5 by 6: 5 * 6 = 30"
+      - "Divide 8 by 2: 8 / 2 = 4"
+      - "Find the remainder when 9 is divided by 3: 9 % 3 = 0"
+      - "Calculate the square root of 16: sqrt(16) = 4"
+      - "Simplify the expression (a+b)/(c-d): (a+b)/(c-d)"
+      - "Factor out the common factor of 2 from 4x + 6y: 2(2x + 3y)"
+      - "Solve for x in the equation 3x - 7 = 2x + 5: x = 12"
+      - "Graph the line y = 2x + 3"
+      - "Approximate pi to three decimal places: 3.142"
+      - "Find the derivative of f(x) = sin(x): f'(x) = cos(x)"
+      - "Integrate g(x) = x^2 over the interval [0, 1]: g(1) - g(0) = 1/3"
+      - "Calculate the determinant of the matrix A = [[2, 3], [4, 5]]: det(A) = 2*5 - 3*4 = -2"
+      - "Solve the system of equations Ax = b: x = [-5, 10]"
+      - "Calculate the sum of the first n natural numbers using the formula Sn = n*(n+1)/2: sum(n=1 to 5) = 15"

model.safetensors.index.json ADDED Viewed

	@@ -0,0 +1 @@

+ {"metadata": {"mergekit_version": "0.0.3.2"}, "weight_map": {"model.embed_tokens.weight": "model-00001-of-00003.safetensors", "model.norm.weight": "model-00001-of-00003.safetensors", "lm_head.weight": "model-00001-of-00003.safetensors", "model.layers.0.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.1.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.2.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.3.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.4.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.5.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.6.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.7.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.8.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.9.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.10.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.11.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.12.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.13.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.14.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.15.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.16.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.17.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.18.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.19.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.20.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.21.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.22.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.23.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.24.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.25.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.26.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.27.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.28.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.29.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.30.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.31.input_layernorm.weight": "model-00001-of-00003.safetensors", "model.layers.0.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.0.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.1.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.1.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.2.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.2.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.3.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.3.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.4.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.4.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.5.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.5.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.6.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.6.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.7.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.7.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.8.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.8.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.9.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.9.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.10.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.10.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.11.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.11.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.12.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.12.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.13.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.13.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.14.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.14.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.15.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.15.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.16.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.16.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.17.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.17.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.18.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.18.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.19.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.19.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.20.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.20.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.21.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.21.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.22.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.22.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.23.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.23.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.24.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.24.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.25.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.25.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.26.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.26.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.27.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.27.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.28.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.28.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.29.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.29.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.30.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.30.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.31.block_sparse_moe.experts.0.w3.weight": "model-00001-of-00003.safetensors", "model.layers.31.block_sparse_moe.experts.1.w3.weight": "model-00001-of-00003.safetensors", "model.layers.0.block_sparse_moe.experts.0.w2.weight": "model-00001-of-00003.safetensors", "model.layers.0.block_sparse_moe.experts.1.w2.weight": "model-00001-of-00003.safetensors", "model.layers.1.block_sparse_moe.experts.0.w2.weight": "model-00001-of-00003.safetensors", "model.layers.1.block_sparse_moe.experts.1.w2.weight": "model-00001-of-00003.safetensors", "model.layers.2.block_sparse_moe.experts.0.w2.weight": "model-00001-of-00003.safetensors", "model.layers.2.block_sparse_moe.experts.1.w2.weight": "model-00001-of-00003.safetensors", "model.layers.3.block_sparse_moe.experts.0.w2.weight": "model-00001-of-00003.safetensors", "model.layers.3.block_sparse_moe.experts.1.w2.weight": "model-00001-of-00003.safetensors", "model.layers.4.block_sparse_moe.experts.0.w2.weight": "model-00001-of-00003.safetensors", "model.layers.4.block_sparse_moe.experts.1.w2.weight": "model-00001-of-00003.safetensors", "model.layers.5.block_sparse_moe.experts.0.w2.weight": "model-00001-of-00003.safetensors", "model.layers.5.block_sparse_moe.experts.1.w2.weight": "model-00001-of-00003.safetensors", "model.layers.6.block_sparse_moe.experts.0.w2.weight": "model-00001-of-00003.safetensors", "model.layers.6.block_sparse_moe.experts.1.w2.weight": "model-00001-of-00003.safetensors", "model.layers.7.block_sparse_moe.experts.0.w2.weight": "model-00001-of-00003.safetensors", "model.layers.7.block_sparse_moe.experts.1.w2.weight": "model-00001-of-00003.safetensors", "model.layers.8.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.8.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.9.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.9.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.10.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.10.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.11.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.11.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.12.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.12.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.13.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.13.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.14.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.14.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.15.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.15.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.16.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.16.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.17.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.17.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.18.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.18.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.19.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.19.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.20.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.20.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.21.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.21.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.22.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.22.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.23.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.23.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.24.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.24.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.25.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.25.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.26.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.26.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.27.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.27.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.28.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.28.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.29.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.29.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.30.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.30.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.31.block_sparse_moe.experts.0.w2.weight": "model-00002-of-00003.safetensors", "model.layers.31.block_sparse_moe.experts.1.w2.weight": "model-00002-of-00003.safetensors", "model.layers.0.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.0.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.1.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.1.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.2.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.2.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.3.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.3.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.4.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.4.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.5.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.5.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.6.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.6.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.7.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.7.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.8.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.8.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.9.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.9.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.10.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.10.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.11.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.11.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.12.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.12.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.13.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.13.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.14.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.14.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.15.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.15.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.16.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.16.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.17.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.17.block_sparse_moe.experts.1.w1.weight": "model-00002-of-00003.safetensors", "model.layers.18.block_sparse_moe.experts.0.w1.weight": "model-00002-of-00003.safetensors", "model.layers.18.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.19.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.19.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.20.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.20.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.21.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.21.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.22.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.22.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.23.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.23.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.24.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.24.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.25.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.25.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.26.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.26.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.27.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.27.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.28.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.28.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.29.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.29.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.30.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.30.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.31.block_sparse_moe.experts.0.w1.weight": "model-00003-of-00003.safetensors", "model.layers.31.block_sparse_moe.experts.1.w1.weight": "model-00003-of-00003.safetensors", "model.layers.0.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.1.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.2.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.3.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.4.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.5.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.6.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.7.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.8.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.9.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.10.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.11.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.12.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.13.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.14.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.15.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.16.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.17.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.18.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.19.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.20.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.21.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.22.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.23.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.24.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.25.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.26.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.27.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.28.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.29.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.30.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.31.post_attention_layernorm.weight": "model-00003-of-00003.safetensors", "model.layers.0.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.1.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.2.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.3.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.4.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.5.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.6.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.7.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.8.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.9.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.10.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.11.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.12.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.13.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.14.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.15.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.16.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.17.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.18.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.19.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.20.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.21.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.22.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.23.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.24.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.25.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.26.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.27.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.28.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.29.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.30.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.31.self_attn.q_proj.weight": "model-00003-of-00003.safetensors", "model.layers.0.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.1.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.2.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.3.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.4.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.5.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.6.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.7.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.8.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.9.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.10.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.11.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.12.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.13.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.14.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.15.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.16.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.17.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.18.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.19.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.20.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.21.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.22.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.23.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.24.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.25.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.26.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.27.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.28.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.29.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.30.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.31.self_attn.k_proj.weight": "model-00003-of-00003.safetensors", "model.layers.0.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.1.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.2.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.3.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.4.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.5.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.6.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.7.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.8.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.9.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.10.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.11.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.12.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.13.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.14.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.15.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.16.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.17.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.18.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.19.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.20.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.21.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.22.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.23.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.24.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.25.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.26.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.27.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.28.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.29.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.30.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.31.self_attn.v_proj.weight": "model-00003-of-00003.safetensors", "model.layers.0.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.1.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.2.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.3.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.4.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.5.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.6.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.7.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.8.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.9.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.10.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.11.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.12.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.13.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.14.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.15.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.16.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.17.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.18.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.19.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.20.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.21.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.22.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.23.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.24.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.25.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.26.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.27.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.28.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.29.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.30.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.31.self_attn.o_proj.weight": "model-00003-of-00003.safetensors", "model.layers.0.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.1.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.2.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.3.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.4.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.5.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.6.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.7.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.8.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.9.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.10.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.11.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.12.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.13.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.14.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.15.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.16.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.17.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.18.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.19.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.20.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.21.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.22.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.23.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.24.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.25.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.26.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.27.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.28.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.29.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.30.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors", "model.layers.31.block_sparse_moe.gate.weight": "model-00003-of-00003.safetensors"}}

original_repo_url.txt ADDED Viewed

	@@ -0,0 +1 @@


1	+ https://huggingface.co/macadeliccc/laser-dolphin-mixtral-2x7b-dpo

output.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:f14437974d59c393670725390a17db1b1e868897caee114b4b89d474f40cbce6
+size 5885402576

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,29 @@

+{
+  "additional_special_tokens": [
+    "<unk>",
+    "<s>",
+    "</s>"
+  ],
+  "bos_token": {
+    "content": "<s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "</s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": "<s>",
+  "unk_token": {
+    "content": "<unk>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer.model ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:dadfd56d766715c61d2ef780a525ab43b8e6da4de6865bda3d95fdef5e134055
+size 493443

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,46 @@

+{
+  "add_bos_token": true,
+  "add_eos_token": false,
+  "added_tokens_decoder": {
+    "0": {
+      "content": "<unk>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "1": {
+      "content": "<s>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "2": {
+      "content": "</s>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "additional_special_tokens": [
+    "<unk>",
+    "<s>",
+    "</s>"
+  ],
+  "bos_token": "<s>",
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "</s>",
+  "legacy": true,
+  "model_max_length": 1000000000000000019884624838656,
+  "pad_token": "<s>",
+  "sp_model_kwargs": {},
+  "spaces_between_special_tokens": false,
+  "tokenizer_class": "LlamaTokenizer",
+  "unk_token": "<unk>",
+  "use_default_system_prompt": true
+}