EpistemeAI
/

Fireball-12B-v1.13a-philosophers

Text Generation

text-generation-inference

Model card Files Files and versions Community

legolasyiu commited on Aug 29, 2024

Commit

6b8d960

·

verified ·

1 Parent(s): fe865a0

Update README.md

Files changed (1) hide show

README.md +35 -7

README.md CHANGED Viewed

@@ -88,13 +88,41 @@ f"""Below is an instruction that describes a task. \
 > ```
 If you want to use Hugging Face `transformers` to generate text, you can do something like this.
 ```py
-from transformers import AutoModelForCausalLM, AutoTokenizer
-model_id = "EpistemeAI2/Fireball-12B-v1.13a-philosophers"
-tokenizer = AutoTokenizer.from_pretrained(model_id)
-model = AutoModelForCausalLM.from_pretrained(model_id)
-inputs = tokenizer("Hello my name is", return_tensors="pt")
-outputs = model.generate(**inputs, max_new_tokens=20)
-print(tokenizer.decode(outputs[0], skip_special_tokens=True))
 ```
 ## Accelerator mode:
 ```py

 > ```
 If you want to use Hugging Face `transformers` to generate text, you can do something like this.
 ```py
+# Import necessary libraries
+from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
+import torch
+# Load the tokenizer
+tokenizer = AutoTokenizer.from_pretrained("EpistemeAI2/Fireball-12B-v1.13a-philosophers")
+quantization_config = BitsAndBytesConfig(load_in_4bit=True)
+# Load the model with 4-bit quantization (no need to use .to() later)
+model = AutoModelForCausalLM.from_pretrained(
+    "EpistemeAI2/Fireball-12B-v1.13a-philosophers",
+    quantization_config=quantization_config,
+    device_map="auto"  # Automatically map model to devices
+)
+# Define the input text
+input_text = "What is the difference between inductive and deductive reasoning?,"
+# Tokenize the input text
+input_ids = tokenizer.encode(input_text, return_tensors="pt")
+# Ensure the input tensors are moved to the correct device
+# Use the first parameter of the model to get the device it's on
+input_ids = input_ids.to(model.device)
+# Generate text using the model
+output_ids = model.generate(input_ids, max_length=100, num_return_sequences=1)
+# Decode the generated tokens to text
+output_text = tokenizer.decode(output_ids[0], skip_special_tokens=True)
+# Print the output
+print(output_text)
 ```
 ## Accelerator mode:
 ```py