anemll commited on May 7

Commit

508ce38

verified ·

1 Parent(s): 0976fe1

Upload folder using huggingface_hub

Browse files

Files changed (28) hide show

.DS_Store +0 -0
.gitattributes +1 -0
README.md +147 -0
chat.py +893 -0
chat_full.py +976 -0
config.json +4 -0
meta.yaml +23 -0
nemo__FFN_PF_chunk_01of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_02of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_03of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_04of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_05of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_06of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_07of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_08of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_09of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_10of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_11of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_12of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_13of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_14of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_15of16.mlmodelc.zip +3 -0
nemo__FFN_PF_chunk_16of16.mlmodelc.zip +3 -0
nemo__embeddings.mlmodelc.zip +3 -0
nemo__lm_head.mlmodelc.zip +3 -0
prefill.py +644 -0
tokenizer.json +3 -0
tokenizer_config.json +2063 -0

.DS_Store ADDED Viewed

Binary file (14.3 kB). View file

.gitattributes CHANGED Viewed

@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text

 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
+tokenizer.json filter=lfs diff=lfs merge=lfs -text

README.md ADDED Viewed

	@@ -0,0 +1,147 @@

+---
+license: mit
+tags:
+- coreml
+- ANE
+- DeepSeek
+- Apple
+- Apple Neural Engine
+- DeepHermes
+---
+# ANEMLL
+**ANEMLL** (pronounced like "animal") is an open-source project focused on accelerating the porting of Large Language Models (LLMs) to tensor processors, starting with the Apple Neural Engine (ANE).
+The goal is to provide a fully open-source pipeline from model conversion to inference for common LLM architectures running on ANE.
+This enables seamless integration and on-device inference for low-power applications on edge devices, ensuring maximum privacy and security.
+This is critical for autonomous applications, where models run directly on the device without requiring an internet connection.
+For more information, visit the [ANEMLL GitHub repository](https://github.com/anemll/anemll).
+---
+## License
+ANEMLL is licensed under the [MIT License](https://opensource.org/license/mit).
+The model is based on Meta's LLaMA 3.2 and may require a separate license.
+This test model is exclusively for the Meta's LLaMA architecture  converted for CoreML, released before the official launch of the ANEMLL repository and minimal documentation. It is intended for early adopters only who requested an early release.
+---
+## Requirements
+- **macOS Sequoia** with Apple Neural Engine and 8GB RAM or more
+- **CoreML Tools** and **HuggingFace Transformers** libraries
+- **Python 3.9**
+`chat.py` provides a sample inference script.
+`chat_full.py` provides a sample inference script with history and conversation management.
+**Installation**
+1. Download the model from Hugging Face:
+```bash
+# Install required tools
+pip install huggingface_hub
+# Install Git LFS (Large File Support)
+# macOS with Homebrew:
+brew install git-lfs
+# Or Ubuntu/Debian:
+# sudo apt-get install git-lfs
+# Initialize Git LFS
+git lfs install
+# Clone the repository with model files
+git clone https://huggingface.co/anemll/anemll-Llama-3.1-Nemotron-Nano-8B-v1-ctx512_0.3.0
+```
+2. Extract model files:
+```bash
+# Navigate to cloned directory
+cd anemll-Llama-3.1-Nemotron-Nano-8B-v1-ctx512_0.3.0
+# Pull LFS files (model weights)
+git lfs pull
+# Extract CoreML model files
+find . -type f -name "*.zip" -exec unzip {} \;
+```
+3. Install dependencies:
+```bash
+pip install coremltools transformers
+```
+**Coremltools:**
+See coremltools installation guide at https://coremltools.readme.io/v4.0/docs/installation
+**How to Run**
+1. Basic chat interface:
+```bash
+python chat.py --meta ./meta.yaml
+```
+2. Full conversation mode with history:
+```bash
+python chat_full.py --meta ./meta.yaml
+```
+> Note: The first time the model loads, macOS will take some time to place it on the device.
+> Subsequent loads will be instantaneous.
+> Use Ctrl-D to exit, Ctrl-C to interrupt inference.
+**More Info**
+Please check following links for later updates:
+* [GitHub](https://github.com/anemll)
+* [Hugging Face Models](https://huggingface.co/anemll)
+* [Twitter/X](https://x.com/anemll)
+* [Website](https://anemll.com)
+[email protected]
+# anemll-Llama-3.1-Nemotron-Nano-8B-v1-ctx512_0.3.0
+This is a CoreML model converted using ANEMLL for Apple Neural Engine inference.
+## Available Distributions
+### Standard Distribution
+- Contains zipped MLMODELC files
+- Suitable for macOS and development
+### iOS Distribution
+- Contains unzipped MLMODELC files
+- Ready for iOS deployment
+- Includes offline tokenizer support
+## Model Information
+- Context Length: %CONTEXT_LENGTH%
+- Batch Size: %BATCH_SIZE%
+- Number of Chunks: %NUM_CHUNKS%
+## Quick Start
+### Test in iOS/macOS App
+Try our sample Chat-Bot app on TestFlight:
+1. Install TestFlight from App Store
+2. Join beta test: [TestFlight Link](https://testflight.apple.com/join/jrQq1D1C)
+3. App includes a small demo model pre-installed
+4. You can add custom models via HuggingFace URLs
+> [!Note]
+> - The TestFlight app works on both iOS and macOS
+> - Demonstrates proper model integration and provides a reference implementation
+> - iOS requires unzipped MLMODELC files and config.json for offline tokenizer
+> - macOS supports both zipped and unzipped model formats
+```

chat.py ADDED Viewed

	@@ -0,0 +1,893 @@

+# chat.py
+#!/usr/bin/env python3
+# chat.py
+# Copyright (c) 2025 Anemll
+# Licensed under the MIT License
+import argparse
+import os
+import re
+import glob
+from pathlib import Path
+import coremltools as ct
+from transformers import LlamaTokenizer, AutoTokenizer
+import torch
+import torch.nn.functional as F
+import numpy as np
+import queue
+import threading
+import time
+import yaml
+import sys
+# ANSI color codes
+LIGHT_BLUE = "\033[94m"
+DARK_BLUE = "\033[34m"
+LIGHT_GREEN = "\033[92m"
+RESET_COLOR = "\033[0m"
+# Add at top with other constants
+WARMUP_TOKEN_LIMIT = 10  # Maximum tokens to generate during warmup
+class TokenPrinter:
+    """Handles background printing of generated tokens."""
+    def __init__(self, tokenizer):
+        self.tokenizer = tokenizer
+        self.token_queue = queue.Queue()
+        self.stop_event = threading.Event()
+        self.thread = None
+        self.buffer = ""
+        self.lock = threading.Lock()
+        self.thinking = True  # Track if we're still in thinking mode
+        self.decoding_buffer = []  # Buffer for token IDs
+        # Add token counting and timing
+        self.start_time = time.time()
+        self.token_count = 0
+        self.start()
+    def start(self):
+        """Start the printer thread."""
+        if self.thread is None:
+            self.thread = threading.Thread(target=self._print_worker)
+            self.thread.daemon = True
+            self.thread.start()
+    def add_token(self, token_id):
+        """Add a token to the print queue."""
+        if not self.stop_event.is_set():
+            self.token_queue.put(token_id)
+            self.token_count += 1
+    def drain_buffer(self):
+        """Decode token IDs from decoding_buffer in the main thread."""
+        if not self.decoding_buffer:
+            return
+        # Decode all tokens at once in the main thread
+        token_str = self.tokenizer.decode(self.decoding_buffer)
+        self.decoding_buffer.clear()
+        # Store the text in buffer for later saving to file
+        with self.lock:
+            self.buffer += token_str
+        # Color-handling logic
+        if self.thinking and "</think>" in token_str:
+            self.thinking = False
+            parts = token_str.split("</think>")
+            if len(parts) > 0:
+                print(parts[0] + "</think>", end='', flush=True)
+                if len(parts) > 1:
+                    print(LIGHT_BLUE + parts[1], end='', flush=True)
+        else:
+            if not self.thinking:
+                print(LIGHT_BLUE + token_str, end='', flush=True)
+            else:
+                print(token_str, end='', flush=True)
+    def _print_worker(self):
+        """Worker thread that takes token_ids from the queue."""
+        while not self.stop_event.is_set():
+            try:
+                token_id = self.token_queue.get(timeout=0.01)
+                with self.lock:
+                    self.decoding_buffer.append(token_id)
+                self.token_queue.task_done()
+            except queue.Empty:
+                continue
+            except Exception as e:
+                print(f"\nError: Token printer error: {str(e)}")
+                break
+    def stop(self):
+        """Stop the printer thread."""
+        if self.thread and self.thread.is_alive():
+            # Ensure any remaining tokens are processed
+            self.drain_buffer()
+            self.stop_event.set()
+            try:
+                self.thread.join(timeout=1.0)
+            except Exception:
+                pass
+            # Calculate and print tokens/s with shorter format in blue
+            elapsed = time.time() - self.start_time
+            if elapsed > 0 and self.token_count > 0:
+                tokens_per_sec = self.token_count / elapsed
+                print(f"\n{DARK_BLUE}{tokens_per_sec:.1f} t/s{RESET_COLOR}")
+            else:
+                print(RESET_COLOR)  # Reset color at the end
+        return self.buffer
+def parse_model_path(path):
+    """Parse model path and return full path with .mlmodelc or .mlpackage extension."""
+    path = Path(path)
+    # If path exists exactly as specified, return it
+    if path.exists():
+        return str(path)
+    # Try with both extensions
+    candidates = [
+        path,  # Original path
+        path.with_suffix('.mlmodelc'),  # With .mlmodelc
+        path.with_suffix('.mlpackage'),  # With .mlpackage
+        Path(str(path) + '.mlmodelc'),  # Handle case where extension is included
+        Path(str(path) + '.mlpackage')
+    ]
+    # Try all possible paths
+    for candidate in candidates:
+        if candidate.exists():
+            print(f"Found model at: {candidate}")
+            return str(candidate)
+    # If we get here, no valid path was found
+    print("\nError: Model not found. Tried following paths:")
+    for candidate in candidates:
+        print(f"  {candidate}")
+    raise FileNotFoundError(f"Model not found: {path}")
+def parse_ffn_filename(path):
+    """Parse FFN model filename to extract chunk information."""
+    path = Path(path)
+    pattern = r'FFN_PF.*_chunk_(\d+)of(\d+)'
+    match = re.search(pattern, path.name)
+    if match:
+        current_chunk = int(match.group(1))
+        total_chunks = int(match.group(2))
+        return current_chunk, total_chunks
+    return None, None
+def find_all_chunks(base_path):
+    """Find all chunk files matching the base FFN path pattern."""
+    path = Path(base_path)
+    pattern = re.sub(r'_chunk_\d+of\d+', '_chunk_*', str(path))
+    return sorted(glob.glob(pattern))
+def load_model(path, function_name=None):
+    """Load a CoreML model, handling both .mlmodelc and .mlpackage formats."""
+    path = Path(path)
+    compute_unit = ct.ComputeUnit.CPU_AND_NE
+    try:
+        if path.suffix == '.mlmodelc':
+            # For compiled models (.mlmodelc), use CompiledMLModel
+            if function_name:
+                return ct.models.CompiledMLModel(str(path), compute_unit, function_name=function_name)
+            else:
+                return ct.models.CompiledMLModel(str(path), compute_unit)
+        else:
+            # For packages (.mlpackage)
+            if function_name:
+                return ct.models.MLModel(str(path), function_name=function_name)
+            else:
+                return ct.models.MLModel(str(path))
+    except RuntimeError as e:
+        if "valid manifest does not exist" in str(e):
+            print(f"\nError: Could not load compiled model at {path}")
+            print("This might be because:")
+            print("1. The model is not properly compiled")
+            print("2. The model was compiled for a different OS version")
+            print("3. The model needs to be recompiled")
+            print("\nTry using the .mlpackage version instead, or recompile the model.")
+        raise
+def load_metadata(model,args):
+    # Extract metadata and config parameters
+    metadata = {}
+    if hasattr(model, 'user_defined_metadata'):
+        meta = model.user_defined_metadata
+        # Extract key parameters with defaults
+        metadata['context_length'] = int(meta.get('com.anemll.context_length', 512))
+        metadata['state_length'] = int(meta.get('com.anemll.state_length', metadata['context_length']))  # Added state_length
+        metadata['batch_size'] = int(meta.get('com.anemll.batch_size', 64))
+        metadata['lut_bits'] = int(meta.get('com.anemll.lut_bits', 0))
+        metadata['num_chunks'] = int(meta.get('com.anemll.num_chunks', 1))
+        print("\nExtracted Parameters:")
+        print(f"  Context Length: {metadata['context_length']}")
+        print(f"  State Length: {metadata['state_length']}")
+        print(f"  Prefill Batch Size: {metadata['batch_size']}")
+        print(f"  LUT Bits: {metadata['lut_bits']}")
+        print(f"  Number of Chunks: {metadata['num_chunks']}")
+        # Print model info
+        print("\nModel Info:")
+        if 'com.anemll.info' in meta:
+            print(f"  {meta['com.anemll.info']}")
+        if 'com.github.apple.coremltools.version' in meta:
+            print(f"  CoreML Tools: {meta['com.github.apple.coremltools.version']}")
+        # Print model input/output shapes
+        print("\nModel Shapes:")
+        if hasattr(model, 'input_description'):
+            print("  Inputs:")
+            for name, desc in model.input_description.items():
+                print(f"    {name}: {desc}")
+        if hasattr(model, 'output_description'):
+            print("  Outputs:")
+            for name, desc in model.output_description.items():
+                print(f"    {name}: {desc}")
+    else:
+        print("\nWarning: No metadata found in model")
+        # Check if model directory name contains context length pattern (ctxXXX)
+        ctx_len = 512
+        if args.context_length is  None:
+            import re
+            ctx_match = re.search(r'ctx(\d+)', str(args.d))
+            if ctx_match:
+                ctx_len0 = int(ctx_match.group(1))
+                if 512 <= ctx_len0 <= 8096:
+                    ctx_len = ctx_len0
+                    print(f"\nDetected context length {ctx_len} from directory name")
+            else:
+                print(f"\nWarning: No context length found in directory  {ctx_len} from directory name {args.d}")
+        else:
+            ctx_len = args.context_length
+        # Use defaults or values from args
+        metadata['context_length'] = ctx_len
+        metadata['state_length'] = ctx_len
+        # Get batch size from args or use default
+        metadata['batch_size'] = getattr(args, 'batch_size', 64)
+        metadata['lut_bits'] = 4
+        metadata['num_chunks'] = getattr(args, 'num_chunks', 4)
+        print("\nUsing parameters:")
+        print(f"  Context Length: {metadata['context_length']}")
+        print(f"  State Length: {metadata['state_length']}")
+        print(f"  Prefill Batch Size: {metadata['batch_size']}")
+        print(f"  LUT Bits: {metadata['lut_bits']}")
+        print(f"  Number of Chunks: {metadata['num_chunks']}")
+    # Override with values from args if they exist
+    if hasattr(args, 'batch_size') and args.batch_size is not None:
+        metadata['batch_size'] = args.batch_size
+        print(f"\nOverriding batch size from args: {args.batch_size}")
+    if hasattr(args, 'num_chunks') and args.num_chunks is not None:
+        metadata['num_chunks'] = args.num_chunks
+        print(f"\nOverriding num chunks from args: {args.num_chunks}")
+    return metadata
+def load_models(args,metadata):
+    """Load all required models and extract metadata."""
+    print("\nLoading models...")
+    try:
+        # Load embeddings model
+        print("\nLoading embeddings model...")
+        embed_path = parse_model_path(args.embed)
+        print(f"Loading from: {embed_path}")
+        embed_model = load_model(embed_path)
+        print("Embeddings model loaded successfully")
+        metadata = load_metadata(embed_model,args)
+        # Load LM head model
+        print("\nLoading LM head model...")
+        lmhead_path = parse_model_path(args.lmhead)
+        print(f"Loading from: {lmhead_path}")
+        lmhead_model = load_model(lmhead_path)
+        print("LM head model loaded successfully")
+        # Parse FFN path and find chunks if needed
+        print("\nLoading FFN+PREFILL model(s)...")
+        ffn_path = parse_model_path(args.ffn)
+        chunk_no, total_chunks = parse_ffn_filename(ffn_path)
+        ffn_models = []
+        if chunk_no and total_chunks:
+            print(f"\nDetected chunked FFN+PREFILL model ({total_chunks} chunks)")
+            # Find and load all chunks
+            chunk_paths = find_all_chunks(ffn_path)
+            if len(chunk_paths) != total_chunks:
+                raise ValueError(f"Found {len(chunk_paths)} chunks but filename indicates {total_chunks} chunks")
+            for chunk_path in chunk_paths:
+                print(f"\nLoading FFN+PREFILL chunk: {Path(chunk_path).name}")
+                try:
+                    # For chunked models, we need both infer and prefill functions
+                    ffn_models.append({
+                        'infer': load_model(chunk_path, function_name='infer'),
+                        'prefill': load_model(chunk_path, function_name='prefill')
+                    })
+                    print("Chunk loaded successfully")
+                except Exception as e:
+                    print(f"Error loading chunk {chunk_path}: {str(e)}")
+                    raise
+            metadata = load_metadata(ffn_models[0],args)
+        else:
+            print("\nLoading single FFN model...")
+            ffn_models.append(load_model(ffn_path))
+            print("FFN model loaded successfully")
+        return embed_model, ffn_models, lmhead_model, metadata
+    except Exception as e:
+        print(f"\nError loading models: {str(e)}")
+        print("\nPlease ensure all model files exist and are accessible.")
+        print("Expected files:")
+        print(f"  Embeddings: {args.embed}")
+        print(f"  LM Head: {args.lmhead}")
+        print(f"  FFN: {args.ffn}")
+        raise
+# At the top of the file, make this a default path
+def initialize_tokenizer(model_path=None):
+    """Initialize and configure the tokenizer."""
+    try:
+        tokenizer = AutoTokenizer.from_pretrained(
+            str(model_path),
+            use_fast=False,
+            trust_remote_code=True
+        )
+        print("\nTokenizer Configuration:")
+        print(f"Tokenizer type: {type(tokenizer)}")
+        print(f"Tokenizer name: {tokenizer.__class__.__name__}")
+        print(f"Vocabulary size: {len(tokenizer)}")
+        print(f"Model max length: {tokenizer.model_max_length}")
+        if tokenizer.pad_token is None:
+            tokenizer.pad_token = tokenizer.eos_token
+            tokenizer.pad_token_id = tokenizer.eos_token_id
+            print("Set PAD token to EOS token")
+        tokenizer.padding_side = "left"
+        print(f"\nSpecial Tokens:")
+        print(f"PAD token: '{tokenizer.pad_token}' (ID: {tokenizer.pad_token_id})")
+        print(f"EOS token: '{tokenizer.eos_token}' (ID: {tokenizer.eos_token_id})")
+        print(f"BOS token: '{tokenizer.bos_token}' (ID: {tokenizer.bos_token_id})")
+        print(f"UNK token: '{tokenizer.unk_token}' (ID: {tokenizer.unk_token_id})")
+        return tokenizer
+    except Exception as e:
+        print(f"\nError: Failed to load tokenizer from {model_path}")
+        print(f"Error details: {str(e)}")
+        print(f"Error type: {type(e)}")
+        print("\nThis code requires a Llama 3.2 model for chat template functionality.")
+        print("Please provide the path to a Llama 3.2 model directory.")
+        import traceback
+        traceback.print_exc()
+        raise
+def make_causal_mask(length, start):
+    """Create causal attention mask."""
+    mask = np.full((1, 1, length, length), -np.inf, dtype=np.float16)
+    row_indices = np.arange(length).reshape(length, 1)
+    col_indices = np.arange(length).reshape(1, length)
+    mask[:, :, col_indices <= (row_indices + start)] = 0
+    return mask
+def initialize_causal_mask(context_length):
+    """Initialize causal mask for transformer attention."""
+    causal_mask = make_causal_mask(context_length, 0)
+    causal_mask = torch.tensor(causal_mask, dtype=torch.float16)
+    print(f"\nInitialized causal mask for context length {context_length}")
+    return causal_mask
+def run_prefill(embed_model, ffn_models, input_ids, context_pos, context_length, batch_size=64, state=None, causal_mask=None):
+    """Run prefill on the input sequence."""
+    # Use provided causal mask or create one if not provided
+    if causal_mask is None:
+        causal_mask = make_causal_mask(context_length, 0)
+        causal_mask = torch.tensor(causal_mask, dtype=torch.float16)
+    # Process in batches
+    batch_pos = 0
+    while batch_pos < context_pos:
+        batch_end = min(batch_pos + batch_size, context_pos)
+        current_batch_size = batch_end - batch_pos
+        # Get current batch
+        batch_input = input_ids[:, batch_pos:batch_end]
+        # Always pad to full batch size for prefill
+        batch_input = F.pad(
+            batch_input,
+            (0, batch_size - current_batch_size),
+            value=0
+        )
+        # Generate position IDs for full batch size
+        position_ids = torch.arange(batch_size, dtype=torch.int32)  # Changed: Always use full batch size
+        batch_causal_mask = causal_mask[:, :, :batch_size, :]  # Changed: Use full batch size
+        # Run embeddings with proper batch size
+        hidden_states = torch.from_numpy(
+            embed_model.predict({
+                'input_ids': batch_input.numpy(),
+                'batch_size': np.array([batch_size], dtype=np.int32)  # Add batch_size parameter
+            })['hidden_states']
+        )
+        # Run through FFN chunks with state
+        for ffn_model in ffn_models:
+            if isinstance(ffn_model, dict):
+                inputs = {
+                    'hidden_states': hidden_states.numpy(),  # [1, 64, hidden_size]
+                    'position_ids': position_ids.numpy(),    # [64]
+                    'causal_mask': batch_causal_mask.numpy(), # [1, 1, 64, context_length]
+                    'current_pos': np.array([batch_pos], dtype=np.int32)  # [1]
+                }
+                output = ffn_model['prefill'].predict(inputs, state)
+                hidden_states = torch.from_numpy(output['output_hidden_states'])
+        batch_pos = batch_end
+    return torch.tensor([context_pos], dtype=torch.int32)
+def generate_next_token(embed_model, ffn_models, lmhead_model, input_ids, pos, context_length, state=None, causal_mask=None, temperature=0.0):
+    """Generate the next token."""
+    # Get current token
+    current_token = input_ids[:, pos-1:pos]  # [1, 1]
+    # Run embeddings
+    hidden_states = torch.from_numpy(
+        embed_model.predict({'input_ids': current_token.numpy()})['hidden_states']
+    )  # [1, 1, hidden_size]
+    # Create masks
+    update_mask = torch.zeros((1, 1, context_length, 1), dtype=torch.float16)
+    update_mask[0, 0, pos-1, 0] = 1.0
+    position_ids = torch.tensor([pos-1], dtype=torch.int32)  # [1]
+    # Use provided causal mask or create one if not provided
+    if causal_mask is None:
+        causal_mask_data = make_causal_mask(context_length, 0)
+        single_causal_mask = torch.tensor(causal_mask_data[:, :, pos-1:pos, :], dtype=torch.float16)  # [1, 1, 1, context_length]
+    else:
+        single_causal_mask = causal_mask[:, :, pos-1:pos, :]
+    # Run through FFN chunks with state
+    for ffn_model in ffn_models:
+        if isinstance(ffn_model, dict):
+            inputs = {
+                'hidden_states': hidden_states.numpy(),
+                'update_mask': update_mask.numpy(),
+                'position_ids': position_ids.numpy(),
+                'causal_mask': single_causal_mask.numpy(),
+                'current_pos': position_ids.numpy()
+            }
+            output = ffn_model['infer'].predict(inputs, state)
+            hidden_states = torch.from_numpy(output['output_hidden_states'])
+    # Run LM head
+    lm_output = lmhead_model.predict({'hidden_states': hidden_states.numpy()})
+    # Debug print
+    #print("\nLM Head output keys:", list(lm_output.keys()))
+    # Combine logits1-8 if they exist
+    if 'logits1' in lm_output:
+        # Concatenate all logits parts
+        logits_parts = []
+        for i in range(1, 9):
+            key = f'logits{i}'
+            if key in lm_output:
+                logits_parts.append(torch.from_numpy(lm_output[key]))
+        logits = torch.cat(logits_parts, dim=-1)  # Concatenate along vocab dimension
+    else:
+        # Try output_logits as fallback
+        logits = torch.from_numpy(lm_output['output_logits'])
+    # Apply temperature and sample
+    if temperature > 0:
+        logits = logits / temperature
+        probs = F.softmax(logits[0, -1, :], dim=-1)
+        next_token = torch.multinomial(probs, num_samples=1).item()
+    else:
+        next_token = torch.argmax(logits[0, -1, :]).item()
+    return next_token
+def create_unified_state(ffn_models, context_length):
+    """Create unified KV cache state for transformer."""
+    if isinstance(ffn_models[0], dict):
+        # Use first FFN model's prefill function to create state
+        state = ffn_models[0]['prefill'].make_state()
+        print(f"\nCreated unified transformer state for {len(ffn_models)} chunks")
+        return state
+    else:
+        state = ffn_models[0].make_state()
+        print("\nCreated unified transformer state")
+        return state
+def chat_loop(embed_model, ffn_models, lmhead_model, tokenizer, metadata, state, causal_mask=None, auto_prompt=None, warmup=False, save_file=None):
+    """Interactive chat loop."""
+    context_length = metadata.get('context_length')
+    batch_size = metadata.get('batch_size', 64)
+    if not warmup:
+        print(f"\nUsing context length: {context_length}")
+        print("\nStarting chat session. Press Ctrl+D to exit.")
+        print("Type your message and press Enter to chat.")
+    # Check if tokenizer has chat template and if it works
+    has_chat_template = False
+    try:
+        # Test if chat template works
+        test_messages = [{"role": "user", "content": "test"}]
+        tokenizer.apply_chat_template(test_messages, return_tensors="pt")
+        has_chat_template = True
+        if not warmup:
+            print("\nUsing chat template for prompts")
+    except:
+        if not warmup:
+            print("\nUsing manual formatting for prompts")
+    conversation = []
+    try:
+        while True:
+            try:
+                if not warmup:
+                    print(f"\n{LIGHT_GREEN}You:{RESET_COLOR}", end=' ', flush=True)
+                if auto_prompt is not None:
+                    user_input = auto_prompt
+                    if not warmup:
+                        print(user_input)
+                else:
+                    user_input = input().strip()
+            except EOFError:
+                if not warmup:
+                    print("\nExiting chat...")
+                break
+            if not user_input:
+                continue
+            # Format prompt based on tokenizer capabilities
+            if has_chat_template:
+                messages = [{"role": "user", "content": user_input}]
+                input_ids = tokenizer.apply_chat_template(
+                    messages,
+                    return_tensors="pt",
+                    add_generation_prompt=True
+                ).to(torch.int32)
+            else:
+                # Manual formatting for Llama models without chat template
+                formatted_prompt = f"[INST] {user_input} [/INST]"
+                input_ids = tokenizer(
+                    formatted_prompt,
+                    return_tensors="pt",
+                    add_special_tokens=True
+                ).input_ids.to(torch.int32)
+            context_pos = input_ids.size(1)
+            if not warmup:
+                print(f"\n{LIGHT_BLUE}Assistant:{RESET_COLOR}", end=' ', flush=True)
+            # Initialize token printer
+            token_printer = TokenPrinter(tokenizer)
+            tokens_generated = 0  # Track number of tokens
+            try:
+                # Start prefill timing
+                prefill_start = time.time()
+                # Run prefill with state and causal mask
+                current_pos = run_prefill(
+                    embed_model,
+                    ffn_models,
+                    input_ids,
+                    context_pos,
+                    context_length,
+                    batch_size,
+                    state,
+                    causal_mask
+                )
+                # Calculate prefill timing
+                prefill_time = time.time() - prefill_start
+                prefill_tokens = context_pos  # Number of tokens in input
+                prefill_tokens_per_sec = prefill_tokens / prefill_time if prefill_time > 0 else 0
+                # Generation loop with state
+                input_ids = input_ids
+                pos = context_pos
+                inference_start = time.time()
+                inference_tokens = 0
+                while pos < context_length - 1:
+                    # Generate next token with causal mask
+                    next_token = generate_next_token(
+                        embed_model,
+                        ffn_models,
+                        lmhead_model,
+                        input_ids,
+                        pos,
+                        context_length,
+                        state,
+                        causal_mask
+                    )
+                    # Add token to sequence
+                    if pos < input_ids.size(1):
+                        input_ids[0, pos] = next_token
+                    else:
+                        input_ids = torch.cat([
+                            input_ids,
+                            torch.tensor([[next_token]], dtype=torch.int32)
+                        ], dim=1)
+                    # Add to printer only if not in warmup
+                    if not warmup:
+                        token_printer.add_token(next_token)
+                        token_printer.drain_buffer()
+                    pos += 1
+                    tokens_generated += 1
+                    inference_tokens += 1
+                    # Check limits
+                    if warmup and tokens_generated >= WARMUP_TOKEN_LIMIT:
+                        break
+                    if next_token == tokenizer.eos_token_id:
+                        break
+                # Calculate inference timing
+                inference_time = time.time() - inference_start
+                inference_tokens_per_sec = inference_tokens / inference_time if inference_time > 0 else 0
+                # Get final response and add to conversation
+                if not warmup:
+                    response = token_printer.stop()
+                    # Print timing stats
+                    prefill_ms = prefill_time * 1000  # Convert to milliseconds
+                    print(f"\nPrefill: {prefill_ms:.1f}ms ({prefill_tokens_per_sec:.1f} t/s)")
+                    print(f"Inference: {inference_tokens_per_sec:.1f} t/s")
+                    print(f"Total: Generated {tokens_generated} tokens in {prefill_time + inference_time:.2f}s")
+                    conversation.append({"role": "assistant", "content": response})
+                    # Save response to file if requested
+                    if save_file:
+                        try:
+                            # Add small delay to ensure all tokens are processed
+                            time.sleep(0.5)
+                            # Make sure response ends with EOS token if it's supposed to
+                            if response and not response.endswith("<|eot_id|>") and not response.endswith("</s>"):
+                                if tokenizer.eos_token:
+                                    eos_text = tokenizer.decode([tokenizer.eos_token_id])
+                                    if not response.endswith(eos_text):
+                                        print(f"\n{DARK_BLUE}Adding missing EOS token for consistency{RESET_COLOR}")
+                                        response += eos_text
+                            with open(save_file, 'w') as f:
+                                f.write(response)
+                            print(f"\n{DARK_BLUE}Response saved to file: {save_file}{RESET_COLOR}")
+                        except Exception as e:
+                            print(f"\n{DARK_BLUE}Error saving to file: {str(e)}{RESET_COLOR}")
+                else:
+                    token_printer.stop()  # Clean up without printing stats
+                # Exit after one response in auto_prompt mode
+                if auto_prompt is not None:
+                    break
+            except KeyboardInterrupt:
+                print("\nGeneration interrupted")
+                token_printer.stop()
+                continue
+    except Exception as e:
+        print(f"\nError in chat loop: {str(e)}")
+        import traceback
+        traceback.print_exc()
+def parse_args():
+    parser = argparse.ArgumentParser(description='Chat with CoreML LLaMA, gil resolved  (c) 2025 Anemll')
+    # Add meta.yaml option
+    parser.add_argument('--meta', type=str, help='Path to meta.yaml to load all parameters')
+    # Model paths
+    parser.add_argument('--d', '--dir', type=str, default='.',
+                       help='Directory containing model files (default: current directory)')
+    parser.add_argument('--embed', type=str, required=False,
+                       help='Path to embeddings model (relative to --dir)')
+    parser.add_argument('--ffn', type=str, required=False,
+                       help='Path to FFN model (can be chunked, relative to --dir)')
+    parser.add_argument('--lmhead', type=str, required=False,
+                       help='Path to LM head model (relative to --dir)')
+    parser.add_argument('--tokenizer', type=str, required=False,
+                       help='Path to tokenizer')
+    # Add new argument for auto-generation
+    parser.add_argument('--prompt', type=str,
+                       help='If specified, run once with this prompt and exit')
+    # Add save option
+    parser.add_argument('--save', type=str,
+                       help='Save assistant\'s response to specified file')
+    # Add no-warmup flag
+    parser.add_argument('--nw', action='store_true',
+                       help='Skip warmup phase')
+    # Model configuration
+    parser.add_argument('--context-length', type=int,
+                       help='Context length for the model (default: 512), if not provided, it will be detected from the model directory name ctxNUMBER')
+    parser.add_argument('--batch-size', type=int,
+                       help='Batch size for prefill (default: 64)')
+    args = parser.parse_args()
+    # If meta.yaml is provided, load parameters from it
+    if args.meta:
+        try:
+            with open(args.meta, 'r') as f:
+                meta = yaml.safe_load(f)
+            params = meta['model_info']['parameters']
+            # Set model directory to meta.yaml directory if not specified
+            if not args.d or args.d == '.':
+                args.d = str(Path(args.meta).parent)
+            # Build model paths based on parameters
+            prefix = params.get('model_prefix', 'llama')  # Default to 'llama' if not specified
+            lut_ffn = f"_lut{params['lut_ffn']}" if params['lut_ffn'] != 'none' else ''
+            lut_lmhead = f"_lut{params['lut_lmhead']}" if params['lut_lmhead'] != 'none' else ''
+            lut_embeddings = f"_lut{params['lut_embeddings']}" if params['lut_embeddings'] != 'none' else ''
+            num_chunks = int(params['num_chunks'])
+            # Set model paths if not specified
+            if not args.lmhead:
+                args.lmhead = f'{prefix}_lm_head{lut_lmhead}'
+            if not args.embed:
+                args.embed = f'{prefix}_embeddings{lut_embeddings}'  # Changed from lm_head to embeddings
+            if not args.ffn:
+                args.ffn = f'{prefix}_FFN_PF{lut_ffn}_chunk_01of{num_chunks:02d}'
+            if not args.tokenizer:
+                args.tokenizer = args.d
+            # Set other parameters if not overridden by command line
+            if args.context_length is None:
+                args.context_length = int(params['context_length'])
+            if args.batch_size is None:
+                args.batch_size = int(params['batch_size'])
+            args.num_chunks = num_chunks
+            print(f"\nLoaded parameters from {args.meta}:")
+            print(f"  Context Length: {args.context_length}")
+            print(f"  Batch Size: {args.batch_size}")
+            print(f"  Num Chunks: {args.num_chunks}")
+            print(f"  Models Directory: {args.d}")
+            print(f"  Embeddings: {args.embed}")
+            print(f"  LM Head: {args.lmhead}")
+            print(f"  FFN: {args.ffn}")
+        except Exception as e:
+            print(f"\nError loading meta.yaml: {str(e)}")
+            sys.exit(1)
+    return args
+def main():
+    args = parse_args()
+    # Convert directory to absolute path
+    model_dir = Path(args.d).resolve()
+    if not model_dir.exists():
+        print(f"\nError: Model directory not found: {model_dir}")
+        return 1
+    print(f"\nUsing model directory: {model_dir}")
+    print(f"Context length: {args.context_length}")
+    try:
+        # Update paths to be relative to model directory
+        args.embed = str(model_dir / args.embed)
+        args.ffn = str(model_dir / args.ffn)
+        args.lmhead = str(model_dir / args.lmhead)
+        # Handle tokenizer path separately since it's not relative to model_dir
+        if args.tokenizer is None:
+            args.tokenizer = str(model_dir)
+        if not Path(args.tokenizer).exists():
+            print(f"\nError: Tokenizer directory not found: {args.tokenizer}")
+            return 1
+        args.tokenizer = str(Path(args.tokenizer).resolve())  # Convert to absolute path
+        print(f"Using tokenizer path: {args.tokenizer}")
+        metadata = {}
+        # Load models and extract metadata
+        embed_model, ffn_models, lmhead_model, metadata = load_models(args,metadata)
+        print(f"\nMetadata befor args.context_length: {metadata}")
+        # Override context length from command line if provided
+        if args.context_length is not None:
+            metadata['context_length'] = args.context_length
+            metadata['state_length'] = args.context_length  # Also update state_length
+            print(f"\nOverriding context length from command line: {args.context_length}")
+        print(f"\nMetadata after load_models: {metadata}")
+        # Load tokenizer with resolved path
+        tokenizer = initialize_tokenizer(args.tokenizer)
+        if tokenizer is None:
+            raise RuntimeError("Failed to initialize tokenizer")
+        # Create unified state once
+        state = create_unified_state(ffn_models, metadata['context_length'])
+        # Initialize causal mask once
+        causal_mask = initialize_causal_mask(metadata['context_length'])
+        # Warmup runs to prevent Python GIL issues with CoreML !
+        if not args.nw:
+            for i in range(2):
+                chat_loop(
+                    embed_model=embed_model,
+                    ffn_models=ffn_models,
+                    lmhead_model=lmhead_model,
+                    tokenizer=tokenizer,
+                    metadata=metadata,
+                    state=state,
+                    causal_mask=causal_mask,  # Pass the causal mask
+                    warmup=True,
+                    auto_prompt="who are you?"
+                )
+        # Main run
+        chat_loop(
+            embed_model=embed_model,
+            ffn_models=ffn_models,
+            lmhead_model=lmhead_model,
+            tokenizer=tokenizer,
+            metadata=metadata,
+            state=state,
+            causal_mask=causal_mask,  # Pass the causal mask
+            warmup=False,
+            auto_prompt=args.prompt,
+            save_file=args.save
+        )
+    except Exception as e:
+        print(f"\nError: {str(e)}")
+        import traceback
+        traceback.print_exc()
+        return 1
+    return 0
+if __name__ == "__main__":
+    exit(main())

chat_full.py ADDED Viewed

	@@ -0,0 +1,976 @@

+# chat.py
+#!/usr/bin/env python3
+# chat.py
+# Copyright (c) 2025 Anemll
+# Licensed under the MIT License
+import argparse
+import os
+import re
+import glob
+from pathlib import Path
+import coremltools as ct
+from transformers import LlamaTokenizer, AutoTokenizer
+import torch
+import torch.nn.functional as F
+import numpy as np
+import queue
+import threading
+import time
+import yaml
+import sys
+# ANSI color codes
+LIGHT_BLUE = "\033[94m"
+DARK_BLUE = "\033[34m"
+LIGHT_GREEN = "\033[92m"
+RESET_COLOR = "\033[0m"
+# Add at the top with other constants
+WARMUP_TOKEN_LIMIT = 10  # Maximum tokens to generate during warmup
+THINKING_MODE = False
+THINKING_PROMPT = """You are a deep thinking AI, you may use extremely long chains of thought to deeply consider the problem and deliberate with yourself via systematic reasoning processes to help come to a correct solution prior to answering. You should enclose your thoughts and internal monologue inside <think> </think> tags, and then provide your solution or response to the problem."""
+DEBUG_LEVEL = 0  # Default debug level
+class TokenPrinter:
+    """Handles background printing of generated tokens."""
+    def __init__(self, tokenizer):
+        self.tokenizer = tokenizer
+        self.token_queue = queue.Queue()
+        self.stop_event = threading.Event()
+        self.thread = None
+        self.buffer = ""
+        self.lock = threading.Lock()
+        self.thinking = True  # Track if we're still in thinking mode
+        self.decoding_buffer = []  # Buffer for token IDs
+        # Timing and stats tracking
+        self.start_time = time.time()
+        self.token_count = 0
+        self.prefill_time = 0
+        self.inference_time = 0
+        self.context_pos = 0
+        self.start()
+    def start(self):
+        """Start the printer thread."""
+        if self.thread is None:
+            self.thread = threading.Thread(target=self._print_worker)
+            self.thread.daemon = True
+            self.thread.start()
+    def add_token(self, token_id):
+        """Add a token to the print queue."""
+        if not self.stop_event.is_set():
+            self.token_queue.put(token_id)
+            self.token_count += 1
+    def drain_buffer(self):
+        """Decode token IDs from decoding_buffer in the main thread."""
+        if not self.decoding_buffer:
+            return
+        # Decode all tokens at once in the main thread
+        token_str = self.tokenizer.decode(self.decoding_buffer)
+        self.decoding_buffer.clear()
+        # Color-handling logic
+        if self.thinking and "</think>" in token_str:
+            self.thinking = False
+            parts = token_str.split("</think>")
+            if len(parts) > 0:
+                print(parts[0] + "</think>", end='', flush=True)
+                if len(parts) > 1:
+                    print(LIGHT_BLUE + parts[1], end='', flush=True)
+        else:
+            if not self.thinking:
+                print(LIGHT_BLUE + token_str, end='', flush=True)
+            else:
+                print(token_str, end='', flush=True)
+    def _print_worker(self):
+        """Worker thread that takes token_ids from the queue."""
+        while not self.stop_event.is_set():
+            try:
+                token_id = self.token_queue.get(timeout=0.01)
+                with self.lock:
+                    self.decoding_buffer.append(token_id)
+                self.token_queue.task_done()
+            except queue.Empty:
+                continue
+            except Exception as e:
+                print(f"\nError: Token printer error: {str(e)}")
+                break
+    def stop(self):
+        """Stop the printer thread."""
+        if self.thread and self.thread.is_alive():
+            self.stop_event.set()
+            try:
+                self.thread.join(timeout=1.0)
+            except Exception:
+                pass
+            print(RESET_COLOR)  # Reset color at the end
+        return self.buffer
+    def set_timing(self, prefill_time, inference_time, context_pos):
+        """Set timing information."""
+        self.prefill_time = prefill_time
+        self.inference_time = inference_time
+        self.context_pos = context_pos
+def parse_model_path(path):
+    """Parse model path and return full path with .mlmodelc or .mlpackage extension."""
+    path = Path(path)
+    # If path exists exactly as specified, return it
+    if path.exists():
+        return str(path)
+    # Try with both extensions
+    candidates = [
+        path,  # Original path
+        path.with_suffix('.mlmodelc'),  # With .mlmodelc
+        path.with_suffix('.mlpackage'),  # With .mlpackage
+        Path(str(path) + '.mlmodelc'),  # Handle case where extension is included
+        Path(str(path) + '.mlpackage')
+    ]
+    # Try all possible paths
+    for candidate in candidates:
+        if candidate.exists():
+            print(f"Found model at: {candidate}")
+            return str(candidate)
+    # If we get here, no valid path was found
+    print("\nError: Model not found. Tried following paths:")
+    for candidate in candidates:
+        print(f"  {candidate}")
+    raise FileNotFoundError(f"Model not found: {path}")
+def parse_ffn_filename(path):
+    """Parse FFN model filename to extract chunk information."""
+    path = Path(path)
+    pattern = r'FFN_PF.*_chunk_(\d+)of(\d+)'
+    match = re.search(pattern, path.name)
+    if match:
+        current_chunk = int(match.group(1))
+        total_chunks = int(match.group(2))
+        return current_chunk, total_chunks
+    return None, None
+def find_all_chunks(base_path):
+    """Find all chunk files matching the base FFN path pattern."""
+    path = Path(base_path)
+    pattern = re.sub(r'_chunk_\d+of\d+', '_chunk_*', str(path))
+    return sorted(glob.glob(pattern))
+def load_model(path, function_name=None):
+    """Load a CoreML model, handling both .mlmodelc and .mlpackage formats."""
+    path = Path(path)
+    compute_unit = ct.ComputeUnit.CPU_AND_NE
+    try:
+        if path.suffix == '.mlmodelc':
+            # For compiled models (.mlmodelc), use CompiledMLModel
+            if function_name:
+                return ct.models.CompiledMLModel(str(path), compute_unit, function_name=function_name)
+            else:
+                return ct.models.CompiledMLModel(str(path), compute_unit)
+        else:
+            # For packages (.mlpackage)
+            if function_name:
+                return ct.models.MLModel(str(path), function_name=function_name)
+            else:
+                return ct.models.MLModel(str(path))
+    except RuntimeError as e:
+        if "valid manifest does not exist" in str(e):
+            print(f"\nError: Could not load compiled model at {path}")
+            print("This might be because:")
+            print("1. The model is not properly compiled")
+            print("2. The model was compiled for a different OS version")
+            print("3. The model needs to be recompiled")
+            print("\nTry using the .mlpackage version instead, or recompile the model.")
+        raise
+def parse_args():
+    parser = argparse.ArgumentParser(description='Full Chat with CoreML LLaMA with context window shifting, gil resolved (c) 2025 Anemll')
+    # Add meta.yaml option
+    parser.add_argument('--meta', type=str, help='Path to meta.yaml to load all parameters')
+    # Add existing arguments
+    parser.add_argument('--d', '--dir', type=str, default='.',
+                       help='Directory containing model files (default: current directory)')
+    parser.add_argument('--embed', type=str, required=False,
+                       help='Path to embeddings model (relative to --dir)')
+    parser.add_argument('--ffn', type=str, required=False,
+                       help='Path to FFN model (can be chunked, relative to --dir)')
+    parser.add_argument('--lmhead', type=str, required=False,
+                       help='Path to LM head model (relative to --dir)')
+    parser.add_argument('--tokenizer', type=str, required=False,
+                       help='Path to tokenizer')
+    # Add new argument for auto-generation
+    parser.add_argument('--prompt', type=str,
+                       help='If specified, run once with this prompt and exit')
+    # Add no-warmup flag
+    parser.add_argument('--nw', action='store_true',
+                       help='Skip warmup phase')
+    # Add debug level
+    parser.add_argument('--debug-level', type=int, default=0,
+                       help='Debug level (0=none, 1=print prompts, 2=more verbose)')
+    # Model configuration
+    parser.add_argument('--context-length', type=int,
+                       help='Context length for the model (default: 512), if not provided, it will be detected from the model directory name ctxNUMBER')
+    parser.add_argument('--batch-size', type=int,
+                       help='Batch size for prefill (default: 64)')
+    args = parser.parse_args()
+    # If meta.yaml is provided, load parameters from it
+    if args.meta:
+        try:
+            with open(args.meta, 'r') as f:
+                meta = yaml.safe_load(f)
+            params = meta['model_info']['parameters']
+            # Set model directory to meta.yaml directory if not specified
+            if not args.d or args.d == '.':
+                args.d = str(Path(args.meta).parent)
+            # Build model paths based on parameters
+            prefix = params.get('model_prefix', 'llama')  # Default to 'llama' if not specified
+            lut_ffn = f"_lut{params['lut_ffn']}" if params['lut_ffn'] != 'none' else ''
+            lut_lmhead = f"_lut{params['lut_lmhead']}" if params['lut_lmhead'] != 'none' else ''
+            lut_embeddings = f"_lut{params['lut_embeddings']}" if params['lut_embeddings'] != 'none' else ''
+            num_chunks = int(params['num_chunks'])
+            # Set model paths if not specified
+            if not args.lmhead:
+                args.lmhead = f'{prefix}_lm_head{lut_lmhead}'
+            if not args.embed:
+                args.embed = f'{prefix}_embeddings{lut_embeddings}'  # Changed from lm_head to embeddings
+            if not args.ffn:
+                args.ffn = f'{prefix}_FFN_PF{lut_ffn}_chunk_01of{num_chunks:02d}'
+            if not args.tokenizer:
+                args.tokenizer = args.d
+            # Set other parameters if not overridden by command line
+            if args.context_length is None:
+                args.context_length = int(params['context_length'])
+            if args.batch_size is None:
+                args.batch_size = int(params['batch_size'])
+            args.num_chunks = num_chunks
+            print(f"\nLoaded parameters from {args.meta}:")
+            print(f"  Context Length: {args.context_length}")
+            print(f"  Batch Size: {args.batch_size}")
+            print(f"  Num Chunks: {args.num_chunks}")
+            print(f"  Models Directory: {args.d}")
+            print(f"  Embeddings: {args.embed}")
+            print(f"  LM Head: {args.lmhead}")
+            print(f"  FFN: {args.ffn}")
+        except Exception as e:
+            print(f"\nError loading meta.yaml: {str(e)}")
+            sys.exit(1)
+    return args
+def load_metadata(model,args):
+    # Extract metadata and config parameters
+    metadata = {}
+    if hasattr(model, 'user_defined_metadata'):
+        meta = model.user_defined_metadata
+        # Extract key parameters with defaults
+        metadata['context_length'] = int(meta.get('com.anemll.context_length', 512))
+        metadata['state_length'] = int(meta.get('com.anemll.state_length', metadata['context_length']))  # Added state_length
+        metadata['batch_size'] = int(meta.get('com.anemll.batch_size', 64))
+        metadata['lut_bits'] = int(meta.get('com.anemll.lut_bits', 0))
+        metadata['num_chunks'] = int(meta.get('com.anemll.num_chunks', 1))
+        print("\nExtracted Parameters:")
+        print(f"  Context Length: {metadata['context_length']}")
+        print(f"  State Length: {metadata['state_length']}")
+        print(f"  Prefill Batch Size: {metadata['batch_size']}")
+        print(f"  LUT Bits: {metadata['lut_bits']}")
+        print(f"  Number of Chunks: {metadata['num_chunks']}")
+        # Print model info
+        print("\nModel Info:")
+        if 'com.anemll.info' in meta:
+            print(f"  {meta['com.anemll.info']}")
+        if 'com.github.apple.coremltools.version' in meta:
+            print(f"  CoreML Tools: {meta['com.github.apple.coremltools.version']}")
+        # Print model input/output shapes
+        print("\nModel Shapes:")
+        if hasattr(model, 'input_description'):
+            print("  Inputs:")
+            for name, desc in model.input_description.items():
+                print(f"    {name}: {desc}")
+        if hasattr(model, 'output_description'):
+            print("  Outputs:")
+            for name, desc in model.output_description.items():
+                print(f"    {name}: {desc}")
+    else:
+        print("\nWarning: No metadata found in model")
+        # Check if model directory name contains context length pattern (ctxXXX)
+        ctx_len = 512
+        if args.context_length is  None:
+            import re
+            ctx_match = re.search(r'ctx(\d+)', str(args.d))
+            if ctx_match:
+                ctx_len0 = int(ctx_match.group(1))
+                if 512 <= ctx_len0 <= 8096:
+                    ctx_len = ctx_len0
+                    print(f"\nDetected context length {ctx_len} from directory name")
+            else:
+                print(f"\nWarning: No context length found in directory  {ctx_len} from directory name {args.d}")
+        else:
+            ctx_len = args.context_length
+        # Use defaults or values from args
+        metadata['context_length'] = ctx_len
+        metadata['state_length'] = ctx_len
+        # Get batch size from args or use default
+        metadata['batch_size'] = getattr(args, 'batch_size', 64)
+        metadata['lut_bits'] = 4
+        metadata['num_chunks'] = getattr(args, 'num_chunks', 4)
+        print("\nUsing parameters:")
+        print(f"  Context Length: {metadata['context_length']}")
+        print(f"  State Length: {metadata['state_length']}")
+        print(f"  Prefill Batch Size: {metadata['batch_size']}")
+        print(f"  LUT Bits: {metadata['lut_bits']}")
+        print(f"  Number of Chunks: {metadata['num_chunks']}")
+    # Override with values from args if they exist
+    if hasattr(args, 'batch_size') and args.batch_size is not None:
+        metadata['batch_size'] = args.batch_size
+        print(f"\nOverriding batch size from args: {args.batch_size}")
+    if hasattr(args, 'num_chunks') and args.num_chunks is not None:
+        metadata['num_chunks'] = args.num_chunks
+        print(f"\nOverriding num chunks from args: {args.num_chunks}")
+    return metadata
+def load_models(args,metadata):
+    """Load all required models and extract metadata."""
+    print("\nLoading models...")
+    try:
+        # Load embeddings model
+        print("\nLoading embeddings model...")
+        embed_path = parse_model_path(args.embed)
+        print(f"Loading from: {embed_path}")
+        embed_model = load_model(embed_path)
+        print("Embeddings model loaded successfully")
+        metadata = load_metadata(embed_model,args)
+        # Load LM head model
+        print("\nLoading LM head model...")
+        lmhead_path = parse_model_path(args.lmhead)
+        print(f"Loading from: {lmhead_path}")
+        lmhead_model = load_model(lmhead_path)
+        print("LM head model loaded successfully")
+        # Parse FFN path and find chunks if needed
+        print("\nLoading FFN+PREFILL model(s)...")
+        ffn_path = parse_model_path(args.ffn)
+        chunk_no, total_chunks = parse_ffn_filename(ffn_path)
+        ffn_models = []
+        if chunk_no and total_chunks:
+            print(f"\nDetected chunked FFN+PREFILL model ({total_chunks} chunks)")
+            # Find and load all chunks
+            chunk_paths = find_all_chunks(ffn_path)
+            if len(chunk_paths) != total_chunks:
+                raise ValueError(f"Found {len(chunk_paths)} chunks but filename indicates {total_chunks} chunks")
+            for chunk_path in chunk_paths:
+                print(f"\nLoading FFN+PREFILL chunk: {Path(chunk_path).name}")
+                try:
+                    # For chunked models, we need both infer and prefill functions
+                    ffn_models.append({
+                        'infer': load_model(chunk_path, function_name='infer'),
+                        'prefill': load_model(chunk_path, function_name='prefill')
+                    })
+                    print("Chunk loaded successfully")
+                except Exception as e:
+                    print(f"Error loading chunk {chunk_path}: {str(e)}")
+                    raise
+            metadata = load_metadata(ffn_models[0],args)
+        else:
+            print("\nLoading single FFN model...")
+            ffn_models.append(load_model(ffn_path))
+            print("FFN model loaded successfully")
+        return embed_model, ffn_models, lmhead_model, metadata
+    except Exception as e:
+        print(f"\nError loading models: {str(e)}")
+        print("\nPlease ensure all model files exist and are accessible.")
+        print("Expected files:")
+        print(f"  Embeddings: {args.embed}")
+        print(f"  LM Head: {args.lmhead}")
+        print(f"  FFN: {args.ffn}")
+        raise
+# At the top of the file, make this a default path
+def initialize_tokenizer(model_path=None):
+    """Initialize and configure the tokenizer."""
+    try:
+        tokenizer = AutoTokenizer.from_pretrained(
+            str(model_path),
+            use_fast=False,
+            trust_remote_code=True
+        )
+        print("\nTokenizer Configuration:")
+        print(f"Tokenizer type: {type(tokenizer)}")
+        print(f"Tokenizer name: {tokenizer.__class__.__name__}")
+        print(f"Vocabulary size: {len(tokenizer)}")
+        print(f"Model max length: {tokenizer.model_max_length}")
+        if tokenizer.pad_token is None:
+            tokenizer.pad_token = tokenizer.eos_token
+            tokenizer.pad_token_id = tokenizer.eos_token_id
+            print("Set PAD token to EOS token")
+        tokenizer.padding_side = "left"
+        print(f"\nSpecial Tokens:")
+        print(f"PAD token: '{tokenizer.pad_token}' (ID: {tokenizer.pad_token_id})")
+        print(f"EOS token: '{tokenizer.eos_token}' (ID: {tokenizer.eos_token_id})")
+        print(f"BOS token: '{tokenizer.bos_token}' (ID: {tokenizer.bos_token_id})")
+        print(f"UNK token: '{tokenizer.unk_token}' (ID: {tokenizer.unk_token_id})")
+        return tokenizer
+    except Exception as e:
+        print(f"\nError: Failed to load tokenizer from {model_path}")
+        print(f"Error details: {str(e)}")
+        print(f"Error type: {type(e)}")
+        print("\nThis code requires a Llama 3.2 model for chat template functionality.")
+        print("Please provide the path to a Llama 3.2 model directory.")
+        import traceback
+        traceback.print_exc()
+        raise
+def make_causal_mask(length, start):
+    """Create causal attention mask."""
+    mask = np.full((1, 1, length, length), -np.inf, dtype=np.float16)
+    row_indices = np.arange(length).reshape(length, 1)
+    col_indices = np.arange(length).reshape(1, length)
+    mask[:, :, col_indices <= (row_indices + start)] = 0
+    return mask
+def run_prefill(embed_model, ffn_models, input_ids, current_pos, context_length, batch_size, state, causal_mask):
+    """Run prefill on the input sequence."""
+    #print(f"[DEBUG] Running prefill from 0 to {current_pos}")
+    # Process in batches
+    batch_pos = 0
+    while batch_pos < current_pos:
+        batch_end = min(batch_pos + batch_size, current_pos)
+        current_batch_size = batch_end - batch_pos
+        #print(f"[DEBUG] Prefill batch {batch_pos}-{batch_end} (size={current_batch_size})")
+        # Get current batch
+        batch_input = input_ids[:, batch_pos:batch_end]
+        # Pad to full batch size
+        batch_input = F.pad(
+            batch_input,
+            (0, batch_size - current_batch_size),
+            value=0
+        )
+        # Generate position IDs for this batch
+        position_ids = torch.arange(batch_pos, batch_pos + batch_size, dtype=torch.int32)
+        # Use the pre-initialized causal mask and extract the batch portion
+        batch_causal_mask = causal_mask[:, :, batch_pos:batch_pos + batch_size, :]
+        # Run embeddings
+        hidden_states = torch.from_numpy(
+            embed_model.predict({'input_ids': batch_input.numpy()})['hidden_states']
+        )
+        # Run through FFN chunks
+        for ffn_model in ffn_models:
+            if isinstance(ffn_model, dict):
+                inputs = {
+                    'hidden_states': hidden_states.numpy(),
+                    'position_ids': position_ids.numpy(),
+                    'causal_mask': batch_causal_mask.numpy(),
+                    'current_pos': np.array([batch_pos], dtype=np.int32)
+                }
+                output = ffn_model['prefill'].predict(inputs, state)
+                hidden_states = torch.from_numpy(output['output_hidden_states'])
+        batch_pos = batch_end
+    return torch.tensor([current_pos], dtype=torch.int32)
+def generate_next_token(embed_model, ffn_models, lmhead_model, input_ids, pos, context_length, state, causal_mask, temperature=0.0):
+    """Generate the next token."""
+    # Get current token
+    current_token = input_ids[:, pos-1:pos]
+    # Run embeddings
+    hidden_states = torch.from_numpy(
+        embed_model.predict({'input_ids': current_token.numpy()})['hidden_states']
+    )
+    # Create masks
+    update_mask = torch.zeros((1, 1, context_length, 1), dtype=torch.float16)
+    update_mask[0, 0, pos-1, 0] = 1.0
+    position_ids = torch.tensor([pos-1], dtype=torch.int32)
+    # Use the pre-initialized causal mask and extract the single position portion
+    single_causal_mask = causal_mask[:, :, pos-1:pos, :]
+    # Run through FFN chunks
+    for ffn_model in ffn_models:
+        if isinstance(ffn_model, dict):
+            inputs = {
+                'hidden_states': hidden_states.numpy(),
+                'update_mask': update_mask.numpy(),
+                'position_ids': position_ids.numpy(),
+                'causal_mask': single_causal_mask.numpy(),
+                'current_pos': position_ids.numpy()
+            }
+            output = ffn_model['infer'].predict(inputs, state)
+            hidden_states = torch.from_numpy(output['output_hidden_states'])
+    # Run LM head and get next token
+    lm_output = lmhead_model.predict({'hidden_states': hidden_states.numpy()})
+    if 'logits1' in lm_output:
+        logits_parts = []
+        for i in range(1, 9):
+            key = f'logits{i}'
+            if key in lm_output:
+                logits_parts.append(torch.from_numpy(lm_output[key]))
+        logits = torch.cat(logits_parts, dim=-1)
+    else:
+        logits = torch.from_numpy(lm_output['output_logits'])
+    if temperature > 0:
+        logits = logits / temperature
+        probs = F.softmax(logits[0, -1, :], dim=-1)
+        next_token = torch.multinomial(probs, num_samples=1).item()
+    else:
+        next_token = torch.argmax(logits[0, -1, :]).item()
+    return next_token
+def create_unified_state(ffn_models, context_length):
+    """Create unified KV cache state for transformer."""
+    if isinstance(ffn_models[0], dict):
+        # Use first FFN model's prefill function to create state
+        state = ffn_models[0]['prefill'].make_state()
+        print(f"\nCreated unified transformer state for {len(ffn_models)} chunks")
+        return state
+    else:
+        state = ffn_models[0].make_state()
+        print("\nCreated unified transformer state")
+        return state
+def initialize_causal_mask(context_length):
+    """Initialize causal mask for transformer attention."""
+    causal_mask = make_causal_mask(context_length, 0)
+    causal_mask = torch.tensor(causal_mask, dtype=torch.float16)
+    print(f"\nInitialized causal mask for context length {context_length}")
+    return causal_mask
+def get_user_input():
+    """Get input from user, handling special key combinations."""
+    global THINKING_MODE
+    try:
+        import termios
+        import tty
+        import sys
+        def _getch():
+            fd = sys.stdin.fileno()
+            old_settings = termios.tcgetattr(fd)
+            try:
+                tty.setraw(sys.stdin.fileno())
+                ch = sys.stdin.read(1)
+            finally:
+                termios.tcsetattr(fd, termios.TCSADRAIN, old_settings)
+            return ch
+        buffer = []
+        while True:
+            char = _getch()
+            # Debug: print the character code
+            print(f"\nKey pressed: {repr(char)} (hex: {hex(ord(char))})")
+            # Check for Enter key
+            if char == '\r' or char == '\n':
+                print()  # Move to next line
+                input_text = ''.join(buffer)
+                # Check if the command is /t
+                if input_text == '/t':
+                    THINKING_MODE = not THINKING_MODE
+                    print(f"Thinking mode {'ON' if THINKING_MODE else 'OFF'}")
+                    buffer = []  # Clear buffer
+                    print(f"\n{LIGHT_GREEN}You{' (thinking)' if THINKING_MODE else ''}:{RESET_COLOR}", end=' ', flush=True)
+                    continue
+                return input_text
+            # Handle backspace
+            if char == '\x7f':  # backspace
+                if buffer:
+                    buffer.pop()
+                    sys.stdout.write('\b \b')  # Erase character
+                    sys.stdout.flush()
+                continue
+            # Handle Ctrl-C
+            if char == '\x03':  # Ctrl-C
+                print("^C")
+                raise KeyboardInterrupt
+            # Print character and add to buffer
+            sys.stdout.write(char)
+            sys.stdout.flush()
+            buffer.append(char)
+    except ImportError:
+        # Fallback for systems without termios
+        return input("> ")
+def chat_loop(embed_model, ffn_models, lmhead_model, tokenizer, metadata, state, causal_mask, auto_prompt=None, warmup=False):
+    """Interactive chat loop."""
+    global THINKING_MODE
+    global DEBUG_LEVEL
+    context_length = metadata.get('context_length')
+    batch_size = metadata.get('batch_size', 64)
+    if not warmup:
+        print(f"\nUsing context length: {context_length}")
+        print("\nStarting chat session. Press Ctrl+D to exit.")
+        print("Type your message and press Enter to chat. Use /t to toggle thinking mode.")
+        print(f"Thinking mode is {'ON' if THINKING_MODE else 'OFF'}")
+    # Keep track of conversation history
+    conversation = []
+    try:
+        while True:
+            try:
+                if not warmup:
+                    print(f"\n{LIGHT_GREEN}You{' (thinking)' if THINKING_MODE else ''}:{RESET_COLOR}", end=' ', flush=True)
+                if auto_prompt is not None:
+                    user_input = auto_prompt
+                    if not warmup:
+                        print(user_input)
+                else:
+                    user_input = input().strip()
+            except EOFError:
+                if not warmup:
+                    print("\nExiting chat...")
+                break
+            if not user_input:
+                continue
+            # Handle /t command
+            if user_input == "/t":
+                THINKING_MODE = not THINKING_MODE
+                print(f"Thinking mode {'ON' if THINKING_MODE else 'OFF'}")
+                continue
+            # Add user message to conversation
+            conversation.append({"role": "user", "content": user_input})
+            # Format using chat template with full history
+            if THINKING_MODE:
+                # Add thinking prompt to system message
+                conversation_with_thinking = [{"role": "system", "content": THINKING_PROMPT}] + conversation
+                base_input_ids = tokenizer.apply_chat_template(
+                    conversation_with_thinking,
+                    return_tensors="pt",
+                    add_generation_prompt=True
+                ).to(torch.int32)
+                # Print full prompt if debug level >= 1
+                if DEBUG_LEVEL >= 1 and not warmup:
+                    print(f"\n{DARK_BLUE}Debug: Full prompt with thinking:{RESET_COLOR}")
+                    print(tokenizer.decode(base_input_ids[0]))
+            else:
+                base_input_ids = tokenizer.apply_chat_template(
+                    conversation,
+                    return_tensors="pt",
+                    add_generation_prompt=True
+                ).to(torch.int32)
+                # Print full prompt if debug level >= 1
+                if DEBUG_LEVEL >= 1 and not warmup:
+                    print(f"\n{DARK_BLUE}Debug: Full prompt:{RESET_COLOR}")
+                    print(tokenizer.decode(base_input_ids[0]))
+            # Check if we need to trim history
+            while base_input_ids.size(1) > context_length - 100:  # Leave room for response
+                # Remove oldest message pair (user + assistant)
+                if len(conversation) > 2:
+                    conversation = conversation[2:]  # Remove oldest pair
+                    base_input_ids = tokenizer.apply_chat_template(
+                        conversation,
+                        return_tensors="pt",
+                        add_generation_prompt=True
+                    ).to(torch.int32)
+                else:
+                    # If only current message remains and still too long, truncate
+                    base_input_ids = base_input_ids[:, -context_length//2:]
+                    break
+            context_pos = base_input_ids.size(1)
+            # Pad sequence to context_size
+            input_ids = F.pad(
+                base_input_ids,
+                (0, context_length - context_pos),
+                value=0
+            )
+            if not warmup:
+                print(f"\n{LIGHT_BLUE}Assistant:{RESET_COLOR}", end=' ', flush=True)
+            # Initialize token printer and collect response
+            token_printer = TokenPrinter(tokenizer)
+            response_tokens = []
+            generation_start_time = time.time()
+            try:
+                # Run prefill on entire context
+                current_pos = run_prefill(
+                    embed_model,
+                    ffn_models,
+                    input_ids,
+                    context_pos,
+                    context_length,
+                    batch_size,
+                    state,
+                    causal_mask
+                )
+                #print(f"\n[DEBUG] After initial prefill - current_pos: {current_pos}")
+                # Generation loop
+                pos = context_pos
+                tokens_generated = 0
+                inference_start = time.time()  # Start inference timing
+                while True:
+                    # Check if we need to shift window
+                    if pos >= context_length - 2:
+                        # Calculate shift to maintain full batches
+                        batch_size = metadata.get('batch_size', 64)
+                        # Calculate max batches that fit in context
+                        max_batches = context_length // batch_size
+                        desired_batches = max(1, max_batches - 2)  # Leave room for new tokens
+                        new_size = min(desired_batches * batch_size, context_length - batch_size)
+                        # Create shifted input_ids
+                        tmp = torch.zeros((1, context_length), dtype=torch.int32)
+                        tmp[:,0:new_size] = input_ids[:,pos-new_size:pos]
+                        input_ids = tmp
+                        # Reset state and run prefill
+                        # keep the same state
+                        #state = create_unified_state(ffn_models, context_length)
+                        current_pos = run_prefill(
+                            embed_model,
+                            ffn_models,
+                            input_ids,
+                            new_size,  # Prefill the entire shifted content
+                            context_length,
+                            batch_size,
+                            state,
+                            causal_mask
+                        )
+                        # Start generating from the next position
+                        pos = new_size  # Don't back up, continue from where we left off
+                        #print(f"\n[DEBUG] After shift - next token will be at pos {pos}")
+                        #print(f"[DEBUG] Context before next token: {tokenizer.decode(input_ids[0, pos-40:pos])}")
+                        window_shifted = True
+                    # Generate next token
+                    next_token = generate_next_token(
+                        embed_model,
+                        ffn_models,
+                        lmhead_model,
+                        input_ids,
+                        pos,
+                        context_length,
+                        state,
+                        causal_mask
+                    )
+                    # Add token
+                    input_ids[0, pos] = next_token
+                    if not warmup:
+                        token_printer.add_token(next_token)
+                        token_printer.drain_buffer()
+                    response_tokens.append(next_token)
+                    pos += 1
+                    tokens_generated += 1
+                    # In warmup mode, limit tokens
+                    if warmup and tokens_generated >= WARMUP_TOKEN_LIMIT:
+                        break
+                    if next_token == tokenizer.eos_token_id:
+                        break
+                inference_time = time.time() - inference_start  # Calculate inference time
+                # Add assistant response to conversation
+                response_text = token_printer.stop()
+                conversation.append({"role": "assistant", "content": response_text})
+                # Print stats only if not in warmup
+                if not warmup:
+                    total_time = time.time() - generation_start_time
+                    prefill_time = total_time - inference_time
+                    inference_tokens_per_sec = len(response_tokens) / inference_time if inference_time > 0 else 0
+                    prefill_ms = prefill_time * 1000
+                    prefill_tokens_per_sec = context_pos / prefill_time if prefill_time > 0 else 0
+                    print(f"{DARK_BLUE}{inference_tokens_per_sec:.1f} t/s, "
+                          f"TTFT: {prefill_ms:.1f}ms ({prefill_tokens_per_sec:.1f} t/s), "
+                          f"{len(response_tokens)} tokens{RESET_COLOR}")
+                if auto_prompt is not None:
+                    break
+            except KeyboardInterrupt:
+                if not warmup:
+                    print("\nGeneration interrupted")
+                token_printer.stop()
+                continue
+    except Exception as e:
+        if not warmup:
+            print(f"\nError in chat loop: {str(e)}")
+            import traceback
+            traceback.print_exc()
+def main():
+    args = parse_args()
+    global DEBUG_LEVEL
+    DEBUG_LEVEL = args.debug_level
+    # Convert directory to absolute path
+    model_dir = Path(args.d).resolve()
+    if not model_dir.exists():
+        print(f"\nError: Model directory not found: {model_dir}")
+        return 1
+    print(f"\nUsing model directory: {model_dir}")
+    print(f"Context length: {args.context_length}")
+    try:
+        # Update paths to be relative to model directory
+        args.embed = str(model_dir / args.embed)
+        args.ffn = str(model_dir / args.ffn)
+        args.lmhead = str(model_dir / args.lmhead)
+        # Handle tokenizer path separately since it's not relative to model_dir
+        if args.tokenizer is None:
+            args.tokenizer = str(model_dir)
+        if not Path(args.tokenizer).exists():
+            print(f"\nError: Tokenizer directory not found: {args.tokenizer}")
+            return 1
+        args.tokenizer = str(Path(args.tokenizer).resolve())  # Convert to absolute path
+        print(f"Using tokenizer path: {args.tokenizer}")
+        metadata = {}
+        # Load models and extract metadata
+        embed_model, ffn_models, lmhead_model, metadata = load_models(args,metadata)
+        print(f"\nMetadata befor args.context_length: {metadata}")
+        # Override context length from command line if provided
+        if args.context_length is not None:
+            metadata['context_length'] = args.context_length
+            metadata['state_length'] = args.context_length  # Also update state_length
+            print(f"\nOverriding context length from command line: {args.context_length}")
+        print(f"\nMetadata after load_models: {metadata}")
+        # Load tokenizer with resolved path
+        tokenizer = initialize_tokenizer(args.tokenizer)
+        if tokenizer is None:
+            raise RuntimeError("Failed to initialize tokenizer")
+        # Create unified state once
+        state = create_unified_state(ffn_models, metadata['context_length'])
+        # Initialize causal mask once
+        causal_mask = initialize_causal_mask(metadata['context_length'])
+        # Warmup runs to prevent Python GIL issues with CoreML !
+        if not args.nw:
+            for i in range(2):
+                chat_loop(
+                    embed_model=embed_model,
+                    ffn_models=ffn_models,
+                    lmhead_model=lmhead_model,
+                    tokenizer=tokenizer,
+                    metadata=metadata,
+                    state=state,  # Pass the state
+                    causal_mask=causal_mask,  # Pass the causal mask
+                    warmup=True,
+                    auto_prompt="who are you?"
+                )
+        # Main run
+        chat_loop(
+            embed_model=embed_model,
+            ffn_models=ffn_models,
+            lmhead_model=lmhead_model,
+            tokenizer=tokenizer,
+            metadata=metadata,
+            state=state,  # Pass the state
+            causal_mask=causal_mask,  # Pass the causal mask
+            warmup=False,
+            auto_prompt=args.prompt
+        )
+    except Exception as e:
+        print(f"\nError: {str(e)}")
+        import traceback
+        traceback.print_exc()
+        return 1
+    return 0
+if __name__ == "__main__":
+    exit(main())

config.json ADDED Viewed

	@@ -0,0 +1,4 @@

+{
+  "tokenizer_class": "LlamaTokenizer",
+  "model_type": "llama"
+}

meta.yaml ADDED Viewed

	@@ -0,0 +1,23 @@

+model_info:
+  name: anemll-Llama-3.1-Nemotron-Nano-8B-v1-ctx512
+  version: 0.3.0
+  description: |
+    Demonstarates running Llama-3.1-Nemotron-Nano-8B-v1 on Apple Neural Engine
+    Context length: 512
+    Batch size: 64
+    Chunks: 16
+  license: MIT
+  author: Anemll
+  framework: Core ML
+  language: Python
+  parameters:
+    context_length: 512
+    batch_size: 64
+    lut_embeddings: none
+    lut_ffn: none
+    lut_lmhead: none
+    num_chunks: 16
+    model_prefix: nemo_
+    embeddings: nemo__embeddings.mlmodelc
+    lm_head: nemo__lm_head.mlmodelc
+    ffn: nemo__FFN_PF.mlmodelc

nemo__FFN_PF_chunk_01of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:edcdbad76faf4ee458a263f3a1aa8b820e0d0ac620bb0c2c5c12cc11be308af5
+size 694852699

nemo__FFN_PF_chunk_02of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:0c1da673a44893fa22cff3f4bb99d346d2d5c27d2830b9006f776e904821e62d
+size 692340458

nemo__FFN_PF_chunk_03of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e7510980bd7a1cb45dafe1a17d8702bf3555410c1dd49c5c4801241e61da7801
+size 692351023

nemo__FFN_PF_chunk_04of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:33edbb186cdd228ea5f148edeb5c71814b2626c17b36fb2bf3ea05bf40bc9833
+size 692319786

nemo__FFN_PF_chunk_05of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:f5b56cbdc705a7e7eefcca36e19161972409546471c0c1023c88d6a0221a5b7c
+size 692485163

nemo__FFN_PF_chunk_06of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:c87dab3c4386b599f4962b03516a24146bcbabaf905230df385dbd21951f5c35
+size 692446777

nemo__FFN_PF_chunk_07of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:4328e1068b54482fee75bcca4d767b9764481f2ac2a4b32efefcd412eb5054bc
+size 692547980

nemo__FFN_PF_chunk_08of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:89285c2ac7bd614edd851db8dbee2289195b09eca211f96ac399a97edcbfc349
+size 692524197

nemo__FFN_PF_chunk_09of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:69984aec9d126550479685f93c02222ee3f621cab5d5bf618175cd4467d30a2f
+size 692290432

nemo__FFN_PF_chunk_10of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:eb6fed8348c34a4dcb94e9e6274a3d4ce13c1cf55c4ea8064201a0f7e60e9a1f
+size 691996314

nemo__FFN_PF_chunk_11of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:347ef4dfb8720c5a80ec921171b4d0c7ec4af078c381f4cd0f87a79c557f110d
+size 691894392

nemo__FFN_PF_chunk_12of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:75ec0763472248d6abb9eb7ff50ff0bb17e76918d8cd0cf90d6b169c31d7c8c2
+size 691735112

nemo__FFN_PF_chunk_13of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:10f9728404778e4a13dc7f69976acf75630074bb9baec63bd06b2d4ca5a18979
+size 691731488

nemo__FFN_PF_chunk_14of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:6fb577e954585168f14b0ed51387c3aaa3fc5248b03f1a096f4af57ecad07ed8
+size 691858182

nemo__FFN_PF_chunk_15of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:ad7cc2817d8089078e15cae9e83aeffd0813a159af79a025214de1deba6946ec
+size 692063899

nemo__FFN_PF_chunk_16of16.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:744ca5f50e7bc118d7a7781cd163019bc78490a84d0fbf727a61faef21a55912
+size 692765141

nemo__embeddings.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e3f84f3a73085f9046ab5747964c27540052a25e08b5c0f43d955558c9164128
+size 807306441

nemo__lm_head.mlmodelc.zip ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:2707f57a46d48a972c0b4bc1a3b56196c6f451d0f0a65a2c6a3443746ae080e7
+size 807604851

prefill.py ADDED Viewed

	@@ -0,0 +1,644 @@

+#!/usr/bin/env python3
+# prefill.py
+# Copyright (c) 2025 Anemll
+# Licensed under the MIT License
+import argparse
+import os
+import re
+import glob
+from pathlib import Path
+import coremltools as ct
+from transformers import AutoTokenizer
+import torch
+import torch.nn.functional as F
+import numpy as np
+import time
+import yaml
+import sys
+# ANSI color codes
+LIGHT_BLUE = "\033[94m"
+DARK_BLUE = "\033[34m"
+LIGHT_GREEN = "\033[92m"
+RESET_COLOR = "\033[0m"
+def parse_model_path(path):
+    """Parse model path and return full path with .mlmodelc or .mlpackage extension."""
+    path = Path(path)
+    # If path exists exactly as specified, return it
+    if path.exists():
+        return str(path)
+    # Try with both extensions
+    candidates = [
+        path,  # Original path
+        path.with_suffix('.mlmodelc'),  # With .mlmodelc
+        path.with_suffix('.mlpackage'),  # With .mlpackage
+        Path(str(path) + '.mlmodelc'),  # Handle case where extension is included
+        Path(str(path) + '.mlpackage')
+    ]
+    # Try all possible paths
+    for candidate in candidates:
+        if candidate.exists():
+            print(f"Found model at: {candidate}")
+            return str(candidate)
+    # If we get here, no valid path was found
+    print("\nError: Model not found. Tried following paths:")
+    for candidate in candidates:
+        print(f"  {candidate}")
+    raise FileNotFoundError(f"Model not found: {path}")
+def parse_ffn_filename(path):
+    """Parse FFN model filename to extract chunk information."""
+    path = Path(path)
+    pattern = r'FFN_PF.*_chunk_(\d+)of(\d+)'
+    match = re.search(pattern, path.name)
+    if match:
+        current_chunk = int(match.group(1))
+        total_chunks = int(match.group(2))
+        return current_chunk, total_chunks
+    return None, None
+def find_all_chunks(base_path):
+    """Find all chunk files matching the base FFN path pattern."""
+    path = Path(base_path)
+    pattern = re.sub(r'_chunk_\d+of\d+', '_chunk_*', str(path))
+    return sorted(glob.glob(pattern))
+def load_model(path, function_name=None):
+    """Load a CoreML model, handling both .mlmodelc and .mlpackage formats."""
+    path = Path(path)
+    compute_unit = ct.ComputeUnit.CPU_AND_NE
+    try:
+        if path.suffix == '.mlmodelc':
+            # For compiled models (.mlmodelc), use CompiledMLModel
+            if function_name:
+                return ct.models.CompiledMLModel(str(path), compute_unit, function_name=function_name)
+            else:
+                return ct.models.CompiledMLModel(str(path), compute_unit)
+        else:
+            # For packages (.mlpackage)
+            if function_name:
+                return ct.models.MLModel(str(path), function_name=function_name)
+            else:
+                return ct.models.MLModel(str(path))
+    except RuntimeError as e:
+        if "valid manifest does not exist" in str(e):
+            print(f"\nError: Could not load compiled model at {path}")
+            print("This might be because:")
+            print("1. The model is not properly compiled")
+            print("2. The model was compiled for a different OS version")
+            print("3. The model needs to be recompiled")
+            print("\nTry using the .mlpackage version instead, or recompile the model.")
+        raise
+def load_metadata(model, args):
+    # Extract metadata and config parameters
+    metadata = {}
+    if hasattr(model, 'user_defined_metadata'):
+        meta = model.user_defined_metadata
+        # Extract key parameters with defaults
+        metadata['context_length'] = int(meta.get('com.anemll.context_length', 512))
+        metadata['state_length'] = int(meta.get('com.anemll.state_length', metadata['context_length']))
+        metadata['batch_size'] = int(meta.get('com.anemll.batch_size', 64))
+        metadata['lut_bits'] = int(meta.get('com.anemll.lut_bits', 0))
+        metadata['num_chunks'] = int(meta.get('com.anemll.num_chunks', 1))
+        print("\nExtracted Parameters:")
+        print(f"  Context Length: {metadata['context_length']}")
+        print(f"  State Length: {metadata['state_length']}")
+        print(f"  Prefill Batch Size: {metadata['batch_size']}")
+        print(f"  LUT Bits: {metadata['lut_bits']}")
+        print(f"  Number of Chunks: {metadata['num_chunks']}")
+    else:
+        print("\nWarning: No metadata found in model")
+        # Check if model directory name contains context length pattern (ctxXXX)
+        ctx_len = 512
+        if args.context_length is None:
+            import re
+            ctx_match = re.search(r'ctx(\d+)', str(args.d))
+            if ctx_match:
+                ctx_len0 = int(ctx_match.group(1))
+                if 512 <= ctx_len0 <= 8096:
+                    ctx_len = ctx_len0
+                    print(f"\nDetected context length {ctx_len} from directory name")
+            else:
+                print(f"\nWarning: No context length found in directory, using default {ctx_len}")
+        else:
+            ctx_len = args.context_length
+        # Use defaults or values from args
+        metadata['context_length'] = ctx_len
+        metadata['state_length'] = ctx_len
+        # Get batch size from args or use default
+        metadata['batch_size'] = getattr(args, 'batch_size', 64)
+        metadata['lut_bits'] = 4
+        metadata['num_chunks'] = getattr(args, 'num_chunks', 4)
+        print("\nUsing parameters:")
+        print(f"  Context Length: {metadata['context_length']}")
+        print(f"  State Length: {metadata['state_length']}")
+        print(f"  Prefill Batch Size: {metadata['batch_size']}")
+        print(f"  LUT Bits: {metadata['lut_bits']}")
+        print(f"  Number of Chunks: {metadata['num_chunks']}")
+    # Override with values from args if they exist
+    if hasattr(args, 'batch_size') and args.batch_size is not None:
+        metadata['batch_size'] = args.batch_size
+        print(f"\nOverriding batch size from args: {args.batch_size}")
+    if hasattr(args, 'num_chunks') and args.num_chunks is not None:
+        metadata['num_chunks'] = args.num_chunks
+        print(f"\nOverriding num chunks from args: {args.num_chunks}")
+    return metadata
+def load_models(args, metadata):
+    """Load all required models and extract metadata."""
+    print("\nLoading models...")
+    try:
+        # Load embeddings model
+        print("\nLoading embeddings model...")
+        embed_path = parse_model_path(args.embed)
+        print(f"Loading from: {embed_path}")
+        embed_model = load_model(embed_path)
+        print("Embeddings model loaded successfully")
+        metadata = load_metadata(embed_model, args)
+        # Load FFN model(s)
+        print("\nLoading PREFILL functionality only...")
+        ffn_path = parse_model_path(args.ffn)
+        chunk_no, total_chunks = parse_ffn_filename(ffn_path)
+        ffn_models = []
+        if chunk_no and total_chunks:
+            print(f"\nDetected chunked model with {total_chunks} chunks")
+            # Find and load all chunks
+            chunk_paths = find_all_chunks(ffn_path)
+            if len(chunk_paths) != total_chunks:
+                raise ValueError(f"Found {len(chunk_paths)} chunks but filename indicates {total_chunks} chunks")
+            for chunk_path in chunk_paths:
+                print(f"\nLoading PREFILL function from chunk: {Path(chunk_path).name}")
+                try:
+                    # For prefill testing, we only need the prefill function
+                    prefill_model = load_model(chunk_path, function_name='prefill')
+                    ffn_models.append(prefill_model)
+                    print("Chunk loaded successfully (prefill only)")
+                except Exception as e:
+                    print(f"Error loading chunk {chunk_path}: {str(e)}")
+                    raise
+            metadata = load_metadata(ffn_models[0], args)
+        else:
+            print("\nLoading single model (prefill functionality only)...")
+            ffn_models.append(load_model(ffn_path))
+            print("Model loaded successfully")
+        return embed_model, ffn_models, metadata
+    except Exception as e:
+        print(f"\nError loading models: {str(e)}")
+        print("\nPlease ensure all model files exist and are accessible.")
+        print("Expected files:")
+        print(f"  Embeddings: {args.embed}")
+        print(f"  FFN: {args.ffn}")
+        raise
+def initialize_tokenizer(model_path=None):
+    """Initialize and configure the tokenizer."""
+    try:
+        tokenizer = AutoTokenizer.from_pretrained(
+            str(model_path),
+            use_fast=False,
+            trust_remote_code=True
+        )
+        print("\nTokenizer Configuration:")
+        print(f"Tokenizer type: {type(tokenizer)}")
+        print(f"Tokenizer name: {tokenizer.__class__.__name__}")
+        print(f"Vocabulary size: {len(tokenizer)}")
+        if tokenizer.pad_token is None:
+            tokenizer.pad_token = tokenizer.eos_token
+            tokenizer.pad_token_id = tokenizer.eos_token_id
+            print("Set PAD token to EOS token")
+        tokenizer.padding_side = "left"
+        return tokenizer
+    except Exception as e:
+        print(f"\nError: Failed to load tokenizer from {model_path}")
+        print(f"Error details: {str(e)}")
+        raise
+def make_causal_mask(length, start):
+    """Create causal attention mask."""
+    mask = np.full((1, 1, length, length), -np.inf, dtype=np.float16)
+    row_indices = np.arange(length).reshape(length, 1)
+    col_indices = np.arange(length).reshape(1, length)
+    mask[:, :, col_indices <= (row_indices + start)] = 0
+    return mask
+def initialize_causal_mask(context_length):
+    """Initialize causal mask for transformer attention."""
+    causal_mask = make_causal_mask(context_length, 0)
+    causal_mask = torch.tensor(causal_mask, dtype=torch.float16)
+    print(f"\nInitialized causal mask for context length {context_length}")
+    return causal_mask
+def run_prefill(embed_model, ffn_models, input_ids, context_pos, context_length, batch_size=64, state=None, causal_mask=None):
+    """Run prefill on the input sequence."""
+    # Use provided causal mask or create one if not provided
+    if causal_mask is None:
+        causal_mask = make_causal_mask(context_length, 0)
+        causal_mask = torch.tensor(causal_mask, dtype=torch.float16)
+    # Process in batches
+    batch_pos = 0
+    while batch_pos < context_pos:
+        batch_end = min(batch_pos + batch_size, context_pos)
+        current_batch_size = batch_end - batch_pos
+        # Get current batch
+        batch_input = input_ids[:, batch_pos:batch_end]
+        # Always pad to full batch size for prefill
+        batch_input = F.pad(
+            batch_input,
+            (0, batch_size - current_batch_size),
+            value=0
+        )
+        # Generate position IDs for full batch size
+        position_ids = torch.arange(batch_size, dtype=torch.int32)
+        batch_causal_mask = causal_mask[:, :, :batch_size, :]
+        # Run embeddings with proper batch size
+        hidden_states = torch.from_numpy(
+            embed_model.predict({
+                'input_ids': batch_input.numpy(),
+                'batch_size': np.array([batch_size], dtype=np.int32)
+            })['hidden_states']
+        )
+        # Run through FFN chunks with state
+        for ffn_model in ffn_models:
+            # Handle both direct model and dictionary formats
+            if isinstance(ffn_model, dict) and 'prefill' in ffn_model:
+                # For backward compatibility with dictionary format
+                prefill_model = ffn_model['prefill']
+            else:
+                # Direct access for models loaded with function_name='prefill'
+                prefill_model = ffn_model
+            inputs = {
+                'hidden_states': hidden_states.numpy(),
+                'position_ids': position_ids.numpy(),
+                'causal_mask': batch_causal_mask.numpy(),
+                'current_pos': np.array([batch_pos], dtype=np.int32)
+            }
+            output = prefill_model.predict(inputs, state)
+            hidden_states = torch.from_numpy(output['output_hidden_states'])
+        batch_pos = batch_end
+    return torch.tensor([context_pos], dtype=torch.int32)
+def create_unified_state(ffn_models, context_length):
+    """Create unified KV cache state for transformer."""
+    if hasattr(ffn_models[0], 'make_state'):
+        # Direct access for models loaded with 'prefill' function_name
+        state = ffn_models[0].make_state()
+        print(f"\nCreated unified transformer state for {len(ffn_models)} chunks")
+        return state
+    else:
+        # Fallback for dictionary-based models (for backward compatibility)
+        if isinstance(ffn_models[0], dict) and 'prefill' in ffn_models[0]:
+            state = ffn_models[0]['prefill'].make_state()
+            print(f"\nCreated unified transformer state for {len(ffn_models)} chunks")
+            return state
+        else:
+            state = ffn_models[0].make_state()
+            print("\nCreated unified transformer state")
+            return state
+def test_prefill_speed(embed_model, ffn_models, tokenizer, batch_size, context_length, num_test_tokens, num_runs=20, test_single_chunk=True):
+    """Test prefill speed with sample token sequences."""
+    print(f"\n{LIGHT_GREEN}Testing prefill speed for {num_test_tokens} tokens (using internal batch size {batch_size}){RESET_COLOR}")
+    print(f"Running {num_runs} iterations for warmup and measurement")
+    # Create sample input sequence of exactly num_test_tokens tokens
+    sample_text = "This is a test sequence. " * ((num_test_tokens + 4) // 5) # Ensure enough text
+    input_ids = tokenizer(sample_text, return_tensors="pt").input_ids.to(torch.int32)
+    # Trim or pad to exactly num_test_tokens tokens
+    if input_ids.size(1) > num_test_tokens:
+        input_ids = input_ids[:, :num_test_tokens]
+    elif input_ids.size(1) < num_test_tokens:
+        pad_length = num_test_tokens - input_ids.size(1)
+        input_ids = F.pad(input_ids, (0, pad_length), value=tokenizer.pad_token_id)
+    print(f"Sample input sequence length: {input_ids.size(1)} tokens")
+    # Test with all chunks first
+    print(f"\n{LIGHT_BLUE}Testing with all chunks ({len(ffn_models)} chunks){RESET_COLOR}")
+    # Create unified state
+    state_all_chunks = create_unified_state(ffn_models, context_length)
+    # Initialize causal mask
+    causal_mask = initialize_causal_mask(context_length)
+    # Run prefill multiple times for warmup and testing
+    all_chunks_times = []
+    for i in range(num_runs):
+        # Reset state for each run
+        if i == 0:
+            print("\nStarting warmup runs...")
+        elif i == num_runs // 2:
+            print("\nWarmup complete, starting measurement runs...")
+        start_time = time.time()
+        # Run prefill
+        run_prefill(
+            embed_model,
+            ffn_models,
+            input_ids,
+            input_ids.size(1),  # context_pos
+            context_length,
+            batch_size, # Internal batching within run_prefill
+            state_all_chunks,
+            causal_mask
+        )
+        elapsed = time.time() - start_time
+        all_chunks_times.append(elapsed)
+        # Print progress
+        if i < num_runs // 2:  # Warmup phase
+            print(f"Warmup run {i+1}/{num_runs//2}: {elapsed:.4f}s ({batch_size/elapsed:.1f} tokens/s)")
+        else:  # Measurement phase
+            print(f"Run {i+1-num_runs//2}/{num_runs//2}: {elapsed:.4f}s ({batch_size/elapsed:.1f} tokens/s)")
+    # Calculate and report statistics for all chunks (excluding warmup runs)
+    all_chunks_measurement_times = all_chunks_times[num_runs // 2:]
+    all_chunks_avg_time = sum(all_chunks_measurement_times) / len(all_chunks_measurement_times)
+    all_chunks_min_time = min(all_chunks_measurement_times)
+    all_chunks_max_time = max(all_chunks_measurement_times)
+    all_chunks_tokens_per_sec = num_test_tokens / all_chunks_avg_time # Use num_test_tokens for speed calculation
+    print(f"\n{LIGHT_BLUE}All Chunks Prefill Speed Results:{RESET_COLOR}")
+    print(f"Number of Chunks: {len(ffn_models)}")
+    print(f"Test Tokens: {num_test_tokens} tokens")
+    print(f"Internal Batch Size: {batch_size} tokens")
+    print(f"Context Size: {context_length} tokens")
+    print(f"Average Time: {all_chunks_avg_time:.4f}s")
+    print(f"Min Time: {all_chunks_min_time:.4f}s")
+    print(f"Max Time: {all_chunks_max_time:.4f}s")
+    print(f"Average Speed: {all_chunks_tokens_per_sec:.1f} tokens/second")
+    print(f"Best Speed: {num_test_tokens / all_chunks_min_time:.1f} tokens/second") # Use num_test_tokens
+    # Test with single chunk if requested and if multiple chunks exist
+    single_chunk_tokens_per_sec = 0
+    if test_single_chunk and len(ffn_models) > 1:
+        print(f"\n{LIGHT_BLUE}Testing with single chunk (first chunk only){RESET_COLOR}")
+        # Create a list with only the first chunk
+        single_chunk_model = [ffn_models[0]]
+        # Create unified state for single chunk
+        state_single_chunk = create_unified_state(single_chunk_model, context_length)
+        # Run prefill multiple times for single chunk
+        single_chunk_times = []
+        for i in range(num_runs):
+            if i == 0:
+                print("\nStarting single chunk warmup runs...")
+            elif i == num_runs // 2:
+                print("\nSingle chunk warmup complete, starting measurement runs...")
+            start_time = time.time()
+            # Run prefill with single chunk
+            run_prefill(
+                embed_model,
+                single_chunk_model,
+                input_ids,
+                input_ids.size(1),  # context_pos
+                context_length,
+                batch_size, # Internal batching within run_prefill
+                state_single_chunk,
+                causal_mask
+            )
+            elapsed = time.time() - start_time
+            single_chunk_times.append(elapsed)
+            # Print progress
+            if i < num_runs // 2:  # Warmup phase
+                print(f"Single chunk warmup run {i+1}/{num_runs//2}: {elapsed:.4f}s ({batch_size/elapsed:.1f} tokens/s)")
+            else:  # Measurement phase
+                print(f"Single chunk run {i+1-num_runs//2}/{num_runs//2}: {elapsed:.4f}s ({batch_size/elapsed:.1f} tokens/s)")
+        # Calculate and report statistics for single chunk
+        single_chunk_measurement_times = single_chunk_times[num_runs // 2:]
+        single_chunk_avg_time = sum(single_chunk_measurement_times) / len(single_chunk_measurement_times)
+        single_chunk_min_time = min(single_chunk_measurement_times)
+        single_chunk_max_time = max(single_chunk_measurement_times)
+        single_chunk_tokens_per_sec = num_test_tokens / single_chunk_avg_time # Use num_test_tokens
+        print(f"\n{LIGHT_BLUE}Single Chunk Prefill Speed Results:{RESET_COLOR}")
+        print(f"Test Tokens: {num_test_tokens} tokens")
+        print(f"Internal Batch Size: {batch_size} tokens")
+        print(f"Context Size: {context_length} tokens")
+        print(f"Average Time: {single_chunk_avg_time:.4f}s")
+        print(f"Min Time: {single_chunk_min_time:.4f}s")
+        print(f"Max Time: {single_chunk_max_time:.4f}s")
+        print(f"Average Speed: {single_chunk_tokens_per_sec:.1f} tokens/second")
+        print(f"Best Speed: {num_test_tokens / single_chunk_min_time:.1f} tokens/second") # Use num_test_tokens
+        # Calculate overhead per chunk
+        if len(ffn_models) > 1:
+            chunk_overhead = (all_chunks_avg_time - single_chunk_avg_time) / (len(ffn_models) - 1)
+            print(f"\n{LIGHT_GREEN}Chunk Overhead Analysis:{RESET_COLOR}")
+            print(f"Single Chunk Time: {single_chunk_avg_time:.4f}s")
+            print(f"All Chunks Time ({len(ffn_models)} chunks): {all_chunks_avg_time:.4f}s")
+            print(f"Additional Time Per Chunk: {chunk_overhead:.4f}s")
+            print(f"Overhead Percentage: {(all_chunks_avg_time/single_chunk_avg_time - 1)*100:.1f}%")
+    return all_chunks_tokens_per_sec, single_chunk_tokens_per_sec
+def parse_args():
+    parser = argparse.ArgumentParser(description='Test prefill speed with CoreML LLaMA models (c) 2025 Anemll')
+    # Add meta.yaml option
+    parser.add_argument('--meta', type=str, help='Path to meta.yaml to load all parameters')
+    # Model paths
+    parser.add_argument('--d', '--dir', type=str, default='.',
+                       help='Directory containing model files (default: current directory)')
+    parser.add_argument('--embed', type=str, required=False,
+                       help='Path to embeddings model (relative to --dir)')
+    parser.add_argument('--ffn', type=str, required=False,
+                       help='Path to FFN model (can be chunked, relative to --dir)')
+    parser.add_argument('--tokenizer', type=str, required=False,
+                       help='Path to tokenizer')
+    # Test configuration
+    parser.add_argument('--batch-size', type=int,
+                       help='Batch size for prefill test (default: 64)')
+    parser.add_argument('--runs', type=int, default=20,
+                       help='Number of test runs (default: 20)')
+    parser.add_argument('--context-length', type=int,
+                       help='Context length for the model')
+    parser.add_argument('--no-single-chunk', action='store_true',
+                       help='Disable single chunk testing')
+    parser.add_argument('--test-tokens', type=int,
+                       help='Number of tokens to use for the speed test (default: batch_size)')
+    parser.add_argument('--test-full-context', action='store_true',
+                       help='Test prefill speed using the full context length (overrides --test-tokens)')
+    args = parser.parse_args()
+    # If meta.yaml is provided, load parameters from it
+    if args.meta:
+        try:
+            with open(args.meta, 'r') as f:
+                meta = yaml.safe_load(f)
+            params = meta['model_info']['parameters']
+            # Set model directory to meta.yaml directory if not specified
+            if not args.d or args.d == '.':
+                args.d = str(Path(args.meta).parent)
+            # Build model paths based on parameters
+            prefix = params.get('model_prefix', 'llama')
+            lut_ffn = f"_lut{params['lut_ffn']}" if params['lut_ffn'] != 'none' else ''
+            lut_embeddings = f"_lut{params['lut_embeddings']}" if params['lut_embeddings'] != 'none' else ''
+            num_chunks = int(params['num_chunks'])
+            # Set model paths if not specified
+            if not args.embed:
+                args.embed = f'{prefix}_embeddings{lut_embeddings}'
+            if not args.ffn:
+                args.ffn = f'{prefix}_FFN_PF{lut_ffn}_chunk_01of{num_chunks:02d}'
+            if not args.tokenizer:
+                args.tokenizer = args.d
+            # Set other parameters if not overridden by command line
+            if args.context_length is None:
+                args.context_length = int(params['context_length'])
+            if args.batch_size is None:
+                args.batch_size = int(params['batch_size'])
+            args.num_chunks = num_chunks
+            print(f"\nLoaded parameters from {args.meta}:")
+            print(f"  Context Length: {args.context_length}")
+            print(f"  Batch Size: {args.batch_size}")
+            print(f"  Num Chunks: {args.num_chunks}")
+            print(f"  Models Directory: {args.d}")
+            print(f"  Embeddings: {args.embed}")
+            print(f"  FFN: {args.ffn}")
+        except Exception as e:
+            print(f"\nError loading meta.yaml: {str(e)}")
+            sys.exit(1)
+    return args
+def main():
+    args = parse_args()
+    # Use default batch size if not specified
+    if args.batch_size is None:
+        args.batch_size = 64
+        print(f"\nUsing default batch size: {args.batch_size}")
+    # Convert directory to absolute path
+    model_dir = Path(args.d).resolve()
+    if not model_dir.exists():
+        print(f"\nError: Model directory not found: {model_dir}")
+        return 1
+    print(f"\nUsing model directory: {model_dir}")
+    try:
+        # Update paths to be relative to model directory
+        args.embed = str(model_dir / args.embed)
+        args.ffn = str(model_dir / args.ffn)
+        # Handle tokenizer path separately
+        if args.tokenizer is None:
+            args.tokenizer = str(model_dir)
+        if not Path(args.tokenizer).exists():
+            print(f"\nError: Tokenizer directory not found: {args.tokenizer}")
+            return 1
+        args.tokenizer = str(Path(args.tokenizer).resolve())
+        print(f"Using tokenizer path: {args.tokenizer}")
+        # Load models and extract metadata
+        metadata = {}
+        embed_model, ffn_models, metadata = load_models(args, metadata)
+        # Override context length from command line if provided
+        if args.context_length is not None:
+            metadata['context_length'] = args.context_length
+            metadata['state_length'] = args.context_length
+            print(f"\nOverriding context length from command line: {args.context_length}")
+        # Load tokenizer
+        tokenizer = initialize_tokenizer(args.tokenizer)
+        if tokenizer is None:
+            raise RuntimeError("Failed to initialize tokenizer")
+        # Determine number of tokens for the test
+        if args.test_full_context:
+            num_test_tokens = metadata['context_length']
+            print(f"\nTesting with full context length: {num_test_tokens} tokens")
+        elif args.test_tokens is not None:
+            num_test_tokens = args.test_tokens
+            print(f"\nTesting with specified tokens: {num_test_tokens} tokens")
+        else:
+            num_test_tokens = args.batch_size # Default to batch size
+            print(f"\nTesting with default tokens (batch size): {num_test_tokens} tokens")
+        # Ensure test tokens do not exceed context length
+        if num_test_tokens > metadata['context_length']:
+            print(f"\nWarning: Requested test tokens ({num_test_tokens}) exceed context length ({metadata['context_length']}).")
+            print(f"Clamping test tokens to context length.")
+            num_test_tokens = metadata['context_length']
+        # Run prefill speed test
+        test_prefill_speed(
+            embed_model=embed_model,
+            ffn_models=ffn_models,
+            tokenizer=tokenizer,
+            batch_size=args.batch_size, # Pass original batch_size for run_prefill internal logic
+            context_length=metadata['context_length'],
+            num_test_tokens=num_test_tokens, # Pass the number of tokens to actually test
+            num_runs=args.runs,
+            test_single_chunk=not args.no_single_chunk
+        )
+    except Exception as e:
+        print(f"\nError: {str(e)}")
+        import traceback
+        traceback.print_exc()
+        return 1
+    return 0
+if __name__ == "__main__":
+    exit(main())

tokenizer.json ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:6b9e4e7fb171f92fd137b777cc2714bf87d11576700a1dcd7a399e7bbe39537b
+size 17209920

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,2063 @@

+{
+  "added_tokens_decoder": {
+    "128000": {
+      "content": "<|begin_of_text|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128001": {
+      "content": "<|end_of_text|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128002": {
+      "content": "<|reserved_special_token_0|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128003": {
+      "content": "<|reserved_special_token_1|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128004": {
+      "content": "<|finetune_right_pad_id|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128005": {
+      "content": "<|reserved_special_token_2|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128006": {
+      "content": "<|start_header_id|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128007": {
+      "content": "<|end_header_id|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128008": {
+      "content": "<|eom_id|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128009": {
+      "content": "<|eot_id|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128010": {
+      "content": "<|python_tag|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128011": {
+      "content": "<|reserved_special_token_3|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128012": {
+      "content": "<|reserved_special_token_4|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128013": {
+      "content": "<|reserved_special_token_5|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128014": {
+      "content": "<|reserved_special_token_6|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128015": {
+      "content": "<|reserved_special_token_7|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128016": {
+      "content": "<|reserved_special_token_8|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128017": {
+      "content": "<|reserved_special_token_9|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128018": {
+      "content": "<|reserved_special_token_10|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128019": {
+      "content": "<|reserved_special_token_11|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128020": {
+      "content": "<|reserved_special_token_12|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128021": {
+      "content": "<|reserved_special_token_13|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128022": {
+      "content": "<|reserved_special_token_14|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128023": {
+      "content": "<|reserved_special_token_15|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128024": {
+      "content": "<|reserved_special_token_16|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128025": {
+      "content": "<|reserved_special_token_17|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128026": {
+      "content": "<|reserved_special_token_18|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128027": {
+      "content": "<|reserved_special_token_19|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128028": {
+      "content": "<|reserved_special_token_20|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128029": {
+      "content": "<|reserved_special_token_21|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128030": {
+      "content": "<|reserved_special_token_22|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128031": {
+      "content": "<|reserved_special_token_23|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128032": {
+      "content": "<|reserved_special_token_24|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128033": {
+      "content": "<|reserved_special_token_25|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128034": {
+      "content": "<|reserved_special_token_26|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128035": {
+      "content": "<|reserved_special_token_27|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128036": {
+      "content": "<|reserved_special_token_28|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128037": {
+      "content": "<|reserved_special_token_29|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128038": {
+      "content": "<|reserved_special_token_30|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128039": {
+      "content": "<|reserved_special_token_31|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128040": {
+      "content": "<|reserved_special_token_32|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128041": {
+      "content": "<|reserved_special_token_33|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128042": {
+      "content": "<|reserved_special_token_34|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128043": {
+      "content": "<|reserved_special_token_35|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128044": {
+      "content": "<|reserved_special_token_36|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128045": {
+      "content": "<|reserved_special_token_37|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128046": {
+      "content": "<|reserved_special_token_38|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128047": {
+      "content": "<|reserved_special_token_39|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128048": {
+      "content": "<|reserved_special_token_40|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128049": {
+      "content": "<|reserved_special_token_41|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128050": {
+      "content": "<|reserved_special_token_42|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128051": {
+      "content": "<|reserved_special_token_43|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128052": {
+      "content": "<|reserved_special_token_44|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128053": {
+      "content": "<|reserved_special_token_45|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128054": {
+      "content": "<|reserved_special_token_46|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128055": {
+      "content": "<|reserved_special_token_47|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128056": {
+      "content": "<|reserved_special_token_48|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128057": {
+      "content": "<|reserved_special_token_49|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128058": {
+      "content": "<|reserved_special_token_50|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128059": {
+      "content": "<|reserved_special_token_51|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128060": {
+      "content": "<|reserved_special_token_52|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128061": {
+      "content": "<|reserved_special_token_53|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128062": {
+      "content": "<|reserved_special_token_54|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128063": {
+      "content": "<|reserved_special_token_55|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128064": {
+      "content": "<|reserved_special_token_56|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128065": {
+      "content": "<|reserved_special_token_57|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128066": {
+      "content": "<|reserved_special_token_58|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128067": {
+      "content": "<|reserved_special_token_59|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128068": {
+      "content": "<|reserved_special_token_60|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128069": {
+      "content": "<|reserved_special_token_61|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128070": {
+      "content": "<|reserved_special_token_62|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128071": {
+      "content": "<|reserved_special_token_63|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128072": {
+      "content": "<|reserved_special_token_64|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128073": {
+      "content": "<|reserved_special_token_65|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128074": {
+      "content": "<|reserved_special_token_66|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128075": {
+      "content": "<|reserved_special_token_67|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128076": {
+      "content": "<|reserved_special_token_68|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128077": {
+      "content": "<|reserved_special_token_69|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128078": {
+      "content": "<|reserved_special_token_70|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128079": {
+      "content": "<|reserved_special_token_71|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128080": {
+      "content": "<|reserved_special_token_72|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128081": {
+      "content": "<|reserved_special_token_73|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128082": {
+      "content": "<|reserved_special_token_74|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128083": {
+      "content": "<|reserved_special_token_75|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128084": {
+      "content": "<|reserved_special_token_76|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128085": {
+      "content": "<|reserved_special_token_77|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128086": {
+      "content": "<|reserved_special_token_78|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128087": {
+      "content": "<|reserved_special_token_79|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128088": {
+      "content": "<|reserved_special_token_80|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128089": {
+      "content": "<|reserved_special_token_81|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128090": {
+      "content": "<|reserved_special_token_82|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128091": {
+      "content": "<|reserved_special_token_83|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128092": {
+      "content": "<|reserved_special_token_84|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128093": {
+      "content": "<|reserved_special_token_85|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128094": {
+      "content": "<|reserved_special_token_86|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128095": {
+      "content": "<|reserved_special_token_87|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128096": {
+      "content": "<|reserved_special_token_88|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128097": {
+      "content": "<|reserved_special_token_89|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128098": {
+      "content": "<|reserved_special_token_90|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128099": {
+      "content": "<|reserved_special_token_91|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128100": {
+      "content": "<|reserved_special_token_92|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128101": {
+      "content": "<|reserved_special_token_93|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128102": {
+      "content": "<|reserved_special_token_94|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128103": {
+      "content": "<|reserved_special_token_95|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128104": {
+      "content": "<|reserved_special_token_96|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128105": {
+      "content": "<|reserved_special_token_97|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128106": {
+      "content": "<|reserved_special_token_98|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128107": {
+      "content": "<|reserved_special_token_99|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128108": {
+      "content": "<|reserved_special_token_100|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128109": {
+      "content": "<|reserved_special_token_101|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128110": {
+      "content": "<|reserved_special_token_102|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128111": {
+      "content": "<|reserved_special_token_103|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128112": {
+      "content": "<|reserved_special_token_104|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128113": {
+      "content": "<|reserved_special_token_105|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128114": {
+      "content": "<|reserved_special_token_106|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128115": {
+      "content": "<|reserved_special_token_107|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128116": {
+      "content": "<|reserved_special_token_108|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128117": {
+      "content": "<|reserved_special_token_109|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128118": {
+      "content": "<|reserved_special_token_110|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128119": {
+      "content": "<|reserved_special_token_111|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128120": {
+      "content": "<|reserved_special_token_112|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128121": {
+      "content": "<|reserved_special_token_113|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128122": {
+      "content": "<|reserved_special_token_114|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128123": {
+      "content": "<|reserved_special_token_115|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128124": {
+      "content": "<|reserved_special_token_116|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128125": {
+      "content": "<|reserved_special_token_117|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128126": {
+      "content": "<|reserved_special_token_118|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128127": {
+      "content": "<|reserved_special_token_119|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128128": {
+      "content": "<|reserved_special_token_120|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128129": {
+      "content": "<|reserved_special_token_121|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128130": {
+      "content": "<|reserved_special_token_122|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128131": {
+      "content": "<|reserved_special_token_123|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128132": {
+      "content": "<|reserved_special_token_124|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128133": {
+      "content": "<|reserved_special_token_125|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128134": {
+      "content": "<|reserved_special_token_126|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128135": {
+      "content": "<|reserved_special_token_127|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128136": {
+      "content": "<|reserved_special_token_128|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128137": {
+      "content": "<|reserved_special_token_129|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128138": {
+      "content": "<|reserved_special_token_130|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128139": {
+      "content": "<|reserved_special_token_131|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128140": {
+      "content": "<|reserved_special_token_132|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128141": {
+      "content": "<|reserved_special_token_133|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128142": {
+      "content": "<|reserved_special_token_134|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128143": {
+      "content": "<|reserved_special_token_135|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128144": {
+      "content": "<|reserved_special_token_136|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128145": {
+      "content": "<|reserved_special_token_137|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128146": {
+      "content": "<|reserved_special_token_138|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128147": {
+      "content": "<|reserved_special_token_139|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128148": {
+      "content": "<|reserved_special_token_140|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128149": {
+      "content": "<|reserved_special_token_141|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128150": {
+      "content": "<|reserved_special_token_142|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128151": {
+      "content": "<|reserved_special_token_143|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128152": {
+      "content": "<|reserved_special_token_144|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128153": {
+      "content": "<|reserved_special_token_145|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128154": {
+      "content": "<|reserved_special_token_146|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128155": {
+      "content": "<|reserved_special_token_147|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128156": {
+      "content": "<|reserved_special_token_148|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128157": {
+      "content": "<|reserved_special_token_149|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128158": {
+      "content": "<|reserved_special_token_150|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128159": {
+      "content": "<|reserved_special_token_151|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128160": {
+      "content": "<|reserved_special_token_152|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128161": {
+      "content": "<|reserved_special_token_153|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128162": {
+      "content": "<|reserved_special_token_154|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128163": {
+      "content": "<|reserved_special_token_155|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128164": {
+      "content": "<|reserved_special_token_156|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128165": {
+      "content": "<|reserved_special_token_157|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128166": {
+      "content": "<|reserved_special_token_158|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128167": {
+      "content": "<|reserved_special_token_159|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128168": {
+      "content": "<|reserved_special_token_160|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128169": {
+      "content": "<|reserved_special_token_161|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128170": {
+      "content": "<|reserved_special_token_162|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128171": {
+      "content": "<|reserved_special_token_163|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128172": {
+      "content": "<|reserved_special_token_164|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128173": {
+      "content": "<|reserved_special_token_165|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128174": {
+      "content": "<|reserved_special_token_166|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128175": {
+      "content": "<|reserved_special_token_167|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128176": {
+      "content": "<|reserved_special_token_168|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128177": {
+      "content": "<|reserved_special_token_169|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128178": {
+      "content": "<|reserved_special_token_170|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128179": {
+      "content": "<|reserved_special_token_171|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128180": {
+      "content": "<|reserved_special_token_172|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128181": {
+      "content": "<|reserved_special_token_173|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128182": {
+      "content": "<|reserved_special_token_174|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128183": {
+      "content": "<|reserved_special_token_175|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128184": {
+      "content": "<|reserved_special_token_176|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128185": {
+      "content": "<|reserved_special_token_177|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128186": {
+      "content": "<|reserved_special_token_178|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128187": {
+      "content": "<|reserved_special_token_179|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128188": {
+      "content": "<|reserved_special_token_180|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128189": {
+      "content": "<|reserved_special_token_181|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128190": {
+      "content": "<|reserved_special_token_182|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128191": {
+      "content": "<|reserved_special_token_183|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128192": {
+      "content": "<|reserved_special_token_184|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128193": {
+      "content": "<|reserved_special_token_185|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128194": {
+      "content": "<|reserved_special_token_186|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128195": {
+      "content": "<|reserved_special_token_187|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128196": {
+      "content": "<|reserved_special_token_188|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128197": {
+      "content": "<|reserved_special_token_189|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128198": {
+      "content": "<|reserved_special_token_190|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128199": {
+      "content": "<|reserved_special_token_191|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128200": {
+      "content": "<|reserved_special_token_192|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128201": {
+      "content": "<|reserved_special_token_193|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128202": {
+      "content": "<|reserved_special_token_194|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128203": {
+      "content": "<|reserved_special_token_195|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128204": {
+      "content": "<|reserved_special_token_196|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128205": {
+      "content": "<|reserved_special_token_197|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128206": {
+      "content": "<|reserved_special_token_198|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128207": {
+      "content": "<|reserved_special_token_199|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128208": {
+      "content": "<|reserved_special_token_200|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128209": {
+      "content": "<|reserved_special_token_201|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128210": {
+      "content": "<|reserved_special_token_202|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128211": {
+      "content": "<|reserved_special_token_203|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128212": {
+      "content": "<|reserved_special_token_204|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128213": {
+      "content": "<|reserved_special_token_205|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128214": {
+      "content": "<|reserved_special_token_206|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128215": {
+      "content": "<|reserved_special_token_207|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128216": {
+      "content": "<|reserved_special_token_208|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128217": {
+      "content": "<|reserved_special_token_209|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128218": {
+      "content": "<|reserved_special_token_210|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128219": {
+      "content": "<|reserved_special_token_211|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128220": {
+      "content": "<|reserved_special_token_212|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128221": {
+      "content": "<|reserved_special_token_213|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128222": {
+      "content": "<|reserved_special_token_214|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128223": {
+      "content": "<|reserved_special_token_215|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128224": {
+      "content": "<|reserved_special_token_216|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128225": {
+      "content": "<|reserved_special_token_217|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128226": {
+      "content": "<|reserved_special_token_218|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128227": {
+      "content": "<|reserved_special_token_219|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128228": {
+      "content": "<|reserved_special_token_220|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128229": {
+      "content": "<|reserved_special_token_221|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128230": {
+      "content": "<|reserved_special_token_222|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128231": {
+      "content": "<|reserved_special_token_223|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128232": {
+      "content": "<|reserved_special_token_224|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128233": {
+      "content": "<|reserved_special_token_225|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128234": {
+      "content": "<|reserved_special_token_226|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128235": {
+      "content": "<|reserved_special_token_227|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128236": {
+      "content": "<|reserved_special_token_228|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128237": {
+      "content": "<|reserved_special_token_229|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128238": {
+      "content": "<|reserved_special_token_230|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128239": {
+      "content": "<|reserved_special_token_231|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128240": {
+      "content": "<|reserved_special_token_232|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128241": {
+      "content": "<|reserved_special_token_233|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128242": {
+      "content": "<|reserved_special_token_234|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128243": {
+      "content": "<|reserved_special_token_235|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128244": {
+      "content": "<|reserved_special_token_236|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128245": {
+      "content": "<|reserved_special_token_237|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128246": {
+      "content": "<|reserved_special_token_238|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128247": {
+      "content": "<|reserved_special_token_239|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128248": {
+      "content": "<|reserved_special_token_240|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128249": {
+      "content": "<|reserved_special_token_241|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128250": {
+      "content": "<|reserved_special_token_242|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128251": {
+      "content": "<|reserved_special_token_243|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128252": {
+      "content": "<|reserved_special_token_244|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128253": {
+      "content": "<|reserved_special_token_245|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128254": {
+      "content": "<|reserved_special_token_246|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "128255": {
+      "content": "<|reserved_special_token_247|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "bos_token": "<|begin_of_text|>",
+  "chat_template": "{%- if messages[0]['role'] == 'system' -%}{%- set system_message = messages[0]['content'] | trim -%}{%- set messages = messages[1:] -%}{%- else -%}{%- set system_message = '' -%}{%- endif -%}{%- if tools is not none -%}{{- '<|begin_of_text|><|start_header_id|>system<|end_header_id|>' + '\n\n' + system_message -}} {{- '\n\n' if system_message else '' -}} {{- '<AVAILABLE_TOOLS>[' -}} {% for t in tools %}{{- (t.function if t.function is defined else t) | tojson() -}}{{- ', ' if not loop.last else '' -}}{%- endfor -%} {{- ']</AVAILABLE_TOOLS>' -}} {{- '<|eot_id|>' -}}{%- else -%}{{- '<|begin_of_text|><|start_header_id|>system<|end_header_id|>' + '\n\n' + system_message + '<|eot_id|>' -}}{%- endif -%}{%- for message in messages -%}{%- if (message['role'] in ['user', 'tool']) != (loop.index0 % 2 == 0) -%}{{- raise_exception('Conversation roles must alternate between user/tool and assistant') -}}{%- elif message['role'] == 'user' -%}{{- '<|start_header_id|>user<|end_header_id|>' + '\n\n' + message['content'] | trim + '<|eot_id|>' -}}{%- elif message['role'] == 'tool' -%}{%- set tool_response = '<TOOL_RESPONSE>[' + message['content'] | trim + ']</TOOL_RESPONSE>' -%}{{- '<|start_header_id|>user<|end_header_id|>' + '\n\n' + tool_response + '<|eot_id|>' -}}{%- elif message['role'] == 'assistant' and message.get('tool_calls') is not none -%}{%- set tool_calls = message['tool_calls'] -%}{{- '<|start_header_id|>assistant<|end_header_id|>' + '\n\n' + '<TOOLCALL>[' -}}{%- for tool_call in tool_calls -%}{{ '{' + '\"name\": \"' + tool_call.function.name + '\", \"arguments\": ' + tool_call.function.arguments | tojson + '}' }}{%- if not loop.last -%}{{ ', ' }}{%- else -%}{{ ']</TOOLCALL>' + '<|eot_id|>' }}{%- endif -%}{%- endfor -%}{%- elif message['role'] == 'assistant' -%}{{- '<|start_header_id|>assistant<|end_header_id|>' + '\n\n' + message['content'] | trim + '<|eot_id|>' -}}{%- endif -%}{%- endfor -%}{%- if add_generation_prompt -%}{{ '<|start_header_id|>assistant<|end_header_id|>' + '\n\n' }}{%- endif -%}",
+  "clean_up_tokenization_spaces": true,
+  "eos_token": "<|eot_id|>",
+  "extra_special_tokens": {},
+  "model_input_names": [
+    "input_ids",
+    "attention_mask"
+  ],
+  "model_max_length": 131072,
+  "tokenizer_class": "PreTrainedTokenizerFast"
+}