matejulcar commited on
Commit
5b36f40
1 Parent(s): cac6701

Added fast tokenizer

Browse files
README.md CHANGED
@@ -9,10 +9,9 @@ Load in transformers library with:
9
  ```
10
  from transformers import AutoTokenizer, AutoModelForMaskedLM
11
 
12
- tokenizer = AutoTokenizer.from_pretrained("EMBEDDIA/est-roberta", use_fast=False)
13
  model = AutoModelForMaskedLM.from_pretrained("EMBEDDIA/est-roberta")
14
  ```
15
- **NOTE**: it is currently *critically important* to add `use_fast=False` parameter to tokenizer if using transformers version 4+ (prior versions have `use_fast=False` as default) By default it attempts to load a fast tokenizer, which might work (ie. not result in an error), but not correctly, as there is no current support for fast tokenizers for Camembert-based models.
16
 
17
  # Est-RoBERTa
18
  Est-RoBERTa model is a monolingual Estonian BERT-like model. It is closely related to French Camembert model https://camembert-model.fr/. The Estonian corpora used for training the model have 2.51 billion tokens in total. The subword vocabulary contains 40,000 tokens.
 
9
  ```
10
  from transformers import AutoTokenizer, AutoModelForMaskedLM
11
 
12
+ tokenizer = AutoTokenizer.from_pretrained("EMBEDDIA/est-roberta")
13
  model = AutoModelForMaskedLM.from_pretrained("EMBEDDIA/est-roberta")
14
  ```
 
15
 
16
  # Est-RoBERTa
17
  Est-RoBERTa model is a monolingual Estonian BERT-like model. It is closely related to French Camembert model https://camembert-model.fr/. The Estonian corpora used for training the model have 2.51 billion tokens in total. The subword vocabulary contains 40,000 tokens.
special_tokens_map.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"bos_token": "<s>", "eos_token": "</s>", "unk_token": "<unk>", "sep_token": "</s>", "pad_token": "<pad>", "cls_token": "<s>", "mask_token": {"content": "<mask>", "single_word": false, "lstrip": true, "rstrip": false, "normalized": true}, "additional_special_tokens": ["<s>NOTUSED", "</s>NOTUSED"]}
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"bos_token": "<s>", "eos_token": "</s>", "sep_token": "</s>", "cls_token": "<s>", "unk_token": "<unk>", "pad_token": "<pad>", "mask_token": {"content": "<mask>", "single_word": false, "lstrip": true, "rstrip": false, "normalized": true, "__type": "AddedToken"}, "additional_special_tokens": ["<s>NOTUSED", "</s>NOTUSED"], "special_tokens_map_file": null, "name_or_path": "EMBEDDIA/est-roberta", "sp_model_kwargs": {}, "tokenizer_class": "CamembertTokenizer", "model_max_length": 512, "do_lower_case": false}