matejulcar
commited on
Commit
•
5b36f40
1
Parent(s):
cac6701
Added fast tokenizer
Browse files- README.md +1 -2
- special_tokens_map.json +1 -0
- tokenizer.json +0 -0
- tokenizer_config.json +1 -0
README.md
CHANGED
@@ -9,10 +9,9 @@ Load in transformers library with:
|
|
9 |
```
|
10 |
from transformers import AutoTokenizer, AutoModelForMaskedLM
|
11 |
|
12 |
-
tokenizer = AutoTokenizer.from_pretrained("EMBEDDIA/est-roberta"
|
13 |
model = AutoModelForMaskedLM.from_pretrained("EMBEDDIA/est-roberta")
|
14 |
```
|
15 |
-
**NOTE**: it is currently *critically important* to add `use_fast=False` parameter to tokenizer if using transformers version 4+ (prior versions have `use_fast=False` as default) By default it attempts to load a fast tokenizer, which might work (ie. not result in an error), but not correctly, as there is no current support for fast tokenizers for Camembert-based models.
|
16 |
|
17 |
# Est-RoBERTa
|
18 |
Est-RoBERTa model is a monolingual Estonian BERT-like model. It is closely related to French Camembert model https://camembert-model.fr/. The Estonian corpora used for training the model have 2.51 billion tokens in total. The subword vocabulary contains 40,000 tokens.
|
|
|
9 |
```
|
10 |
from transformers import AutoTokenizer, AutoModelForMaskedLM
|
11 |
|
12 |
+
tokenizer = AutoTokenizer.from_pretrained("EMBEDDIA/est-roberta")
|
13 |
model = AutoModelForMaskedLM.from_pretrained("EMBEDDIA/est-roberta")
|
14 |
```
|
|
|
15 |
|
16 |
# Est-RoBERTa
|
17 |
Est-RoBERTa model is a monolingual Estonian BERT-like model. It is closely related to French Camembert model https://camembert-model.fr/. The Estonian corpora used for training the model have 2.51 billion tokens in total. The subword vocabulary contains 40,000 tokens.
|
special_tokens_map.json
ADDED
@@ -0,0 +1 @@
|
|
|
|
|
1 |
+
{"bos_token": "<s>", "eos_token": "</s>", "unk_token": "<unk>", "sep_token": "</s>", "pad_token": "<pad>", "cls_token": "<s>", "mask_token": {"content": "<mask>", "single_word": false, "lstrip": true, "rstrip": false, "normalized": true}, "additional_special_tokens": ["<s>NOTUSED", "</s>NOTUSED"]}
|
tokenizer.json
ADDED
The diff for this file is too large to render.
See raw diff
|
|
tokenizer_config.json
ADDED
@@ -0,0 +1 @@
|
|
|
|
|
1 |
+
{"bos_token": "<s>", "eos_token": "</s>", "sep_token": "</s>", "cls_token": "<s>", "unk_token": "<unk>", "pad_token": "<pad>", "mask_token": {"content": "<mask>", "single_word": false, "lstrip": true, "rstrip": false, "normalized": true, "__type": "AddedToken"}, "additional_special_tokens": ["<s>NOTUSED", "</s>NOTUSED"], "special_tokens_map_file": null, "name_or_path": "EMBEDDIA/est-roberta", "sp_model_kwargs": {}, "tokenizer_class": "CamembertTokenizer", "model_max_length": 512, "do_lower_case": false}
|