Dean Byrne's picture
🔄 In a Training Loop

Dean Byrne PRO

Quazim0t0

AI & ML interests

DaisyChainAI🌼 / SmallLM's / San Francisco / Open Source

Recent Activity

liked a model about 13 hours ago
opencerebral/Boris-1.7-D60M-n30M
reacted to FlameF0X's post with 🔥 about 13 hours ago
Hello HuggingFace! I would like to share a small preview of a language model that I have been experiencing for a while. In the first attached image there is a small sample of the Myosotis-1, an attempt to make a small 100m parameter flagship model that is built on my bizarre architecture that is somewhat similar to an S4/S5 model with WKV added. I call it FWKV (Feed-Forward WKV). The model is currently still in training because of the nature of RNN-like models. It can also be seen that the model has insane prompt processing and token generation speed (evaluation done on a 2x Titan XP); even for its small size, some similar Transformer models do struggle to get the same results without custom kernels (some Transformer models can achieve this level of throughput on cheap hardware). In the second image, you can see checkpoint 20k of the model in its next token prediction state (this means it can't chat), ranking in the top 100 on https://huggingface.co/spaces/AxiomicLabs/Open_SLM_Leaderboard (the results have not been submitted since the model is not done training). - Why not just use Transformers? Have you seen any pure non-Transformers SLMs besides RWKV and Mamba? - Should you expect this project to become the next LFM or another very fast language model thing on some Raspberry Pi? No, the model is still an experiment; it's very sensible and prone to collapse (by the time of this post, it can be seen in the 1st image). - Should you use it? Maybe not yet; the architecture itself is still very "naive"—that's how I could call it at its current level. If you just want to play with it and see what you could do or how fast the model is on your hardware, then you can do it. Once the training is finished and I feel satisfied with the model next token prediction (the base model) and "assistants" (the instruction-tuned model) capabilities, I will make open weights at https://huggingface.co/FWKV with full support of the 🤗 Transformers.
reacted to FlameF0X's post with 👍 about 13 hours ago
Hello HuggingFace! I would like to share a small preview of a language model that I have been experiencing for a while. In the first attached image there is a small sample of the Myosotis-1, an attempt to make a small 100m parameter flagship model that is built on my bizarre architecture that is somewhat similar to an S4/S5 model with WKV added. I call it FWKV (Feed-Forward WKV). The model is currently still in training because of the nature of RNN-like models. It can also be seen that the model has insane prompt processing and token generation speed (evaluation done on a 2x Titan XP); even for its small size, some similar Transformer models do struggle to get the same results without custom kernels (some Transformer models can achieve this level of throughput on cheap hardware). In the second image, you can see checkpoint 20k of the model in its next token prediction state (this means it can't chat), ranking in the top 100 on https://huggingface.co/spaces/AxiomicLabs/Open_SLM_Leaderboard (the results have not been submitted since the model is not done training). - Why not just use Transformers? Have you seen any pure non-Transformers SLMs besides RWKV and Mamba? - Should you expect this project to become the next LFM or another very fast language model thing on some Raspberry Pi? No, the model is still an experiment; it's very sensible and prone to collapse (by the time of this post, it can be seen in the 1st image). - Should you use it? Maybe not yet; the architecture itself is still very "naive"—that's how I could call it at its current level. If you just want to play with it and see what you could do or how fast the model is on your hardware, then you can do it. Once the training is finished and I feel satisfied with the model next token prediction (the base model) and "assistants" (the instruction-tuned model) capabilities, I will make open weights at https://huggingface.co/FWKV with full support of the 🤗 Transformers.
View all activity

Organizations

Build Small Hackathon's profile picture DaisyChainAI's profile picture Neural Verified's profile picture