My series of fully open, state-of-the-art small mixture-of-experts models.
🔄 In a Training Loop
aquilesfd
aquilesfd
AI & ML interests
research
Recent Activity
repliedto Banaxi-Tech's post 1 day ago
We're announcing our BananaMind 2.1 model series!
The models will include:
- BananaMind 2.1 Nano: 10M parameters with 60B tokens.
- BananaMind 2.1 Lite: 25M parameters with 40B tokens.
- BananaMind 2.1 Flash: 50M parameters with 55B tokens.
- BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens.
These model will use a multi tower architecture (like https://huggingface.co/BananaMind/BananaMind-2.1-Unified) with some more architectural changes.
BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one!
We're currently training some experimental models based on this architecture to see its scaling!
Follow us:
https://huggingface.co/BananaMind
@Banaxi-Tech
@vovaRL
@DedeProGames
https://huggingface.co/bananamind-research-community reacted to Banaxi-Tech's post with 🔥 1 day ago
We're announcing our BananaMind 2.1 model series!
The models will include:
- BananaMind 2.1 Nano: 10M parameters with 60B tokens.
- BananaMind 2.1 Lite: 25M parameters with 40B tokens.
- BananaMind 2.1 Flash: 50M parameters with 55B tokens.
- BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens.
These model will use a multi tower architecture (like https://huggingface.co/BananaMind/BananaMind-2.1-Unified) with some more architectural changes.
BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one!
We're currently training some experimental models based on this architecture to see its scaling!
Follow us:
https://huggingface.co/BananaMind
@Banaxi-Tech
@vovaRL
@DedeProGames
https://huggingface.co/bananamind-research-community liked a model 3 days ago
SupraLabs/Supra2-Medium-Base