AI & ML interests

None defined yet.

AtAndDev 
posted an update about 3 hours ago
view post
Post
219
@Banaxi-Tech stop hiding my comments. AND STOP STEALING PAPERS AND SPREADING MISINFORMATION.
your BGA blog is a copy of NSA (deepseek, 2025) branded under your name. literally the same top16 selected blocks, 512 local window, router over block summaries, all you did was change block size from 64 to 128.
you didnt cite NSA once but you put a “please cite BGA” bibtex at the bottom.
i commented under your post and said that there is no way that you can support claims like: “The Accuracy Should BE WAy better than DSA but untested yet.” you didnt run a single experiment. and the 256x isnt from BGA, its just n/2k with k=2048 so the exact same k DSA uses. if opus wrote this for you, at least read it before posting.
i commented again after you hid my comment despite it having constructive and correct feedback and you hid that too. and again.
you can hide the truth and just try to get hf post likes..... but is it really the thing that needs to be done? do you really want to take papers and make them yours while barely even changing the params?

admitting your mistakes and doing something about them needs humbleness, intelligence, humanness.
i encourage you to admit your mistakes and try to do better next time (at least read what blog your ai wrote or do proper experiments to back your stuff up).
  • 34 replies
·
AtAndDev 
posted an update about 1 month ago
view post
Post
2893
SPECK 2 IS ALREADY OUT: specklabs/Speck2-140M

Pretrained on 4x more tokens than the previous releases (20b vs 5b).
Instruct tuned versions are coming soon.
Very interesting models are coming soon too (hint: super long context).

Thanks for everyone supporting!
  • 3 replies
·
AtAndDev 
posted an update about 1 month ago
view post
Post
210
SPECK1.5 IS COMING SOON!
Same 5B token budget but much better corpus quality.

Also getting a ton of downloads, thanks for everyone downloading and liking <3

specklabs
AtAndDev 
posted an update about 2 months ago
view post
Post
160
NEW SPECK UPDATES:

Just hit #14 and #15 with out FIRST models on Open SLM Leaderboard. The models were trained on 5B tokens, while competing with similarly sized models trained on more than 6-20x the data.

A new base model Speck1.5-140M being trained right now on a higher quality corpus and will be released soon.
SpeckChat3 is coming very soon with 1 million samples, specifically designed to post train small base models.

Also, just to clarify stuff, we will NOT release anything that is NOT MIT licensed EVER. Openness is needed in small language research.

Thanks to everyone supporting the project, and stay tuned for new releases!
AtAndDev 
posted an update about 2 months ago
view post
Post
1894
SPECK UPDATES:
1 New instruct model tuned on top of Speck1-140M: specklabs/Speck1-140M-Instruct
2 Instruction tuning datasets
2 GGUFs

Much more coming soon:
Speck1.1-140M-Instruct that is post trained on SpeckChat2 will be coming very soon
New base model Speck1.5-140M is coming with a much higher quality corpus

Thanks to everyone who is already supporting the project, and stay tuned for new releases!
  • 3 replies
·
AtAndDev 
posted an update about 2 months ago
view post
Post
2130
FIRST SPECK MODEL RELEASED:
specklabs/Speck1-140M

new models coming very soon (both instruct and much better models), with much much higher training scale as i am getting marenostrum5 access soon!
we will be looking at 100b-2t token budgets :)
  • 4 replies
·
Nymbo 
posted an update 2 months ago
view post
Post
2437
Anthropic gave me six months of Claude Max 20x through the Claude for Open Source program, granted based on my Hugging Face work. Thank you
Anthropic
for supporting open source.

So far I've been pointing it at Markdown Minimap, an Obsidian plugin that adds a scrollable IDE-style minimap to your notes. This week I've been clearing a backlog of user-reported issues on it, with Claude often handling them end to end.

https://github.com/Nymbo/Markdown-Minimap — issues and PRs welcome.
Nymbo 
posted an update 3 months ago
view post
Post
6127
Introducing Inflect-v2, two exceptionally small, open-weight English TTS models at just 3.9M and 9.3M parameters. Both generate speech multiple times faster than real-time on CPU. Despite their size, Inflect-v2 delivers quality that is competitive with much larger lightweight TTS systems, including KittenTTS, Piper, and Supertonic-3.

CPU, CUDA, PyTorch, and ONNX are supported. Apache 2.0.

See it for yourselves:
owensong/Inflect-Micro-v2
owensong/Inflect-Nano-v2

Try the Demos:
Nymbo/Inflect-TTS (unlimited CPU usage)
owensong/Inflect-v2 (ultra-fast ZeroGPU usage)
  • 6 replies
·
Nymbo 
posted an update 7 months ago
view post
Post
7999
We should really have a release date range slider on the /models page. Tired of "trending/most downloaded" being the best way to sort and still seeing models from 2023 on the first page just because they're embedded in enterprise pipelines and get downloaded repeatedly. "Recently Created/Recently Updated" don't solve the discovery problem considering the amount of noise to sift through.

Slight caveat: Trending actually does have some recency bias, but it's not strong/precise enough.
  • 3 replies
·