Andres Marafioti's picture

Andres Marafioti

andito

AI & ML interests

Multimodal models, VLM and TTS

Recent Activity

reacted to merve's post with ๐Ÿค— 1 day ago
So many open releases at Hugging Face past week ๐Ÿคฏ recapping all here โคต๏ธ https://huggingface.co/collections/merve/march-21-releases-67dbe10e185f199e656140ae ๐Ÿ‘€ Multimodal > Mistral AI released a 24B vision LM, both base and instruction FT versions, sota ๐Ÿ”ฅ (OS) > with IBM we released SmolDocling, a sota 256M document parser with Apache 2.0 license (OS) > SpatialLM is a new vision LM that outputs 3D bounding boxes, comes with 0.5B (QwenVL based) and 1B (Llama based) variants > SkyWork released SkyWork-R1V-38B, new vision reasoning model (OS) ๐Ÿ’ฌ LLMs > NVIDIA released new Nemotron models in 49B and 8B with their post-training dataset > LG released EXAONE, new reasoning models in 2.4B, 7.8B and 32B > Dataset: Glaive AI released a new reasoning dataset of 22M+ examples > Dataset: NVIDIA released new helpfulness dataset HelpSteer3 > Dataset: OpenManusRL is a new agent dataset based on ReAct framework (OS) > Open-R1 team released OlympicCoder, new competitive coder model in 7B and 32B > Dataset: GeneralThought-430K is a new reasoning dataset (OS) ๐Ÿ–ผ๏ธ Image Generation/Computer Vision > Roboflow released RF-DETR, new real-time sota object detector (OS) ๐Ÿ”ฅ > YOLOE is a new real-time zero-shot object detector with text and visual prompts ๐Ÿฅน > Stability AI released Stable Virtual Camera, a new novel view synthesis model > Tencent released Hunyuan3D-2mini, new small and fast 3D asset generation model > ByteDance released InfiniteYou, new realistic photo generation model > StarVector is a new 8B model that generates svg from images > FlexWorld is a new model that expands 3D views (OS) ๐ŸŽค Audio > Sesame released CSM-1B new speech generation model (OS) ๐Ÿค– Robotics > NVIDIA released GR00T, new robotics model for generalized reasoning and skills, along with the dataset *OS ones have Apache 2.0 or MIT license
liked a model 4 days ago
HuggingFaceM4/idefics-80b-instruct
liked a model 6 days ago
HuggingFaceTB/SmolLM2-135M
View all activity

Organizations

Hugging Face's profile picture HuggingFaceM4's profile picture Huggingface Projects's profile picture Hugging Face H4's profile picture Hugging Face OSS Metrics's profile picture Hugging Face Smol Models Research's profile picture MLX Community's profile picture Distillation Hugs's profile picture Argilla Warehouse's profile picture Hugging Face FineVideo's profile picture smol-explorers's profile picture Hugging Face Science's profile picture Open R1's profile picture Smolvencoder's profile picture

andito's activity

published an article about 1 month ago
view article
Article

SmolVLM2: Bringing Video Understanding to Every Device

By orrzohar and 6 others โ€ข
โ€ข 215
published an article 2 months ago
view article
Article

SmolVLM Grows Smaller โ€“ Introducing the 250M & 500M Models!

By andito and 2 others โ€ข
โ€ข 165
published an article 4 months ago
view article
Article

SmolVLM - small yet mighty Vision Language Model

By andito and 4 others โ€ข
โ€ข 227
published an article 5 months ago
view article
Article

Deploying Speech-to-Speech on Hugging Face

By andito and 3 others โ€ข
โ€ข 39
published an article 6 months ago
view article
Article

FineVideo: behind the scenes

By mfarre and 5 others โ€ข
โ€ข 30
published an article 8 months ago
view article
Article

LAVE: Zero-shot VQA Evaluation on Docmatix with LLMs - Do We Still Need Fine-Tuning?

By danaaubakirova and 1 other โ€ข
โ€ข 17
published an article 8 months ago
view article
Article

Docmatix - a huge dataset for Document Visual Question Answering

By andito and 1 other โ€ข
โ€ข 72
published an article 9 months ago
view article
Article

Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models

By andito and 2 others โ€ข
โ€ข 190