ICCV2023

community

Activity Feed Request to join this org

AI & ML interests

None defined yet.

Recent Activity

frankzydou authored a paper 5 days ago

Align3R: Aligned Monocular Depth Estimation for Dynamic Videos

frankzydou authored a paper 5 days ago

ProTracker: Probabilistic Integration for Robust and Accurate Point Tracking

frankzydou authored a paper 5 days ago

SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation

View all activity

ICCV2023's activity

AdinaY

posted an update 1 day ago

Post

1091

New models from Qwen 🔥

Qwen3-Embedding and Qwen3-Reranker Series just released on the hub by
Alibaba Qwen team.

✨ 0.6B/ 4B/ 8B with Apache2.0
✨ Supports 119 languages 🤯
✨ Top-tier performance: Leading the MTEB multilingual leaderboard！

Reranker:
Qwen/qwen3-reranker-6841b22d0192d7ade9cdefea
Embedding:
Qwen/qwen3-embedding-6841b2055b99c44d9a4c371f

AdinaY

posted an update 1 day ago

Post

1138

OpenAudio S1-mini 🔊 a new OPEN multilingual TTS model trained on 2M+ hours of data, by FishAudio

fishaudio/openaudio-s1-mini

✨ Supports 14 languages
✨ 50+ emotions & tones
✨ RLHF-optimized
✨ Special effects: laughing, crying, shouting, etc.

1 reply

Ljzycmd

authored a paper 2 days ago

ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

Paper • 2506.03107 • Published 3 days ago • 1

AdinaY

posted an update 3 days ago

Post

931

AReaL-boba² 🔥 A fully async RL system by Ant Research & Tsinghua.

Paper: AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning (2505.24298)
Model:
inclusionAI/areal-boba-2-683f0e819ccb7bb2e1b2f2d5

✨ 8B/14B/32B models, datasets & paper – all on the hub
✨ 2.77× faster training
✨ Native Agentic RL support

AdinaY

posted an update 4 days ago

Post

830

SynLogic 🧠 logical reasoning model & dataset by MiniMax.

MiniMaxAI/synlogic-6836c3246fca0277657ff032

✨ 3 models: 7B/32B/ Mix-3-32B (MIT license)
✨ Dataset: 35 verifiable logic tasks (Sudoku, Cipher, Arrow Maze etc.)
✨ RL training with auto-verifiable rewards
✨ Generalizes to math without explicit math training
✨ +6 pts on BBEH, +9.5 on KOR-Bench vs baselines

AdinaY

posted an update 4 days ago

Post

1615

Video-XL-2 🔥 long video understanding model by BAAI & Shanghai Jiaotong University

BAAI/Video-XL-2

✨ Apache 2.0
✨ Handles up to 10,000+ frames on a single GPU
✨ 2048-frame encoding in just 12s
✨ Efficient Chunk-based Prefilling & Bi-granularity KV decoding

AdinaY

posted an update 5 days ago

Post

2113

May highlights from China’s open source ecosystem 🔥

zh-ai-community/may-2025-open-works-from-the-chinese-community-681a3494145f2914dc679b7c

✨ DeepSeek dropped R1 updates
- Both R1 & 8B distralled smol model

✨ Bytedance goes big on open source:
- BAGEL, Dolphin, Seedcoder, Dream0...

✨ Multimodal is on fire!
- HuyuanCustom / HunyuanVideo-Avatar / HunyuanPortrait
- MiniMax: SynLogic / Orsta-7B
- Xiaomi: MiMo VL
- Alibaba Wan: Wan2.1-VACE
- OpenGVlab: ZeroGUI
- StepFun: ACE-Step-v1/Step1X-3D

✨ Specialized models/datasets excels
- Alibaba Qwen: World PM 72B
- BAAI:RobotBrain (MLLM for robotic)
- HiThink Research: BizFinBench (dataset)
- OpenBMB: Ultra FineWeb (dataset)
- Bilibili: Index-anisora (Anime/ACG)
- Skywork:Matrix-Game (game)

More awesome releases: Alibaba QwenLong-L1-32B, SkyWork OR1, OpenS2V-5M etc...

AdinaY

posted an update 8 days ago

Post

509

MiMo-VL 🔥 smol & mighty vision language model by Xiaomi

XiaomiMiMo/mimo-vl-68382ccacc7c2875500cd212

✨ 7B with RL & SFT
✨ Native resolution ViT for fine grained perception
✨ MORL = smarter alignment across perception, grounding & reasoning

AdinaY

posted an update 10 days ago

Post

2634

🔥 New benchmark & dataset for Subject-to-Video generation

OPENS2V-NEXUS by Pekin University

✨ Fine-grained evaluation for subject consistency
BestWishYsh/OpenS2V-Eval
✨ 5M-scale dataset:
BestWishYsh/OpenS2V-5M
✨ New metrics – automatic scores for identity, realism, and text match

2 replies

AdinaY

posted an update 10 days ago

Post

2245

HunyuanVideo-Avatar 🔥 another image to video model byTencent Hunyuan

tencent/HunyuanVideo-Avatar

✨Emotion-controlled, high-dynamic avatar videos
✨Multi-character support with separate audio control
✨Works with any style: cartoon, 3D, real face, while keeping identity consistent

AdinaY

posted an update 10 days ago

Post

1908

HunyuanPortrait 🔥 video model by Tencent Hunyuan team.

HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation (2503.18860)
tencent/HunyuanPortrait

✨Portrait animation from just one image + a video prompt
✨Diffusion-based, implicit motion control
✨Superior temporal consistency & detail

dbaranchuk

authored a paper 11 days ago

Alchemist: Turning Public Text-to-Image Data into Generative Gold

Paper • 2505.19297 • Published 12 days ago • 73

AdinaY

posted an update 11 days ago

Post

2824

Orsta 🔥 vision language models trained with V-Triune, a unified reinforcement learning system by MiniMax AI

One-RL-to-See-Them-All/one-rl-to-see-them-all-6833d27abce23898b2f9815a

✨ 7B & 32B with MIT license
✨ Masters 8 visual tasks: math, science QA, charts, puzzles, object detection, grounding, OCR, and counting
✨ Uses Dynamic IoU rewards for better visual understanding
✨Strong performance in visual reasoning and perception

AdinaY

posted an update 11 days ago

Post

2086

QwenLong-L1🔥 long-context reasoning model by Alibaba Tongyi Zhiwen team.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning (2505.17667)
Tongyi-Zhiwen/QwenLong-L1-32B

✨ 32B & Apache 2.0
✨ Outperforms OpenAI-o3-mini & Qwen3-235B-A22B
✨ Trained on a unique 1.6K DocQA RL dataset spanning math, logic & multi-hop reasoning

AdinaY

posted an update 17 days ago

Post

2784

ByteDance is absolutely cooking lately🔥

BAGEL 🥯 7B active parameter open multimodal foundation model by Bytedance Seed team.

ByteDance-Seed/BAGEL-7B-MoT

✨ Apache 2.0
✨ Outperforms top VLMs (Qwen2.5-VL & InternVL-2.5)
✨ Mixture-of-Transformer-Experts + dual encoders
✨ Trained on trillions of interleaved tokens

AdinaY

posted an update 18 days ago

Post

2402

Dolphin 🔥 A multimodal document image parsing model from ByteDance
, built on an analyze-then-parse paradigm.

ByteDance/Dolphin

✨ MIT licensed
✨ Handles text, tables, figures & formulas via:
- Reading-order layout analysis
- Parallel parsing with smart prompts

AdinaY

posted an update 19 days ago

Post

576

Index-AniSora 🎬 an open anime video model released by Bilibili

👉https://huggingface.co/IndexTeam/Index-anisora

✨ Apache2.0
✨ Supports many 2D styles: anime, manga, VTubers, and more
✨ Fine control over characters and actions with smart masking

AdinaY

posted an update 19 days ago

Post

1841

Data quality is the new frontier for LLM performance.

Ultra-FineWeb 📊 a high-quality bilingual dataset released by OpenBMB

openbmb/Ultra-FineWeb

✨ MIT License
✨ 1T English + 120B Chinese tokens
✨ Efficient model-driven filtering

2 replies

AdinaY

posted an update 21 days ago

Post

1931

2 cool papers from Alibaba Qwen team featured in today’s Daily Papers🔥

✨ WorldPM: Scaling Human Preference Modeling
WorldPM: Scaling Human Preference Modeling (2505.10527)
✨ Parallel Scaling Law for Language Models
Parallel Scaling Law for Language Models (2505.10475)

AdinaY

posted an update 24 days ago

Post

2679

Skywork-VL Reward🔥A multimodal reward model for both understanding & reasoning tasks, released by Skywork 昆仑万物-天工

Paper: Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning (2505.07263)
Model: Skywork/Skywork-VL-Reward-7B

✨ 7B
✨ Trained on large scale, high-quality preference data
✨ SOTA on VL-RewardBench + boosts reasoning via MPO

AI & ML interests

Recent Activity

Team members 209

ICCV2023's activity