view article Article SigLIP 2: A better multilingual vision language encoder +1 ariG23498, merve, qubvel-hf • Feb 21, 2025 • 227
TIPSv2 Collection TIPSv2 foundational vision-language models. Webpage: https://gdm-tipsv2.github.io/ • 9 items • Updated Jul 21 • 44
Core ML Model Zoo Collection PyTorch models converted to Core ML for on-device inference on iPhone, iPad and Mac. • 46 items • Updated Aug 3 • 1
view article Article Fine-Tune a Semantic Segmentation Model with a Custom Dataset tobiasc, nielsr • Mar 17, 2022 • 36
NVIDIA OmniDreams Collection NVIDIA OmniDreams model checkpoints and sample datasets. • 3 items • Updated 27 days ago • 10
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams Paper • 2601.02281 • Published Jan 5 • 33
RoMa v2: Harder Better Faster Denser Feature Matching Paper • 2511.15706 • Published Nov 19, 2025 • 9
EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion Paper • 2507.16535 • Published Jul 22, 2025 • 23
Probing the 3D Awareness of Visual Foundation Models Paper • 2404.08636 • Published Apr 12, 2024 • 14
VideoChat-R1 Collection VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning • 4 items • Updated Sep 28, 2025 • 8
Cosmos-Tokenizer1 Collection ⚠️ This collection is archived. 👉 https://huggingface.co/collections/nvidia/cosmos3 • 22 items • Updated 27 days ago • 44
BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation Paper • 2407.17952 • Published Jul 25, 2024 • 32