We’re excited to release Pebble-25M and Pebble-25M-Chat!
Both models use our 3:1 Mamba2/Transformer hybrid architecture and were pretrained on 25B tokens. Pebble-25M-Chat was then further fine-tuned on an additional 250M tokens from smol-smoltalk, following the same approach used for the Pebble-10M models.
Hey I've updated my Hugging Face text generation model search space B-Sides. I always wanted more from HF's model search, so I built one.
I went deeper than the model card, embedding the relevant .json and .py files so you can search for models with custom kernels or exotic imports and specific architecture shapes. You can narrow it down to quantization types and training stacks. If you want to search it and it's not available just make a community post and I would gladly make each query more detailed with an update.
So far over 428,440 text generation models are catalogued with more coming weekly.