Dataset and models for transforming LFM2 2.6B into a Tic Tac Toe master using RL Environments. Free course: https://t.ly/4jIFq
Stefano Fiorucci PRO
anakin87
AI & ML interests
Language Models: orchestration, post-training, GRPO, synthetic data...
Contributing to Haystack LLM framework 🏗️
Recent Activity
updated a Space 11 days ago
deepset/autoquizzer upvoted an article 18 days ago
Can you train a model on Simon Willison's deeply unscientific pelican benchmark?