appvoid
AI & ML interests
Recent Activity
Organizations
We are releasing Limen0.2B, a 222.5M-parameter base language model developed as a research platform for efficient pretraining and superword tokenization at smaller scales.
Limen0.2B was trained from scratch on 50B tokens and uses a compact 16K BoundlessBPE vocabulary. The project explores whether SuperBPE-style tokenization can remain effective in a substantially smaller model and vocabulary regime than those examined in earlier large-scale experiments.
The model also combines a deep-and-narrow transformer design with Exclusive Self-Attention, grouped-query attention, and tied embeddings. Its compact vocabulary reduces the embedding footprint and leaves a larger share of the parameter budget available to the transformer layers.
Despite its relatively modest training budget, Limen0.2B achieves competitive results for its scale across the reported language understanding, commonsense reasoning, and grammatical evaluation tasks. Comparisons with other compact models are provided as context rather than strict rankings, as their training data, token budgets, architectures, and evaluation settings differ.
The release includes the model weights, implementation, training configuration, checkpoint progression, and evaluation results, all under Apache 2.0.
UniversalComputingResearch/Limen0.2B
Technical feedback, independent evaluations, and further experiments with the model and tokenizer are welcome.
Cool stuff right there! Keep it up
It measures model performance on a variety of different tasks:
Language Completion
Common sense too
World Knowledge
Context Tracking
Quantitative
Logical Reasoning
Code Completion
Each has a different score and 1 overall score.
Submit your own model:
BananaMind/BananaMindBench-Leaderboard
Check it out:
BananaMind/BananaMindBench-Leaderboard
CPU, CUDA, PyTorch, and ONNX are supported. Apache 2.0.
See it for yourselves:
owensong/Inflect-Micro-v2
owensong/Inflect-Nano-v2
Try the Demos:
Nymbo/Inflect-TTS (unlimited CPU usage)
owensong/Inflect-v2 (ultra-fast ZeroGPU usage)
because its my own and im currently training it
you should do more of that magic you did with hellaswag on BananaMind-2-Medium
BananaMind 2 pro... ok i know
Let's Go!!!
One is Kimi-K3, which I have heard of briefly. What's the other one? I hope it's a video generation model
You guessed right! One of them is the best biggest model. The other one is the best smallest.
Following
Hey
following
mine are all <300M
Great, following already!
Surely I do!
following
is there anything you have always wanted to achieve? that idea you know you want to be true? is possible if any
keep working ๐ช
achieve maximum benchmark score per param
lol
you're talking about robust performance and not benchmaxxing, right? right?
is there anything you have always wanted to achieve? that idea you know you want to be true? is possible if any
keep working ๐ช
BananaMind 2 Nano is the smallest member of the BananaMind 2.0 family โ a 10M-parameter language model that shows how much you can squeeze out of a tiny footprint. It uses the family's digit-isolated tokenizer, so it keeps solid arithmetic despite its size, and it's small enough to run just about anywhere.
Trained on 30B tokens in about a day on a single RTX 5070 Ti (16GB), 4096-token context.
Benchmarks:
Average 35.77
ARC Easy 36.20
PIQA 55.98
ARC Challenge 23.38
HellaSwag 27.50
That 35.77 average edges out Pythia-31M (~34.79) at roughly a third the parameters.
Released under Apache 2.0 on Hugging Face: BananaMind/BananaMind-2-Nano โ weights, tokenizer, and config included.
following you!
ok
I want to follow people that also make super cool small models so comment here to follow you