·
AI & ML interests
None yet
Recent Activity
reacted to wop's post with 🔥 3 days ago https://huggingface.co/bench-labs/GCTokenizer-v1 , a multilingual tokenizer which does not require a training corpus
https://huggingface.co/bench-labs developed **GCTokenizer-v1**, which is a multi-lingual tokenizer
Available in four sizes: 32K, 65K, 131K and 262K tokens "S, M, L, XL"
It utilizes an encoding scheme which allows it to handle characters in any language around the world
General (multi lingual)
Consensus (from multiple model tokenizers consensus)
Tokenizer
We included an implementation script too,
built like BPE- it can encode arbitrary text, most of the time, efficiently View all activity Organizations