Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
S's picture

S PRO

aeon37
2 1 20
Abramech's profile picture Talle2-X's profile picture hppdqdq's profile picture
ยท

AI & ML interests

None yet

Recent Activity

reacted to wop's post with ๐Ÿ”ฅ 4 days ago
https://huggingface.co/bench-labs/GCTokenizer-v1 , a multilingual tokenizer which does not require a training corpus https://huggingface.co/bench-labs developed **GCTokenizer-v1**, which is a multi-lingual tokenizer Available in four sizes: 32K, 65K, 131K and 262K tokens "S, M, L, XL" It utilizes an encoding scheme which allows it to handle characters in any language around the world General (multi lingual) Consensus (from multiple model tokenizers consensus) Tokenizer We included an implementation script too, built like BPE- it can encode arbitrary text, most of the time, efficiently
View all activity

Organizations

ML intern explorers's profile picture

aeon37 's datasets

None public yet
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs