Emo

Suggest emoji faster than you can type.

Multilingual on-device emoji suggestion.

Type a word or a sentence and get the emoji that fits. Tuned for to-dos, calendar entries, notes, and message drafts across 22 languages (including CJK, Arabic, Thai, Hindi, and more). The whole thing, model and tokenizer, is small (5MB on Apple via Core ML, 11MB via LiteRT elsewhere) and runs in well under 2ms on device.

"Dentist appointment" → 🦷 · "réserver un vol pour Tokyo" → ✈️ · "犬の散歩" → 🐕 · "จองโรงแรม" → 🏨

Try it

Platforms iOS, macOS, tvOS, visionOS, Android, Linux, Windows, Browser, Node
Languages 22
Weights v0.7.0

Install

Swift (requirements)

.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.1.0")

Then add the Emo product to your target.

Kotlin (requirements)

implementation("ai.desertant:emo:3.1.0")

JavaScript (requirements)

npm i @desert-ant-labs/emo @litertjs/core   # browser
npm i @desert-ant-labs/emo                  # Node, prebuilt native core

Files

File Format Size Contents
emo.tflite LiteRT / TFLite (int8) 10.2MB Fixed-window n-gram + masked semantic inputs, softmax probabilities output; runs on Android, Linux, Node, and the web (bundled by default in the Kotlin SDK; downloaded on demand by the JavaScript SDK)
emo.mlmodelc Compiled Core ML 4.6MB Mixed 4-/8-bit-palettized transformer, ready to load on Apple platforms (used by the Swift SDK)
emo_tokenizer.bin Pruned unigram tokenizer 750KB 48k SentencePiece pieces + scores; token ids = semantic-table rows
emo_meta.json JSON tiny emoji labels + n-gram hashing / fixed-window config the runtime needs
emo.pt PyTorch checkpoint 48MB Full-precision weights + semantic table + tokenizer (for retraining / other runtimes)

Older revisions (tags v0.6.0 and earlier) carry Emo.mlmodelc and emo.safetensors for SDK versions that predate the unified cross-platform migration.

Architecture

A compact two-stream classifier - no large encoder, just a tiny transformer over the semantic tokens:

  • Lexical stream: script-aware character/word n-grams (Latin, Han·Kana, Hangul jamo, Devanagari clusters, SE-Asian, …) hashed into a fixed multi-hash signed embedding table. Its size is independent of the number of languages.
  • Semantic stream: a frozen multilingual static embedding (Model2Vec potion-multilingual-128M, distilled from BAAI bge-m3), PCA-reduced to 128 dims and vocab-pruned to the 48k tokens that matter for the 22 target languages. Gives cross-lingual generalization and handles out-of-vocabulary words. The matching 750KB unigram tokenizer ships alongside (emo_tokenizer.bin).
  • Semantic pooling: a small 2-layer transformer encoder runs over the semantic token sequence, then an attention pool - order-aware, so it composes phrases and idioms instead of averaging tokens.
  • Head: a small MLP fusing the two streams into a softmax over a curated vocabulary of 812 everyday emojis (the emojis that actually come up most across the training phrases). Trained with n-gram dropout so the head relies on the semantic stream, which is what makes it generalize across languages.

Inputs and outputs

  • Input: a plain text string. Best on short, intent-oriented text.
  • Output: a probability distribution over the 812-emoji vocabulary; take the top-1 (or top-k). Optimized for top-1 relevance.

Languages

English, Spanish, Portuguese, French, German, Italian, Dutch, Russian, Polish, Turkish, Arabic, Chinese (Simplified & Traditional), Japanese, Korean, Hindi, Indonesian, Thai, Vietnamese, Ukrainian, Swedish, Danish, Czech.

Limitations

  • Tuned for short, intent-oriented text; long-form text produces noisier suggestions.
  • Emoji semantics are imprecise; near-ties at the top of the ranking are expected.
  • Per-language quality varies; lower-resource languages in the set are somewhat weaker.

Built on

  • minishlab/potion-multilingual-128M (MIT): semantic embedding stream (PCA-reduced, vocab-pruned derivative) + tokenizer lineage.
  • BAAI/bge-m3 (MIT): teacher the static embedding was distilled from.
  • Model2Vec (MIT): static-embedding distillation method.
  • Unicode CLDR emoji annotations: multilingual keyword grounding in the training data.

License

Desert Ant Labs Source-Available License. Free for most apps, and a commercial license is required at scale. Full terms are at the link. Licensing: licensing@desertant.com.

See THIRD_PARTY_NOTICES.md.

Citation

@software{emo_2026,
  title  = {Emo: Multilingual on-device emoji suggestion},
  author = {Desert Ant Labs},
  year   = {2026},
  url    = {https://huggingface.co/desert-ant-labs/emo},
}

© 2026 Desert Ant Labs · https://desertant.com

Downloads last month
22,654
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using desert-ant-labs/emo 1