Models / Madmon
Embeddings

Madmon

Multilingual text embeddings specialized for Arabic dialects, Berber languages, and other underrepresented language families. Designed with deep focus on the languages frontier models ignore.

Vector Size
768 dimensions
Focus
Underrepresented languages
Best For
Search and retrieval
Input
Multilingual text
Access
Sawalni API

Purpose

Optimized for Underrepresented Languages

Madmon creates useful semantic representations for multilingual search, retrieval, classification, and analysis in complex linguistic environments.

It is designed with a focus on:

  • Arabic dialects — Darija, Egyptian, Gulf, Tunisian, Levantine, and more
  • Berber languages — Tamazight, Kabyle, Tachelhit in Latin, Arabic, and Tifinagh scripts
  • Other low-resource families — African, Southeast Asian, and indigenous languages

Semantic Search

Search across multilingual corpora. A Darija query retrieves relevant MSA, French, or English documents through shared semantic space.

Content Classification

Classify Arabic dialect content, social media posts, and user-generated text with high-quality representations.

Cross-Lingual Retrieval

Match content across scripts and languages. Arabizi, Arabic-script, and Latin-script content mapped to the same embedding space.

Clustering & Analytics

Group social media content by dialect, topic, or sentiment. Discover patterns in multilingual datasets.

Try Madmon now

Generate 768-dimensional semantic representations for multilingual search and content workflows.

Open Playground