Back to Projects
WikiLLM
February 2026
LLMLow-Resource LanguagesWikipediaTokenizerDatasets
A family of compact, locally-runnable language models trained on curated Wikipedia data with custom tokenizers per language family. Designed to set a reproducible open baseline for low-resource LLM training with publishable evaluation benchmarks.