Shekar
A high-performance Persian NLP library providing tokenization, normalization, POS tagging, NER, embeddings, spell checking, sentiment analysis, and dependency parsing, all in one package.
Libraries, datasets, and models built to advance Persian AI, all freely available under permissive licenses.
A high-performance Persian NLP library providing tokenization, normalization, POS tagging, NER, embeddings, spell checking, sentiment analysis, and dependency parsing, all in one package.
A large-scale open Persian speech dataset collected via community crowdsourcing. Version 6.0 holds 62,279 recordings totalling 99.02 hours of native Persian speech with predefined speaker-disjoint splits, for ASR, TTS, and representation learning.
The FineWeb of Persian: a quality-filtered web corpus of 6.25M documents from 14 sources for LLM pretraining, with synthetic queries for embedding and retrieval training.
Open Persian text embedding models for semantic search, retrieval, clustering and similarity. Includes Noql, a 12M-parameter Matryoshka model scoring 58.21 on FaMTEB.