Open Source

Projects

Libraries, datasets, and models built to advance Persian AI, all freely available under permissive licenses.

Python Library

Shekar

A high-performance Persian NLP library providing tokenization, normalization, POS tagging, NER, embeddings, spell checking, sentiment analysis, and dependency parsing, all in one package.

Normalization Tokenization POS & NER Embeddings
Speech Dataset

Neyshekar

A large-scale open Persian speech dataset collected via community crowdsourcing. Version 6.0 holds 62,279 recordings totalling 99.02 hours of native Persian speech with predefined speaker-disjoint splits, for ASR, TTS, and representation learning.

ASR TTS 99+ hours CC0 1.0
Text Dataset

Shiraz

The FineWeb of Persian: a quality-filtered web corpus of 6.25M documents from 14 sources for LLM pretraining, with synthetic queries for embedding and retrieval training.

LLM Pretraining Retrieval 6.25M docs ODC-BY 1.0