Open Source

Projects

Libraries, datasets, and models built to advance Persian AI, all freely available under permissive licenses.

Python Library

Shekar

A high-performance Persian NLP library providing tokenization, normalization, POS tagging, NER, embeddings, spell checking, sentiment analysis, and dependency parsing, all in one package.

Normalization Tokenization POS & NER Embeddings
Speech Dataset

Neyshekar

A large-scale open Persian speech dataset collected via community crowdsourcing. Version 5.0 holds 50,026 recordings totalling 79.22 hours of native Persian speech, for ASR, TTS, and representation learning.

ASR TTS 79+ hours CC0 1.0