All posts

Persian text transliteration with Shekar library

Shekar now supports Iranian Persian ↔ Tajikistan Persian transliteration: converting text between the Perso-Arabic script of Iranian Persian and the Cyrillic script of Tajikistan Persian. We fine-tuned Google’s ByT5-small model and integrated it into the Shekar Python library, making both directions available through a simple interface. The model is available on Hugging Face.

Why transliteration between Iranian Persian and Tajikistan Persian matters

A familiar language can become difficult to read when it appears in an unfamiliar alphabet. Iranian Persian is written in Perso-Arabic script, while Tajikistan Persian uses Cyrillic. These closely related varieties of Persian share much, but readers who know only one script face a barrier to the other’s writing. This is the problem that motivated the ParsText research.

Script conversion can help students explore texts, readers approach literature and news, and developers make documents searchable across alphabets. For Shekar, it is another way to make Persian language technology useful across communities.

Transliteration changes the writing system. It does not guarantee that vocabulary, regional expressions, or style will be adapted for another audience. A converted text may still benefit from an editor familiar with the target variety.

How the ByT5 model handles two scripts

A letter-by-letter substitution can only go so far. Iranian Persian usually leaves short vowels unwritten, so converting into the Cyrillic script of Tajikistan Persian requires choosing sounds that are not explicitly shown in the source. Context matters.

ByT5 reads text as bytes—the basic units used to store digital text—rather than relying on a fixed vocabulary of word pieces. We fine-tuned it on paired Iranian Persian and Tajikistan Persian text so it could learn patterns across both scripts. One model handles both directions. The model card describes the architecture and intended uses.

The dataset behind the model

Our model card identifies the training data as ParsText, a parallel corpus used here for Iranian Persian and Tajikistan Persian. A parallel corpus places corresponding text in the two scripts side by side, giving the model examples to learn from.

The training notebook loads a merged file called all_data.tsv, removes empty pairs and pairs with either side longer than 200 characters, and standardizes letters and spacing in Iranian Persian with Shekar. Each remaining pair becomes two examples: Iranian Persian → Tajikistan Persian and Tajikistan Persian → Iranian Persian.

The model card reports 751,650 directional examples after preparation, equivalent to 375,825 pairs before doubling, with a 90% training, 5% validation, and 5% test split. These are example counts, not a count of unique sentences.

There is a provenance detail to distinguish: Rayyan Merchant and Kevin Tang’s original 2024 ParsText paper describes 2,813 sentence pairs collected from blogs and news. The larger merged file used here is not documented source by source in the notebook, so its reported size should not be attributed to that original release.

Results: how well does it work?

The model card reports these results on 37,582 test examples. Higher chrF++ and exact-match scores are better; lower character error rates are better.

Reported ByT5 test results
DirectionExampleschrF++Character error rateExact match
Iranian Persian → Tajikistan Persian18,70587.904.67%42.3%
Tajikistan Persian → Iranian Persian18,87791.762.83%64.9%
Overall37,58289.683.82%53.7%

chrF++ measures similarity to the reference text using character and word patterns; 89.68 is not “89.68% accuracy.” Character error rate counts the character edits needed to match the reference. Exact match asks a stricter question: was the entire output identical to the reference?

The outputs can therefore be close to their references while fewer match perfectly. Iranian Persian → Tajikistan Persian is the harder direction, consistent with the need to infer unwritten vowels. Names, unfamiliar loanwords, and text outside the training material still need care.

How to read these numbers: the notebook splits the data after creating both directions. A pair in the test set can therefore have its reverse in training. These are reported example-level results, not a test on entirely unseen text pairs. The evaluation also uses four decoding beams, while Shekar defaults to one; the card does not provide a separate benchmark for the quantized library version.

Try Iranian Persian ↔ Tajikistan Persian transliteration in Shekar

Install the Shekar library with pip install shekar, then use the two direction-specific classes:

from shekar import FarsiToTajik, TajikToFarsi

to_tajik = FarsiToTajik()
to_persian = TajikToFarsi()

print(to_tajik("ایران مادر است!"))
# Эрон модар аст!

print(to_persian("Донишгоҳи Теҳрон"))
# دانشگاه تهران

These examples come from the model card. Shekar uses a compressed ONNX version that runs on a CPU without PyTorch. For longer documents, process short sentences individually: training examples were limited to 200 characters per side.

The model is released under the MIT license. Explore the weights and model card, follow the library usage guide, or read the training notebook to see how it was built. We hope it helps more readers and developers work across the two scripts.

Cite Shekar

If you use Shekar in your research, please cite the library’s paper:

“Amirivojdan, A. (2025). Shekar: A Python Toolkit for Persian Natural Language Processing. Journal of Open Source Software, 10(114), 9128. DOI: 10.21105/joss.09128.”