Semantic sovereignty in large language models: Auditing the representation of persianate and Shi‘i meanings across languages

Document Type : Original Research Paper

Authors

1 PhD Candidate in Applied Linguistics, Hakim Sabzevari University, Sabzevar, Iran

2 Assistant Professor, Department of English Language Teaching, Gonbad KavoosBranch, Islamic Azad University, Gonbad Kavoos, Iran

10.22034/spektrum.2026.587146.1069
Abstrakt
Large language models increasingly translate, summarize, and explain traditions they have not inherited. This study develops semantic sovereignty as a framework for testing whether multilingual models preserve Persianate and Shi‘i meanings or relocate them into dominant Anglophone interpretive regimes. A bounded cross-lingual mixed-method audit examined 120 purposively selected units from Rumi, Hafez, Ferdowsi, and Shi‘i hermeneutics. Four model systems, blinded during expert scoring, were tested under Persian-only, English-only, cultural-context, and expert-context conditions, with two runs per prompt, yielding 3,840 outputs. A three-member disciplinary panel applied a 0–4 Semantic Preservation Scale and a multi-label error taxonomy; mutually exclusive percentages summarize each output’s adjudicated dominant error. Across 16 model-by-domain cells, the mean score was 2.53/4.00 (SD = 0.38). Rumi was frequently psychologized into global spirituality, Hafez was over-clarified, Ferdowsi was depoliticized into leadership language, and Shi‘i concepts were flattened into generic religious vocabulary. Expert-context prompts produced the strongest descriptive means, but displacement persisted. Because product identities, exact access dates, complete raw transcripts, and output-level scores were not retained, the analysis is descriptive and bounded rather than inferential or product-reproducible. Multilingual AI should therefore be evaluated not only for linguistic performance or cultural bias, but for hermeneutic accountability: preserving metaphor, doctrine, ambiguity, register, authority, and civilizational memory. The study contributes a culturally grounded audit framework for identifying subtle semantic loss across languages.

Keywords

Subjects

Abaskohi, A., Baruni, S., Masoudi, M., Abbasi, N., Babalou, M. H., Edalat, A., Kamahi, S., Mahdizadeh Sani, S., Naghavian, N., Namazifard, D., Sadeghi, P., & Yaghoobzadeh, Y. (2024). Benchmarking large language models for Persian: A preliminary study focusing on ChatGPT. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 2189–2203). ELRA and ICCL.
Abdollahpour Sangchi, F., Rahnamaei, H., Asgariyazdi, A., & Rezaee, M. (2025). Artificial intelligence and digital hermeneutics: Data bias, algorithmic ethics, and social implications. Spektrum Iran, 38(2), 213–242.
Akbari, R. (2025). Rumi’s conceptions of meaning in life: Reading the Seven Sermons (Majāles-e Sabʿa) toward wisdom pedagogy in adult education. Spektrum Iran, 38(1), 1–29.
Akhavan, M., Ameli, S. R., Rahgozar, M., & Shahghasemi, E. (2025). Dual-spacization of intelligence: A theoretical retroduction of the socialization of artificial intelligence in meaning construction. Spektrum Iran, 38(2), 1–30.
Alak, A. I. (2023). The Islamic humanist hermeneutics: Definition, theory and application. Islam and Christian–Muslim Relations.
Azadibougar, O., & Patton, S. (2015). Coleman Barks’ versions of Rumi in the USA. Translation and Literature, 24(2), 172–189.
Bender, E. M., & Friedman, B. (2018). Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6, 587–604.
Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 5185–5198). Association for Computational Linguistics.
Blodgett, S. L., Barocas, S., Daumé III, H., & Wallach, H. (2020). Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 5454–5476). Association for Computational Linguistics.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems (Vol. 33, pp. 1877–1901). Curran Associates.
Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. In S. A. Friedler & C. Wilson (Eds.), Proceedings of the 1st Conference on Fairness, Accountability, and Transparency (Proceedings of Machine Learning Research, Vol. 81, pp. 77–91). PMLR.
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., Ye, W., Zhang, Y., Chang, Y., Yu, P. S., Yang, Q., & Xie, X. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3), Article 39.
Davidson, O. M. (1994). Poet and hero in the Persian Book of Kings. Cornell University Press.
Davis, D. (2006). Shahnameh: The Persian Book of Kings. Viking.
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 4171–4186). Association for Computational Linguistics.
Durmus, E., Nguyen, K., Liao, T. I., Schiefer, N., Askell, A., Bakhtin, A., Chen, C., Hatfield-Dodds, Z., Hernandez, D., Joseph, N., Lovitt, L., McCandlish, S., Sikder, O., Tamkin, A., Thamkul, J., Kaplan, J., Clark, J., & Ganguli, D. (2023). Towards measuring the representation of subjective global opinions in language models. arXiv.
Farahani, M., Gharachorloo, M., Farahani, M., & Manthouri, M. (2021). ParsBERT: Transformer-based model for Persian language understanding. Neural Processing Letters, 53, 3831–3847.
Gadamer, H.-G. (2004). Truth and method (J. Weinsheimer & D. G. Marshall, Trans.; 2nd rev. ed.). Continuum. (Original work published 1960)
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86–92.
Havsson, M., & Sajjadi, M. (2025). Iranian digital discourse, affective alignments, and the geopolitics of AI. Spektrum Iran, 38(2), 187–212.
Ingenito, D. (2018). Hafez’s “Shirāzi Turk”: A geopoetical approach. Iranian Studies, 51(6), 901–930.
Khashabi, D., Cohan, A., Shakeri, S., Hosseini, P., Pezeshkpour, P., Alikhani, M., Aminnaseri, M., Bitaab, M., Brahman, F., Ghazarian, S., Gheini, M., Kabiri, A., Karimi Mahabadi, R., Memarrast, O., Mosallanezhad, A., Noury, E., Raji, S., Rasooli, M. S., Sadeghi, S., … Yaghoobzadeh, Y. (2021). ParsiNLU: A suite of language understanding challenges for Persian. Transactions of the Association for Computational Linguistics, 9, 1147–1162.
Kia, M. (2020). Persianate selves: Memories of place and origin before nationalism. Stanford University Press.
Lewis, F. D. (2015). Guest editor’s introduction: The Shahnameh of Ferdowsi as world literature. Iranian Studies, 48(3), 313–321.
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Ré, C., Acosta-Navas, D., Hudson, D. A., … Wu, Y. (2023). Holistic evaluation of language models. Transactions on Machine Learning Research.
Mavani, H. (2013). Religious authority and political thought in Twelver Shi‘ism: From Ali to post-Khomeini. Routledge.
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), Article 115.
Meraji Oskuie, S. (2025). Gender construction in anthropomorphizing generative AI: An interplay of society and technology. Spektrum Iran, 38(2), 293–324.
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220–229). Association for Computing Machinery.
Naghmeh-Abbaspour, B., Mahadi, T. S. T., & Zulkali, I. (2021). Ideological manipulation of translation through translator’s comments: A case study of Barks’ translation of Rumi’s poetry. Pertanika Journal of Social Sciences & Humanities, 29(3), 1831–1851.
Nasr, S. H. (2006). Islamic philosophy from its origin to the present: Philosophy in the land of prophecy. State University of New York Press.
Noble, S. U. (2018). Algorithms of oppression: How search engines reinforce racism. New York University Press.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems (Vol. 35, pp. 27730–27744). Curran Associates.
Qin, L., Chen, Q., Zhou, Y., Chen, Z., Li, Y., Liao, L., Li, M., Che, W., & Yu, P. S. (2025). A survey of multilingual large language models. Patterns, 6(1), Article 101118.
Ricoeur, P. (1981). Hermeneutics and the human sciences: Essays on language, action and interpretation (J. B. Thompson, Ed. & Trans.). Cambridge University Press.
Said, E. W. (1978). Orientalism. Pantheon Books.
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., & Hashimoto, T. (2023). Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202, pp. 29971–30004). PMLR.
Sazmand, B., & Mozaffari Falarti, M. (2024). Universal messages of peace: The enduring legacy of Jalal ad-Din Muhammad Balkhi (widely known as Rumi or Mawlana) and Hakim Abul-Qasim Ferdowsi in global discourse. Spektrum Iran, 37(2), 83–104.
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., Kluska, A., Lewkowycz, A., Agarwal, A., Power, A., Ray, A., Warstadt, A., Kocurek, A. W., Safaya, A., Tazarv, A., … Wu, Z. (2023). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research.
Tao, Y., Viberg, O., Baker, R. S., & Kizilcec, R. F. (2024). Cultural bias and cultural alignment of large language models. PNAS Nexus, 3(9), Article pgae346.
Totaro, M. W., Gheisi, L., & Shahghasemi, E. (2025). Affective asymmetries in AI: Sentiment bias between English and Persian in harmonized LLM pipelines. Spektrum Iran, 38(2), 143–157.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems (Vol. 30, pp. 5998–6008). Curran Associates.
Xu, Y., Hu, L., Zhao, J., Qiu, Z., Xu, K., Ye, Y., & Gu, H. (2025). A survey on multilingual large language models: Corpora, alignment, and bias. Frontiers of Computer Science, 19(11), Article 191101.
Yarshater, E. (1988). Hafez I: An overview. In Encyclopaedia Iranica.
Zhao, H., Chen, H., Yang, F., Liu, N., Deng, H., Cai, H., Wang, S., Yin, D., & Du, M. (2024). Explainability for large language models: A survey. ACM Transactions on Intelligent Systems and Technology, 15(2), Article 20.
Zhao, W., Mondal, D., Tandon, N., Dillion, D., Gray, K., & Gu, Y. (2024). WorldValuesBench: A large-scale benchmark dataset for multi-cultural value awareness of language models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 17696–17714). ELRA and ICCL.
Zhu, S., Supryadi, S., Xu, S., Sun, H., Pan, L., Cui, M., Du, J., Jin, R., Branco, A., & Xiong, D. (2024). Multilingual large language models: A systematic survey. arXiv.