Article In: Terminology: Online-First Articles
From readable to executable
A governance stress test of translation variants under generative ai in WHO traditional Chinese medicine terminology standards
This content is being prepared for publication; it may be subject to changes.
Abstract
Traditional terminology standards often assume that trained users can
resolve coexisting translation variants through contextual judgement. As generative
artificial intelligence (genAI) and large language models (LLMs) increasingly mediate
terminology reuse, this raises questions about whether standards encode sufficiently
explicit term-selection information for auditable automated selection. This study uses
ChatGPT-4o in a prompt-testing procedure to examine whether the WHO’s 2007 and 2022
Traditional Chinese Medicine (TCM) terminology standards provide sufficient guidance for
AI-mediated selection among coexisting English variants. Each sampled term was tested
using the same three-prompt sequence: constrained variant selection, brief justification,
and a within-session shift from general medical use to possible use in a WHO terminology
standard. The model consistently followed the single-choice restriction. Its rationales
clustered into four observable justification patterns. Most responses to the third prompt
stated that the initial choice would remain unchanged; these responses are reported only
as statements produced under this specific sequence. The findings suggest that terminology
standards designed for human readability may not provide sufficiently explicit guidance
for AI-mediated variant selection. The study therefore proposes a governance stress-test
procedure and a minimal set of candidate executable fields for term selection, subject to
expert validation and controlled testing.
Article outline
- 1.Introduction
- 2.Literature review
- 2.1Generative AI in translation and terminology work
- 2.2Translation variation
- 2.3WHO TCM terminology standards
- 2.4LLMs in terminology-resource diagnosis
- 2.5Conceptual framework
- 3.Methodology
- 3.1Research design
- 3.2Data source and collection
- 3.3Sample selection
- 3.4Experimental setting and procedural metadata
- 3.5Prompt design
- 3.6Interaction structure and output recording
- 3.7Handling of potential non-compliant responses
- 3.8Coding and reliability procedure for Q2 rationales
- 3.9Analysis of stated continuity and revision (Q3)
- 3.10Methodological boundaries and validity considerations
- 4.Results
- 4.1Q1 compliance and Q3 outcomes
- 4.2Recurrent justification patterns in Q2 rationales
- 4.3Distribution of Q2 rationale categories across domain branches
- 4.4Revisions in response to Q3
- 5.Discussion
- 6.Conclusion
- Acknowledgements
- Author queries
References
References (53)
Altakhaineh, Abdel Rahman Mitib, Ghazi Ayed Alghathian, and Mashal Mufleh Jarrah. 2025. “A Comparative Study of Accuracy in Human vs. AI Translation of Legal Documents into Arabic.” International Journal of Language & Law (JLL) 14: 63–80.
Bartsch, Henning, Ole Jorgensen, Domenic Rosati, Jason Hoelscher-Obermaier, and Jacob Pfau. 2023. “Self-Consistency of Large Language Models under Ambiguity.” In Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, 89–105. Singapore: Association for Computational Linguistics.
Bender, Emily M., Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?” In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623. New York: Association for Computing Machinery.
Bommasani, Rishi, Drew A. Hudson, Ehsan Adeli, et al. 2021. “On the Opportunities and Risks of Foundation Models.” arXiv:2108.07258.
Bowker, Lynne. 2015. “Terminology and Translation.” In Handbook of Terminology, vol. 1, ed. by Hendrik J. Kockaert and Frieda Steurs, 304–323. Amsterdam: John Benjamins.
Braun, Virginia, and Victoria Clarke. 2006. “Using Thematic Analysis in Psychology.” Qualitative Research in Psychology 3 (2): 77–101.
Cabré, M. Teresa. 2003. “Theories of Terminology: Their Description, Prescription and Explanation.” Terminology 9 (2): 163–199.
Cantos-Gómez, Pascual. 2025. “AI in Specialised Translation: Terminology, Quality and the Evolving Role of Human Translators.” Ibérica 50: 19–44.
Chen, Christine L., Yue Dong, Claudia Castillo-Zambrano, et al. 2025. “A Systematic Multimodal Assessment of AI Machine Translation Tools for Enhancing Access to Critical Care Education Internationally.” BMC Medical Education 25 (1): 1022.
De Schryver, Gilles-Maurice. 2023. “Generative AI and Lexicography: The Current State of the Art Using ChatGPT.” International Journal of Lexicography 36 (4): 355–87.
Freixa, Judit. 2006. “Causes of Denominative Variation in Terminology: A Typology Proposal.” Terminology 12 (1): 51–77.
Fu, Linling, and Lei Liu. 2024. “What Are the Differences? A Comparative Study of Generative Artificial Intelligence Translation and Human Translation of Scientific Texts.” Humanities and Social Sciences Communications 11: 1236.
Gergel, Peter, Oľga Wrede, Daša Munková, and Lucia Benková. 2025. “Maschinelle und menschliche Übersetzung im Vergleich am Beispiel des Entwurfs des ersten ungarischen Zivilgesetzbuches (Deutsch–Slowakisch) [A Comparison of Machine and Human Translation Using the Draft of the First Hungarian Civil Code as an Example (German–Slovak)].” Fachsprache 47 (1–2): 62–87.
Giampieri, Patrizia. 2025. “Assessing the Quality of AI and MT in Legal Translation.” Altre Modernità, no. 33: 143–159.
Hasler, Eva, Adrià de Gispert, Gonzalo Iglesias, and Bill Byrne. 2018. “Neural Machine Translation Decoding with Terminology Constraints.” In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2, 506–512. New Orleans: Association for Computational Linguistics.
Heinisch, Barbara. 2025. “Next-Gen Terminology: Transforming Terminology Work with Large Language Models.” Across Languages and Cultures 26 (S): 64–80.
Hendy, Amr, Mohamed Abdelrehim, Amr Sharaf, et al. 2023. “How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation.” arXiv:2302.09210.
Hernández Fresno, Elena, and María Teresa Ortego Antón. 2025. “Hacia una traducción automática inclusiva: La intersección entre inteligencia artificial, terminología LGTBIQ+ y sesgo de género [Towards inclusive machine translation: The Intersection of Artificial Intelligence, LGBTQIA+ Terminology, and Gender Bias].” ELUA: Estudios de Lingüística. Universidad de Alicante 44: 125–145.
Hou, Yuehui. 2025. “Pruning Translation of Logical and Accidental Polysemy in Traditional Chinese Medicine Terminology.” Terminology 31 (2): 311–334.
Hsu, Elisabeth. 2000. “Spirit (Shen), Styles of Knowing, and Authority in Contemporary Chinese Medicine.” Culture, Medicine and Psychiatry 24 (2): 197–229.
Hu, Weilin, Wei Lin, Qizheng Li, Xiaona Yu, and Chengyan Zhu. 2025. “Textile AI-Enhanced Translation System Based on Mapping Probability and In-Context Learning.” ACM Transactions on Asian and Low-Resource Language Information Processing 24 (7): 73.
International Organization for Standardization. 2018. ISO/TR 20694:2018: A Typology of Language Registers. Geneva: International Organization for Standardization.
. 2019. ISO 30042:2019: Management of Terminology Resources — TermBase eXchange (TBX). Geneva: International Organization for Standardization.
. 2022. ISO 12620–1:2022: Management of Terminology Resources — Data Categories — Part 1: Specifications. Geneva: International Organization for Standardization.
Jacovi, Alon, and Yoav Goldberg. 2020. “Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?” In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 4198–4205. Association for Computational Linguistics.
Kasneci, Enkelejda, Kathrin Seßler, Stefan Küchemann, et al. 2023. “ChatGPT for Good? On Opportunities and Challenges of Large Language Models for Education.” Learning and Individual Differences 103: 102274.
Lai, Han, Aaron Lee Moore, Xiaodan Yan, and Weihong Li. 2025. “A Qualitative Analysis of WHO’s International Standard Terminologies on TCM.” Chinese Medicine and Culture 8 (1): 86–95.
Landis, J. Richard, and Gary G. Koch. 1977. “The Measurement of Observer Agreement for Categorical Data.” Biometrics 33 (1): 159–174.
Li, Biwei. 2023. “Conceptual Deviation in Terminology Translation: A Case Study on Translating COVID-19 Terminology in Multilingual News Media.” Terminology 29 (2): 351–382.
Li, Ruoxin, and Yajun Li. 2022. “A Review of Standardization of TCM Terminology Translation.” Lecture Notes on Language and Literature 5 (3): 26–31.
Lu, Yao, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. “Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity.” In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8086–8098. Dublin: Association for Computational Linguistics.
McCrae, John P., Julia Bosque-Gil, Jorge Gracia, Paul Buitelaar, and Philipp Cimiano. 2017. “The OntoLex-Lemon Model: Development and Applications.” In Electronic Lexicography in the 21st Century: Lexicography from Scratch. Proceedings of the eLex 2017 Conference, ed. by Iztok Kosem, Carole Tiberius, Miloš Jakubíček, Jelena Kallas, Simon Krek, and Vít Baisa, 587–597. Brno: Lexical Computing CZ s.r.o. [URL]
Miles, Alistair, and Sean Bechhofer, eds. 2009. SKOS Simple Knowledge Organization System Reference. W3C Recommendation, August 18. World Wide Web Consortium. [URL]
Moneus, Ahmed Mohammed, and Yousef Sahari. 2024. “Artificial Intelligence and Human Translation: A Contrastive Study Based on Legal Texts.” Heliyon 10 (6): e28106.
Pang, Jianhui, Fanghua Ye, Derek Fai Wong, et al. 2025. “Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models.” Transactions of the Association for Computational Linguistics 13: 73–95.
Pezeshkpour, Pouya, and Estevam Hruschka. 2024. “Large Language Models Sensitivity to the Order of Options in Multiple-Choice Questions.” In Findings of the Association for Computational Linguistics: NAACL 2024, 2006–2017. Mexico City: Association for Computational Linguistics.
Post, Matt, and David Vilar. 2018. “Fast Lexically Constrained Decoding with Dynamic Beam Allocation for Neural Machine Translation.” In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1, 1314–1324. New Orleans: Association for Computational Linguistics.
Rao, Pavithra, Lauren M. McGee, and Casey A. Seideman. 2024. “A Comparative Assessment of ChatGPT vs. Google Translate for the Translation of Patient Instructions.” Journal of Medical Artificial Intelligence 7: 11.
Raus, Rachele, and Michela Tonti. 2025. “Intelligence artificielle, corpus et diversité linguistique: enjeux et perspectives. Introduction.” Langages 237 (1): 7–20.
Reinke, Uwe. 2013. “State of the Art in Translation Memory Technology.” Translation: Computation, Corpora, Cognition 3 (1): 27–48.
Sclar, Melanie, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024. “Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I Learned to Start Worrying about Prompt Formatting.” In The Twelfth International Conference on Learning Representations. [URL]
Temmerman, Rita. 2000. Towards New Ways of Terminology Description: The Sociocognitive Approach. Amsterdam: John Benjamins.
Turpin, Miles, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. “Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.” Advances in Neural Information Processing Systems 36: 74952–74965.
Vidal Sabanés, Laia, and Iria da Cunha. 2025. “AI as a Resource for the Clarification of Medical Terminology: An Analysis of Its Advantages and Limitations.” Terminology 31 (1): 37–71.
Warburton, Kara. 2018. “Terminology Resources in Support of Global Communication.” In The Human Factor in Machine Translation, ed. by Sin-wai Chan, 118–136. London: Routledge.
. 2025. “Terminology in the Age of AI.” MultiLingual, March 7. [URL]
World Health Organization. 2007. WHO International Standard Terminologies on Traditional Medicine in the Western Pacific Region. Manila: WHO Regional Office for the Western Pacific.
. 2022. WHO International Standard Terminologies on Traditional Chinese Medicine. Geneva: World Health Organization.
Xu, Qihe. 2023. “WHO International Standard Terminologies on Traditional Chinese Medicine: Use in Context, Creatively.” Integrative Medicine in Nephrology and Andrology 10 (2): e00029.
Ye, Xiao, and Hongxia Zhang. 2017. “A History of Standardization in the English Translation of Traditional Chinese Medicine Terminology.” Journal of Integrative Medicine 15 (5): 344–350.
Zhou, Tong, and Jinghui Wang. 2025. “Embodied Empathy in Translation Studies: Enhancing Global Readers’ Cognitive and Emotional Engagement with Translations of Traditional Chinese Medicine Terminology.” Frontiers in Psychology 16: 1618531.