Accuracy and Readability of Large Language Models’ Responses to Frequently Asked Questions on Social Assistance


Tepe H. T.

Journal of the Society for Social Work and Research, cilt.17, sa.2, ss.217-240, 2026 (SSCI, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 17 Sayı: 2
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1086/740347
  • Dergi Adı: Journal of the Society for Social Work and Research
  • Derginin Tarandığı İndeksler: Social Sciences Citation Index (SSCI), Scopus, IBZ Online, Psycinfo, Sociology Source Ultimate (EBSCO)
  • Sayfa Sayıları: ss.217-240
  • Anahtar Kelimeler: access to information, artificial intelligence and large language models, digital inclusion, social assistance, Türkiye
  • Bilecik Şeyh Edebali Üniversitesi Adresli: Evet

Özet

Objective: Artificial-intelligence-powered large language models (LLMs) are reshaping information access and increasingly serve as digital sources for public-service-related content, including social assistance. However, concerns remain regarding the accuracy and readability of the information LLMs generate, particularly for vulnerable populations. This study examines the accuracy and readability of responses produced by ChatGPT-3.5, ChatGPT-4o, Gemini, and Microsoft Copilot to frequently asked questions about social assistance programs in Türkiye. Method: Using 20 questions selected from the official website of the Ministry of Family and Social Services, LLM-generated responses were collected in June 2024 and evaluated with a researcher-developed Likert-type scale for accuracy and the Ateşman Readability Formula for readability. Results: The models showed limited performance in both domains, with ChatGPT-3.5 producing the lowest averages in accuracy and readability. Gemini achieved the highest readability scores, whereas Microsoft Copilot performed best in terms of accuracy. ChatGPT-4o demonstrated moderate performance across both metrics. I found no significant relationship between accuracy and readability. Conclusions: These results highlight the need to assess LLMs not only for their technical accuracy but also for their ability to produce accessible and comprehensible content, particularly in the context of public information on social assistance.