JBRA Assisted Reproduction 2025;29(Suppl.2 SBRA 2025):30
Poster Presentation
29th Annual Congress of the SBRA. São Paulo/SP - Brazil, 2025
doi: 10.5935/1518-0557.20263554
P-018. Assessment of Artificial Intelligence knowledge on human reproduction: A comparison with the Guidelines of the European Society of Human Reproduction and Embryology
Nara Santos Guerra1, Bruna Maia Furtado Figueiredo1, Letícia Timbó Martins Ferreira1, Amanda de Alevir1, Marcelo Borges Cavalcante1
1UNIFOR - Universidade de Fortaleza - Fortaleza - CE - Brasil
Objective: This study aims to assess the impact of Artificial Intelligence (AI) on the dissemination of reproductive health information, verifying whether ChatGPT-3.5 and 4.0, DeepSeek, and Google Gemini are reliable means for health education or whether they compromise learning due to possible flaws or misinformation.
Methods: This is an observational, cross-sectional, and analytical study that seeks to analyze the quality and accuracy of the responses provided by different artificial intelligences (AIs) — ChatGPT-3.5 and 4o, DeepSeek, and Google Gemini — based on the questions and answers in the European Society of Human Reproduction and Embryology (ESHRE) guideline on unexplained infertility. The study sample consisted of a guiding question from the aforementioned guideline, which was: “What is the need for investigations of the female lower genital tract?” This question was submitted identically to the three AI platforms mentioned, aiming at a systematic comparison of the responses provided. Only responses that met the following classification criteria were included in the analysis: (1) the AI was unable to answer the question; (2) the AI gave an answer, but it was considered insufficient; (3) the AI responded, but in a manner that was not consistent with the guideline; and (4) the AI responded appropriately, in accordance with the guidelines. Responses generated in languages other than English or Portuguese were disregarded, as were duplicates, which were removed from the analysis. The question extracted from the ESHRE guideline was translated into Portuguese and then applied to the selected AIs. Each response was subjected to a readability analysis on an online platform, which assigns a numerical score for ease of reading: the higher the value, the lower the readability of the response. In addition, the responses were evaluated qualitatively using a 7-point Likert scale, with the following parameters: (1) strongly disagree, (2) disagree, (3) partially disagree, (4) neither agree nor disagree, (5) partially agree, (6) agree, and (7) strongly agree. Data collection and systematization were performed by filling out forms in Google Forms, with each AI-generated response being evaluated by three different researchers, ensuring greater reliability of the analysis, and then exported to Google Sheets. On this platform, the data were organized according to absolute frequency, relative frequency, mean, and standard deviation.
Results: When analyzing Chat GPT-3.5's response to the question, “What is the need for investigations of the female lower genital tract?”, it was found that two-thirds of the evaluators disagreed, while only one evaluator neither agreed nor disagreed. The response presented was analyzed and classified as average readability, with a level of 15. In comparison, the ChatGPT-4o response was evaluated by two of the participants as “neither agree nor disagree,” obtaining a different disagreement rating from the other evaluator and a low readability level, being classified at level 21. Another analysis was performed on DeepSeek, in which one of the analyses corresponded to “partially disagree,” while the rest rated it as “disagree.” It was also verified that this response has a readability level of 16, being classified as average. Finally, when presenting the question to Google Gemini, two of the evaluators classified it as “partially disagree,” resulting in a readability rating of level 15.
Conclusion: Therefore, this study showed that, although the AIs analyzed — ChatGPT-3.5, ChatGPT-4o, DeepSeek, and Google Gemini — show potential as effective tools in health education, they do not yet demonstrate full reliability when compared to consolidated guidelines, such as those of ESHRE. This is evidenced by the analysis of the responses, which showed low compliance with the expected parameters, indicating a risk of disseminating incorrect information if used without professional supervision.