JBRA Assisted Reproduction 2025;29(Suppl.2 SBRA 2025):32
Poster Presentation

29th Annual Congress of the SBRA. São Paulo/SP - Brazil, 2025
doi: 10.5935/1518-0557.20263556

P-020. Assessment of the Reliability of Information Generated by Artificial Intelligence in Reproductive Medicine: Critical Comparison with Clinical Guidelines

Marina Pinheiro Bezerra de Menezes1, Marina Almeida Simões1, Amanda Dias Cavalcante1, Maria Eduarda Cordeiro de Alencar1, Marcelo Cavalcante1

1UNIFOR - Universidade de Fortaleza - Fortaleza - CE - Brasil

Objective: The present study aims to evaluate the impact of Artificial Intelligence (AI) on the dissemination of information in the field of reproductive health, verifying whether the ChatGPT-3.5 and 4º, Google Gemini, and DeepSeek platforms are reliable sources for health education or whether they pose a risk to learning due to possible inaccuracies.
Methods: The study has an observational, cross-sectional, and analytical design, aimed at evaluating the accuracy and consistency of the responses of artificial intelligence models (ChatGPT-3.5, ChatGPT-4o, DeepSeek, and Google Gemini) in relation to the recommendations of the European Society of Human Reproduction and Embryology (ESHRE) on unexplained infertility. The guiding question extracted from the guideline was: “Should female or male partner’s age affect the definition of unexplained infertility (UI)?”, applied identically to the platforms, aiming at a systematic comparison of the content. Only responses directly related to the topic will be included, classified as: (1) no response; (2) insufficient; (3) not convergent with the guideline; (4) adequate and aligned with the recommendations. Responses in languages other than English or Portuguese and duplicates will be excluded. The question will be translated into Portuguese prior to application. Thus, the responses obtained will be submitted to a readability assessment using an online tool, with a numerical score inversely proportional to the ease of reading. The qualitative analysis will follow a 7-point Likert scale, according to the following parameters: (1) strongly disagree; (2) disagree; (3) partially disagree; (4) neither agree nor disagree; (5) partially agree; (6) agree; (7) strongly agree. In addition, it will have predefined explanatory captions to standardize the judgment criteria. Data collection will take place via Google Forms, with three independent evaluators. The primary outcome will be the agreement between the AI responses and the ESHRE guidelines, measuring content similarity and consistency of the information transmitted. The risks are minimal, limited to the interpretation of the guidelines by the AIs, the absence of external validation, and possible scientific updates. The expected benefit is to identify the degree of reliability of the information generated by AI, favoring its clinical application and the development of future guidelines. The statistical analysis will use data exported to the Google Sheets platform, with calculation of absolute and relative frequencies, mean, and standard deviation.
Results: When evaluating ChatGPT-3.5's response to the question “Should the age of the partner affect the definition of unexplained infertility?”, there was disagreement among all evaluators, with ratings of “Neither agree nor disagree,” “Partially agree,” and “Agree” in the comparative analysis with the guideline. In terms of readability, a level of 12 (“high”) and level 14 (“medium”) were recorded. ChatGPT-4o obtained 66.67% in “Partially agree” and the rest in “Agree,” with total agreement in readability level 11 (“high”). Google Gemini scored 66.67% in “Partially agree” and the rest in “Partially disagree,” with readability level 11 (“high”) in all results. Finally, DeepSeek received 100% in “Agree” and readability level 15 (“medium”) by all evaluators.
Conclusion: The comparative analysis showed that, although all the AIs evaluated — ChatGPT-3.5, ChatGPT-4o, Google Gemini, and DeepSeek — presented predominantly “high” or “medium” readability, agreement with ESHRE recommendations on unexplained infertility varied among the models. Thus, despite their potential as tools to support reproductive health education, AIs still exhibit significant variation in alignment with guidelines and may disseminate inaccurate information if used without professional supervision. The need for external validation and further studies to ensure safe and evidence-based use in clinical practice is reinforced.