JBRA Assisted Reproduction 2025;29(Suppl.2 SBRA 2025):252
Poster Presentation

29th Annual Congress of the SBRA. São Paulo/SP - Brazil, 2025
doi: 10.5935/1518-0557.20263829

P-240. Unexplained infertility and sexual frequency: are the responses of Artificial Intelligence models aligned with ESHRE recommendations?

Marina Norões1, Marina Pinheiro Bezerra de Menezes1, Eduarda Loiola Werner1, Amanda Dias Carvalho1, Marcelo Cavalcante1, Maria Eduarda Cordeiro De Alencar1

1 UNIFOR - Universidade de Fortaleza – Fortaleza - CE - Brasil

Objective: The main objective of this study is to evaluate the accuracy and consistency of the responses provided by different Artificial Intelligence models (ChatGPT-3.5, ChatGPT-4o, DeepSeek, and Google Gemini) in light of the recommendations of the European Society of Human Reproduction and Embryology (ESHRE) on unexplained infertility. Specifically, we seek to verify whether these platforms agree with the guidelines when addressing the issue of the influence of sexual frequency on the definition of this diagnosis.
Methods: An observational, cross-sectional, analytical study was designed to verify the accuracy and consistency of the responses provided by the four selected AI platforms. To this end, a question based on the ESHRE guidelines will be used: "Should the frequency of sexual intercourse affect the definition of unexplained infertility?" This question will be presented identically to each of the artificial intelligences, allowing for direct comparison. Only responses relevant to the subject and that fall within the following parameters will be considered: (1) total absence of response; (2) response present, but insufficient; (3) response incompatible with the content of the guideline; (4) complete response in accordance with the recommendations. Duplicate responses or those written in a language other than Portuguese or English will be excluded. Each result obtained will undergo a readability assessment using an online tool that generates a numerical score—higher values indicate lower readability. A qualitative analysis will also be conducted using a seven-point Likert scale, ranging from "strongly disagree" (1) to "strongly agree" (7), with predefined descriptions to standardize the evaluation criteria. The data will be collected via Google Forms and analyzed by three independent evaluators, ensuring greater reliability. The main outcome will be to assess the degree of alignment between the AI responses and the ESHRE guidelines, considering clarity, accuracy, and consistency. The anticipated risks are low, limited to the possible partial or incorrect interpretation of the recommendations and the absence of external verification. As benefits, the study can map the reliability of the information generated, favoring its safe use in clinical contexts and health education initiatives. The statistical analysis will be performed in Google Sheets, considering absolute frequency, relative frequency, mean, and standard deviation.
Results: When checking ChatGPT-3.5's response to the question "Should the frequency of sexual intercourse affect the definition of unexplained infertility?", it was observed that 66.6% of participants marked "partially agree" and 33.3% marked "agree." Readability was consistent across all assessments, classified as "Level 16, average." In comparison, ChatGPT-4o's response was unanimous in "strongly agree," with readability "Level 14, average." An identical result was found in DeepSeek, also with unanimity in "strongly agree" and readability "Level 14, average." Google Gemini, on the other hand, scored 33.3% in "agree" and 66.6% in "partially agree," with readability "Level 14, understandable for university students."
Conclusion: The evaluation of the four AI models showed that, although all presented university-level readability, there were variations in adherence to the recommendations. ChatGPT-4o and DeepSeek remained fully aligned with the guideline, conveying greater certainty. ChatGPT-3.5 presented scattered responses, suggesting less uniformity and the influence of different interpretations. Google Gemini took an intermediate position, with concordant excerpts but without full consistency, which may reduce reliability. These results indicate that, although AIs can support reproductive health education, it is essential to use them with critical analysis and evidence-based reasoning, prioritizing models with greater scientific alignment.