JBRA Assist. Reprod. 2025;29(Suppl 1):15-16
POSTER PRESENTATION
doi: 10.5935/1518-0557.20250071
1Cegyr, Eugin Group - Research department (Buenos Aires, Argentina)
Objective:
To compare the predictive accuracy of domain-specific AI algorithms, general AI models, and infertility specialists in forecasting IVF outcomes.
To evaluate the sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV) of domain-specific across all methods in IVF outcome prediction.
To assess the predictive accuracy of a general AI model in comparison to domain-specific AI algorithms and fertility specialists.
To determine the ability of infertility specialists to predict IVF outcomes and compare their performance with AI-based models.
Methods: Reliable prediction of IVF outcomes is essential for managing patient expectations and guiding treatment decisions. Oocyte age is a key prognostic factor but insufficient for precise outcome prediction. AI algorithms offer the potential to improve accuracy by integrating multiple predictive variables. This prospective cohort study included 206 patients undergoing IVF with their own oocytes between October 2023 and March 2024. Predictive performance of domain-specific AI algorithms (ExpectMore, SART, CDC), a general AI model (ChatGPT), and infertility specialists were compared. Key metrics included sensitivity, specificity, area under the curve (AUC), and positive predictive value. PPV was analyzed as the proportion of true negatives among predicted negatives. Optimal thresholds were determined before analysis.
Results: All methods were more accurate in predicting that a woman would not get pregnant (low chances of success) than predicting that she would (high chances of success). PPV emerged as the most reliable metric, reflecting the proportion of cases with true negative results among all cases predicted as negative. ExpectMore (Se: 27%, Sp: 95%, AUC: 0.61) and SART (Se: 31%, Sp: 93%, AUC: 0.62) demonstrated similar performance, with superior PPVs (93% and 92%, respectively) compared to CDC (Se: 41%, Sp: 84%, AUC: 0.62, PPV: 86%), ChatGPT (Se: 37%, Sp: 76%, AUC: 0.55, PPV: 78%, p<0.01), and specialists (Se: 30%, Sp: 81%, AUC: 0.57, PPV: 78%, p<0.01). Both senior and junior specialists performed worse than domain-specific AI tools, with no significant differences based on experience (kappa=0.67).
Conclusion: Domain-specific AI algorithms could complement clinical expertise by more effectively identifying poor prognosis cases, thereby improving decision-making, patient counseling, and expectation management. However, both the analyzed algorithms and specialist assessments still overlook important predictive variables. Future research should focus on addressing these gaps and enhancing the accuracy of predicting positive IVF outcomes, ultimately leading to more tailored and effective treatment strategies. These findings require validation with larger, more diverse patient populations to confirm generalizability.

Table 1. Prediction parameters for Poor Prognosis of Cumulative Pregnancy Rate.

Table 2. Comparison Between AI and Expert Physicians in Predicting Poor Prognosis of Cumulative Pregnancy Rate.