Introduction
Artificial intelligence (AI) tools such as ChatGPT are increasingly used in healthcare, but their effectiveness in providing accurate patient education remains uncertain. In a prior study, a retrieval-augmented model trained on urogynecologic patient education materials outperformed standard ChatGPT in blinded assessments. To further validate its utility under “real-world” conditions where experts independently pose questions, we evaluated the model’s response quality and usability for common pelvic floor disorder queries.
Methods
We developed a retrieval-augmented ChatGPT model using the American Urogynecologic Society’s (AUGS) patient education leaflets. Pelvic floor disorder specialists were recruited via AUGS forums, email, and community outreach. Participants submitted patient-representative questions and rated model responses across six domains—accuracy, clarity, relevance, completeness, usefulness, and provision of references—using the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool (5-point Likert scale) [1]. Usability was measured using the 10-item System Usability Scale (SUS, score range 0–100) [2]. Participants also provided qualitative feedback. Descriptive statistics and qualitative content analysis were performed.
Results
Thirteen specialists entered the survey, with 11 completing evaluations of 22 model responses. Most respondents were physicians (85%) with diverse experience levels, from fellows to >21 years in practice, and varied familiarity with AI. The median QAMAI total score was 28 (IQR 24–29.8), reflecting moderate to high response quality. Usability was rated highly, with a mean SUS score of 81.8±11.2.
Qualitative feedback highlighted three main areas for improvement: (1) consistent source citation, (2) greater clinical depth, and (3) more comprehensive answers beyond simplified overviews (Table 1). Respondents praised the model’s ease of use, patient-friendly tone, and integration of AUGS links but noted occasional inaccuracies and insufficient detail, reinforcing the need for clinician oversight in patient education.
Conclusions
A retrieval-augmented ChatGPT model grounded in AUGS patient education materials produced high-quality, highly usable responses to patient questions in urogynecology. Enhancements in citation practices and clinical depth are needed to strengthen its role as a patient education tool. Findings support AI’s potential to complement, but not replace, provider-patient communication in pelvic floor care.
Table
Table 1. Participant-identified areas for improvement in model responses.
| Theme | Respondent Quotes |
|---|---|
| Suggested changes or improvements to model responses | |
| Inconsistent source citation | “When I have seen other AI tools, there is a way to see where they came up with their information, and with this one, there is no information about how it came up with this information.” “Providing sources and references would be helpful - there are FDA pages that specifically address risks of mesh for prolapse and mesh for incontinence. separating mid-urethral slings from other mesh for prolapse would also be an improvement.” “Have the model provide references without prompting for its responses, example ‘Some studies suggest that urethral bulking has minimal impact on sexual function, ...’ what studies? please cite. I could continue to probe it and see what it would respond but it would be nice if it provides citation without prompting.” |
| Lack of clinical depth | “Useful information but not comprehensive or detailed.” “There is no mention of biofeedback or loperamide or advanced management strategies for fecal incontinence, like sacral neuromodulation or eclipse vaginal bowel control system, but it is a good first pass of information for a patient looking for more info.” “Consider including something about the small risk of blood transfusion. Also maybe the rate of mesh complications with this procedure.” |
| Over-simplified response | “This answer provides a great general content review but does require additional interaction with a urogynecologist for a comprehensive and detailed review.” “Not as complete as my typical counseling’s.” “I found this answer to be incredibly oversimplified but helpful as an introduction for patient orientation.” |
References
1. Vaira LA, Lechien JR, Abbate V, Allevi F, Audino G, Beltramini GA, Bergonzani M, Boscolo-Rizzo P, Califano G, Cammaroto G, Chiesa-Estomba CM, Committeri U, Crimi S, Curran NR, di Bello F, di Stadio A, Frosolini A, Gabriele G, Gengler IM, Lonardi F, Maglitto F, Mayo-Yáñez M, Petrocelli M, Pucci R, Saibene AM, Saponaro G, Tel A, Trabalzini F, Trecca EMC, Vellone V, Salzano G, De Riu G (2024) Validation of the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool: a new tool to assess the quality of health information provided by AI platforms. Eur Arch Otorhinolaryngol 281 (11):6123-6131. doi:10.1007/s00405-024-08710-0
2. Brooke J (1995) SUS: A quick and dirty usability scale. Usability Eval Ind 189