Introduction
Accurate discharge prediction is critical for health care delivery, affecting key factors in healthcare operations including scheduling, bed utilization, capacity management, and patient satisfaction1. While large language models (LLMs) have demonstrated effectiveness in healthcare operational tasks, including length of stay prediction and readmission forecasting,2 their performance can vary based on operational factors such as hospital-specific clinical workflows, primary service team preferences, and cultural differences in care. This study analyzed the ability of a LLM to predict hospital discharge among patients hospitalized at a large academic medical center, identified performance barriers for LLM success, and developed optimization strategies for improving future LLM performance.
Methods
Same-day discharge predictions were generated using the Llama 3.3-70B model for a randomized dataset of 1,000 inpatients at a tertiary academic medical center hospitalized from January 1, 2024 through December 31, 2024. Two-step prompting was used to generate a binary discharge prediction (1=discharge, 0=no discharge) and a clinical explanation to support this prediction. Clinical notes from the previous day through 6am on prediction day were provided to the model. To characterize errors made by the model, a qualitative analysis was performed on 202 prediction errors (24 false positives, 178 false negatives) by four trained reviewers. Reviewers examined clinical documentation, model predictions, and model reasoning outputs to identify discordance patterns and extract representative examples.
Results
False positives (n=24) occurred due to inadequate knowledge of specialty-specific discharge protocols, failure to account for patient preferences despite clinical readiness, and reliance on sparse or outdated documentation. False negatives (n=178) resulted from excessive conservatism in prediction thresholds, overemphasis on early hospitalization events, inability to distinguish inpatient versus outpatient management requirements, and overweighting surgical complexity rather than current clinical stability.
Conclusions
LLMs have the potential to improve discharge prediction accuracy, but require integration of institution and specialty-specific clinical knowledge to improve reasoning capabilities. Service-specific interventions including targeted prompting for recovery protocols, improved distinction between inpatient and outpatient care needs, and better incorporation of current patient status over historical complexity may optimize model performance. These improvements could enhance resource allocation efficiency and support safer discharge decisions across inpatient populations.
References
Barnes S, Hamrock E, Toerper M, et al. Real-time prediction of inpatient length of stay for discharge prioritization. J Am Med Inform Assoc. 2016;23(e1):e2-e10. doi:10.1093/jamia/ocv106
Jiang LY, Liu XC, Nejatian NP, et al. Health system-scale language models are all-purpose prediction engines. Nature. 2023;619(7969):357-362. doi:10.1038/s41586-023-06160-y