Skip to main content
Literature Review

Literature Review: Racial and Social Bias in Artificial Intelligence for Healthcare

Author: Alexandria Balde (University of Michigan)

  • Literature Review: Racial and Social Bias in Artificial Intelligence for Healthcare

    Literature Review

    Literature Review: Racial and Social Bias in Artificial Intelligence for Healthcare

    Author:

Abstract

Artificial intelligence (AI), its machine learning applications, and predictive models are increasingly used in healthcare to streamline processes and improve patient outcomes. However, numerous limitations and imperfections have led to racially and socially biased results, along with discriminatory actions against underrepresented patients. The overburdening and shortage of healthcare professionals and the increased need for cost-effective solutions have led to the use of AI to automate healthcare activities such as administrative tasks, prognoses, therapies, and diagnoses. However, there are unresolved instances of prejudices, misdiagnoses, and mistreatment based on race and intersectional minorities. The lack of diverse and representative data for training AI models contributes to biased outcomes. Additionally, the regulation of AI devices in medicine is incomplete, and the many design flaws do not address the shortcomings of AI. This continues to exacerbate existing health disparities by perpetuating unequal access to quality care and resources. In the short term, professionals must approach AI cautiously to avoid overreliance, and temporary measures can also be taken to increase individual patient precision medicine literacy and improve representation in genetic information by addressing travel and language access to clinical studies. In the long term, stricter protocols must be enforced to ensure compliance, specifically monitoring AI in all stages of production to combat bias and the multiple weaknesses of AI. The training data also must ensure that a comprehensive representation of racial and ethnic groups helps mitigate bias in machine learning models.

Keywords: Artificial Intelligence, Healthcare, Clinical Decision-making, Discrimination, Bias

How to Cite:

Balde, A., (2026) “Literature Review: Racial and Social Bias in Artificial Intelligence for Healthcare”, University of Michigan Undergraduate Research Journal 18: 12. doi: https://doi.org/10.3998/umurj.9828

54 Views

18 Downloads

Published on
2026-06-04

Peer Reviewed

Current Ramifications of AI in Patient Care

The deployment of AI, language learning models (LLMs), and predictive models or algorithms has become increasingly prevalent in medical settings. Technology is now used to automate administrative tasks, suggest testing, and shortcut diagnoses (2). However, there are many limitations and imperfections in the training and development aspects that can produce bias which actively perpetuates health disparities for underrepresented communities. For example, in November 2023, two patients’ families filed a lawsuit against private provider UnitedHealth when their automatic AI model falsely denied their insurance, and the lawsuit noted their model has been known to have a 90% error rate (13). Another study found UnitedHealth’s model deeming African American patients’ health needs as “less than” that of white patients, thus resulting in only 17.7% of African American patients being determined to need critical care, in contrast with the more accurate number of 46.5% (21, p. 221).

Numerous healthcare systems also rely on algorithms to organize specialized care, such as high-risk care management that supports the patient with greater medical attention via a “risk score.” Another nationally used algorithm concluded that African American patients ranked in the same risk score, meaning similar assessment of health, as white patients were still significantly sicker (14, pp. 2–3). And any algorithms that uses cost to causally predict health outcomes is not without unaddressed racial bias. African Americans have been historically economically disadvantaged and are more likely to be uninsured than white Americans (9). As such, African Americans have been underrecognized for AI-backed targeted health programs, and the in-place algorithm’s faulty logic perpetuates racism and unequal accessibility to needed programs and resources.

Overburdened Practitioners and Overreliance on AI

The advent of AI in healthcare followed the Affordable Care Act (ACA)’s significant shaping of modern United States healthcare, allowing more people to qualify for and use care coverage. However, the increased patient population is met with a shortage of physicians; even before the enactment of the ACA, the Bureau of Health Professions stated that 7000 physicians were needed for under-resourced communities alone. After the ACA, the current projection of needed physicians is tenfold (10). The shortage of physicians leads to the overburdening of current professionals and the increased need for aid. And as a solution, AI’s predictive mechanisms allow physicians to organize health management and offload workloads while reducing the cost. Research reports the use of AI can reduce the average 13 hours per week that physicians and staff spend, and automate 50–75% of authorizations (2). Healthcare systems also benefit from AI’s role in automating administrative tasks like prior authorizations, to approve patients’ coverage for procedures, medication, and other healthcare costs. However, patient-categorizing tasks that can be automated, such as targeted care-management programs, can lead to heavy reliance on algorithms to identify patients (14, pp. 1–2). Of which, those algorithmic predictions are without the input of the providers or systems using it. LLMs and their generative text are also being used in electronic health records documentation such as Microsoft’s SlicerDicer, which is used at Stanford Health Care, UC San Diego Health, and UW Health in Wisconsin. And on the patient side, LLMs have also been incorporated in answering questions in various specialties such as cardiology, oncology, and anesthesiology. However, research has also shown that LLMs return erroneous, racially biased answers, such as the scientifically refuted use of race to predict kidney function via the glomerular filtration rate equation; moreover, models have been shown to fabricate equations completely (16).

Alongside overburdening of tasks, a global shortage of radiologists has led to a national need for automation in radiology. A previous study of the CheXNet deep learning detection models showcased that instances of AI are applicable in recognizing pneumonia in screenings (18). Yet another deep learning model utilizing the same neural network, CheXpert, along with two other public datasets, was shown to consistently and selectively underdiagnose disparate racial and gender populations, notably Hispanic female patients (19). In radiology specifically, diagnosis in triage is a deciding factor in clinical priority and lower accessibility to receive the necessary medical attention; thus, screening misdiagnosis can lead to fatal, unnoticed, and preventable consequences. Another study presenting a group of Californian computer-aided clinical detection tools, R2 ImageChecker M100 for mammograms, showed that the second-reader group that relied on the tool was less sensitive to image abnormalities than the group who screened without the aid (3). Thus, if AI and machinery are overreliant in diagnosing patients, improper machine conclusions could lead to misinformed physicians and other readers and misdiagnoses.

Historic Exclusion in Algorithm Training Models

It is fundamental to understanding that AI predictions are wholly the product of the data it was built upon: That is to say, AI will reflect the trends, bias, discrimination, inclusion, or lack-there-of, of the datasets that were used to inform it. Any racial bias and discrimination in data only exacerbates existing health disparities and can perpetuate unequal access to quality care and resources. The genome-wide association study (GWAS), for example, is a dataset that has made remarkable strides in disease etiology and detection labels for cancers and other diseases, which are the basis for genotype algorithms and precision medicine. Significant innovations have been made in genetic markers and disease association and drug mechanisms predictions; yet 70% of GWAS makeup is made by those of European ancestry, and only 2% are represented by African ancestry (20, p. 28). The disproportionate representation ultimately leads scientists, doctors, and any AI trained on GWAS to ignore genetic diversity, which can lead to unequal health outcomes. As an example, GWAS connected a genetic mutation to cystic fibrosis, which accounted for more than 70% of the cases in Europe, yet left other underrepresented populations, mainly African Americans, underdiagnosed at only 29% of cases even though there are distinct genetic markers for cystic fibrosis for that population (20, p. 26). Another study showed a pharmacogenetic genotype-based algorithm consistently over-prescribed medication for African American patients, which led to preventable, uncontrolled bleeding (7, p. 79). Moreover, because disproportionate participation of European ancestry provided the basis for research, that could lead to continual favor and perpetuation of that distribution (4, p. 259). And while there is no inherent biological and genetic difference among different races, genetics do significantly vary endemically. Thus genetic association testing, pharmacologic studies, and precision medicine treatments built from GWAS are designed with incredible disadvantage for underrepresented populations, and there are a myriad of barriers to increasing inclusivity. Research has shown that minority populations have limited financial and linguistic access to specialty care which can serve as an entry to participating in health research (15, pp. 3–4). There is also generational abuse and exploitation of minorities in research (8, p. 243); and substantial genetic literature is written by white male authors, and even in collaboration with authors from low- and middle-income countries, the latter are underrepresented as first and last authors (8, p. 244). Clearly, these barriers are long withstanding and have led to historic exclusion and manipulation of certain communities in data documentation and collection, and in order for AI to not repeat these offensive patterns, much work needs to be done.

Insufficient Legislation of New Medical Technologies

There is also a clear gap of established policies to govern the development of AI in healthcare, and if laws for the emerging technology remain underdeveloped, that will threaten anti-discriminatory patient protection. The Food and Drug Administration (FDA), released a framework for Software as Medical Devices (SaDMs)—which are commercial software available to the public and healthcare providers—that requires transparency and validity to meet qualifications. Yet this model is insufficient as AI in healthcare is still largely unregulated and has fewer obligations to comply with when compared to drug prescriptions (21, p. 224). FDA-approved chest X-ray diagnostic devices have already been proven to perform worse at different sites, and some even show disparity among African American and white patients (23, p. 583). None of the high-risk SaDMs were scrutinized under prospective studies nor had a control comparison, and most devices did not report multi-site testing (23, pp. 582–583). Which poses an urgent risk as single-site assessment only focuses on homologous data and can lead to vulnerabilities and bias (11). It has been supposed that the homologous data training could be a result of regional bias, or ‘vernacular medicine’, and most importantly it may not expand to broader communities (1, p. 3). Another effort the FDA has piloted is regulation for Clinical Decision Support software (CDS) which prohibits AI from making independent recommendations, thus preventing systematic discrimination. However, tools can circumvent the regulation by labeling as a “Non-Device CDS” and documenting directions and rationale in non-medical terminology (21, pp. 227–228).

There is opportunity to increase jurisdiction in administrative technology, including insurance and authorization AI use, which can contribute heavily to prejudices against patients. To cover discrimination caused by AI, like that of the UnitedHealth case, there needs to be non-discriminatory legislation. There is potential legal grounding in the Affordable Care Act Section 1557 that prohibits the providers and plan entity discrimination, specifically including decisions based on clinical algorithms (5). But legislation like Title VII has yet to be applied, thus grievances will continue to be inadequately addressed legally as there is no “constitutional property right to receive” or object of grievance (21, pp. 231–232).

Future Directions

Short Term: The healthcare professional must understand the limitations and inconsistencies that AI tools produce before they reference them, and temporary measures can also be taken to increase individual patient representation in genetic information. It has been shown that it is essential to have routine genetic information and socioeconomic status screening to address inaccessibility adequately (12). During an appointment for example, the patient can input genetic ancestry and socioeconomic status in their electronic health record, so AI and genetic tools can base their suggestions on that. There is also the importance of patient genetic transparency and literacy. When the patient receives a diagnosis or prescription via AI, providing comprehensive and clear reasoning behind the diagnosis is especially imperative. Therefore, the physician can mitigate negligence, and the patient can interact and actively correct misguided conclusions.

Immediate steps can also be taken to address the socioeconomic barriers to clinical research participation, such as genotype research. Research has shown that participation in studies is increased when the researcher shares a similar cultural background with them or can communicate in their language (15, p. 5). Providers can establish that by recruiting genetic counselors who share similar cultural backgrounds to the patients they are treating, thus creating a more inclusive social and community context. Fixing the stunted pipeline to specialty care for socioeconomic minorities can also be approached with travel support, food, and childcare support during a clinical study. There already exists government acknowledgment and monetary support for these types of effort such as the Revitalization Act which instructs that cost considerations cannot be the reason for the exclusion of minorities (15, p. 6).

Long Term: Diversifying and expanding studies to include African ancestry and other racial minorities both advances scientific understanding of diseases and mitigates bias in the data and AI models. Research shows that a more extended population history and greater genetic diversity of African populations could be helpful in exploring the holistic evolution of Fragile X syndrome and other Mendelian disorders, which have been well-documented yet not well-developed in European populations (17). One such endeavor is the “All of Us” Research Program, which targets enrolling a more diverse population, and it has been reported in a survey that more than 75% of participants are from underrepresented backgrounds in research and 45% are from racial and ethnic minorities (6). Furthermore, it is not just the gathering of genomic data that is important to collect, surveys also need to include socioeconomic factors like race, gender, and income. These all impact social determinants of health, and using that to create a more holistic picture of the patient leads to more robust training models. The survey collection process should be made flexible and free, or have other means of support such as offering subsidies for participants if the research requires equipping them with devices to collect biometric data. It is also necessary to remove language barriers, which can be implemented similarly to the “All of Us” Research Program which is available in both English and Spanish at approachable reading levels. And intentional recruitment at diverse centers that treat patients with different economic backgrounds is necessary to identify the participants that might not have easy access to education and healthcare.

Alongside proper data aggregation and collection, stricter protocols must enforce compliance and oversight of AI development. The FDA’s framework of SaDMs and other regulations need to require external site testing of the AI before validation, therefore exposing possible bias against underrepresented population groups. In contrast to the unspecified safety assurances of the FDA, an alternative is the European Union Medical Device Regulation which defines a more specific requirement for manufacturers to maintain post-market monitoring systems (22, p. 738). That showcases how measures that can be taken to supervise AI more directly. There is a great safety risk for a design without post-market surveillance and pre-market in-depth documentation of the methods behind AI algorithms. The fruition of more transparency fosters more trust and mitigates healthcare professionals and staff from turning a blind eye to the actual workings of the AI used in their work setting. Such practices of AI validation throughout its lifetime could be created based on the FDA’s Total Product Life Cycle (TPLC) model for traditional medical devices. One study proposes the addition of standard equity metrics to each of the technology’s phases, beginning in the conception phase to quantify bias related to historical access and intended use through population-achieved sensitivity (22, pp. 738–739). Another study suggests that during the design and development phase, the training data needs to be scrutinized, and it should be susceptible to retrospective use of older datasets (1, pp. 4–5). It also should include an investigation of the algorithms’ reference standard, which may consist of bias in predicting health outcomes. Finally, the access phase should require comprehensive monitoring that also considers bias in the reporting of the AI tool throughout its market use (1, p. 5). In monitoring AI in healthcare throughout all of its life cycle: training, modeling, and predicting, active and interpretable measures are taken to eliminate bias and prejudice in the technology.

References

1. Abràmoff, M. D., Tarver, M. E., Loyo‐Berríos, N., Trujillo, S., Char, D., Obermeyer, Z., Eydelman, M. B., & Maisel, W. H. (2023). Considerations for addressing bias in artificial intelligence for health equity. Npj Digital Medicine, 6(1). https://doi.org/10.1038/s41746-023-00913-9https://doi.org/10.1038/s41746-023-00913-9

2. AI ushers in next-gen prior authorization in healthcare. (2022, April 19). McKinsey & Company. https://www.mckinsey.com/industries/healthcare/our-insights/ai-ushers-in-next-gen-prior-authorization-in-healthcarehttps://www.mckinsey.com/industries/healthcare/our-insights/ai-ushers-in-next-gen-prior-authorization-in-healthcare

3. Alberdi, E., Povyakalo, A., Strigini, L., & Ayton, P. (2004). Effects of incorrect computer-aided detection (CAD) output on human decision-making in mammography. Academic Radiology, 11(8), 909–918. https://doi.org/10.1016/j.acra.2004.05.012https://doi.org/10.1016/j.acra.2004.05.012

4. Bentley, A. R., Callier, S., & Rotimi, C. N. (2017). Diversity and inclusion in genomic research: why the uneven progress? Journal of community genetics, 8(4), 255–266. https://doi.org/10.1007/s12687-017-0316-6https://doi.org/10.1007/s12687-017-0316-6

5. Cary, M. P., Zink, A., Wei, S., Olson, A., Yan, M., Senior, R., Bessias, S., Gadhoumi, K., Jean-Pierre, G., Wang, D., Ledbetter, L., Economou-Zavlanos, N. J., Obermeyer, Z., & Pencina, M. J. (2023). Mitigating racial and ethnic bias and Advancing health equity in Clinical Algorithms: A scoping review. Health Affairs, 42(10), 1359–1368. https://doi.org/10.1377/hlthaff.2023.00553https://doi.org/10.1377/hlthaff.2023.00553

6. Denny JC, Rutter JL, Goldstein DB, Philippakis A, Smoller JW, Jenkins G, Dishman E: The “All of Us” research program. The New England Journal of Medicine, 381(7), 668–76. https://doi.org/10.1056/nejmsr1809937https://doi.org/10.1056/nejmsr1809937

7. Drozda, K., Wong, S. S., Patel, S., Bress, A. P., Nutescu, E. A., Kittles, R. A., & Cavallari, L. H. (2015). Poor warfarin dose prediction with pharmacogenetic algorithms that exclude genotypes important for African Americans. Pharmacogenetics and Genomics, 25(2), 73–81. https://doi.org/10.1097/fpc.0000000000000108https://doi.org/10.1097/fpc.0000000000000108

8. Fatumo, S., Chikowore, T., Choudhury, A., Ayub, M., Martin, A. R., & Kuchenbaecker, K. (2022). A roadmap to increase diversity in genomic studies. Nature Medicine, 28(2), 243–250. https://doi.org/10.1038/s41591-021-01672-4https://doi.org/10.1038/s41591-021-01672-4

9. Herman, J. (2022, April 26). Racism, inequality, and health care for African Americans. The Century Foundation. https://tcf.org/content/report/racism-inequality-health-care-african-americans/https://tcf.org/content/report/racism-inequality-health-care-african-americans/

10. Holtzman, J. (2013, October 21). The Physician Shortage and the Future of the Affordable Care Act: The Coverage without Care Conundrum. Stanford Journal of Public Health. https://web.stanford.edu/group/sjph/cgi-bin/sjphsite/the-physician-shortage-and-the-future-of-the-affordable-care-act-the-coverage-without-care-conundrum/https://web.stanford.edu/group/sjph/cgi-bin/sjphsite/the-physician-shortage-and-the-future-of-the-affordable-care-act-the-coverage-without-care-conundrum/

11. Kaushal, A., Altman, R. B., & Langlotz, C. P. (2020). Geographic distribution of US cohorts used to train deep learning algorithms. JAMA, 324(12), 1212. https://doi.org/10.1001/jama.2020.12067https://doi.org/10.1001/jama.2020.12067

12. Landry, L., Ali, N., Williams, D. R., Rehm, H. L., & Bonham, V. L. (2018). Lack of diversity in genomic databases is a barrier to translating precision medicine research into practice. Health Affairs, 37(5), 780–785. https://doi.org/10.1377/hlthaff.2017.1595https://doi.org/10.1377/hlthaff.2017.1595

13. Napolitano, E. (2023, November 21). UnitedHealth uses faulty AI to deny elderly patients medically necessary coverage, lawsuit claims. CBS News. https://www.cbsnews.com/news/unitedhealth-lawsuit-ai-deny-claims-medicare-advantage-health-insurance-denials/https://www.cbsnews.com/news/unitedhealth-lawsuit-ai-deny-claims-medicare-advantage-health-insurance-denials/

14. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019b). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342https://doi.org/10.1126/science.aax2342

15. Oh, S. S., Galanter, J., Thakur, N., Pino‐Yanes, M., Barceló, N. E., White, M. J., De Bruin, D. M., Greenblatt, R. M., Bibbins‐Domingo, K., Wu, A. H., Borrell, L. N., Gunter, C., Powe, N. R., & Burchard, E. G. (2015a). Diversity in clinical and biomedical research: a promise yet to be fulfilled. PLOS Medicine, 12(12), e1001918. https://doi.org/10.1371/journal.pmed.1001918https://doi.org/10.1371/journal.pmed.1001918

16. Omiye, J. A., Lester, J., Spichak, S., Rotemberg, V., & Daneshjou, R. (2023). Large language models propagate race-based medicine. Npj Digital Medicine, 6(1). https://doi.org/10.1038/s41746-023-00939-zhttps://doi.org/10.1038/s41746-023-00939-z

17. Peprah, E., Xu, H., Tekola‐Ayele, F., & Royal, C. (2014). Genome-Wide Association Studies in Africans and African Americans: Expanding the framework of the genomics of human traits and disease. Public Health Genomics, 18(1), 40–51. https://doi.org/10.1159/000367962https://doi.org/10.1159/000367962

18. Rajpurkar, P. (2017, November 14). CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning. arXiv.org. https://arxiv.org/abs/1711.05225https://arxiv.org/abs/1711.05225

19. Seyyed-Kalantari, L., Zhang, H., McDermott, M. B. A., Chen, I. Y., & Ghassemi, M. (2021). Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nature Medicine, 27(12), 2176–2182. https://doi.org/10.1038/s41591-021-01595-0https://doi.org/10.1038/s41591-021-01595-0

20. Sirugo, G., Williams, S. M., & Tishkoff, S. A. (2019). The Missing Diversity in Human Genetic Studies. Cell, 177(1), 26–31. https://doi.org/10.1016/j.cell.2019.02.048https://doi.org/10.1016/j.cell.2019.02.048

21. Takshi S. (2021). Unexpected Inequality: Disparate-Impact From Artificial Intelligence in Healthcare Decisions. Journal of Law and Health, 34(2), 215–251.

22. Vokinger, K. N., & Gasser, U. (2021). Regulating AI in medicine in the United States and Europe. Nature Machine Intelligence, 3(9), 738–739. https://doi.org/10.1038/s42256-021-00386-zhttps://doi.org/10.1038/s42256-021-00386-z

23. Wu, E. H., Wu, K., Daneshjou, R., Ouyang, D., Ho, D. E., & Zou, J. (2021a). How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals. Nature Medicine, 27(4), 582–584. https://doi.org/10.1038/s41591-021-01312-xhttps://doi.org/10.1038/s41591-021-01312-x