1. The Meaning of ‘Racist’
1.1. Conceptual Inflation
We live in an era of inflation. According to Ipsos’s “What Worries the World” survey from March 2024, inflation is the “number one global concern overall” (Ipsos 2024). Some theorists and pundits worry that in addition to the devaluation of our currency, some of our most important concepts are also undergoing a worrisome process of inflation. In 1989, sociologist Robert Miles used the expression conceptual inflation to critique the process whereby the concept of racism “has been redefined to refer to a wider range of phenomena” (Miles 1989/2003: 58). A decade later, Miles’s conceptual inflation critique of ‘racism’ was taken up by philosophers. In 1997, J. L. A. Garcia endorsed Miles’s critique and connected it to concerns in popular political commentary: “the editors of the neoconservative opinion journal First Things … warn that if we do not tighten up our usage, then they will consign it to the same dustbin to which they have already consigned the terms ‘sexism’ and ‘homophobia’” (Garcia 1997: 5). In 2002, Lawrence Blum echoed and elaborated on Miles’s critique:
To name an act or a person “racism” or “racist” is particularly severe condemnation. But the terms are in danger of losing their moral force, for they have been subject to conceptual inflation … thus inhibiting honest interracial exchange. (Blum 2002: 31)
In 2019, in the heterodox online magazine Quillette, Spencer Case made the same critique as Blum, but with ‘racism’ a target alongside ‘sexism’ and ‘colonialism’:
When speakers expand the reference of a word in order to attach its associations to new things, they dilute the associations of the original word. Just as printing too much paper currency diminishes the value of the currency, concept inflation degrades the rhetorical effect of inflated words and phrases.
In 2022, linguist John McWhorter expressed concern about “the lexical mission creep which has led to how confusingly we use the word ‘racism’ now” in The New York Times, and Elizabeth Anderson affirmed a 2018 editorial in the Pittsburgh Post-Gazette that (in an echo of the First Things editorial from 1997) complained “about the excessively wide scope of the term ‘racist’: ‘For if every person who speaks inelegantly, or from a position of privilege, or ignorance, or expresses an idea we dislike, or happens to be a white male, is a racist, the term is devoid of meaning’.” (Anderson 2022: 77, quoting Burris 2018).
Anderson took the concern expressed by the editorial seriously and argued that, given its current meaning, ‘racist’ should be reserved for referring to only the “most hateful and vile individuals,” and not be used at all to communicate concerns about problematic racial phenomena (2022: 82; see also Anderson 2010: 48 and Blum 2004: 77). The excessively wide scope of the term ‘racist’—that is, when it is applied beyond the most hateful and vile individuals—is problematic, according to Anderson, because “outside of academic and progressive circles, almost no one talks like this” and “such wide scope … often leaves addressees in the dark about what is wrong” (2022: 82).
The charge of conceptual inflation has also been levied against a wide range of terms besides ‘racist’ and ‘racism,’ including “human rights” and “rule of law,” “capitalism” and “socialism,” and even “antiracism” (see Liao and Hansen forthcoming: §1). There is also a related expression of concept creep, coined by psychologist Nick Haslam, that has been used to critique “psychology’s expansionary redefinition of negative phenomena [that] arguably reflects a liberal social agenda,” which targets terms such as ‘abuse,’ ‘bullying,’ ‘trauma,’ ‘mental disorder,’ ‘addiction,’ and ‘prejudice’ (Haslam 2016: 1). However, as the brief history above illustrates, ‘racism’ and ‘racist’ remain prominent targets of conceptual inflation critiques across the political spectrum.
The conceptual inflation critique makes a normative judgment (inflation is bad) on the back of a descriptive claim, that the meaning of a term in ordinary language has inflated over time. Critics of conceptual inflation have different reasons for thinking that it is bad, and there are advocates of conceptual change who hold that certain kinds of inflation are good. But our goal in this essay is not to tarry with the normative; we are interested in trying to find evidence that supports the descriptive claim that conceptual inflation is in fact happening. After all, if it turns out that there is no evidence that ‘racist’ is undergoing inflation, then the debate over whether and why inflation is bad or good loses its urgency. The first step in gathering evidence of whether ‘racist’ is undergoing inflation is to be more precise about what conceptual inflation itself is.
1.2. Operationalizing ‘Conceptual Inflation’
The descriptive component of the conceptual inflation critique can be precisified into three claims that are in principle empirically testable. First, as made explicit by Miles (“refer to a wider range”) and Case (“expand the reference”), an important part of the use of ‘racism’ and ‘racist’ is what these terms refer to, and there is a concern that the extensions of these terms have expanded over time. Second, as made explicit by Blum (“losing their moral force”) and Case (“degrades the rhetorical effect”), an important part of the use of ‘racism’ and ‘racist’ is the evaluative force they have when they are used in moral criticism, and there is a concern that the moral intensity of ‘racism’ and ‘racist’ as terms of criticism has decreased over time. Let us name these two claims:
| extension expansion | The reference class of a term in ordinary language has increased. |
| intensity bleaching | The moral force of a term in ordinary language has decreased. |
Moreover, these two aspects of meaning change seem to be connected. Both Blum and Case explicitly contend that it is extension expansion that causes intensity bleaching, but other critics are not as explicit. To explicate this alleged relationship between the two aspects of a term’s meaning, but without the causal commitment, let us name this third claim:
| inverse correlation | When the reference class of a term in ordinary language increases, the moral force of that term decreases. |
On our understanding, these three claims—extension expansion, intensity bleaching, and inverse correlation—are central to the concerns about conceptual inflation that have been raised over the last three decades, even if it is not the case that every single critic explicitly endorses all three claims. The aim of our studies is to assess the veracity of the descriptive component of the conceptual inflation critique, as captured by these three claims. An advantage of operationalizing in this way is that it also allows us to assess other ways of understanding conceptual inflation. For example, there is a neighboring way of understanding conceptual inflation that one might endorse, namely the idea that what constitutes conceptual inflation is the combination of extension expansion with no intensity bleaching while making no claim about the causal relationship between extension and intensity. We will discuss this neighboring way of understanding conceptual inflation in relation to our empirical findings in §5.3.
To put the descriptive question in our precisified terms: Has ‘racist’ undergone extension expansion and intensity bleaching, as the critics have been claiming over the last three decades? Economists do not track currency inflation with mere vibes; they use measurements such as the consumer price index. For example, using CPI, it is observable that there has been considerable US dollar inflation over the last three decades or so: $1.00 in 1989 had the purchasing power of $2.54 in 2024 (U.S. Bureau of Labor Statistics 2024). Has ‘racist’ undergone a similar transformation, as critics of conceptual inflation contend? How could that be measured? Before we plunge into the normative debate, we should gather systematic empirical evidence about whether the purported conceptual inflation of ‘racist’ has in fact been happening.
To be clear, there are two dimensions to critics’ worries about extension expansion, and in this paper we will only gather systematic empirical evidence about one of them. A critic might be concerned that the reference of an expression is being vertically expanded, insofar as it is being applied to more instances of a category. Our empirical investigation focuses on this dimension, on ‘racist’ as applied to persons. However, a critic might also be concerned about horizontal extension expansion, insofar as it is being applied to more categories. (See Tse & Haslam 2024 for the vertical/horizontal distinction used to measure “concept creep” in psychological concepts.) For example, Miles (1989/2003: 66–71) objects to applying the concept of racism to institutions and actions rather than individuals and beliefs. He cites the 1967 UNESCO statement on race as the root of how “the concept of racism was expanded in meaning to include not only beliefs but, more importantly, all actions, individual and institutional, which had the consequence of sustaining or increasing the subordination of ‘black’ people” (Miles 1989/2003: 67). In making this objection, Miles exemplifies a common narrative nowadays that ‘racism’ and ‘racist’ are historically individualist, and the institutionalist sense is a recent invention.
While our present empirical investigation is not on the horizontal dimension of conceptual inflation, we do want to note at the outset that Miles’s narrative is itself a recent invention. The earliest use of ‘racism’ recorded in the Oxford English Dictionary, in 1902, reads to us as more institutionalist than individualist: “[s]egregating any class or race of people apart from the rest of the people kills the progress of the segregated people or makes their growth very slow. Association of races and classes is necessary to destroy racism and classism” (quoted in Ichikawa ms). Similarly, in the Corpus of Historical American English (COHA; Davies 2010), which contains more than 475 million words of text from the 1820s–2010s, the first clear use of ‘racist’ as an institutionalist adjective (“black men were tried in racist courts”; 1942) appears earlier than the first clear use of ‘racist’ as an individualist adjective (“You racist little prick”; 1948). In COHA, there are at least as many, if not more, instances of using ‘racist’ to talk about institutions as there are instances of using ‘racist’ to talk about individuals before 1967, such as: “his racist goverriment [sic]” (1949); “the Army does not intend to, abolish its racist quota system or its segregation” (1950); “… provisions of the immigration laws of the United States because they would introduce racist concepts” (1956); “South Africa’s racist apartheid policy” (1961); “entirely peaceful demonstration by Moslem Algerians in Paris who were protesting a racist curfew” (1961). We don’t take this brief survey of corpus evidence to establish that ‘racist’ has not undergone horizontal extension expansion, but we think it should shift the burden of proof onto those theorists who think that it has (for more discussion of the lack of evidence of horizontal extension expansion, see Liao and Hansen [forthcoming: §4]). As we will explain in §5.2, we think that worries about horizontal extension expansion, like worries about conceptual inflation more generally, are examples of a linguistic recency illusion (Zwicky 2005; van der Meulen 2022), whereby people judge that a phenomenon is a more recent development than it in fact is.
1.3. Measuring Change
One form of systematic empirical evidence that might be used to provide a factual foundation for worries about conceptual inflation comes from increased frequency of the use of ‘racist’ over time (see Figure 1).
Figure 1: Relative frequency in academic abstracts (solid line) and news media (broken line) of expressions that are claimed to be undergoing conceptual inflation, based on the Semantic Scholar Open Research Corpus and selected international newspapers (Rozado 2022: fig. 5).
It is true that, in academia and media, the words ‘racism’ and ‘racist’—alongside ‘sexism’ and ‘sexist’ and ‘homophobia’ and ‘homophobic’—are used more frequently now than they were three decades ago (Rozado 2022). And increased frequency of use can be a cause of semantic bleaching (Sweetser 1988), the process by which expressions come to lose specific features of their meaning: “The mechanism behind bleaching is habituation: a stimulus loses its impact if it occurs very frequently” (Moder 2007: 339). However, increased frequency does not itself entail semantic bleaching or extension expansion. There are many possible reasons for a term’s increased frequency of use that are consistent with unchanged meaning: for example, increased frequency could be the result of increased uptake of the expression in the community, changing pragmatic norms about the conversational appropriateness of using the expression, or increases in the prevalence of the phenomenon itself that the expression is used to describe. Evidence of changing frequency of use is therefore only one piece of the puzzle of showing a term’s changing extension or force.
Another form of systematic empirical evidence that might be used to support the conceptual inflation critique comes from a closer inspection of linguistic corpora that goes beyond mere frequency of use. It is now commonplace for linguists, philosophers, and legal scholars to use linguistic corpora to investigate the meaning of linguistic expressions (for example: Vetter 2014; Andow 2015; Lee & Mouritsen 2017; Goldfarb 2017; Pinillos & Nichols 2018; Sytsma et al. 2019; Hansen et al. 2021; Tobia 2020; Zahorec et al. 2023). In the debate about the meaning of ‘racist,’ we used contemporary corpora to respond to the worry that ‘expressions condemning oppression’ (the adjectives ‘racist,’ ‘sexist’ and ‘homophobic’) have been inflated to the point that they could no longer express nuanced moral condemnation (Liao and Hansen 2023). But in that study we could not directly address questions about whether these expressions had changed meaning over time because the available historical corpora, such as COHA (Davies 2010), don’t have enough examples of ‘racist’ to draw firm conclusions. While linguistic corpora are in principle a good form of systematic empirical evidence for resolving the controversy about the meaning change of ‘racism’ and ‘racist,’ existing corpora don’t provide solid evidence either way.
Finally, another form of systematic empirical evidence that might be used to support the conceptual inflation critique comes from surveying people about their uses of ‘racist.’ In an ideal world, we would conduct a real time longitudinal study that tracked the way a group of speakers over time or different samples from a population at different times use ‘racist’ and see if it shows signs of meaning change. Unfortunately, it is too late for us now to start asking people questions about their use of ‘racist’ three decades ago. The General Social Survey, which is normally an excellent resource for studying change over time, only has one question in the vicinity of our interest—“What are your personal feelings about people who believe whites are racially superior to all other races?”—and it was only asked once, in 1990 (Davern et al. 1972–2024). But we can adopt a technique used by sociolinguists to study linguistic change that relies on a proxy for real change through time: the apparent time construct.
Apparent time contrasts with real time. In a real time study, the linguistic behavior of a cohort of speakers is measured at different stages of their lives, or samples of different speakers from a population are taken over time and compared. In an apparent time study, the linguistic behavior of speakers of different ages is sampled at a single time, on the assumption that “linguistic differences among different generations of a population (apparent-time differences) would mirror actual diachronic developments in the language (real-time linguistic changes)” (Bailey 2004: 313). This assumption depends on the idea that post-adolescence, speakers’ “vernaculars remain stable throughout the course of an adult lifetime” (Bailey 2004: 320). Studying the way older cohorts use language is a window through which we can view historical patterns of use, namely those that were prevalent when the cohort’s linguistic habits were being formed. By looking through that window at the current linguistic behavior of speakers of different ages, we get a view of linguistic variation over time.
The apparent time approach to studying linguistic change is widespread in sociolinguistics. It was used in some of the earliest studies of variation, like Louis Gauchat’s (1905) study of sound change in a Swiss village, and William Labov’s (1963) study of phonetic variation among residents on Martha’s Vineyard. And there is evidence, from real-time studies of linguistic behavior, supporting the idea that post-adolescent vernaculars are stable (see Bailey 2004: §2.2 for a survey of the evidence).
One use of the apparent time approach to study lexical change comes from sociolinguists Sali Tagliamonte and Katharina Pabst’s (2020) study of adjectives of “highly positive evaluation”—expressions like ‘amazing,’ ‘cool,’ ‘brilliant,’ and ‘wonderful.’ She shows that the use of these expressions varies by the age of speakers in Toronto and York (U.K.). For example, ‘cool’ is a common adjective in Toronto among younger speakers, but not older speakers, and ‘brilliant’ has a similar pattern of use among younger and older British speakers.
In Figure 2, we can see the rise of ‘cool’ to become the preferred adjective of highly positive evaluation among younger speakers in Toronto, and the relative decline of ‘great’ and ‘wonderful.’ Of course, younger speakers can choose to use ‘wonderful’ as a term of highly positive evaluation—though in so doing they would sound old-fashioned, just as older speakers can say that things are ‘cool’ if they want to sound more hip—or cringe (Tagliamonte & Pabst 2020: 24). ‘Cool’ itself may sound dated to today’s young people, since the data Tagliamonte and Pabst are relying on only goes up to 1999. At the risk of sounding hopelessly middle-aged, kids these days may prefer to express highly positive evaluation by using ‘dope,’ ‘fire,’ or, as a young guest at a pop-up restaurant recently said to one of the authors, ‘gas.’
Figure 2: Distribution of Adjectives of Highly Positive Evaluation by Decade of Birth (in Toronto) (Tagliamonte & Pabst 2020: fig. 2).
But—we want to be clear up front before applying the apparent time method in this study—the method has not frequently been used to study semantic change. One exception is Jean-Phillipe Magué (2006), who conducted an apparent time study of the French expression ‘maison’ using a semantic field approach that asked participants to judge the semantic similarity of various words to ‘maison.’ Magué did find significant differences between the younger (median age 21) and older (median age 56) cohorts he recruited in the semantic similarity rating for certain expressions (such as ‘chateau’ and ‘immeuble’), and he shows “a correlation between this semantic variation and the age of speakers which is, under the apparent time hypothesis, the synchronic manifestation of a change in progress” (2006: 233). But Magué expresses caution about how much weight to put on this finding, because he points out that unlike with studies of sound change, the “apparent time hypothesis has never been verified for semantics, and thus we cannot exclude that speakers modify their semantic structure of the semantic field as they get older.” That kind of diachronic linguistic change is known as “age-grading,” where older speakers use an expression differently than younger speakers not because community use has changed, but because “generational differences…repeat themselves from one decade to the next, and teenagers end up speaking just like their parents” (Boberg 2004: 257)
An example of an age-graded change is the greater prevalence of the pronunciation of the letter “z” as “zee,” rather than “zed,” among young people in Southern Ontario, a difference which does not persist as speakers in that region age. That change in pronunciation has been explained as an effect of children learning to pronounce the letter “z” as “zee” from hearing the “alphabet song” on the American children’s TV show Sesame Street. As the Canadian speakers age, they drop the American pronunciation in favor of the standard “zed” (Chambers 2002: ch. 4; though see Boberg 2004 for conflicting evidence from Montreal English). For our purposes, distinguishing generational change in the use of a term like ‘racist,’ which would be evidence of dynamic conceptual inflation, from an age-graded effect, which would only indicate that older speakers use the expression differently than younger speakers while overall use is stable, will only be important if there are in fact differences in use between younger and older cohorts.
Another interesting possible pattern of linguistic change is discussed by Charles Boberg (2004), namely generational change plus individual change among older speakers: “it is not impossible for older people to adapt their speech to the new patterns they hear around them” (252). For example, among speakers of English in Montreal, Boberg finds evidence that not only has the term “Chesterfield” fallen completely out of favor among 1999 teenagers as a term used to refer to “an upholstered piece of furniture that seats three people in a row,” in contrast to 1999 “grandparents,” 30% of whom still prefer it to “couch.” This might look like a very clear apparent time effect, but using a real-time study comparing Canadian speakers from 1972 and 1999, Boberg shows that there is a massive drop-off in the use of “Chesterfield” among 1999 “parents” in comparison with their 1972 teenage counterparts, from around 60% to around 15%. That change indicates that the rapid decline in popularity of “Chesterfield” is not just a generational difference, but is accelerated by individual changes among the “parent” cohort. Determining whether linguistic change is purely the result of generational change, purely an age graded effect, or a combination of factors would require combining real and apparent time methods: “Evidence of late adoption [of lexical changes], in fact, can only emerge from studies that marshal both kinds of data” (Boberg 2004: 266). One way of combining real and apparent time methods for measuring meaning change is to run an apparent time study on a word for which there is consensus that its meaning has indeed changed. Justyna Robinson (2012) does this with the word “gay,” showing that the change in the word’s meaning shows up in differences in the way older and younger speakers apply the term and how they explain its meaning. Older speakers persist in using the word with its “happy” meaning, while younger speakers do not use it that way, using it to mean “homosexual,” and the youngest cohort additionally use it as a “general term of disapproval” meaning something like “lame” (Robinson 2012: 47).
Unfortunately, the previous generation of researchers didn’t experiment on the meaning of terms like ‘racist’ that have been accused of undergoing conceptual inflation, and unlike ‘gay,’ ‘racist’ is not an uncontroversial case of a term undergoing change of meaning, so combining a real time study with an apparent time study of ‘racist’ is not yet possible. But we can look for prima facie evidence of linguistic change using the apparent time method and see if anything turns up, and thereby lay the foundation for future real time studies that can begin to answer these questions with more confidence.
1.4 Overview of Studies
We conducted two preregistered studies and one non-preregistered replication that aimed to evaluate each of the descriptive components of the conceptual inflation critique:
| extension expansion | Has the extension of ‘racist’ expanded? |
| intensity bleaching | Has the intensity of ‘racist’ decreased? |
| inverse correlation | Is there an inverse correlation between the extension and intensity of ‘racist’? |
In all three studies, we used an apparent time technique, recruiting participants of varying ages so that we could evaluate whether different age cohorts interpret ‘racist’ as having different extensions and different intensities.
In Study 1, we also (i) compared ‘racist’ with thin moral terms (‘disagreeable,’ ‘terrible,’ and ‘worthless’) in order to get a non-politically contested benchmark against which to compare the extension and intensity of ‘racist’; (ii) we examined what happened to the extension and intensity of ‘racist’ when it was combined with degree modifiers like ‘slightly,’ ‘moderately,’ and ‘extremely,’ in order to evaluate the claim we made in Liao and Hansen (2023) that these modifiers allow uses of ‘racist’ to express nuanced moral condemnations; and (iii) we evaluated proposals made by critics of conceptual inflation to replace ‘racist’ with expressions that are alleged to provide better expressive resources for criticizing racial ills.
In Study 2 and Study 3, we compared ‘racist’ with ‘queer,’ a “reclaimed” slur that is generally regarded to have undergone meaning change. We also exploratorily examined other terms that we suspected would be judged differently according to participants’ age: ‘slutty’ and ‘neurodivergent.’
To summarize our main findings: we did not find any evidence that ‘racist’ has undergone intensity bleaching in any of our three studies. Moreover, in Study 1, we found that the intensity of ‘racist’ is greater than ‘terrible,’ and on par with ‘worthless,’ one of the most intense thin moral terms of criticism. That is evidence that ‘racist’ is still regarded as an extreme form of criticism. We found some evidence in Study 2 (but not Study 1 or Study 3) that ‘racist’ has undergone some extension expansion: the younger cohort of participants apply ‘racist’ to a larger group of people than the older cohort does. In Study 2 and Study 3, we found that ‘queer’ has undergone both extension expansion and intensity bleaching, but we found evidence that ‘racist’ has undergone extension expansion only in Study 2; and that the changes that ‘queer’ has undergone are significantly greater than those that ‘racist’ has undergone. Finally, we did not find any evidence, in any of our three studies, that the extension and intensity of ‘racist’ are correlated. This pattern of findings should be surprising to anyone who believes that ‘racist’ is undergoing conceptual inflation.
2. Study 1: Demographically Representative Sample
In our first study, we investigated the meaning of ‘racist’ with a sample that is demographically representative of the overall population of the USA with respect to age, race, and gender. To understand the current meaning of ‘racist,’ we compare it against a range of thin moral terms (§2.2) and against a range of modified forms. To track the meaning change of ‘racist’ with the apparent time construct, we examine its perceived extension and intensity across participants of different ages, keeping in mind a potential difference between participants of different races (§2.4).
2.1. Methods
In a preregistered study, 417 U.S.-based participants were recruited using Prolific’s “representative sample” function, which aims to match the sample to the demographic distribution of the overall U.S. population with respect to age, race, and gender. (All material, data, and analysis can be found at https://doi.org/10.17605/OSF.IO/GY6CN.) The median time to complete the study was 3 minutes and 4 seconds, and participants were paid $1 on completion. A final sample of 240 participants (Mage = 45.23, SDage = 16.08; 46.67% men, 49.58% women, 0.42% non-binary; 72.92% white, 25.00% nonwhite) was obtained after the preregistered exclusion criteria were applied. (The pattern of statistical inferences is the same with and without exclusions, with one exception to be noted. All statistical tests are confirmatory—that is, based on the preregistration—rather than exploratory, unless explicitly noted otherwise.) No other demographic information was collected.
Participants completed the study in Qualtrics and provided informed consent before proceeding. There was no experimental manipulation and all participants were asked all questions. To start, participants were asked, in counterbalanced randomized order, the following two questions about the term ‘racist’ on a 0–100 slider with text anchors, each on as an individual block on its own page (Figure 3):
| [extension] | What percentage of people can be reasonably called ‘racist’? [none – all] |
| [intensity] | How bad is it for a person to be called ‘racist’? [not at all – the worst] |
Then, participants were asked, in counterbalanced randomized order, six blocks of questions. Within each block, the questions were presented in randomized order. The six blocks involved the two types of questions shown above—extension and intensity—with, respectively, three sets of three terms substituted in the place of ‘racist’:
| thin moral terms | ‘disagreeable’ / ‘terrible’ / ‘worthless’ |
| ‘racist’ with degree modifiers | ‘slightly racist’ / ‘moderately racist’ / ‘extremely racist’ |
| alternative vocabulary terms | ‘racially ignorant’ / ‘racially insensitive’ / ‘racially unjust’ |
The thin moral terms were drawn from an earlier study on the valence of English words (Warriner et al. 2013) to include ones that have been rated higher than, about the same as, and lower than ‘racist.’ The degree modifiers were selected to provide a range of expressions.
The alternative vocabulary items we investigated come from linguistic recommendations made by Blum (2002b: 209), in light of his concern about conceptual inflation:
To help us avoid the first form of confusion about racism—conceptual inflation—I will suggest a core meaning rooted in the history of its use, that confines “racism” to phenomena deserving of the severest moral condemnation …. Fixing on such a definition should encourage us to make use of the considerable other resources our language affords us for describing and evaluating race-related ills that do not characteristically rise to the level of racism—racial insensitivity, racial conflict, racial injustice, racial ignorance, racial discomfort, and others.
While Blum holds that these expressions should just be used to supplement our use of ‘racist,’ Anderson (2022: 82) advocates replacement: “it would be wise … to use other terms to describe moral problems concerning race.” In addition to giving us words to express weaker forms of moral condemnation, these theorists also hope that this vocabulary can improve communicative clarity.
Finally, participants were asked demographic questions about age, race, and gender following a recommendation on inclusive practices for collecting demographic questions (Hughes et al 2016).
2.2. ‘Racist’
What percentage of people can be reasonably called ‘racist’? Participants gave a wide range of responses to this question (M = 32.33, SD = 21.44). The response distribution is significantly different from a normal distribution, but still approximately normal (W = 0.941, p < 0.001). ‘Racist’ is applied neither exclusively nor indiscriminately. On average, participants neither reserved the term for the extremely narrow group of “neo-Nazis who consciously endorse particularly hateful beliefs and attitudes toward members of a racial group,” as Anderson claims about popular moral discourse, nor applied it to the extremely wide group of “every person who speaks inelegantly, or from a position of privilege, or ignorance, or expresses an idea we dislike, or happens to be a white male,” as Pittsburgh Post-Gazette catastrophized about conceptual inflation.
How bad is it for a person to be called ‘racist’? Again, participants gave a wide range of responses to this question (M = 72.71, SD = 25.87). The response distribution is significantly different from a normal distribution, but still approximately normal (W = 0.862, p < 0.001). ‘Racist’ still possesses a strong moral force. On average, participants interpreted it as a “severe condemnation,” as Blum claimed.
While there is some value to these initial findings, the numbers themselves remain difficult to interpret in isolation. For example, one might ask with respect to the extension question: Do our participants, on average, really think that 32.33% of people can be reasonably called ‘racist’? This difficulty is especially salient since extant research on survey methods shows that people are, to say the least, not great at estimating the size of groups. For example, a 2022 YouGov study showed that Americans systematically overestimate the size of minority groups—for example, estimating that 21% of the population is transgender when the reality is 1%—and systematically underestimate the size of majority groups—for example, estimating that 58% of the population are Christian when the reality is 70% (Orth 2022). To be clear, this systematic unreliability with estimation is not only a problem for us, but also a problem—perhaps an even bigger problem—for anyone who eschews empirical methods and makes claims about the meaning of terms based on their own impressions of how they are used. To avoid putting too much argumentative weight on these numbers in isolation, we want to interpret them against the backdrop of other terms.
2.3. ‘Racist’ versus Thin Moral Terms
How does ‘racist’ compare with other terms of moral condemnation? We can address this question by drawing on extant research on the valence of English expressions. Amy Beth Warriner and colleagues (2013) had participants rate the emotional valence of 13,915 English expressions on a scale from 9 (completely happy, pleased, satisfied, contented, hopeful) to 1 (unhappy, annoyed, unsatisfied, melancholic, despaired, or bored). For example, the top-rated expression is ‘vacation’ (mean valence of 8.53) and the lowest rated expression is ‘pedophile’ (mean valence of 1.26). From this large list of expressions, we chose three thin moral terms from the negative end of the valence scale: ‘disagreeable’ (3.29), ‘terrible’ (2.1), and ‘worthless’ (1.89). Our aim is to use these terms as benchmarks with which to compare the meaning of ‘racist.’
Since the distribution of responses to the extension question about the term ‘racist’ violates the normality assumption of our original planned parametric statistical tests, we report results from applying the nonparametric Friedman test to compare across ‘racist’ and the set of thin moral terms, and Nemenyi test for pairwise comparisons between all terms. These are non-parametric equivalents of much more familiar ANOVA and planned comparison t-tests. The pattern of statistical inferences remains the same with parametric and nonparametric tests, since linear regression is quite robust to normality violation (Schmidt & Finan 2018).
Comparing their perceived extensions, there is a statistically significant difference across ‘racist’ and the set of thin moral terms (χ2(3) = 333.64, p < 0.001). With pairwise comparisons, there are statistically significant differences between all pairs of terms. As the graph shows (Figure 4), the perceived extension of ‘racist’ is narrower than that of ‘disagreeable’ (M = 40.98, SD = 20.31), and wider than that of ‘terrible’ (M = 22.06, SD = 17.83) and ‘worthless’ (M = 12.69, SD = 16.89).
Figure 4: Comparing the extension of ‘racist’ with thin moral terms. (These graphs have three layers of information. First, the colored dots are in a sinaplot, which represents the individual data points in a way that shows their density and distribution. Second, the boxplot shows the quartile cutoffs and the bold horizontal line indicates the median. Third, the diamond indicates the mean.)
There are a wide range of views about the extension of ‘racist,’ but the term is not meaningless. The statistical results confirm our initial impression: In ordinary language, ‘racist’ neither refers exclusively nor indiscriminately. Our participants thought there are fewer people reasonably called ‘racist’ than those reasonably called ‘disagreeable.’ However, our participants also thought that there are more people reasonably called ‘racist’ than those reasonably called ‘terrible’ and ‘worthless’: that is, the label of ‘racist’ is not reserved only for the most hateful and vile individuals.
In terms of perceived intensity, there is a statistically significant difference across ‘racist’ and the set of thin moral terms (χ2(3) = 374.93, p < 0.001). With pairwise comparisons, there are statistically significant differences between all pairs of terms, except ‘racist’ and ‘worthless’ (p = 0.96). As planned, we conducted a follow-up equivalence test to interpret this statistically nonsignificant result (Lakens, Scheel, & Isager 2018). Using the conventional “small” effect as our equivalence bounds (d = 0.2; Cohen 1988), there is no difference between the perceived intensity of ‘racist’ and ‘worthless’ (t(239) = –1.7, p = 0.04). As the graph shows (Figure 5), the perceived intensity of ‘racist’ is stronger than that of ‘disagreeable’ (M = 29.93, SD = 22.52) and ‘terrible’ (M = 55.45, SD = 27.76), and about the same as ‘worthless’ (M = 69.97, SD = 27.47).
As with the perceived extension of ‘racist,’ there are a wide range of views on the perceived intensity of ‘racist.’ But, again, that does not mean the term is meaningless. The thin moral terms also have a wide range of perceived intensity. And the statistical results confirm that in ordinary language, ‘racist’ expresses very intense moral criticism. Our participants thought of it as on par with ‘worthless’—which, to our ears (and according to Warriner et al. 2013) is one of the worst ‘thin’ moral terms that a human being can be called. Our participants also thought ‘racist’ is a worse thing to be called than ‘terrible’ or ‘disagreeable.’ Being called ‘racist’ remains a particularly severe moral condemnation.
2.4. ‘Racist’ with Degree Modifiers versus Alternative Vocabulary Terms
As mentioned, there is some affirmative evidence from corpora, where examples of speakers modifying ‘racist’ to weaken (‘slightly racist’) or strengthen (‘extremely racist’) its moral force can be found (Liao & Hansen 2023). Our approach allows us to directly investigate how such degree modifiers change the perceived extension and intensity of ‘racist.’
Comparing their perceived extensions, there is a statistically significant difference across bare ‘racist’ and the set of its modified forms (χ2(3) = 236.02, p < 0.001). With pairwise comparisons, there are statistically significant differences between all pairs of terms, except ‘racist’ and ‘moderately racist’ (p = 0.344). As planned, we conducted a follow-up equivalence test to interpret this statistically nonsignificant result: using the conventional “small” effect as our equivalence bounds (d = 0.2), we cannot conclude that there is no difference between ‘racist’ and ‘moderately racist’ (t(239) = –1.1, p = 0.13). As the graph shows (Figure 6), the perceived extension of ‘racist’ is narrower than that of ‘slightly racist’ (M = 39.58, SD = 25.09), in the vicinity of ‘moderately racist’ (M = 30.23, SD = 20.08), and wider than ‘extremely racist’ (M = 19.48, SD = 20.72).
Comparing their perceived intensity, there is a statistically significant difference across bare ‘racist’ and the set of its modified forms (χ2(3) = 453.5, p < 0.001). With pairwise comparisons, there are statistically significant differences between all pairs of terms. As the graph shows (Figure 7), the perceived intensity of ‘racist’ is stronger than that of ‘slightly racist’ (M = 51.61, SD = 29.97) and ‘moderately racist’ (M = 64.88, SD = 26.27), and weaker than that of ‘extremely racist’ (M = 81.44, SD = 24.00).
The statistical results confirm our observations in Liao and Hansen (2023) based on linguistic corpora that there are resources available in ordinary language that enable ‘racist’ to express a range of moral condemnation. Degree modifiers can indeed strengthen and weaken the moral force of ‘racist’: Our participants thought that ‘slightly racist’ is weaker than bare ‘racist’ and that ‘extremely racist’ is stronger than bare ‘racist.’ Moreover, degree modifiers can also narrow and widen the reference class of ‘racist’: Our participants thought that ‘slightly racist’ picks out more people than bare ‘racist’ and that ‘extremely racist’ picks out fewer people than bare ‘racist.’ Given these patterns, we can also conclude that degree modifiers expand the expressive range of ‘racist’ both upwards and downwards. Ordinary language enables speakers to draw distinctions, even when ‘racist’ takes wide scope.
By contrast, theorists who advocate for a narrow-scope conception of ‘racist’ typically recommend adding related, but distinct, expressions that refer to race-related ills that do not rise to the level of severity that they associate with ‘racism’ and ‘racist.’ For example, according to Blum (2002a: 206–207), a 9-year-old white soccer player who says “‘Boy, pass the ball over here’ to one of his black teammates” is racially insensitive but not racist; someone who “exhibits culpable ignorance about racial matters” is racially ignorant but not racist; and a policy that has unintended negative consequences for Black people is racially unjust but not racist. Along the same lines, Anderson (2022: 82) recommends that we “use other terms to describe moral problems concerning race” instead of ‘racist.’ How do these alternative vocabulary terms—‘racially insensitive,’ ‘racially ignorant,’ and ‘racially unjust’—compare to ‘racist’ in terms of extension and intensity? One of the concerns with the wide scope use of ‘racist’ is that it “often leaves addressees in the dark about what is wrong.” Are alternative vocabulary terms more communicatively clear than ‘racist’?
Comparing their perceived extensions, there is a statistically significant difference across bare ‘racist’ and the set of alternative vocabulary terms (χ2(3) = 143.82, p < 0.001). With pairwise comparisons, there are statistically significant differences between all pairs of terms, except between ‘racist’ and ‘racially unjust’ (p = 0.79) and between ‘racially ignorant’ and ‘racially insensitive’ (p = 0.88). As planned, we conducted a follow-up equivalence test to interpret the statistically nonsignificant result: using the conventional “small” effect as our equivalence bounds (d = 0.2), there is no difference between the perceived extension of ‘racist’ and ‘racially unjust’ (t(239) = 2.57, p < 0.01). As the graph shows (Figure 8), the perceived extension of ‘racist’ is narrower than that of ‘racially ignorant’ (M = 45.30, SD = 25.50) and ‘racially insensitive’ (M = 43.00, SD = 24.29), and about the same as ‘racially unjust’ (M = 33.09, SD = 22.99).
Comparing their perceived intensity, there is a statistically significant difference across bare ‘racist’ and the set of alternative vocabulary terms (χ2(3) = 262.6, p < 0.001). With pairwise comparisons, there are statistically significant differences between all pairs of terms, except between ‘racially ignorant’ and ‘racially insensitive’ (p = 0.087). As the graph shows (Figure 9), the perceived intensity of ‘racist’ is stronger than that of ‘racially insensitive’ (M = 44.50, SD = 28.46), ‘racially ignorant’ (M = 49.69, SD = 27.23), and ‘racially unjust’ (M = 56.20, SD = 28.36).
The statistical results show that these alternative vocabulary terms do expand the range of extension and intensity that speakers express beyond what is available with bare ‘racist.’ If the goal of the critics of conceptual inflation is to recommend vocabulary that allows for a wider reference class and a weaker moral force, then these terms do accomplish that goal. However, if their goal is to recommend vocabulary that improves communicative clarity, these terms do not accomplish that. The variations in participants’ understandings of these terms’ extension and intensity, as measured by standard deviations, are about the same as or greater than those of ‘racist.’
We can compare the expressive range of the two options on the table, wide-scope ‘racist’ with degree modifiers versus narrow-scope ‘racist’ with alternative vocabulary terms. To operationalize, we can take the mean difference between the terms at the extremes of each set (‘slightly racist’ vs. ‘extremely racist’ and ‘racially ignorant’ vs. ‘racially unjust’). For the intensity measures, the expressive range for ‘racist’ with degree modifiers is statistically significantly wider (t(239) = 8.391, p < 0.001). For the extension measures, the expressive range for degree modifiers is also statistically significantly wider (t(239) = –4.105, p < 0.001). However, we want to emphasize that—unlike Anderson—we do not see the choice as exclusive. Natural language provides speakers with many types of expressive resources for making moral criticisms, and we should take advantage of those resources to say what we mean. Alternative vocabulary terms have their place, but so does ‘racist,’ especially in combination with degree modifiers.
2.5. ‘Racist’ over Apparent Time
After situating the extension and intensity of ‘racist’ against the backdrop of other terms, we can now assess the two central claims of the conceptual inflation critique. Has ‘racist’ undergone extension expansion? Has ‘racist’ undergone intensity bleaching? As mentioned, we will use the apparent time construct to answer these questions. That is, we will examine how perceived extension and perceived intensity of ‘racist’ vary by participant age. Since some critics (such as Anderson) mention race—specifically, white versus nonwhite—as another potential source of variation in how these terms are perceived, we will also take into account how perceived extension and perceived intensity of ‘racist’ vary by participant race. We did so on the assumption that white people and nonwhite people might perceive both the extension and intensity of the term differently.
We fitted a model that predicts the perceived extension of ‘racist’ with participant age and race. As the graph shows (Figure 10), the model’s explanatory power is weak (R2 = 0.02). There is a statistically significant relationship with race (beta = –18.41, 95% CI [–36.36, –0.46], p = 0.044). (However, this relationship is statistically not significant in the sample without the preregistered exclusions.) There is no statistically significant relationship with age (beta = –0.32, 95% CI [–0.69, 0.05], p = 0.093) or the age-race interaction (beta = 0.35, 95% CI [–0.07, 0.77], p = 0.101). We also fitted a model that predicts the perceived intensity of racist with participant age and race. As the graph shows (Figure 11), the model’s explanatory power is very weak (R2 = 0.01). There is no statistically significant relationship with age (beta = –0.30, 95% CI [–0.76, 0.17], p = 0.210) or race (beta = –16.91, 95% CI [–39.23, 5.41], p = 0.138), or the age-race interaction (beta = 0.45, 95% CI [–0.07, 0.97], p = 0.089).
We also exploratorily examined the age variable by itself. For the perceived extension of ‘racist,’ the relationship is negative, statistically not significant, and tiny (r = –0.05, 95% CI [–0.17, 0.08], p = 0.463). For the perceived intensity of ‘racist,’ the relationship with age is positive, statistically not significant, and tiny (r = 0.04, 95% CI [–0.09, 0.16], p = 0.556). In isolation, there are also no statistically significant relationships between either perceived extension or intensity of ‘racist’ and, respectively, race and gender. (Since there are not similar studies on this topic, we have given conventional labels to effect sizes following Cohen 1988, Field 2013, and Funder & Ozer’s 2019 recommendations.)
The statistical results do not show that there has been extension expansion or intensity bleaching with ‘racist.’ Study 1 does not reveal any evidence of different linguistic behavior with ‘racist’ with respect to apparent time. Moreover, regardless of the conventional threshold of statistical significance, the effects with age are tiny in their magnitudes. That is, even if there were changes to the meaning of ‘racist’ over time, these changes would be hard to notice without statistical tools.
What about the alleged inverse correlation between extension expansion and intensity bleaching, the intuitive idea that as the extension of a term increases, its moral intensity decreases? When we designed this study, we simply assumed—like some of the critics of conceptual inflation—that the claim of inverse correlation was true. However, we began to suspect that our assumption was faulty when we compared ‘racist’ with other terms and saw that the comparative patterns with extension and intensity do not always match. For example, remember that ‘racist’ is about as strong as ‘worthless’ in its moral force (Figure 5), but picks out more people than ‘worthless’—indeed, more people than ‘terrible’—in its reference class (Figure 4). (A referee asked whether this finding is evidence that participants are not paying sufficient attention to the prompts. We embedded an attention check throughout Study 1 whereby we excluded anyone from our analysis who did not respond to the modified expressions ‘slightly racist,’ ‘moderately racist,’ and ‘extremely racist’ in the expected way, namely giving ‘slightly racist’ lower intensity and greater extension ratings than ‘moderately racist’ and ‘extremely racist.’ We therefore think that the fact that intensity and extension ratings come apart cannot be explained in terms of lack of attention.) Surprisingly to us, as the graph shows (Figure 12), there was no statistically significant correlation between individual participants’ perceived extension and intensity of ‘racist’ (tau = 0.055, p = 0.891). Even when outliers were excluded, there was still no statistically significant correlation (tau = –0.002, p = 0.48). Contrary to some critics’ (and our) assumption, it is not true that perceiving a wider reference class correlates with perceiving a weaker moral force of ‘racist.’
2.6. Summary
We precisified the critique of conceptual inflation into three claims: extension expansion, intensity bleaching, and inverse correlation. In Study 1, we found no evidence for any of these three claims. Despite critics’ outcries over the last three decades, Study 1 turned up no support for ‘racist’ having undergone a transformation that is comparable to the case of currency inflation.
3. Study 2: Comparison with Reclaimed Slurs
In Study 1 we only found an absence of evidence for conceptual inflation, and not evidence for an absence of conceptual inflation. Are we using the right tools and looking in the right place to find evidence of conceptual inflation? Maybe our questions are too coarse and the participants’ responses are too noisy for us to uncover evidence of conceptual inflation, or maybe public debates about what ‘racist’ should and should not be applied to have pushed older speakers’ use of the term into alignment with younger speakers so that what is in fact a real-time change does not show up in our apparent time study.
One way to demonstrate that our methods are capable of finding evidence of meaning change is to test them on unrelated terms that we have good reason to think have changed their meaning. If our approach doesn’t detect meaning changes with these terms, that would raise concerns about our methods. A particularly promising domain to test our methods on concerns the reclamation of slurs, a topic that has received philosophical attention in recent years (Anderson 2018; Cepollaro & López de Sa 2022; Jeshion 2020; Popa-Wyatt 2020; Ritchie 2017). In particular, we think that the reclaimed terms ‘queer’ (Brontsema 2004) and ‘slut’ (Herbert 2015), as well as the replacement term ‘neurodivergent,’ are promising examples of terms that have undergone recent extension expansion and intensity bleaching. Moreover, it seems plausible that inverse correlation between extension expansion and intensity bleaching holds for reclaimed slurs, though the causal relationship is the opposite of what is alleged with conceptual inflation: if activists’ work to reduce the moral badness of these terms has been successful, then that opens the door to wider use—including, for example, adding oneself to the reference class. Philosophers investigating the “implementation challenge” in conceptual engineering (Cappelen & Plunkett 2020), which concerns methods for bringing about change in concepts or word meanings, should be interested in whether extension expansion brings about or correlates with intensity bleaching, because whether such a relationship exists could affect what strategies are most effective for bringing about change in the meaning of expressions.
We ran a second study that investigates the meaning of ‘racist’ in comparison with reclaimed slurs. Instead of using a demographically representative sample, in this study we wanted to make any meaning change that exists as vivid as possible by specifically recruiting participants from two different age groups that are three decades apart: people 30 years old and under (the under-30s) and people 60 years old and over (the over-60s). We chose these groups because conceptual inflation worries about ‘racist’ have been around for at least 30 years; if the worries are on target, then it should be relatively easy to find a difference between age groups that are at least 30 years apart.
3.1. Methods
In a preregistered study, 395 U.S.-based participants were recruited using Prolific. (All material, data, and analysis can be found at https://doi.org/10.17605/OSF.IO/8D7NK.) The median time to complete the study was 1 minute and 54 seconds for the under-30s, and 3 minutes 4 seconds for the over-60s. Participants were paid $1 on completion. No exclusion criterion was applied. We intentionally recruited participants in two age groups, and did not attempt to achieve representativeness on other demographic variables. 8 participants (2.03%) preferred to not disclose their age in the questionnaire. The remaining participants were divided into two planned groups: the under-30s (n = 194, 49.11%; Mage = 25.13, SDage = 3.45) and the over-60s (n = 193, 48.86%; Mage = 65.54, SDage = 4.86). (There exist small discrepancies between participants’ age as reported on Prolific, which we used for recruitment screening, versus in Qualtrics, which we used for the questionnaire. We included 1 36-year-old in the under-30s group and 2 59-year-olds in the over-60s group.) In addition to race and gender, in this study we also collected information on political orientation (0–100 scale: the left – the right) and education attainment (some high school / some undergraduate / some postgraduate). Since we did not attempt to achieve representativeness, there exist other demographic differences between the under-30s (46.3% men, 48.4% women, 5.3% non-binary; 54.5% white, 44.85% nonwhite; Mpolitical = 27.63, SDpolitical = 25.82; 15.5% some high school, 69.1% some undergraduate, 15.5% some postgraduate) and over-60s (33.9% men, 65.6% women, 0.5% non-binary; 89.6% white, 10.36% nonwhite; Mpolitical = 35.78, SDpolitical = 31.96; 9.3% some high school, 66.3% some undergraduate, 24.4% some postgraduate).
Participants completed the study in Qualtrics and provided informed consent before proceeding. There was no experimental manipulation and all participants were asked all questions. Participants were presented with two main blocks of four questions in counterbalanced randomized order. Within each block, participants were asked one of the following two types of questions on a 0–100 slider with text anchors:
| [extension] | What percentage of people can be reasonably called ‘racist’? [none – all] |
| [intensity] | How bad is it for a person to be called ‘racist’? [not at all – the worst] |
Within each block, participants were asked, in counterbalanced randomized order, about the target term ‘racist,’ and the terms ‘queer,’ ‘slutty,’ and ‘neurodivergent.’ (In particular, as preregistered, our main comparison is between ‘racist’ and ‘queer,’ since we expected ‘queer’ to be most likely to show a change of meaning over time.) Finally, participants were asked demographic questions about age, race, gender, political orientation, and educational attainment, following a recommendation on inclusive practices for collecting demographic questions wherever possible (Hughes et al 2016).
3.2. ‘Racist’ over Apparent Time
Has ‘racist’ undergone extension expansion or intensity bleaching? Once again, we used the apparent time construct to answer these questions. We conducted Welch two-sample t-tests to compare the responses of our two age cohorts.
As Figure 13 shows, for perceived extension, the difference between under-30s (M = 38.77, SD = 21.78) and over-60s (M = 33.26, SD = 21.07) was statistically significant and small (t(384.70) = –2.53, p = 0.012; Cohen’s d = –0.26, 95% CI [–0.46, –0.06]). As Figure 14 shows, for perceived intensity, the difference between under-30s (M = 76.86, SD = 25.58) and over-60s (M = 75.07, SD = 29.41) was statistically not significant and tiny (t(377.19) = –0.64, p = 0.525; Cohen’s d = –0.06, 95% CI [–0.26, 0.13]). As planned, we conducted a follow-up equivalence test to interpret this statistically nonsignificant result concerning perceived intensity: Using the conventional “small” effect as our equivalence bounds (d = 0.2), we cannot conclude that there was no difference between the two groups (t(377.19) = 1.33, p = 0.09).
The statistical results indicate that there has been some extension expansion with ‘racist,’ but do not indicate that there has been intensity bleaching. Unlike Study 1, we did find some evidence of different linguistic behavior with ‘racist’ across our two age cohorts. Contrary to our prediction based on Study 1, older participants do use ‘racist’ to pick out a narrower reference class compared to younger participants. However, as predicted based on Study 1, there remained no evidence that older participants use ‘racist’ with a stronger moral force compared to younger participants. If the apparent time approach is a legitimate proxy for real time change in use (keeping in mind the serious caveats about this inference discussed in §1.3), then we do have some reason to believe that there has been a small extension expansion of the meaning of ‘racist,’ but we still do not have reason to believe that there has been intensity bleaching.
What about inverse correlation? In an exploratory test in Study 1, we were surprised to find no correlation between individual participants’ perceived extension and intensity of ‘racist.’ In Study 2, we preregistered a one-tailed nonparametric correlation test to re-examine this relationship. Once again, as the graph shows (Figure 13), there was no correlation between individual participants’ perceived extension and intensity of ‘racist’ (tau = 0.02, p = 0.677). As planned, we conducted a follow-up equivalence test to interpret this statistically nonsignificant result: using the conventional “small” effect as our equivalence bounds (d = 0.2), the relationship between participants’ perceived extension and intensity of ‘racist’ was no different from none at all (p < 0.001).
3.3. ‘Racist’ in Comparison
How does ‘racist’ compare to a reclaimed slur like ‘queer’? We asked this question because we thought the comparison would be helpful as a demonstration that our methods are sufficiently sensitive to detect meaning change if it does in fact exist. However, since we did find expansion extension with ‘racist’ in Study 2, this comparison can also provide a benchmark for the magnitude of change.
As expected, we found evidence for different linguistic behavior with ‘queer’ with respect to apparent time. As Figure 14 shows, for perceived extension, the difference between under-30s (M = 24.85, SD = 18.65) and over-60s (M = 17.54, SD = 16.14) was statistically significant and small (t(377.76) = –4.12, p < 0.001; Cohen’s d = –0.42, 95% CI [–0.62, –0.22]). As the graph shows (Figure 15), for perceived intensity, the difference between under-30s (M = 23.20, SD = 28.76) and over-60s (M = 35.32, SD = 34.62) was statistically significant and small (t(371.76) = 3.74, p < 0.001; Cohen’s d = 0.38, 95% CI [0.18, 0.58]).
However, since the difference between statistical significance and statistical non-significance is not itself statistically significant (Gelman & Stern 2006), we also conducted MANOVAs that crossed ‘racist’ versus ‘queer’ with over-30s versus under-60s to compare the two terms’ change over apparent time. For perceived extension, there was a statistically significant small difference between ‘racist’ and ‘queer’ in the difference between under-30s and over-60s (F(1, 385) = 7.11, p < 0.001; partial eta2 = 0.04). For perceived intensity, there was a statistically significant small difference between ‘racist’ and ‘queer’ in the difference between under-30s and over-60s (F(1, 385) = 9.07, p < 0.001; partial eta2 = 0.05).
We chose ‘queer’ as the primary comparison since, despite the precarity of all linguistic reclamation projects, its meaning change is arguably the least controversial. By contrast, we have reservations about ‘slutty’ since its reclaimed status remains controversial and there may be a gendered response pattern, and we also have reservations about ‘neurodivergent’ because it is too recent and specifically coined to replace an older term in the process of deprecation. That said, we found the same pattern of responses with these terms too. When we conducted MANOVAs that crossed ‘racist’ versus ‘slutty’ with over-30s versus under-60s, we also found a statistically significant small difference with perceived extension (F(1, 385) = 6.45, p < 0.001; partial eta2 = 0.03) and with perceived intensity (F(1, 385) = 7.22, p < 0.001; partial eta2 = 0.04). When we conducted MANOVAs that crossed ‘racist’ versus ‘neurodivergent’ with over-30s versus under-60s, we found a statistically significant small difference with perceived extension (F(1, 385) = 9.99, p < 0.001; partial eta2 = 0.05) but no statistically significant difference with perceived intensity (F(1, 385) = 2.93, p = 0.054; partial eta2 = 0.02).
The statistical results show that our methods are sensitive enough to detect meaning change, given the reasonable assumption that ‘queer’ has undergone some change in meaning over the past three decades (and perhaps ‘slutty’ and ‘neurodivergent’ too). Moreover, while we found some evidence of extension expansion with ‘racist,’ this change is smaller in magnitude compared to the extension expansion with ‘queer’ (as well as ‘slutty’ and ‘neurodivergent’). At the same time, while we found no evidence for intensity bleaching with ‘racist,’ there has been intensity bleaching with ‘queer’ (as well as ‘slutty’), presumably from reclamation efforts.
3.4. Summary
We precisified the critique of conceptual inflation into three claims: extension expansion, intensity bleaching, and inverse correlation. In Study 2, with ‘racist,’ we found evidence for a small effect of extension expansion (in discordance with Study 1), but still no evidence for intensity bleaching or inverse correlation (in concordance with Study 1). Through comparison with reclaimed slurs, especially ‘queer,’ we found that ‘racist’ does not show the clear profile of meaning change that less controversial examples do.
4. Study 3: Replicate Central Elements of Study 1 and Study 2
4.1. Methods
A referee suggested that we run a version of Study 2 using a demographically representative sample, as we did in Study 1. Doing so would allow us to use the same measure of conceptual inflation for ‘racist’ across different studies, treating age as a continuous variable rather than comparing the results of participants under 30 years old and those over 60 years old. With that in mind, we ran a third study that used the same experimental materials and design as Study 2 but recruited a demographically representative sample of the U.S. population as we did in Study 1. Since this study aims to replicate findings from Study 1 and Study 2, we did not conduct a separate preregistration.
398 U.S.-based participants were recruited using Prolific’s “representative sample” function (Mage = 45.50, SDage = 15.49; 48.5% men, 49.1% women, 2.4% non-binary; 62.6% white, 27.4% nonwhite). The median time to complete the study was 2 minutes and 47 seconds, and participants were paid $1 on completion. Experimental materials and design were identical as those used in Study 2 (see §3.1). (All material, data, and analysis can be found at https://doi.org/10.17605/OSF.IO/8VWRZ.)
4.2. ‘Racist’ over Apparent Time
As in Study 1, we fitted a model that predicts the perceived extension of ‘racist’ with participant age and race. The model’s explanatory power is weak (R2 = 0.05). There is no statistically significant relationship with race (beta = –3.56, 95% CI [–17.73, 10.60], p = 0.622), or with age (beta = 0.07, 95% CI [–0.18, 0.32], p = 0.584) or the age-race interaction (beta = –0.14, 95% CI [–0.45, 0.17], p = 0.382). We also fitted a model that predicts the perceived intensity of racist with participant age and race. The model’s explanatory power is very weak (R2 = 0.01). There is no statistically significant relationship with age (beta = 0.11, 95% CI [–0.22, 0.44], p = 0.510) or race (beta = 5.58, 95% CI [–12.74, 23.89], p = 0.551), or the age-race interaction (beta = –0.02, 95% CI [–0.42, 0.38], p = 0.907).
We also examined the age variable by itself. For the perceived extension of ‘racist,’ the relationship is negative, statistically not significant, and very small (r = –0.06, 95% CI [–0.16, 0.04], t(387) = –1.23, p = 0.220) (see Figure 16). For the perceived intensity of ‘racist,’ the relationship with age is positive, statistically not significant, and very small (r = 0.07, 95% CI [–0.03, 0.17], t(387) = 1.35, p = 0.176) (see Figure 17). This replicates what we found in Study 1: we did not find evidence in Study 3 of either extension expansion or intensity bleaching with ‘racist’ over apparent time.
There was also no statistically significant correlation between individual participants’ perceived extension and intensity of ‘racist’ (tau = 0.08, p = 0.989). With eqb at ±0.02, the correlation is indistinguishable from zero (p < 0.001). This replicates what we found in both Study 1 and Study 2.
4.3. ‘Queer’ over Apparent Time
As with ‘racist,’ we examined whether participant age predicts perceived extension and intensity of ‘queer.’ For the extension measure, the relationship with age is negative, statistically significant, and medium (r = –0.20, 95% CI [–0.30, –0.11], t(387) = –4.11, p < .001) (see Figure 18). For the intensity measure, the relationship with age is positive, statistically significant, and small (r = 0.10, 95% CI [2.45e–03, 0.20], t(387) = 2.01, p = 0.045) (see Figure 19).
This replicates what we found in Study 2, but with age as a continuous variable: we did find evidence for extension expansion and intensity bleaching with ‘queer’ over apparent time.
4.4. Comparing ‘Racist’ and ‘Queer’ over Apparent Time
Once again, since the difference between statistical significance and statistical non-significance is not itself statistically significant, we need to directly compare the age correlations for ‘racist’ and ‘queer.’ There are a few ways for doing this, and we are reporting the results from Hittner, May, and Silver’s (2003) method. Contrary to Study 2 (albeit with a different sample and analysis), we did not find a significant difference between the age variation with respect to the intensity of ‘racist’ and ‘queer’ (z = –0.498, p = 0.619). Congruent with Study 2, we did find a significant difference between the age variation with respect to the extension of ‘racist’ and ‘queer’ (z = 2.502, p = 0.012).
However, when we compared the under-30s with the over-60s in this sample, as we did in Study 2, we did replicate its results. For perceived extension, there was a statistically significant small difference between ‘racist’ and ‘queer’ in the difference between under-30s and over-60s (F(1, 170) = 5.18, p = 0.007; partial eta2 = 0.06). For perceived intensity, there was a statistically significant small difference between ‘racist’ and ‘queer’ in the difference between under-30s and over-60s (F(1, 170) = 3.07, p = 0.049; partial eta2 = 0.04).
4.5. Summary of Study 3
Study 3 confirmed the big picture findings of Study 1 and Study 2, while adding a few interesting complications. We replicated two main findings of Study 1: We did not find any evidence of extension expansion or intensity bleaching of ‘racist’ over apparent time, and we found no correlation between perceived intensity and extension of ‘racist’ (which would be expected if there were a causal relation between extension expansion and intensity bleaching). And we replicated our findings in Study 2, that while there is clear evidence for ‘queer’ having undergone extension and intensity expansion over apparent time, there is no evidence that ‘racist’ has done so to the same extent. Specifically, contrary to Study 2 but congruent with Study 1, we did not find evidence for extension expansion of ‘racist’ over apparent time. Furthermore, when comparing correlations over the entire demographically representative sample, we replicated our finding in Study 2 that the difference between the differences of perceived extension for ‘racist’ and ‘queer’ over apparent time is significant, though we did not replicate our finding that the difference between the differences of perceived intensity over apparent time is significant. However, when comparing the same differences between differences with under-30s vs. over-60s, as we did in Study 2, we did replicate both findings.
5. General Discussion
While our Study 2 did find evidence that ‘racist’ has undergone some extension expansion (over-60s applied it more restrictively than the under-30s), we did not find evidence of any change in either extension or intensity in Study 1 or Study 3. And across our three studies we did not find any evidence that ‘racist’ has either of the other two features that are typically combined in worries about conceptual inflation: intensity bleaching and inverse correlation of extension and intensity. If the conceptual inflation of ‘racist’ is as rampant and problematic as critics have proposed, our results should be surprising. What explains the mismatch between the frequently expressed belief that ‘racist’ is undergoing conceptual inflation and our results?
5.1. What about Politics?
After completing Study 1, a participant volunteered a comment to us about the politics they perceive behind conflicting claims about what counts as ‘racist’:
I just complete[d] your study on being called “racist” and I have an issue with the study’s design. Would not being called any degree of racist depend on the person doing the name-calling? Nowadays, too many people resort to wrongly calling others racist when they are losing their argument. The term is thrown at others so commonly that the word has become almost meaningless. Race is used by too [m]any politicians to divide people, a wedge issue to get votes. Democrats use it frequently to make up for lack of a legitimate platform.
Our results in Study 1 show that this participant is just wrong that ‘racist’ has become “almost meaningless”: it is as intense a term of criticism as ‘worthless,’ and it is not applied indiscriminately; its extension is judged to be somewhere between the extension of ‘terrible’ and ‘disagreeable.’ But what about the participant’s claim that people with different politics make different judgments about the meaning of ‘racist’? In Study 2 and Study 3, we included a demographic question about political orientation so we could begin to evaluate whether judgments about the intensity and extension of ‘racist’ correlate with people’s self-identified politics. In both Study 2 and 3, we did not find evidence that the perceived intensity of ‘racist’ correlates with political orientation. In Study 2, the correlation between intensity and political orientation is negative, statistically not significant, and tiny (r = –0.04, 95% CI [–0.14, 0.06], t(383) = –0.71, p = 0.479); in Study 3, the correlation between intensity and political orientation is negative, statistically not significant, and very small (r = –0.07, 95% CI [–0.17, 0.03], t(381) = –1.31, p = 0.191). In Study 2, we did find evidence that there is a correlation between extension and political orientation; people reporting a political orientation on the right tend to judge that ‘racist’ has a narrower extension than people on the left: the relationship between extension and political affiliation is negative, statistically significant, and small (r = –0.18, 95% CI [–0.28, –0.09], t(383) = –3.67, p < .001) (Figure 20). But in Study 3 we did not find evidence of a correlation between the perceived extension of ‘racist’ and political affiliation: the relationship is positive, statistically not significant, and tiny (r = 0.01, 95% CI [–0.09, 0.11], t(381) = 0.24, p = 0.812) (Figure 21).
As our participant observed, what counts as ‘racist’ is a contentious political issue, so it’s not surprising that we found some evidence that the meaning of ‘racist’ correlates with political affiliation. (What is more surprising is that we found evidence of such a correlation only in Study 2 and not Study 3). But does this tell us anything about conceptual inflation? It shows us that there is some evidence of systematic variation in how people judge the extension of ‘racist,’ but not that the meaning of ‘racist’ is changing over time. As we will discuss in the next section, there is reason to suspect that people notice this kind of variation and mistake it for evidence of linguistic change.
5.2. Is Conceptual Inflation an Illusion?
A recent study investigated the universal human tendency, documented since Livy’s History of Rome (written between 27 and 9 BC), to think that one’s own historical era is undergoing moral decline (Mastroianni & Gilbert 2023). Mastroianni and Gilbert find that while there is widespread and persistent belief in moral decline, responses to 107 survey questions “administered to 4,483,136 people across a 55-year span from 1965 to 2020” showed that “people’s reports of the current morality of their contemporaries were stable over time” (4), suggesting that belief in moral decline is an illusion (1). They give a two factor explanation for this illusion: First, “human beings are especially likely to seek and attend to negative information about others,” information amply provided by media with a focus on current negative events (5); and second, “when people recall positive and negative events from the past, the negative events are more likely to be forgotten, more likely to be misremembered as their opposite, and more likely to have lost their emotional impact” (5). Those two factors, when taken together, explain why people underestimate the morality of their contemporaries and overestimate the morality of their historical predecessors, generating the illusion of moral decline. Similarly, worries about the decline of English have been around almost as long as the language itself (Shariatmadari 2020: ch.1), and are just as illusory as perennial beliefs about moral decline, because there is no evidence of actual linguistic decline. Could beliefs about the conceptual inflation of ‘racist’ as standardly understood be similarly illusory?
One combination of factors that might explain false beliefs about conceptual inflation is (a) there is variation in people’s judgments about both the extension and the intensity of ‘racist,’ including variation in extension that correlates with political affiliation (which we found evidence of in Study 2); (b) the frequency of uses of ‘racist’ has recently increased (Rozado 2022); and (c) there is evidence that people are susceptible to a “recency illusion” about linguistic facts, whereby they mistake a linguistic phenomenon that they have recently noticed for a phenomenon that is genuinely new (Zwicky 2005; van der Meulen 2022). This combination of factors would produce a situation in which there is no underlying change in the meaning of ‘racist,’ but because uses of ‘racist’ have become more frequent, people are more aware of variation in the way people use the term, and because of the recency illusion, they misinterpret their increasing awareness of that variation as a change in the meaning of the term. That combination of factors would explain why there is an illusion of conceptual inflation.
5.3. An Alternative Conception of Conceptual Inflation
The most explicit accounts of what conceptual inflation is (see, for example, the accounts from Blum and Case given above) involve the three components that we have been evaluating throughout our studies: (i) extension expansion, (ii) intensity bleaching, and (iii) the causal claim that extension expansion causes intensity bleaching. Our findings cast doubt on the existence of that standard conception of conceptual inflation. But there is a neighboring way of understanding conceptual inflation that our findings in Study 2 lend some weak support to, namely the idea that what constitutes conceptual inflation is extension expansion with no intensity bleaching—and no claim about the causal relationship between (i) and (ii). Our Study 2 did find evidence of (i) and did not find evidence of (ii), so it could be taken to lend some support to this alternative conception of conceptual inflation (though note that the difference between age cohorts we found was small, and we did not find evidence of this difference in Study 3). On this alternative conception, what makes conceptual inflation bad is not that the moral force of ‘racist’ is being diluted, but that it is not being diluted while the extension of the concept expands. That means more people will find themselves in the target area of a powerful form of moral criticism. Is that a problem? In this paper we are not joining debates about whether (and if so, why) conceptual inflation is bad. Our central goal is to operationalize conceptual inflation so that we can gather evidence about whether or not it is taking place. One benefit of operationalization is that it clarifies the fact that there are different ways of understanding conceptual inflation; on the standard interpretation, we did not find evidence that it was taking place, but if it is understood as extension expansion without intensity bleaching, then we did find some equivocal evidence that it is taking place.
6. Conclusion, Limitations, and Suggestions for Future Research
Our primary aim in this paper is to improve on all the existing claims by philosophers and social critics who allege that ‘racist’ is undergoing conceptual inflation, by showing different ways that amorphous claim can be made more precise and picking one of those precisifications and looking for evidence that it is occurring. In this subsection, we highlight some of the limitations of this approach.
We want to be clear that we don’t take ourselves to show that ‘racist’ has not undergone any meaning change; only that one focused way of looking for changing use of the term did not find evidence of one way of understanding conceptual inflation. We think that particular way of understanding conceptual inflation is important and worth our attention, and if conceptual inflation is in fact taking place as alleged, then we would expect to find this kind of evidence of the changing use of the term (as we did for ‘queer’ and ‘slutty’). But this is just a first step in a systematic investigation of this purported phenomenon. One natural next step would be to focus on the horizontal dimension of conceptual inflation we discuss in §1.2; in addition to the corpus evidence we present which indicates that early uses of ‘racist’ were already being applied horizontally to institutional categories beyond individuals, an experimental investigation of the horizontal dimension of conceptual inflation could be modeled on the apparent time study of the meaning change of the word ‘gay’ in Robinson (2012). The format that Robinson used involved asking these questions:
Q1. “Who or what is X” (where “gay” was the target expression)
Q2. “Why is Y X?” where “Y” was the answer to Q1.
Robinson then coded the responses to evaluate whether participants of different ages were applying ‘gay’ to different categories.
As we discuss in §1.3, there haven’t been comparisons of apparent time studies of meaning with real time studies of the kind that validate the use of the apparent time method in studying phonological, syntactic, and lexical change (Bailey 2004). That leaves one assumption of the method in need of additional support. Moreover, there are deep difficulties in any attempt to identify meaning change by investigating linguistic behavior or language users’ beliefs.
Suppose we did find an apparent time difference in how people use the term ‘racist.’ There would be lots of potential explanations of that difference: it could be due to there simply being more (or fewer) racists around, or changes in various pragmatic norms concerning the use of the term. We made these observations urging caution about interpreting Rozado’s data that shows the increasing frequency of terms like ‘racist’ as evidence of changing meaning (Figure 1).
In the same way, there are lots of explanations for not finding differences in use that are compatible with meaning in fact changing. If, for example, the number of extreme racists goes down, but the threshold for what counts as a racist also drops, then the overall proportion of racist people could stay the same over time. That might count (on some ways of understanding meaning) as a change in meaning (if changing thresholds counts as a change in meaning) but without a noticeable change in responses to our question “What % of people can reasonably be called ‘racist’?” So we can’t conclude from there being no evidence of change in use that the meaning has remained stable over time.
This limitation of our investigation—that there are deep difficulties in drawing conclusions about meaning change (or stability of meaning) from evidence of language use—is equally a worry for our opponents, those who claim (on the basis of anecdotal evidence) that ‘racist’ is undergoing conceptual inflation. They haven’t tried to gather systematic evidence that use is in fact changing (aside from evidence of changing frequency), let alone done any work to try to exclude those competing explanations of what might account for any observed changes in use. Even with this limitation, we believe in the value of systematic empirical investigation of changing uses of language, and we have demonstrated one—but by no means the only—way to empirically study changing use.
It is frustrating that there aren’t earlier quantitative studies of ‘racist’ and other contested expressions that would make a straightforward real time study of meaning change possible. The apparent time method relies on assumptions, like the constancy of the use of expressions over the post-adolescent lifetimes of individual speakers, that may turn out not to hold for ‘racist.’ Any conclusions we draw about the presence or absence of conceptual inflation on the basis of our apparent time studies must therefore remain tentative. Our main goals in this study are to make the phenomenon people are worried about when they worry about ‘conceptual inflation’ more precise and do better at searching for evidence of it than the existing claims that it is taking place—not to settle the question whether ‘racist’ is undergoing conceptual inflation. Settling that question will require many different experimental and observational approaches. If the only result of our study is that people come to see how inadequate the empirical support is for the standard claims that conceptual inflation is taking place, then we will be happy; though we would be even happier if we encourage people to run other studies using different techniques to examine the phenomenon from different directions. But our studies provide the best evidence we currently have, and they take the first steps beyond the impressions of those who claim to have noticed the conceptual inflation of ‘racist.’ And they constitute the first part of a future real time study for which our data will, eventually, be a historical record of how ‘racist’ was once used in the hard-to-remember era of the mid-2020s.
Acknowledgements
Thanks to Heather Burnett, Kathryn Francis, Ana Gantman, Aidan Gray, Jonathan Ichikawa, Judy Sein Kim, Joshua Knobe, Quill Kukla, Matthew Lindauer, Eliot Michaelson, Ethan Nowak, Stephanie Solt, Alejandro Vesga, Chun-Ping Yen, and Tomasz Zyglewicz for comments and discussion. Audiences at the Words Workshop, Berlin Sociolinguistics Workshop, Meaning and Reality in Social Context II Conference, Yale Experimental Philosophy Lab, CUNY PsyPhi Lab, Princeton University Center for Human Values Race Group, Cal State Long Beach, Siena College, the UK X-Phi Triangle, the 7th PLM in Prague, University College Dublin, University of Puget Sound, and Western Washington University gave us very helpful comments. Nat Hansen gratefully acknowledges support from the Alexander von Humboldt Foundation and from the University of Reading for a University Research Fellowship. Shen-yi Liao gratefully acknowledges support from University of Puget Sound research fund and from Princeton University’s University Center of Human Values Visiting Faculty Fellowship.
References
Anderson, Elizabeth (2010). The Imperatives of Integration. Princeton University Press.
Anderson, Elizabeth (2022). Can We Talk? Communicating Moral Concern in an Era of Polarized Politics. Journal of Practical Ethics, 10(1), 67–92.
Anderson, Luvell (2018). Calling, Addressing, and Appropriation. In David Sosa (Ed.), Bad Words: Philosophical Perspectives on Slurs (6–28). Oxford University Press.
Andow, James (2015). How “Intuition” Exploded. Metaphilosophy, 46(2), 189–212.
Bailey, Guy (2004). Real and Apparent Time. In J.K. Chambers, Peter Trudgill, and Natalie Schelling-Estes (Eds.), The Handbook of Language Variation and Change (312–332). Blackwell.
Blum, Lawrence (2002a). ‘I’m Not a Racist, But…’: The Moral Quandaries of Race. Cornell University Press.
Blum, Lawrence (2002b). Racism: What It Is and What It Isn’t. Studies in Philosophy and Education, 21(3), 203–218.
Blum, Lawrence (2004). What Do Accounts of ‘Racism’ Do? In Michael P. Levine and Tamas Pataki (Eds.), Racism in Mind (56–77). Cornell University Press.
Boberg, Charles (2004). Real and Apparent Time in Language Change: Late Adoption of Changes in Montreal English. American Speech, 79(3), 250–269.
Brontsema, Robin (2004). A Queer Revolution: Reconceptualizing the Debate Over Linguistic Reclamation. Colorado Research in Linguistics, 17(1), 1–17.
Burris, Keith (2018, January 15). Reason as Racism: An Immigration Debate Gets Derailed. Pittsburgh Post Gazette. https://www.post-gazette.com/opinion/editorials/2018/01/15/Reason-as-racism-An-immigration-debate-gets-derailed/stories/201801150024
Cappelen, Herman and David Plunkett (2020). A Guided Tour of Conceptual Engineering and Conceptual Ethics. In Alexis Burgess, Herman Cappelen, and David Plunkett (Eds.), Conceptual Engineering and Conceptual Ethics (1–26). Oxford University Press.
Case, Spencer (2019). The Boy Who Inflated the Concept of ‘Wolf.’ Quillette. Retrieved from https://quillette.com/2019/02/14/the-boy-who-inflated-the-concept-of-wolf/
Cepollaro, Bianca and Dan López de Sa (2022). Who Reclaims Slurs? Pacific Philosophical Quarterly, 103(3), 606–619.
Chambers, J. K. (2002). Sociolinguistic Theory: Linguistic Variation and Its Social Significance. Blackwell.
Cohen, Jacob (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Routledge.
Davern, Michael, Rene Bautista, Jeremy Freese, Pamela Herd, and Stephen L. Morgan; General Social Survey 1972–2024. https://gssdataexplorer.norc.org/variables/4087/vshow. Accessed December 19, 2024.
Davies, Mark (2010). The Corpus of Historical American English (COHA): 400 million words, 1810–2009. Available online at https://www.english-corpora.org/coha/
Field, Andy (2013). Discovering Statistics Using IBM SPSS Statistics. SAGE.
Funder, David C. and Daniel J. Ozer (2019). Evaluating effect size in psychological research: Sense and nonsense. Advances in Methods and Practices in Psychological Science, 2(2), 156–168.
Garcia, Jorge L. A. (1997). Current Conceptions of Racism: A Critical Examination of Some Recent Social Philosophy. Journal of Social Philosophy, 28(2), 5–42.
Gauchat, Louis (1905). L’Unite Phonetique Dans Le Patois D’Une Commune. Max Niemeyer Verlag.
Gelman, Andrew and Hal Stern (2006). The Difference Between ‘Significant’ and ‘Not Significant’ Is Not Itself Statistically Significant. The American Statistician, 60(4), 328–331.
Goldfarb, Neal (2017). A Lawyer’s Introduction to Meaning in the Framework of Corpus Analysis. Brigham Young University Law Review, 6(6), 1359–1417.
Hansen, Nat, J. D. Porter, and Kathryn Francis (2021). A Corpus Study of ‘Know’: On the Verification of Philosophers’ Frequency Claims about Language. Episteme, 18(2), 242–268.
Haslam, Nick (2016). Concept Creep: Psychology’s Expanding Concepts of Harm and Pathology. Psychological Inquiry, 27(1), 1–17.
Herbert, Cassie (2015). Precarious Projects: The Performative Structure of Reclamation. Language Sciences, 52, 131–138.
Hittner, James B., Kim May, and N. Clayton Silver (2003). A Monte Carlo Evaluation of Tests for Comparing Dependent Correlations. The Journal of General Psychology, 130(2), 149–68.
Hughes, Jennifer L., Abigail A. Camden, and Tenzin Yangchen (2016). Rethinking and Updating Demographic Questions: Guidance to Improve Descriptions of Research Samples. Psi Chi Journal of Psychological Research, 21(3), 138–151.
Ichikawa, Jonathan (manuscript). “How Racist Is Racist?”
Ipsos (2024). What Worries the World, March 2024. https://www.ipsos.com/en-nl/what-worries-world-march-2024 (accessed April 17, 2024).
Jeshion, Robin (2020). Pride and Prejudiced. Grazer Philosophische Studien, 97(1), 106–137.
Labov, William (1963). The Social Motivation of a Sound Change. Word, 19(3), 273–309.
Lakens, Daniël, Anne M. Scheel, and Peder M. Isager (2018). Equivalence Testing for Psychological Research: A Tutorial. Advances in Methods and Practices in Psychological Science, 1(2), 259–269.
Lee, Thomas R. and Stephen C. Mouritsen (2017). Judging Ordinary Meaning. The Yale Law Journal, 127(4), 788–879.
Liao, Shen-yi and Nat Hansen (2023). ‘Extremely Racist’ and ‘Incredibly Sexist’: An Empirical Response to the Charge of Conceptual Inflation. Journal of the American Philosophical Association, 9(1), 72–94.
Liao, Shen-yi and Nat Hansen (forthcoming). “Conceptual Inflation,” EurAmerica.
Magué, Jean-Phillipe (2006). Semantic Changes in Apparent Time. 32nd Annual Meeting of the Berkeley Linguistics Society, 227–235.
Mastroianni, Adam and Daniel T. Gilbert (2023). The Illusion of Moral Decline. Nature, 618, 782–789.
McWhorter, John (2022, November 15). When ‘Racism’ Is Not Really Racism. The New York Times. Retrieved from https://www.nytimes.com/2022/11/15/opinion/racism-systemic-structural.html
Miles, Robert (1989/2003). Racism. Routledge.
Moder, Carol L. (2007). Mechanisms of Change in Grammaticalization: The Role of Frequency. In Joan Bybee (Ed.), Frequency of Use and the Organization of Language (336–358). Oxford University Press.
Orth, Taylor (2022). From Millionaires to Muslims, Small Subgroups of the Population Seem Much Larger to Many Americans. Retrieved May 18, 2024 from https://today.yougov.com/politics/articles/41556-americans-misestimate-small-subgroups-population
Pinillos, Ángel and Shaun Nichols (2018). Skepticism and the Acquisition of Knowledge. Mind & Language, 33(4), 397–414.
Popa-Wyatt, Mihaela (2020). Reclamation: Taking Back Control of Words. Grazer Philosophische Studien, 97(1), 159–176.
Ritchie, Katherine (2017). Social Identity, Indexicality, and the Appropriation of Slurs. Croatian Journal of Philosophy, 17(2), 155–180.
Robinson, Justyna A. (2012). A Gay Paper: Why Should Sociolinguistics Bother with Semantics? English Today, 28(4), 38–54.
Rozado, David (2022). Themes in Academic Literature: Prejudice and Social Justice. Academic Questions, 35(2), 16–29.
Shariatmadari, David (2020). Don’t Believe a Word: The Surprising Truth about Language. W.W. & Norton Co.
Sweetser, Eve E. (1988). Grammaticalization and Semantic Bleaching. Proceedings of the Fourteenth Annual Meeting of the Berkeley Linguistics Society, 389–405.
Sytsma, Justin, Roland Bluhm, Pascale Willemsen, and Kevin Reuter (2019). Causation Attributions and Corpus Analysis. In Eugen Fischer and Mark Curtis (Eds.), Methodological Advances in Experimental Philosophy (209–238). Bloomsbury.
Tagliamonte, Sali A. and Katharina Pobst (2020). A Cool Comparison: Adjectives of Positive Evaluation in Toronto, Canada and York, England. Journal of English Linguistics, 48(1), 3–30.
Tobia, Kevin (2020). Testing Ordinary Meaning: An Experimental Assessment of What Dictionary Definitions and Linguistic Usage Data Tell Legal Interpreters. Harvard Law Review, 134(2), 1–62.
Tse, Jesse S. Y. and Nick Haslam (2024). Broad Concepts of Mental Disorder Predict Self-Diagnosis. SSM Mental Health, 6, 1–8.
U.S. Bureau of Labor Statistics (2024). Inflation Calculator. https://www.bls.gov/data/inflation_calculator.htm
van der Meulen, Marten (2022). Are We So Illuded? Recency and Frequency Illusions in Dutch Prescriptivism. Languages, 7(42), 1–18.
Vetter, Barbara. (2014). Dispositions without Conditionals. Mind, 123(489), 129–156.
Warriner, Amy Beth, Victor Kuperman, and Marc Brysbaert (2013). Norms of Valence, Arousal, and Dominance for 13,915 English Lemmas. Behavior Research Methods, 45, 1191–1207.
Zahorec, Mike, Robert Bishop, Nat Hansen, John Schwenkler, and Justin Sytsma (2023). No Modification without Aberration? The Use of Linguistic Corpora in Ordinary Language Philosophy. In David Bordonaba-Plou (Ed.), Experimental Philosophy of Language: Perspectives, Methods, and Prospects (121–149). Springer.
Zwicky, Arnold (2005, August 7). Just Between Dr. Language and I. Language Log. Retrieved April 30, 2024, from http://itre.cis.upenn.edu/~myl/languagelog/archives/002386.html





















