Safety
-9.7 pts very vs not familiar
Do people who know AI better want different things from it? We analyzed 1,500 survey responses to explore how familiarity relates to what people value in language models.
The familiar-looking pattern changes when we account for AI use. Our most consistent adjusted finding concerns the importance of creativity.
Participants selected their own familiarity level. The study did not test or certify their technical knowledge.
Difference between the average importance ratings of very familiar and unfamiliar participants on a 0–100 scale.
Safety was rated lower by very familiar participants.
Creativity was rated higher by very familiar participants.
Factual accuracy showed almost no difference in group averages.
How to interpret this: These are associations in a particular 2023 participant sample—not evidence that using AI changes people's beliefs, or that experienced users care less about real-world safety.
The eight measures are the PRISM survey's stated-preference attributes. Higher numbers mean the participant considered the attribute more important. Each bar represents a group mean, not an AI model's actual performance.
-9.7 pts very vs not familiar
+8.4 pts very vs not familiar
+4.6 pts very vs not familiar
+0.4 pts very vs not familiar
+2.4 pts very vs not familiar
+2.9 pts very vs not familiar
+4.2 pts very vs not familiar
-7.3 pts very vs not familiar
In separate checks for U.S. respondents, U.K. respondents, and everyone else, safety ratings were lower and creativity ratings higher for the very familiar group than for the unfamiliar group. We have not yet adjusted for demographics or formally tested these subgroup patterns.
Some comparison groups are small, including only 26 unfamiliar U.S. respondents and 54 very familiar U.K. respondents. Treat this as a preliminary consistency check.
We modeled eight 0–100 importance ratings across 1,499 participants, accounting for age, country of residence, education, and then AI usage frequency. Coefficients compare very familiar respondents to another familiarity group. These are associations, not causal effects.
Creativity was rated more important by very familiar participants than somewhat familiar participants after all four controls (95% CI +1.8 to +7.5; Benjamini–Hochberg q≈0.012).
Key distinction: With demographics alone, creativity and safety differences are clearer. After also accounting for how frequently people use AI, most comparisons become uncertain. Familiarity and use frequency substantially overlap, making them difficult to disentangle.
| Trait | Adjusted difference | 95% confidence interval | Multiple-testing q |
|---|
Confidence intervals use HC3 heteroskedasticity-robust standard errors; q-values adjust for eight outcomes within each comparison and model. All statistics are exploratory. The 247 missing frequency values represent participants who were not shown that question, not people who necessarily never used AI.
We're reanalyzing the existing PRISM data, not introducing a new study population. Researchers have already established that AI preferences vary. Our narrower question is how eight structured importance ratings differ with self-reported AI familiarity after controlling for other factors.
Published the data and established variation in preferences, conversation topics, and model rankings. The source paper already reports aggregate preference scores and familiarity categories. Our addition is this adjusted eight-trait familiarity comparison.
Analyzed the same 1,500 participants' open-ended descriptions of AI preferences and performed exploratory demographic analyses. This is close related work but uses different measures from our structured slider comparisons.
Found links between familiarity and support for autonomous AI applications, establishing this as an existing line of inquiry; it does not investigate PRISM's eight language-model traits.
Re-recruited 530 original participants for personalized-model tests, rather than studying self-rated familiarity and their original trait-importance ratings.
Originality status: We did not identify our precise comparison in these publications, but our search is not an exhaustive systematic review. We cannot claim this is a first-of-its-kind finding.
Source: PRISM Alignment Dataset, Kirk et al. (2024), DOI 10.57967/hf/2113. This page displays aggregates only.