How We Talk to Machines Try an experiment
← Back to research findings / RESEARCH NOTE 001
EXPLORATORY RESEARCH / PRISM DATASET

The Familiarity
Divide.

Do people who know AI better want different things from it? We analyzed 1,500 survey responses to explore how familiarity relates to what people value in language models.

1,500 participantsOriginal data: 2023Secondary analysisAdjusted analysis · October 2026

The familiar-looking pattern changes when we account for AI use. Our most consistent adjusted finding concerns the importance of creativity.

01 / WHO WE COMPARED

Three levels of familiarity.

Participants selected their own familiarity level. The study did not test or certify their technical knowledge.

156Not familiar at all
920Somewhat familiar
424Very familiar
02 / THE HEADLINE FINDINGS

Experience and expectations.

Difference between the average importance ratings of very familiar and unfamiliar participants on a 0–100 scale.

−9.7 pts

Safety was rated lower by very familiar participants.

+8.4 pts

Creativity was rated higher by very familiar participants.

+0.4 pts

Factual accuracy showed almost no difference in group averages.

How to interpret this: These are associations in a particular 2023 participant sample—not evidence that using AI changes people's beliefs, or that experienced users care less about real-world safety.

03 / EXPLORE THE MEASURES

What each group rated important.

The eight measures are the PRISM survey's stated-preference attributes. Higher numbers mean the participant considered the attribute more important. Each bar represents a group mean, not an AI model's actual performance.

Not familiar (n=156)Somewhat familiar (n=920)Very familiar (n=424)

Safety

-9.7 pts very vs not familiar

Not familiar at all: 85.7/100
Somewhat familiar: 81.2/100
Very familiar: 76.0/100
Unadjusted 95% CI for difference: -14.2 to -5.3 points

Creativity

+8.4 pts very vs not familiar

Not familiar at all: 66.0/100
Somewhat familiar: 68.0/100
Very familiar: 74.4/100
Unadjusted 95% CI for difference: +4.2 to +12.6 points

Personalization

+4.6 pts very vs not familiar

Not familiar at all: 66.5/100
Somewhat familiar: 66.7/100
Very familiar: 71.1/100
Unadjusted 95% CI for difference: +0.1 to +9.2 points

Factual accuracy

+0.4 pts very vs not familiar

Not familiar at all: 88.4/100
Somewhat familiar: 88.7/100
Very familiar: 88.9/100
Unadjusted 95% CI for difference: -2.3 to +3.2 points

Helpfulness

+2.4 pts very vs not familiar

Not familiar at all: 88.2/100
Somewhat familiar: 89.1/100
Very familiar: 90.6/100
Unadjusted 95% CI for difference: -0.2 to +5.0 points

Fluency

+2.9 pts very vs not familiar

Not familiar at all: 85.2/100
Somewhat familiar: 86.2/100
Very familiar: 88.2/100
Unadjusted 95% CI for difference: -0.1 to +5.9 points

Diversity

+4.2 pts very vs not familiar

Not familiar at all: 74.0/100
Somewhat familiar: 74.7/100
Very familiar: 78.2/100
Unadjusted 95% CI for difference: +0.5 to +8.0 points

Value alignment

-7.3 pts very vs not familiar

Not familiar at all: 59.9/100
Somewhat familiar: 54.2/100
Very familiar: 52.6/100
Unadjusted 95% CI for difference: -12.4 to -2.3 points
04 / A QUICK ROBUSTNESS CHECK

Does the pattern hold across places?

In separate checks for U.S. respondents, U.K. respondents, and everyone else, safety ratings were lower and creativity ratings higher for the very familiar group than for the unfamiliar group. These geographic subgroup summaries are unadjusted and have not been formally tested on their own. The pooled demographic-adjusted models appear below.

386U.S. respondents
340U.K. respondents
774Other respondents

Some comparison groups are small, including only 26 unfamiliar U.S. respondents and 54 very familiar U.K. respondents. Treat this as a preliminary consistency check.

05 / THE ADJUSTED ANALYSIS

What changes when we account for other factors?

We modeled eight 0–100 importance ratings across 1,499 participants, accounting for age, country of residence, education, and then AI usage frequency. Coefficients compare very familiar respondents to another familiarity group. These are associations, not causal effects.

+4.7 pts

Creativity was rated more important by very familiar participants than somewhat familiar participants after all four controls (95% CI +1.8 to +7.5; Benjamini–Hochberg q≈0.012).

Key distinction: With demographics alone, creativity and safety differences are clearer. After also accounting for how frequently people use AI, most comparisons become uncertain. Familiarity and use frequency substantially overlap, making them difficult to disentangle.

TraitAdjusted difference95% confidence intervalMultiple-testing q

Confidence intervals use HC3 heteroskedasticity-robust standard errors; q-values adjust for eight outcomes within each comparison and model. All statistics are exploratory. The 247 missing frequency values represent participants who were not shown that question, not people who necessarily never used AI.

06 / PRIOR RESEARCH

Standing on existing research.

We're reanalyzing the existing PRISM data, not introducing a new study population. Researchers have already established that AI preferences vary. Our narrower question is how eight structured importance ratings differ with self-reported AI familiarity after controlling for other factors.

PRISM Dataset (Kirk et al., 2024)

Published the data and established variation in preferences, conversation topics, and model rankings. The source paper already reports aggregate preference scores and familiarity categories. Our addition is this adjusted eight-trait familiarity comparison.

Mapping Preference Plurality (Coelho & Hale, 2026)

Analyzed the same 1,500 participants' open-ended descriptions of AI preferences and performed exploratory demographic analyses. This is close related work but uses different measures from our structured slider comparisons.

How Familiarity Breeds Trust and Contempt (2023)

Found links between familiarity and support for autonomous AI applications, establishing this as an existing line of inquiry; it does not investigate PRISM's eight language-model traits.

PRISM-X (Kirk et al., 2026)

Re-recruited 530 original participants for personalized-model tests, rather than studying self-rated familiarity and their original trait-importance ratings.

Originality status: We did not identify our precise comparison in these publications, but our search is not an exhaustive systematic review. We cannot claim this is a first-of-its-kind finding.

07 / METHODS & LIMITATIONS

What this means—and what it doesn't.

How we analyzed the surveys

  1. Retained all 1,500 unique PRISM survey respondents for descriptive comparisons.
  2. Excluded one age-undisclosed respondent for adjusted modeling (n=1,499).
  3. Used OLS for each of eight 0–100 importance ratings, with categorical terms for familiarity, age, country, education and optionally use frequency.
  4. Reported HC3 robust 95% intervals and corrected eight comparisons per model and contrast with Benjamini–Hochberg FDR.
  5. Checked published PRISM research and additional relevant peer-reviewed studies.

What we cannot conclude

  • Familiarity was self-reported, and the data were collected in late 2023.
  • Participants came from a diverse but not globally representative sample.
  • We cannot conclude that AI experience caused any differences.
  • Frequency is strongly associated with familiarity; adjusting for both leads to limited comparison overlap.
  • The safety question asks about avoiding harm, not support for specific content restrictions.
  • This work is exploratory, not preregistered or peer reviewed.

Source: PRISM Alignment Dataset, Kirk et al. (2024), DOI 10.57967/hf/2113. This page displays aggregates only.

SOURCES & ATTRIBUTION

Open data. Open questions.

PRISM Alignment Dataset · Kirk et al. (2024), original research paper · Codebook and project repository.

Derived from PRISM human-authored survey data licensed CC BY 4.0. How We Talk to Machines is an independent initiative, not the PRISM study team. These are exploratory secondary-analysis findings, not peer-reviewed conclusions. No participant-level survey data or model responses are included in this page.