Are Personality Tests Pseudoscience

TL;DR

Are Personality Tests Pseudoscience? What the Evidence Actually Supports

  • Four criteria separate science from theatre: retest reliability, predictive validity, construct validity, falsifiability.
  • The Big Five passes. The type instruments - MBTI, DISC, Insights - fail on the evidence.
  • When the publishers themselves disclaim prediction, the science question answers itself.
  • Feeling accurately described is the Barnum effect - demonstrated with a recycled horoscope in 1949.

The charge of pseudoscience is serious and specific. It means an instrument looks like science, talks like science, and fails science's own tests. Some personality assessments deserve the label. Others do not. The answer depends on which test, which claim, and - critically - what you are using it for.

Are personality tests pseudoscience? Not all of them, and the blanket charge is lazy. The Big Five model - built and rebuilt across thousands of peer-reviewed studies over six decades - meets every standard criterion for scientific measurement: replicable factor structure, predictive validity for life outcomes, stability on retest. The MBTI, DISC, and the colour-based Jungian instruments occupy different ground: they sort continuous traits into discrete types, produce classifications that are unstable on retest for up to half of takers, and are marketed with technical language that implies a precision their evidence does not support. Whether that makes them pseudoscience or simply poor measurement depends on how tightly you draw the line - but the more important question is whether any of them, good or bad, should be making decisions about people at work.

50% Of MBTI takers classified into a different type within five weeks - the retest instability that type instruments share Pittenger, 1993
.42 Validity of structured interviews for predicting job performance - the highest of any selection method Sackett et al., 2022
4.3/5 Accuracy rating students gave an identical "personal" sketch assembled from horoscopes - the Barnum effect Forer, 1949 / Britannica

What Would Make a Personality Test Scientific?

Before answering whether personality tests are pseudoscience, it helps to agree on what science requires. The criteria are not obscure - they are the same tests any measurement instrument must pass, in any field.

Test-retest reliability. Give the same person the same test twice, with a reasonable gap, and get the same answer. This is the most basic requirement. A thermometer that reads 37°C on Monday and 41°C on Thursday, with nothing having changed, is broken. A personality test that assigns you a different type in five weeks is broken in the same way - unless it explicitly claims to measure something that changes, in which case the change is the finding, not the flaw.

Predictive validity. The test must predict something beyond itself. If a personality assessment claims to tell you how someone will work, it must correlate with how they actually work - job performance, satisfaction, retention, something observable. A test that produces vivid descriptions but predicts nothing is a mirror, not a measure.

Falsifiability. Karl Popper's criterion: a scientific claim must be capable of being proved wrong. If every possible result confirms the theory - if there is no outcome that would count as evidence against it - the theory is not science. When a personality instrument produces a profile that feels accurate regardless of what the profile says, that is a warning sign, not a feature.

Construct validity. The test must measure the thing it claims to measure, and that thing must correspond to real, measurable differences between people - not arbitrary categories imposed on continuous variation. If the boundaries between "types" are drawn through the middle of a normal distribution rather than between genuinely distinct groups, the types are a fiction of the scoring system, not a discovery about people.

These are not unreasonable standards. They are the floor.

Which Tests Pass?

The Big Five (or Five-Factor Model) passes. Its five dimensions - openness, conscientiousness, extraversion, agreeableness, and neuroticism - were derived empirically by factor-analysing thousands of trait descriptions across languages and cultures, starting with the lexical hypothesis work at Brooks Air Force Base in 1961 and replicated hundreds of times since. The factor structure is robust. Conscientiousness predicts job performance across roles; extraversion correlates with leadership emergence and sales outcomes; neuroticism predicts burnout risk. The predictions are modest - personality is never the whole story - but they are real, replicable, and subject to the kind of scrutiny that science demands.

The HEXACO model, developed in 2000 by Ashton and Lee, extends the Big Five with a sixth dimension (honesty-humility) and was shaped by the same empirical process. It passes the same tests.

These are not perfect instruments. The Big Five's validity for selection - using it to decide who gets a job - is weaker than its validity for description. Morgeson and colleagues, in a paper authored by a panel of former editors of Personnel Psychology, described personality tests' predictive validity for job performance as "very low." That is not an attack on the Big Five's scientific status. It is a reminder that a scientifically valid measure can still be the wrong tool for a given job.

A panel of former editors of the field's top journals described personality tests' predictive validity for job performance as "very low" - a reminder that scientific validity and practical usefulness are not the same claim.

- Morgeson et al., 2007, Personnel Psychology

Which Tests Fail?

The type-based instruments - MBTI, DISC, Insights Discovery, and their variants - fail on multiple criteria.

Test-retest reliability. Pittenger's review found up to 50 per cent of MBTI takers classified differently on retest within five weeks. The problem is structural, not a quirk of one instrument: any system that draws a boundary through continuous data will misclassify people near the line, and most people are near the line. This applies to every Jungian type instrument, not just the MBTI.

Construct validity. The MBTI's four dichotomies are imposed on continuous distributions - there are no bimodal clusters at the introvert and extravert poles. People score across a bell curve, and the "type" they receive depends on which side of an arbitrary midpoint they fall. An "INTJ" one point from the boundary is psychologically near-identical to an "ENTJ" one point the other side; the four-letter label says otherwise. DISC-family instruments face the same criticism: independent researchers found the four dimensions were not statistically independent and were better explained as combinations of Big Five traits.

Predictive validity. The MBTI publisher states plainly that the instrument cannot "measure or predict job performance". Wiley, publisher of Everything DiSC, says DISC profiles are "not recommended for pre-employment screening". When the publishers themselves disclaim prediction, the science question answers itself.

"The MBTI assessment is not designed to be used to measure or predict job performance… tell companies who they should hire."

- The Myers-Briggs Company, publisher of the MBTI

The Barnum tell. Forer's 1949 experiment remains the sharpest test of whether a profile is measurement or flattery. He gave 39 students identical personality sketches assembled from an astrology book; they rated the sketch's personal accuracy at 4.3 out of 5. Type-based profiles are careful to describe every category as a strength - nobody is ever told they are difficult - and the language is warm, general, and framed as being about you. The experience of recognition is real. It is not evidence of measurement.

The Sharper Question

"Are personality tests pseudoscience?" treats the field as one thing, and it is not. The Big Five is science. The MBTI is not science, by most reasonable definitions. DISC and the colour models sit in similar territory - instruments built on outdated or untested theory, marketed with the trappings of science (technical manuals, reliability coefficients, professional certifications), producing outputs that feel compelling for reasons the science can explain (the Barnum effect) but the instruments cannot.

But the sharper question is not whether an instrument is scientific. It is what the instrument is being used for.

A scientifically valid personality measure used to screen job applicants is still being misused - the evidence does not support that application, and the field's own leading researchers have said so. A less-than-scientific colour profile used to start a team conversation about working styles may do something genuinely helpful, despite its scientific limitations - because the value is in the conversation, not the measurement.

The problem with the people-testing industry is not that some instruments are unscientific, though some clearly are. It is that the whole architecture - questionnaire, type, normative database, certified interpreter, static report - was built in an era when those were the only tools available, and nobody has questioned whether the architecture itself is the constraint.

What the Appetite Actually Deserves

The appetite behind the personality-testing industry is not pseudoscientific. It is one of the most legitimate impulses in working life: the desire to understand how you work, how your colleagues work, and how to bridge the gap. Two million people a year do not take the MBTI because they are gullible. They take it because nobody has offered them anything better.

People do their best work when they play to their preferences - and preferences need no types, no categories, no certified interpreter, and no claim to predict what someone will become. Measured properly, how you work is a set of positions on continua: each one plain enough that you can read it yourself, each one revisited as you change, each one discussable without judgement. That is not a personality test. It is a different category - measurement in the here and now, not a verdict from the there and then. The ten-minute survey is free for any individual - and it will never call itself science while failing science's own tests.

Are personality tests pseudoscience?

Not all of them. The Big Five model meets scientific criteria: replicable factor structure, predictive validity, stability on retest. Type-based instruments (MBTI, DISC, Insights Discovery) sort continuous traits into discrete categories, produce unstable classifications on retest, and are marketed with technical language that implies a precision their evidence does not support. Whether that constitutes pseudoscience depends on the definition, but the instruments fail multiple standard scientific criteria.

Is the MBTI scientifically valid?

By most reasonable standards, no. Up to 50 per cent of takers receive a different type on retest within five weeks, the four dichotomies are imposed on continuous distributions without bimodal clustering, and the publisher states the instrument cannot predict job performance. The experience of feeling accurately described is explained by the Barnum effect, not by measurement precision.

What makes a personality test scientific?

Four criteria: test-retest reliability (same result on retesting), predictive validity (the score predicts something real), construct validity (the test measures what it claims, and those constructs correspond to real differences), and falsifiability (it must be possible for the theory to be proved wrong).

Are the Big Five personality traits scientific?

Yes. The five-factor structure has been replicated across thousands of studies, languages, and cultures over six decades. Individual traits predict life outcomes - conscientiousness predicts job performance across roles, neuroticism predicts burnout risk. The predictions are modest but real, replicable, and subject to ongoing scrutiny. The Big Five's scientific validity does not, however, mean it should be used for hiring - the predictive validity for job performance specifically is modest.

What is the Barnum effect?

A cognitive bias identified by Bertram Forer in 1949: people rate vague, flattering personality descriptions as highly accurate when they believe the description was written specifically for them. Forer's students rated an identical sketch - assembled from an astrology book - at 4.3 out of 5 for personal accuracy. The effect explains why personality profiles feel right regardless of their measurement quality.


The individual instruments examined: has Myers-Briggs been debunked · how accurate is Insights Discovery · how accurate is a DISC assessment. The wider selection evidence is in personality tests don't predict job performance. The full verdict table is in personality tests at work. What preferences are - and why they carry no judgement - is defined in what are work preferences.


Sources