![]()
Annie Murphy Paul's investigation exposed a $500 million industry built on pseudoscience. Two decades later, the same myths still shape how organisations hire, manage, and misunderstand their people.
In the summer of 2025, I picked up a battered copy of The Cult of Personality Testing by Annie Murphy Paul from a second-hand bookshop. I started reading it that afternoon and couldn't put it down. Not because it was telling me something I didn't already suspect, but because it had the courage to say what so many practitioners in the assessment industry have known for years but rarely articulate publicly: the most popular and commercially successful personality tests in the world are built on foundations that would embarrass the people who use them, if only they knew.
Reading Paul's meticulously researched investigation felt like watching Hans Christian Andersen's tale of the Emperor's New Clothes unfold in real time. The entire court - the Fortune 500 companies, the HR departments, the executive coaches, the career counsellors - all nodding along, admiring the fine garments, while a handful of researchers stand at the back of the crowd trying to point out that the emperor is, in fact, naked.
This article draws on Paul's work to examine why the personality testing industry continues to thrive despite decades of evidence against its core products, what that means for modern workplaces, and why the principles behind Sariio were designed specifically to break this mould.
A $500 Million Industry Built on Faith
The personality testing industry generates an estimated $500 million annually and grows at roughly 8% to 10% per year. Over 2,500 personality questionnaires are commercially available. The Myers-Briggs Type Indicator alone is administered to approximately 3.5 million people every year. The MMPI - the Minnesota Multiphasic Personality Inventory - is taken by an estimated 15 million Americans annually.
These are staggering numbers. And they would be entirely reasonable if the instruments behind them delivered what they promise: a reliable, valid, and actionable picture of who a person is and how they will behave. But Paul's investigation, drawing on decades of peer-reviewed research, reveals something far more uncomfortable. The most widely used personality tests fail the two basic scientific tests that any measurement instrument must pass: reliability (do you get the same result when you measure again?) and validity (are you actually measuring what you claim to measure?).
The industry persists not because the science supports it, but because institutions need a way to process human beings at scale. An "objective" test provides a defensible reason to hire one candidate over another, to promote one manager ahead of their peers, or to assign a struggling child to a particular educational track. The test result becomes a shield against the messy, subjective, politically fraught reality of actually getting to know someone. It is, as Paul describes it, the promise of "an X-ray of the soul."
The MBTI: Comfort Over Rigour
The Myers-Briggs Type Indicator is the most recognisable personality test on the planet. Used by 88% of Fortune 500 companies and administered in 115 countries, it classifies people into 16 personality types based on four dichotomies: Introversion/Extraversion, Sensing/Intuition, Thinking/Feeling, and Judging/Perceiving. The language of MBTI - "I'm an INTJ" or "She's an ENFP" - has become part of everyday workplace conversation, dating profiles, and social media bios.
Paul reveals something that would surprise most of the millions who have taken the test: neither Katharine Cook Briggs nor her daughter Isabel Briggs Myers had any formal training in psychology or psychometrics. They were self-taught enthusiasts who developed their system from an idiosyncratic reading of Carl Jung's psychological theories. Jung himself had reservations about reducing his complex theoretical framework to a simple classification system.
As many as 75% of test-takers receive a different personality type when they retake the Myers-Briggs after as little as five weeks. If a test cannot produce the same result twice, it cannot reliably measure anything.
The reliability data is damning. Research cited by Paul shows that between 39% and 75% of test-takers receive a different four-letter type upon retesting. Some studies place the type-change rate at the lower end of that range; others at the higher end. But even at the most generous estimate, four in ten people become a fundamentally different "personality type" within weeks. If a medical thermometer gave you a different temperature reading 40% of the time, you would throw it away. In personality testing, you build a global industry around it.
The MBTI's predictive validity - its ability to tell you anything useful about how a person will actually perform at work - is equally weak. Research consistently shows very low correlations between MBTI type and job performance, leadership effectiveness, or team outcomes. Some researchers have described the instrument as "irresponsible armchair philosophy" that has been dressed up as science through decades of aggressive marketing by its publisher.
Yet the test endures because, in Paul's analysis, it is "non-judgemental." There are no bad types. Every result is framed positively. This makes it feel safe and affirming - qualities that drive commercial adoption but have nothing to do with scientific accuracy. In some extreme corporate cases, managers have been pressured to conform to a specific "type" or face consequences, turning a tool designed for self-understanding into a mandate for conformity of identity.
The MMPI: A Clinical Tool Applied Where It Was Never Meant to Go
The Minnesota Multiphasic Personality Inventory occupies the opposite end of the spectrum. Where the MBTI is friendly and affirming, the MMPI is clinical and diagnostic. Developed in the 1940s by psychologist Starke Hathaway and neuropsychiatrist J.C. McKinley, the MMPI was designed to identify psychiatric conditions - depression, schizophrenia, paranoia - in clinical populations.
Paul's investigation reveals the extraordinary story of the test's construction. The baseline for "normal" human psychology - the standard against which millions of people would be measured - was established by studying a group known as the "Minnesota Normals." These were not a carefully selected, nationally representative sample. They were, for the most part, relatives and visitors of patients at the University of Minnesota hospitals: predominantly white, rural Minnesotans from a specific time and place. The world's standard of psychological normality was calibrated against a convenience sample from 1940s rural America.
The test's 504 true-or-false items include questions about bowel movements, religious beliefs, family relationships, and sexual thoughts - items that were designed to detect clinical pathology, not to measure workplace suitability. Yet the MMPI has migrated far beyond clinical settings. It is routinely used in pre-employment screening for law enforcement, in child custody disputes, and in forensic evaluations. A tool built to distinguish psychiatric patients from hospital visitors is now being used to decide who gets hired, who keeps their children, and who goes to prison.
The cross-context problem
The MMPI was designed for clinical diagnosis. The MBTI was designed for self-understanding. Neither was designed to predict job performance, team dynamics, or leadership potential. Yet both are used daily for exactly those purposes. This is the core dysfunction that Paul identifies: the personality testing industry takes tools built for one context and applies them to another, with no evidence that the transfer is valid.
The Big Five: Better Science, Still Limited Answers
Academic psychology's response to the MBTI and MMPI has been the "Big Five" model: Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism (OCEAN). Unlike the MBTI's arbitrary types, the Big Five measures personality on a spectrum. Unlike the MMPI, it was built through decades of factor-analytic research rather than clinical convenience. It is, by consensus, the most scientifically defensible framework for personality that mainstream psychology has produced.
But Paul's analysis - and the meta-analytic research she cites - demonstrates that even the Big Five has severe limitations when applied to workplace decisions. The most comprehensive reviews of decades of studies show that Big Five personality traits account for only about 5% to 7% of the variance in job performance. The strongest individual predictor, Conscientiousness, correlates with job performance at just 0.10 to 0.15 - a statistically meaningful but practically tiny relationship.
of job performance variance explained by personality traits
Meta-analytic reviews show personality tests predict performance at roughly the same level as unstructured interviews - a method most experts consider ineffective.
To put this in perspective: knowing someone's Big Five profile tells you almost nothing about how well they will actually do their job. Personality tests, even the best-validated ones, perform at roughly the same predictive level as unstructured interviews - a method that most industrial psychologists consider one of the weakest selection tools available. Yet the testing industry continues to market these instruments as though they offer deep, actionable insight into human potential.
The Human Cost
Paul's most powerful chapters move beyond psychometric statistics to document the real-world consequences of institutional faith in personality testing. The harm falls disproportionately on the people least able to resist it.
Children Sorted Before They Start
In educational settings, children as young as five are subjected to personality assessments that assign them colours, types, or behavioural categories. These labels - intended to help teachers differentiate instruction - often become self-fulfilling prophecies. A child told they are a "blue" (helper) or an "orange" (energiser) begins to see themselves through that one-dimensional lens. Paul argues that this process "flattens and quashes" the natural complexity of a developing identity, replacing genuine curiosity about who a child might become with a pre-assigned category.
Workers Filtered by Algorithm
In hiring, "honesty" and "integrity" tests - many of which are personality tests rebranded - have false positive rates so high that large numbers of honest, capable candidates are rejected because they do not fit a specific psychometric profile. These individuals often have no legal recourse, because "personality" is not a protected class in the way that race or gender are. The result is a hidden layer of discrimination that operates beneath the surface of otherwise lawful hiring practices.
Neurodivergent People Screened Out
The ethical implications are particularly acute for neurodivergent individuals. Tests that measure "soft skills" or "agreeableness" often penalise the specific features of autism, ADHD, and other neurological differences - treating them as personality deficits rather than cognitive variations. A highly skilled candidate who struggles with the performative aspects of an "extraversion" scale may be screened out by an automated algorithm before ever reaching a human interviewer. Paul's work, combined with more recent advocacy, frames this as a form of systemic exclusion that uses scientifically questionable tools to reinforce barriers against vulnerable populations.
Human beings are complicated, contradictory, and changeable across time and place. When we reduce them to a type, a score, or a four-letter code, we miseducate our children, mismanage our companies, and misunderstand ourselves.
The Narrative Alternative
Paul does not end with a critique alone. She points toward a fundamentally different way of understanding people - one rooted in narrative rather than typology. In this view, personality is not a fixed internal essence that can be captured by a questionnaire. It is a story that unfolds over time, shaped by context, relationships, and experience. A person might be assertive in a professional setting where they feel confident, and cautious in a personal relationship where they feel uncertain. They might lead with analytical thinking in one role and shift toward intuitive decision-making after a career change. These shifts are not inconsistencies or measurement errors. They are how human beings actually work.
Research in situationism - the study of how context shapes behaviour - supports this view. Behaviour is driven as much by the environment as by internal traits. The person you are in a Monday morning leadership meeting is not the same person you are at a Friday evening dinner with friends. Any instrument that claims to capture "who you really are" in a single sitting and stamp that result as permanent truth is making a claim that science does not support.
From type to narrative
Paul and the researchers she cites advocate for understanding personality as "an entire rainbow rather than a single colour." When institutions replace personal narratives with psychometric labels, they flatten the very human diversity they claim to value. The question is not "what type are you?" The question is "how do you prefer to work, right now, in this context - and how has that changed?"
Why Sariio Was Built to Break This Mould
I did not build Sariio because the world needed another personality test. I built it because the world needed something fundamentally different from one. Every design decision in the MAPS framework was shaped by the failures that Paul documents so thoroughly - and by the principles that her investigation points toward.
Preferences, not personality
MAPS does not claim to measure who you are. It measures how you prefer to work - right now, in your current role and context. The distinction matters. Personality implies permanence. Preference implies choice, agency, and the capacity for change. You are not an "INTJ" or a "High C" or a score on a neuroticism scale. You are a person with work preferences that evolve as you grow.
Dynamic, not fixed
Traditional tests produce a result and treat it as a settled truth. MAPS is designed for retesting. Take the survey now. Take it again in three months. The shift in your profile is not an error - it is information. It shows how you are developing, adapting, and growing in response to your work environment. This is precisely the narrative understanding of personality that Paul advocates: a living story, not a frozen label.
Context-aware, not context-blind
The MAPS survey asks about work preferences in work contexts. It does not extrapolate from clinical pathology scales. It does not ask about your bowel movements or your relationship with God. It measures twelve specific preference pairs - trust vs wariness, procedures vs options, reflection vs doing - that are directly relevant to how people collaborate, communicate, and deliver in professional settings.
Insight for development, not gatekeeping
Paul's most devastating critique is that personality tests are used to exclude - to filter out job candidates, to sort children, to label people as unsuitable. MAPS is designed for the opposite purpose: to create mutual understanding. When two colleagues compare their MAPS profiles, they see where their preferences align and where they diverge. The goal is not to decide who is "right" but to make the invisible visible, so that collaboration can be intentional rather than accidental.
What Paul's Investigation Means for You
If you are a coach, an HR professional, a team leader, or simply someone who has ever been asked to take a personality test at work, Paul's book is essential reading. It will change how you think about the instruments that institutions use to evaluate people - and it will sharpen your ability to distinguish between tools that serve people and tools that merely sort them.
The personality testing industry survives because most people never ask the questions that Paul asks. What is the test-retest reliability of this instrument? What population was it normed on? What percentage of job performance variance does it actually predict? These are not hostile questions. They are the minimum standards that any measurement tool should meet. When you ask them about the most popular personality tests, the answers are deeply unsatisfying.
The alternative is not to abandon all assessment. It is to demand better. To insist on instruments that measure what they claim to measure, that respect the complexity of human experience, that evolve as people evolve, and that are used to develop rather than to exclude. That is the standard we hold ourselves to at Sariio. It is the standard that Paul's investigation makes urgently necessary.
Key takeaways
- The personality testing industry generates $500 million annually, despite most instruments failing basic tests of scientific reliability and validity
- Up to 75% of MBTI test-takers receive a different personality type upon retesting - the most widely used test in the world cannot produce consistent results
- The MMPI's standard for 'normal' psychology was calibrated against 1940s rural Minnesotans - a convenience sample used to judge millions
- Even the Big Five, academia's best-validated model, explains only 5-7% of the variance in job performance
- Personality tests disproportionately harm children, neurodivergent individuals, and honest job seekers who do not fit narrow psychometric profiles
- The alternative is no assessment - it is a better assessment: preference-based, dynamic, context-aware, and designed for development rather than gatekeeping