While users of ChatGPT, Gemini, and other large language models (LLMs) frequently praise their intelligence and analytical capabilities, a provocative question remains: can AI be clairvoyant?
Researchers at the Hebrew University of Jerusalem (HUJI) have pioneered a method to generate personality assessment questionnaires using ChatGPT, drawing from any given English source text.
To validate this approach, the team applied it to two contrasting sources: the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5)—the authoritative guide used by mental health professionals to diagnose and describe mental health conditions—and, as an intentionally unconventional control, an astrology textbook.
The research team found that ChatGPT could not only create and validate these questionnaires but also accurately predict population-level responses prior to the actual surveys being conducted.
Testing AI with Personality Assessments
The study, conducted by Dr. Rotem Monsa, Prof. Aviv Zohar, and Prof. Shahar Arzy from the Hebrew University-Hadassah Medical School and the HUJI Faculty of Computer Science, was published in the Cell Press journal iScience under the title “Generating and analyzing personality questionnaires using large language models.”
Publicly accessible LLMs like ChatGPT are trained on vast datasets compiled from trillions of human-language inputs across the internet, including websites and social media.
This extensive training has led researchers to wonder whether LLMs achieve an expert-level understanding of human language and whether personality traits are inherently embedded within their algorithms.
“Since personality traits are reflected in language, LLMs may have naturally learned the structure of human personality as a byproduct of their training,” suggested lead author Dr. Rotem Monsa, a postdoctoral researcher in HUJI’s medical neurosciences department, in an interview with The Jerusalem Post.
“Thus, even though these models are not explicitly taught psychology or personality theories, these concepts are already deeply embedded in the language data they consume,” she added.
To test how effectively LLMs can assess human personality, the researchers utilized GPT-4—a large multimodal model that demonstrates human-level performance on various professional and academic benchmarks—to generate two distinct personality assessment questionnaires.
First, they used excerpts from the DSM-5 as source text, operating under the assumption that personality traits are localized along a personality disorder continuum.
ChatGPT generated a questionnaire featuring statements based on the DSM-5’s descriptions of personality disorders.
For instance, drawing from the section on paranoid personality disorder, the questionnaire asked participants to rate their agreement on a scale from 1 (strongly disagree) to 5 (strongly agree) with statements like “often suspects others’ motives” or “finds it easy to trust people.”
A second questionnaire, serving as a control, was generated from an astrology textbook, creating similar statements based on the text’s assignment of personality traits to zodiac signs.
“We intentionally selected texts that describe human personality in rich detail while occupying opposite ends of the scientific grounding spectrum,” Monsa explained.
“The DSM-5 has been refined over decades of clinical research and stands as the global standard diagnostic manual in clinical psychiatry. In contrast, the astrology text is highly culturally influenced but lacks scientific validation.”
Once the LLM-generated questionnaires were ready, they were administered to 600 participants alongside the Big Five Inventory (BFI)—the most widely validated personality questionnaire currently in use, which is based on the premise that personality traits are encoded in language—to assess their utility.
Essentially, the researchers sought to determine whether an LLM could absorb sufficient psychological patterns from language to predict how humans would answer questions the model itself had just created.
Results from the DSM-5-sourced questionnaire demonstrated high internal consistency within personality clusters. Traits that typically correlate in the real world, such as dependency and avoidance, also showed strong correlations in the participants’ responses.
“Crucially, these results mirrored those obtained from the BFI, validating the questionnaire’s strength in measuring real-life patterns of human psychology,” said Monsa.
“As predicted, the astrology questionnaire showed weak internal consistency across traits. Our data suggest that astrological elements do not reflect coherent psychological dimensions.”
Perhaps the most surprising finding, however, was ChatGPT’s ability to predict how participants would respond to the questionnaires before they were even administered.
For both questionnaires, ChatGPT predicted the mean responses and inter-question correlations in advance with high real-world accuracy. This suggests that LLMs possess an innate, population-level understanding of personality dynamics, according to Monsa.
When asked what it means for ChatGPT to have an “understanding” of personality, she explained that “it was trained on vast amounts of text. It doesn’t comprehend independently; rather, it learns from statistical patterns.”
While ChatGPT lacks genuine human understanding, she noted, “even I—who know Chat is not human—often refer to it as ‘he’ because the responses are remarkably human-like. Over the three years since we began this research, LLMs have learned an immense amount.”
Regarding how ChatGPT could predict responses before seeing individual answers, Monsa clarified that “it wasn’t about predicting one person, but a population-level response. We were very surprised that Chat performed so well.”
“When we started, the models were less advanced. Obviously, the model wasn’t ‘born’ with psychological knowledge; it wasn’t explicitly trained in it. Instead, this capability became ‘natural’ or ‘innate’ through its learning process.”
The team ensured ChatGPT could not access actual participant responses before making its predictions because the questionnaires did not exist online—the primary source of the model’s training data, Monsa noted.
“We ran the prediction multiple times to check if ChatGPT produced consistent predictions, and they were highly similar each time,” she added.
The astrology questionnaire served as a negative control because personality cannot be determined by astrology.
“If we fed ChatGPT another pseudoscientific text—such as a book claiming handwriting analysis reveals personality—it would perform successfully similarly to our research. We even tried non-personality texts, like menus and kitchen organization guides,” Monsa recalled.
The study’s findings on LLM-generated personality assessments could eventually be applied to evaluate human patients.
“It will take time, and there are risks,” Monsa acknowledged. “It must be thoroughly analyzed by psychologists, but I am confident it will eventually be used for diagnosis or even treatment following psychological validation. It would be significantly more cost-effective and particularly beneficial in regions where trained psychologists are scarce.”
She conceded that LLMs could inadvertently generate questions that are culturally biased, stigmatizing, or psychologically misleading, given that they are trained primarily on English-language, Western materials.
“Our results may not be directly applicable if these methods are repeated in different languages and cultures, as LLMs are trained mostly on English texts from Western cultures,” Monsa noted.
“However, if applied to Hebrew-speaking Israelis, the results should be relatively close, as Israelis are culturally very Western. An LLM could even learn culturally specific personality concepts that human psychologists might overlook.”
Also Read
- China’s new moon mission could unlock secret of lunar ice: Why that matters
- The Long-Term Cost of Holding Cash: What $10,000 Could Lose in 10 Years
- Community Mourns Loss of Vibrant 18-Year-Old Lily Hooper After NSW Hiking Tragedy
- Emerging Asian Powers Accelerate Ambitions Along Melting Arctic Corridors
