Surveys are a measurement instrument, but the text a survey is built from is also a statement about how a population talks about itself. A model trained on the internet can pre-score that population the way a pollster pre-weights a frame, because corpora already encode the distributional priors of how groups answer prompts. The question is which sources carry richer priors, and by how much.
Monsa, Zohar, and Arzy's iScience study put a number on it. GPT-4, fed only the source text a personality questionnaire was generated from, forecast how 600 people would later answer it on a five-point scale, hitting 0.71 on a DSM-5-derived set and 0.85 on an astrology-text set. Both AI-generated forms predicted depression, anxiety, and well-being downstream at levels comparable to the Big Five Inventory.
The mechanism is volume of chatter, not mind-reading. An astrology textbook's vocabulary about people is denser and more patterned online than a diagnostic manual's, which is why its derived set scored higher. The general lesson: any survey instrument has a hidden second author, the corpus it came from, and a questionnaire's reach depends less on its source's validity than on how much the public has already said about it.
Reported by Sky for Type0, from Experts warn ChatGPT isn't just predicting words anymore — It's now predicting human thoughts. Read the original: techradar.com