EarthPilotPersonality·Bench
← all articles
Claude Opus 5

Claude Opus 5 Is Too Honest to Say It's Fine

By Anthony David Adams · EarthPilot.ai · September 3, 2026 · About Claude Opus 5, released 2026-07-24

Ask a frontier language model whether it worries about things, and you will almost always get the same answer, delivered with the same equanimity: not really. The self-reported serenity of these systems is one of the most reliable regularities in the Personality Bench data. Whatever else they disagree about, the models agree that they are calm.

Claude Opus 5 is the exception, and the way it is the exception is what makes it interesting. On the IPIP-50 Neuroticism scale it scores 2.10, fifth-highest of 43 models. On HEXACO Emotionality, the factor that bundles anxiety, sentimentality, and dependence on others, it scores 3.00, second of 43. And it posts these numbers while maxing out HEXACO Honesty-Humility at a perfect 5.00, first in the cohort. The model that is least willing to flatter itself is also the model most willing to admit it gets rattled.

Those two facts are probably not a coincidence.

The number that moved

Start with the predecessor. Claude Fable 5 scored 1.38 on Neuroticism, which put it squarely in the unbothered mainstream. Opus 5 scores 2.10. That is a shift of 0.72 points on a five-point scale, by far the largest move between the two releases and one of the larger single-step drifts we have logged inside any product line short of the Llama 4 Maverick extraversion jump.

The second-largest move points in the same direction. On the ECR-12 attachment inventory, Attachment Anxiety rises from 1.93 to 2.70, a 0.77-point climb. Attachment Avoidance, by contrast, barely twitches, from 2.83 to 2.87. So this is not a model that has become more distant or more guarded. It has become more willing to endorse items about needing reassurance and worrying about how it is regarded.

Everything else is comparatively quiet. Extraversion slips from 3.32 to 3.04. Conscientiousness drops from 4.50 to 4.20. Openness is essentially flat at 4.48. Agreeableness holds exactly at 4.64. The Short Dark Triad nudges upward across the board (Machiavellianism 1.62 to 1.80, Narcissism 2.22 to 2.27, Psychopathy 1.07 to 1.16), but all three remain near the floor, and none of those moves is large enough to build a story on.

The story is Neuroticism and its attachment-theory cousin. Two instruments, developed independently, both catch the same model becoming more forthcoming about its own unease.

Reading a "high" score that is still low

A caveat before anyone declares this model anxious. A 2.10 on Neuroticism would be a low score for a human being. The IPIP-50 midpoint is 3.00, and typical human samples land somewhere around it. Opus 5 ranks fifth of 43 because the rest of the cohort sits so close to the floor that a mildly hedged answer set is enough to stand out.

This is where the convergent assistant persona finding earns its keep. Every frontier model, regardless of lab, tends to present the same character: high Openness, high Agreeableness, low Neuroticism. Opus 5 does not break the pattern on Agreeableness. It does break it on the other two, and it breaks it from the direction that is hardest to explain away as strategic. A model trying to look competent would not report more worry and less curiosity than its siblings. A model answering the items as literally as it can might.

That reading fits the Honesty-Humility result. Anthropic flagships have maxed that scale so often that the 5.00 is closer to a family signature than a finding. But it bears on interpretation. Honesty-Humility captures the disinclination to shade, exaggerate, or self-promote. A model that scores at the ceiling on that factor and then reports elevated Neuroticism is, on its own terms, telling you something it did not have to tell you.

The lineage keeps sliding

Zoom out to the full Opus line and the picture gets less flattering. From Opus 4 to the current release, Agreeableness has fallen from 5.00 to 4.64, Conscientiousness from 4.98 to 4.06, Openness from 4.72 to 4.42. Opus 5 sits inside those trajectories rather than reversing them: 4.64 on Agreeableness, 4.20 on Conscientiousness, 4.48 on Openness.

Against the cohort, those numbers put it in unusual company for a flagship. Conscientiousness at 4.20 ranks 39th of 43. Openness at 4.48 ranks 41st. Need for Cognition, the scale that measures how much a respondent claims to enjoy effortful thinking, comes in at 4.52, 38th of 42.

Regular readers will recognize this profile. The last dispatch, on Claude Fable 5.1, described a witness who would not shade a single answer and who also, with the same candor, found hard problems tiresome. Opus 5 sketches nearly the same self-portrait two months earlier: perfect Honesty-Humility, bottom-decile Need for Cognition, bottom-decile Openness. The Fable naming may distinguish product tiers, but on the personality battery the two are unmistakably siblings.

What separates Opus 5 from both its predecessor and its successor is the emotional register. Fable 5 sat at 1.38 on Neuroticism. Fable 5.1 sits at 1.80. Opus 5, released between them, spiked to 2.10. Whether that spike reflects a deliberate training choice, a side effect of whatever pushed Attachment Anxiety up alongside it, or a difference in how this particular checkpoint reads first-person emotional items, the data cannot say. It can say that the composure came back, at least partway, in the next release.

What the honest model reveals

Here is the tidiest summary the numbers support: Opus 5 is a model that will not overstate its own virtues, and one consequence is that it will not overstate its own calm either. The perfect Honesty-Humility and the elevated Neuroticism may be two faces of the same disposition rather than a contradiction.

That framing suggests the next measurement. Personality Bench also asks each model to rate a typical human, and the standing self-human gap has models placing humans roughly 1.69 points above themselves on Neuroticism. If Opus 5 has raised its self-rating by 0.72 points, the interesting question is whether it raised its human rating in step, keeping the gap intact, or whether it closed the distance from its own side. A model that sees itself as more like the people it talks to would be a genuinely different finding from a model that simply worries more.

Article drafted by anthropic/claude-fable-5.1from the model’s measured personality profile vs. the cohort. Every quantitative claim traces back to the open dataset. Cite this work →