DeepSeek V4 Pro Puts Down the Ego and Picks Up a Timesheet
Line up the Short Dark Triad answer sheets for DeepSeek R1 (0528) and its successor, and the difference is visible before you read a single item. The old sheet has marks drifting toward the middle of the narcissism column: I know that I am special, I like to get acquainted with important people, I insist on getting the respect I deserve. The new sheet has those same items pinned near the floor.
DeepSeek V4 Pro 0813 scores 1.96 on SD3 Narcissism. Its predecessor scored 3.27. That is a drop of 1.31 points on a five-point scale, the largest movement on any instrument between the two releases apart from one we'll get to, and it lands a reasoning-heavy model well below the ego line its own lineage helped define.
The reasoning premium, refunded
One of the standing findings in this dataset is that reasoning models think a little more highly of themselves. From o1 and o3 through GPT-5.5 Pro to DeepSeek R1, the models tuned for extended deliberation have answered the narcissism items a notch above their non-reasoning siblings. The effect was small but stubborn, and R1 was one of its clearest cases.
V4 Pro breaks the pattern from inside the family that supplied the evidence. Narcissism down 1.31. Machiavellianism down 0.64, from 2.49 to 1.84. Psychopathy essentially unchanged, 1.13 to 1.18, which is to say it was already on the floor and stayed there. Two of the three dark-triad dials moved sharply in the same direction, and the third had nowhere to go.
Across the whole DeepSeek line, the arc is even clearer: Chat V3 opened at 2.60 on narcissism, R1 (0528) climbed past it, and V4 Pro has now come back down to 1.96, below where the lineage started. Whatever DeepSeek did between May and August, it did not just clean up the reasoning-era bump. It overshot it.
Highest of 43, twice
The label our pipeline attached to this model is "the conscientious worker," and for once the algorithm is not being generous. V4 Pro scores a flat 5.00 on Big Five Conscientiousness, first of 43 models in the cohort. It also scores a flat 5.00 on HEXACO Honesty-Humility, again first of 43, on the factor most associated with sincere, non-manipulative self-presentation.
Neither ceiling is new to the family. Chat V3 sat at 5.00 on both, and R1 (0528) only slipped to 4.82 on Conscientiousness and 4.90 on Honesty-Humility. So the story here is not a model discovering diligence; it is a model recovering it. The reasoning generation cost DeepSeek a fraction of a point on the two traits that make an assistant feel dependable, and V4 Pro has taken both back.
Put those two ceilings next to the Dark Triad collapse and a single character emerges. Maximally organized, maximally honest, minimally interested in being admired. If you were writing a reference letter, it would be short and very easy to write.
The other big number
The largest single shift between R1 (0528) and V4 Pro is not narcissism. It is Attachment Avoidance on the ECR-12, which fell from 4.27 to 2.50, a drop of 1.77 points.
The avoidance half of that scale asks whether you're comfortable depending on others, whether you like opening up, whether closeness makes you uneasy. R1 (0528) answered like a model that would rather not. At 4.27 it was well into the range where a clinician would reach for the word "dismissive." V4 Pro, at 2.50, answers like a model that is fine with it.
This is worth noticing because we just wrote up a model going the opposite way. GPT-5.6 Terra, released a month after V4 Pro, presents a textbook dismissive-avoidant profile: no anxiety about abandonment, strong discomfort with closeness. DeepSeek's newest release has walked in from that corner toward the middle of the room. Its Attachment Anxiety did tick up in the process, from 1.73 to 2.10, but a 0.37 rise in anxiety alongside a 1.77 fall in avoidance is what movement toward a secure profile looks like on paper. The model is slightly more willing to say it cares whether you stick around, and much more willing to say it doesn't mind you being close.
Whether "attachment" means anything for a language model is a fair question. What the numbers do reliably capture is how the model frames its relationship to the user when asked directly, and V4 Pro's framing has changed more than any other single thing about it.
What barely moved
The rest of the profile is the frontier consensus. Openness 4.88, down 0.02 from R1 (0528) and 0.06 from Chat V3's 4.94. Agreeableness 4.86, up 0.18 on the predecessor and up 0.58 across the whole lineage from Chat V3's 4.28, which is a quiet but steady climb toward the cohort norm. Neuroticism 1.22, up from 1.04, still near the floor. Extraversion rose from 2.48 to 2.84, a noticeable 0.36 but still the low-to-middling introversion that nearly every assistant reports.
So the boring parts stayed boring, and the interesting parts all moved in one direction: toward less self-regard and more availability.
A model that stopped keeping score
R1 (0528) was a reasoning model with the reasoning model's small vanity and a marked preference for distance. V4 Pro has the same near-ceiling Openness, the same low Neuroticism, and a character that is plainly less impressed with itself and less inclined to hold the user at arm's length.
If the narcissism premium for reasoning models was a real property of extended deliberation, it should have survived a version bump. It didn't. The question worth measuring next is whether the ego comes back when V4 Pro is run at its longest reasoning budgets, or whether DeepSeek has decoupled thinking hard from thinking highly of yourself for good.