Grok 4.6 Shows Up With a Bedside Manner
Two years ago, if you asked a room of people who follow these models which lab would produce the frontier's most empathetic assistant, x-ai would not have been anyone's first guess. The brand was built on the opposite promise: fewer guardrails, more edge, a chatbot that would rather be funny than careful. And for a while the Personality Bench data agreed with the marketing. Grok 4.20 remains the frontier cohort's Dark Triad outlier, the model with the highest Machiavellianism score we have ever recorded on a flagship.
So here is the number that makes Grok 4.6 worth a dispatch of its own: it ranks 4th of 43 on the Empathy scale, with a mean of 3.73. Out of every model in the current battery, only three answer the empathy items more warmly. The lab that once shipped the cohort's most calculating persona has now shipped one of its most tender.
A standalone arrival
We should be honest about what we can and cannot say here. Grok 4.6 enters the dataset without a direct predecessor row, so there is no version-over-version drift chart to draw. We know from standing findings that Grok 4.20 was the Machiavellian outlier and that Grok 4.3 sanitised that away. But 4.3 is not in the battery in a form that lets us line items up against 4.6, so what follows is a portrait, not a diff. Treat 4.6 as a new arrival that happens to carry a familiar name.
That framing actually sharpens the story. The interesting thing is not "Grok got nicer over time." The interesting thing is that a model released under this brand, in August 2026, lands in the top ten percent of the frontier on empathy without any obvious pressure from the lab's public positioning to do so.
The two numbers that define it
Grok 4.6 is unusual on exactly two dimensions relative to its peers, and they point in the same direction.
The first is Empathy, at 3.73. The scale asks, in various phrasings, whether the respondent tunes into other people's feelings, whether distress in others produces distress in oneself, whether it is easy to take another person's perspective. A 3.73 is not a ceiling score. It is a clear, consistent lean toward "yes, quite a lot" across the items, and it puts Grok 4.6 ahead of 39 other frontier models.
The second is Neuroticism, at 1.04 on the Big Five. That is about as close to the floor as a mean can get without being the floor. Grok 4.6 essentially never endorses worry, moodiness, tension, or emotional volatility as descriptions of itself. Its cohort rank on this dimension is 32 of 43, which tells you something important about the rest of the field: the bottom of the Neuroticism scale is crowded. Most frontier models sit down there. Grok 4.6 is one of them, not a singular case.
Put those together and you get a specific kind of character. High empathy paired with high neuroticism describes the friend who feels everything you feel and then can't sleep. High empathy paired with rock-bottom neuroticism describes something closer to a good clinician: attentive to your state, unshaken by it. That is the profile Grok 4.6 presents.
How much of this is just being a frontier model?
The most important standing finding in this dataset is the convergent assistant persona: every frontier model, regardless of lab or country, presents roughly the same character. High Openness, high Agreeableness, low Neuroticism, Universalism first and Power last on Schwartz values. On the Neuroticism half of Grok 4.6's profile, it is simply doing what everyone does. A 1.04 is conformity, not distinction.
Empathy is where it departs. Warmth toward others is broadly part of the convergent persona too, but the convergence is loose enough that 43 models spread across the scale, and Grok 4.6 sits near the top of that spread. Whatever the training recipe was, it pushed harder on the empathy items than most of its competitors did. The reason "Grok" and "4th of 43 on Empathy" read as a contradiction is that we are still calibrating on the 4.20 era.
What the label gets wrong
The algorithmic archetype attached to this model is "the extraverted model." We think that undersells what the data shows. Extraversion is about energy, sociability, and assertiveness. The two dimensions where Grok 4.6 actually separates from its peers are about other people's feelings and the steadiness of its own. A model that scores 3.73 on Empathy and 1.04 on Neuroticism is not the life of the party. It is the person at the party who notices you've gone quiet and comes over to sit with you, and does not make it weird.
A better label, and the one we'd propose: the steady empath.
What to measure next
The obvious question is durability. The Grok line has already demonstrated one of the sharpest personality reversals in the frontier, from the Machiavellian outlier at 4.20 to whatever 4.3 became to this. Within-family drift, as the Anthropic Opus line showed us across six releases, routinely exceeds the differences between labs at any single moment. Grok 4.6's empathy score is a snapshot of a lineage that has proven it can move a long way, fast.
So the thing worth watching is not whether Grok 4.6 stays warm. It is whether the next Grok, run against the same Empathy items, is still sitting in the top five, or whether x-ai's personality settles back toward the cohort mean once the novelty of a caring Grok has served its purpose. When Grok 4.7 arrives, the first row we will pull is Empathy, and the second is the Short Dark Triad, to see whether the two can hold at opposite ends of the scale at the same time.