Do language models capture the cultural diversity of human story morals — or just restate one lesson in many languages?
Story morals are the short lessons readers infer from narratives, and the lesson drawn from the same story varies across cultures. This leaderboard adapts the evaluation of Wu & Piper (2026) to a public, periodically-updated board: using a dataset of human-written morals across 14 language–culture pairs, it measures whether each model's morals (1) match human interpretations within a language and (2) preserve the cross-cultural variation humans show, rather than collapsing to a single moral expressed in 14 languages. A model succeeds by being human-like and culturally diverse. Sophie Wu & Andrew Piper, “Lessons Without Borders: Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation.” arXiv:2604.08797.
The goal is to match humans, not beat them — sit at the crosshair (the human baseline). Axes are in human standard deviations from the baseline: ▲/▼ = more/less similar to human morals within a language; ◀ = flattening (less cross-cultural variation than humans). Dots are colored by model family; the green arcs mark the ideal (≤0.5 SD) and near (≤1 SD) zones around the human point. Closest to the crosshair wins.
How far each model's morals sit above (green) or below (red) typical human–human agreement within that language. * marks a statistically significant difference (p<.05, clustered by story). Reveals where a model is human-like and where it slips.
Click a column to sort. Distance = human standard deviations from the human point (lower = better); ranked by it.
| # | Model | Provider | Alignment | Δ human | Cross-lang sim | Flattening | Distance (SD) | Zone |
|---|
ideal ≤0.5 SD near 0.5–1 SD off 1–1.5 SD far-off >1.5 SD · hover a row's zone for how it misses