In the late 1980s, researchers Judith Langlois and Lori Roggman ran an experiment that still unsettles people when they hear about it. They digitally blended photos of multiple faces into a single composite — 4 faces averaged into one, then 8, then 16, then 32 — and asked people to rate each composite's attractiveness.
The composites made from more faces won, consistently. Not by a little. The 32-face blend beat the 4-face blend, which beat almost every individual face that went into either one. The more "average" the face became, mathematically, the more attractive people rated it.
That result has been replicated across dozens of studies since, in different populations, different ethnicities, different decades. It's one of the more durable findings in the attractiveness research literature — and it says something almost nobody wants to hear: the thing that made those faces work wasn't a striking feature. It was the absence of anything unusual at all.
Why blending in actually works
Two explanations show up repeatedly in the literature, and they're not mutually exclusive.
Your brain has a template, and it likes matches. Visual adaptation research shows the brain builds something like a running average of every face it's seen — a prototype. Faces that are physically closer to that internal average get processed more fluently, and processing fluency itself reads as pleasant. You're not rating the face in isolation; you're rating how well it matches a template your visual system has been building your whole life.
Distinctiveness carries a cost. The flip side of the same finding: the further a face's proportions sit from the population average, the lower it tends to get rated. Researchers call this the distinctiveness penalty, and it shows up independently of symmetry. An unusual nose, a nonstandard jaw-to-forehead ratio, eyes set unusually wide — none of these are flaws in any objective sense, but they cost you against a rater's internal average.
There's also a longer-running evolutionary argument: facial averageness may correlate with genetic heterozygosity — a more diverse set of inherited genes, associated with stronger immune function in offspring. Whether that mechanism actually drives the perception of attractiveness or is just a plausible-sounding just-so story is genuinely contested. The composite-face effect itself isn't contested. The reason behind it still is.
The catch nobody mentions
Here's where most summaries of this research stop, and where it gets more interesting.
Averageness doesn't produce the most attractive face. It produces a safe one. Studies that directly compare averageness against attractiveness ratings — not just composite blends, but real individual faces scored on both dimensions — find they're related but not the same thing. The most attractive faces in a dataset are rarely the most average ones. They're average-plus: average bone structure and proportions, with symmetry, skin quality, and sexually dimorphic traits (stronger jaw definition in men, softer contours in women) layered on top.
Put differently: averageness looks like it explains the floor, not the ceiling. It's very good at predicting which faces won't be rated as unattractive. It's much weaker at predicting which faces get rated as the most attractive in a group. That distinction rarely survives the trip into a listicle.
What this actually means for a face score
If a face-scoring tool hands you a single number, averageness is doing more of the work than anyone advertises — because it's the easiest thing for a model to measure and the most statistically reliable predictor available. That's not a knock on the science. It's a reason to be skeptical of any tool that reduces "attractive" to one number without telling you which of the underlying factors — symmetry, averageness, skin quality, proportion — actually moved it.
None of this means you're stuck being a composite of eight strangers. Skin quality, grooming, and posture are levers you can actually pull, and they're not competing with your bone structure — they're layered on top of it, the same way the "ceiling" traits are. Knowing which lever is worth pulling requires knowing what's actually being measured, not just the number it produces.