AI Has No Taste Because It's Wired to Average Everything
Deep thoughts on AI and aspiration —— ByteDance Deep Think Circle
Open any AI content generation platform and you’ll notice something familiar: the output keeps growing, but it increasingly feels like it all came from the same person. The color palettes are those few combinations, the copy has that tone, the designs look like they rolled off the same assembly line. This phenomenon has a name: AI slop.
Most people attribute this to “models can only imitate, they have no creativity.” This explanation sounds plausible but says nothing—it blames the model without explaining why the model behaves this way.
I recently heard an explanation that breaks down the underlying logic cleanly. It came from Thais, founder of Taste Labs, a company dedicated to researching “why AI fails in subjective domains.” Her conclusion boils down to one sentence: AI has no taste because its root cause lies in the training signal—it constantly rewards the average.
Models Only Improve Where You Can Measure
Start with a background observation. We’re used to saying “models are great at writing code,” as if this were an inherent capability. Thais sees it the opposite way: this is actually a property of code itself.
Code can be executed and verified. Write it wrong, run it, and it throws an error; write it right, and tests pass. The feedback signal is clean, definitive, directly usable for training. Design doesn’t work this way. Ask someone to look at a page and they’ll likely say “something feels off,” unable to pinpoint what; ask someone else and their judgment might be completely different. Taste also shifts—what was trendy five years ago might be dated today. But 2 plus 2 equals 4 won’t become 5 just because times change.
She summarizes this difference in one line: capability follows measurability. Where you can measure, models improve; where you can’t, models stand still.
This framework is far more useful than debates about “whether AI has creativity.” It transforms the question from mysticism to structure: subjective domains are hard because “good” has no objective answer—it’s a judgment dependent on people and context that drifts over time. Mixing these two types of problems on one capability map was never going to work.
The Model’s Nature Is to Average
Next comes the sharpest concept she articulates, and the underlying mechanism explaining AI slop.
The way models are trained is essentially to predict “the next most likely content.” In math and code, the most likely answer happens to also be the best answer: 2 plus 2 equals 4 is both the most frequent answer and the only correct one—they coincide. In design and writing, a wide gap separates the most common answer from the best answer.
What truly makes creative work memorable almost always occurs at the distribution’s edge: the ad that deliberately breaks convention, the poster that deviates from norms, the color palette others wouldn’t dare use. Meanwhile, the model’s optimization objective naturally pulls output toward the center, toward “most common.” The result is output increasingly resembling the mean: not bad, but never striking—like something you’ve seen a hundred times before.
And this problem won’t automatically disappear with more parameters. Larger models just perform more precise mean estimation on larger datasets; they won’t spontaneously learn “when to break the rules.” When you complain that AI-generated content lacks soul, the real question to ask is: has the training signal ever rewarded “non-average”?
Decomposing Taste Into Verifiable Problems
Thais offers several actionable paths, with decomposition at the core.
“Is this design good?” can’t be answered, but “Does this page align with brand identity?” is far more concrete. Break down a brand: what colors does it use, what fonts, how is spacing defined, are animations fast or slow, is there texture? Each individual item has right and wrong. An AI-generated page can be scored item by item: Are the colors correct? Are the fonts from the brand guidelines? Does the spacing follow the design system? The value of decomposition is that it transforms a vague “good” into a series of verifiable sub-problems, and these sub-problems themselves can serve as ground truth for training.
There’s another more hidden trap in the data. The common approach is to get a bunch of people to rate things, then average the scores as training signal. But different people have genuinely different tastes—these disagreements are part of the real world. After averaging them out, you don’t get a better answer, just a mixture representing no one’s preference.
Her approach is to build a preference vector for each user type and bind it to the training data. What the model learns isn’t “a unified standard of good design” but “for people with this preference, what counts as good.” She also handles cases where experts disagree: when two experts give opposite assessments of the same design, if it’s a disagreement on style preference, the data is genuinely capturing preference diversity and should be preserved; only when they disagree on objective issues like alignment does it indicate the data itself is flawed.
One-line summary: In subjective domains, data quality far outweighs data quantity. Feed in averages, you get averages out; feed in expert data that preserves genuine diversity, and you might get something with taste.
What This Means for Product Builders
Combining these mechanisms, here are some takeaways you can apply directly.
First, don’t expect “waiting for bigger models” to solve taste problems. First figure out which parts of your product have objective right and wrong, and which are purely subjective. For the parts with right and wrong, confidently let models iterate; for purely subjective parts, either get your hands dirty decomposing “good” into rules, or diligently accumulate high-quality preference data. Feed vague problems to vague signals, and results will inevitably be mediocre.
Second, personalization needs to go deeper. Most products today do “one size fits one” by tagging at the behavior level: you clicked this, I recommend that. Personalization at the aesthetic level is different—it’s understanding a person’s judgment criteria, what they like in what contexts. Build this layer, and what your product can do shifts from “recommending what you might like” to “helping you generate and filter things that match your taste.”
Third, the average itself is also a business. Safe designs that please most people can make money by “never being wrong.” So the real question to answer is: do you want most people to find it acceptable, or a small group to find it stunning? Choose the latter, and you need to start accumulating differentiated taste data now. This path is expensive and slow with no shortcuts, because it directly determines the ceiling of model output.
Key points: AI slop’s root cause is training signals rewarding averages (most common ≠ best); code trains well because it’s verifiable, capability follows measurability; the solution is decomposing into verifiable sub-problems and using preference vectors to preserve disagreements; the average is also a business—first decide whether you want the safety of mass appeal or the brilliance of niche appeal.