On the part of design that AI cannot quite reach
There is an old line attributed to Justice Potter Stewart, who, when asked to define obscenity, gave up and said: I know it when I see it.
Designers have spent the better part of a century saying the same thing about taste, with the same mixture of conviction (and vagueness). We know it when we see it. We cannot quite say what it is. We are nevertheless prepared to defend our position with a confidence the definition does not really support. This is awkward when you are trying to do the work, and it has become considerably more awkward in the last three years, because a system that has read everything ever written about taste, and seen most of what has ever been designed, is now sitting patiently at the other end of your cursor, asking what you would like it to make.
So the question is no longer just what is taste. The question is whether taste is the kind of thing that can, in principle, be learned by something that is not a human.
The thing I keep coming back to is not whether the models are good. They are great! And they are getting better. The thing I keep coming back to is what the goodness is made of, and whether what it is made of is the same thing taste is made of.
It is tempting to pick a side quickly, and with feeling; I am writing this to find out where I land.
“I shall not today attempt further to define the kinds of material I understand to be embraced within that shorthand description; and perhaps I could never succeed in intelligibly doing so. But I know it when I see it, and the motion picture involved in this case is not that.”
The Bull Case (AI gets there)
The bull case is that taste is not magic, it is data, and we have only just started feeding the machine. Every judgement a designer makes is, at some level, a function of inputs they have absorbed: products they have used, interfaces they have studied, the specific corner of the internet they grew up clicking around in. None of this is mystical; it is a corpus. The corpus is large and partly tacit, (but tacit is not the same as inaccessible, and the history of machine learning is mostly the history of things we used to call ineffable being quietly eaten by scale). Chess was ineffable. Protein folding was ineffable. So, until recently, was language itself! The bull case does not require the model to feel anything. It only requires that enough good design, paired with enough signal about which design was ‘good’ and why, eventually produces a system whose outputs are indistinguishable from the work of someone with taste. At which point the distinction stops mattering, the way the distinction between a human grandmaster and a chess engine stopped mattering somewhere around 2005.
The mistake some designers make (and I include myself in this), is assuming outright that taste lives somewhere the data cannot reach. It might not. Taste lives in the artefacts. Every product you have ever felt was right exists somewhere on a server, alongside the products that felt wrong, and increasingly alongside the analytics that tell you which felt better to whom and for how long etc. The signal might be noisy, but so is everything these systems learn from. That is to say, they do not need the ‘taste rule’. They need enough examples of the rule being followed and broken that the shape of the rule begins to emerge in the weights. This is, mechanically speaking, not very different from how a junior designer becomes a senior one, except faster and at a scale no human career can match. A model does not have to understand why one choice is better than another. It only has to have seen enough examples of each, in enough contexts, with enough signal about which contexts rewarded which choice. After enough of that, the right answer becomes a thing the model produces without effort, the way a designer with twenty years of experience produces it without effort, and possibly for the same underlying reason.
I still do not think the strongest version of this case is that frontier models will replace designers. It is that the part of design we have been calling taste may turn out, in retrospect, to have been a much more legible problem than we wanted it to be. Perhaps we wanted it to be ineffable because ineffable is flattering. It made us feel like artisans! But the same was true of the chess grandmaster, whose play was described in mystical terms right up until the moment a search algorithm started consistently winning, at which point said mysticism simply resolved into pattern recognition. Design may be next. The uncomfortable possibility is not that the models will get good at taste. It is that taste, examined honestly, is closer to a skill than a sense, the former of which is just data + compute + time.
The Bear Case (it does not, and cannot, in the way that counts)
The bear case is that taste is not a prediction problem at all, and treating it as one mistakes the surface of the work for the work. A model trained on every interface ever shipped will get very good at producing the centre of that distribution. (Yes, post-training shifts the centre. We will come back to this). This is useful. This is, in fact, most of what most software needs. But the work that matters, the work people remember, is almost always a small and deliberate violation of something the average would never break. The optically-aligned icon nudged half a pixel off-grid because on-grid looked wrong. The sentence that lost two words so the line break would land where the eye wanted it. The whitespace that is technically nothing, and yet somehow feels like the most important thing on the page. (The list goes on and you will have your own examples, I'm sure). A predictor cannot do this, not really, because to do it you have to know the rule, and feel the rule, and then decide that this particular product, in this particular context, is better served by breaking it. That feels like less of a math problem, rather, something closer to character, which is probably the one thing you cannot prompt your way into.
The problem at the heart of the bull case is that the training data does not include the rejected work. Yet every product I admire is, structurally, the negative space around all the things it could have been and is not. None of the major training corpora (as far as I am aware) contain Figma version history, or design Slack threads, or the slew of screens dragged to the trash. The model sees the survivors; it does not see the killing floor.
It also does not see the difference between the survivors. Most design that ships is mediocre. The labs know this, and the response has been preference data: humans who rate and compare and rank. I do not know exactly who the labs are using as their raters, but, from what I can tell, much of the preference-data labour runs through annotation platforms (usually crowdsourced workers paid by the hour), rather than senior practitioners with the eye to reliably flag work above the median. The economics make this unavoidable. The best designers are well compensated to actually design things; nobody is going to pay them what their day rate would cost to sit and rate model outputs by the hundred. And if anyone tried, the work itself is a kind of low-grade torture for someone whose joy is surely in the making.
The frontier labs are no doubt trying to recruit better raters where they can, and this picture may change. But the rater question, in the end, is not even the most salient. Whoever they are, however carefully selected, preference data inherits their ceiling. This is a version of what AI researchers in adjacent contexts have started calling the scalable oversight problem: a system trained by human evaluation cannot exceed the evaluation, and where the evaluation is the bottleneck, the system stalls there. And even if the labs solved this, even if they recruited the best designers in the world to rate every output, the rating itself is the wrong shape for what taste actually does. Taste is not the act of choosing between two finished candidates. Taste is the act of producing a candidate worth choosing in the first place. A rater is making A or B decisions on work that is already in front of them. A designer with taste is making decisions about work that does not yet exist, in conditions that have not yet been specified, against constraints that nobody has written down. A predictor cannot lead. It can only follow extremely well, at extremely high resolution, across extremely many examples. Reducing this to a ranked-pairs problem is, by its nature, throwing away the part where the taste was. So yes, we can rank the outputs of a process. But we cannot rank the process itself, because the process is not really a thing that can be put on a screen and clicked.
Design taste is one of the cleaner cases of this. The part of design quality that defines great work is precisely the part that most evaluators cannot reliably flag. The model will climb to the top of what its evaluators can see, and then it will stop. Past that point, there is nothing in the system pointing the way. And not because the evaluators are not good enough, but because the work the evaluators are evaluating is not the work where taste lives. Which I think is broadly the point this essay has been working toward. You can train a model to do almost anything. You cannot train it to decide, from first principles, what is worth doing. Taste is not a math problem because the math, however good, is being done downstream of where taste lived.
Where I land (for now)
I land ultimately on the bear, though not in a way that feels triumphant. The models will keep getting more capable, more polished, more legibly 'correct'. The part of design that exists below taste is going to become abundant in a way it has never been. None of this is bad. Most software needs more competence than it currently has, and the labs are about to deliver an enormous amount of it.
Yet when I look at the designers whose work I keep coming back to, three things keep showing up. There is deciding, which is the largest and the most upstream: whether to make something at all, whether this feature is worth shipping, whether the direction is the right direction in the first place. There is framing, which sits just below it: given that you are making something, what is the actual question you are answering, what are the constraints, who is this for, what is the real problem under the apparent one? And there is refusing, the myriad of little nos inside the work: the modal cut, the easing dialled back, the accent colour muted etc.
None of these three are things a model can be ranked into. The deciding is upstream of anything the system could measure. The framing is below the surface that gets measured. The refusals are absences, and you cannot rank an absence. The models will get better at producing things; we will get better at deciding what is worth producing. The job has narrowed, but it has also clarified.
The taste, in the end, is still ours.