Key takeaways

  • As co-founder Grace Li tells it, her company started a few weeks before graduation in 2025, with a handful of college friends trying to…
  • There was no substitute for human judgment, they decided, and soon they were brainstorming ways to get honest human feedback at scale.
  • “It was the missing bottleneck for a lot of these models to make improvements in the design space,” Li says.

What happened

As co-founder Grace Li tells it, her company started a few weeks before graduation in 2025, with a handful of college friends trying to make their AI game engine work. The models could make functional games, but none of the games were fun — which raised the interesting question, how can you tell if a game will be fun?

The users tend to be indifferent to which models they’re ranking — as Li puts it, they just want the best output they can get — so their rankings can give critical input to what users really want.

For frontier labs, that’s a service worth paying for, Li says, adding the site is currently generating $60 million in ARR, solidifying its position as a key source of human-led evaluation data for the AI industry. Crucially, users have to log in to get their output, so Intelligence can also track how those tastes change across different continents and over time.

LM Arena, which takes a similar approach to text-based responses, raised $150 million in a Series A in January, just four months after formally launching its paid product. When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Why it matters

There was no substitute for human judgment, they decided, and soon they were brainstorming ways to get honest human feedback at scale. 3 million people around the world. As it turned out, there were lots of AI companies looking for scalable user feedback — and many of them were willing to pay for it.

“It was the missing bottleneck for a lot of these models to make improvements in the design space,” Li says. 9 million seed round led by Index Ventures with participation from Conviction (Sarah Guo and Mike Vernal), A*, Valkyrie, and others. For non-enterprise users, using DesignArena is a lot like using a sophisticated model router.

There’s a Chat-GPT-style window for prompts, with separate dropdowns for websites, images, and a dozen other visual formats. Once you put in the request, format and style, you’ll be presented with a series of “A vs. B” choices until you’ve ranked the handful of outputs from best to worst.

It’s a useful service, but the real value of the platform comes from the enterprise side, where participating models can treat it as a source of endless instant feedback for their media-generating models.

What to watch

Less than a year after launching, Yupp shuttered its doors earlier this year after raising $33M from a16z crypto’s Chris Dixon. 3 million users, but still couldn’t build a sustainable long-term business. Even so, other startups based on human evaluation seem to be thriving.