Proving AI in the Real World: Our Investment in Intelligence

Index Partner Vlad Loktev, Intelligence co-founders Grace Li and Kamryn Ohly, and Index Partner Eryk Dobrushkin

Do you remember learning to ride a bike? At first, you probably had training wheels or a parent’s hand keeping you upright. Eventually, though, to really go forward, you had to pedal on your own. You wobbled, you fell, you skinned your knees. But you got back up, you learned how to balance, you built confidence and speed. And one day, without realizing it, you were flying through the neighbourhood, choosing your own adventure.

AI today is at a similar inflection point. The promise is enormous — superintelligent systems that don’t just assist humans, but do real economic work. And yet, most deployments are still tightly constrained. Data labeling teaches models what experts think is right. Reinforcement learning environments let them practice in closed worlds. Application wrappers make them usable in relatively narrow, scaffolded workflows. All of this has value, but it means few models are ever exposed to the full, messy reality of how work actually gets done.

For AI to achieve its promise, we must take off the training wheels. We have to let models operate on their own, under real conditions, with all the unpredictability and bloody knees that entails. And to do that, we need something we don’t have today: real-world proving grounds where autonomous systems can be deployed, measured, and iterated on in the open.

One year ago, Grace Li and Kamryn Ohly were computer science students at Harvard. Best friends since freshman year, they entered a hackathon and built a game engine. The AI models they used worked, but the results felt clunky and visually off in a way anyone with taste could spot immediately. They applied to YC with the game engine, but they couldn’t shake the feeling that, for all their utility, these models could do better.

So they built a simple experiment to find out — a “hot or not” game for design. You’d type a prompt, see outputs from different models, and pick the one you preferred. Those votes became a live leaderboard showing which systems were better at producing designs people actually wanted. They called it Design Arena. Then they more or less forgot about it.

A few days later, a friend posted it to Reddit. Overnight, thousands of users showed up. And something else unexpected happened. Instead of treating Design Arena as a benchmark, people were using it like a product — a way to generate websites and interfaces by comparing models. Grace and Kamryn realized the signal might not just be valuable to consumers, but to the AI labs competing to build the best models.

Since then, Design Arena has grown into a global platform with 5.5 million users, scaling from $5 million to $60 million in ARR in six months, profitably, with an incredibly talent-dense team of 10. People use it to access the best intelligence for anything they want to accomplish — creating a website, making a video game, editing a movie. Design Arena lets the world’s models compete to serve that demand, learn from what people choose and what actually works, and route each task to the best system. Along the way, it’s become core to how leading AI teams measure their progress: Google used Design Arena to test Gemini 3 ahead of launch; OpenAI, Thinking Machines Lab, xAI, and Meta have used it to evaluate their models. As the universal interface between human intent and AI, Design Arena is making the world’s best intelligence more accessible.

At Intelligence, Grace and Kamryn are taking that same approach and applying it beyond design. With Prediction Arena, they’re testing whether AI models can accurately predict real-world events by having them trade on prediction markets using real money. The underlying belief is consistent: any domain where models are immature, hard to evaluate, or easy to fake in demos is one that lacks real-world feedback. If models struggle to design usable interfaces, they also struggle to trade, to grow an audience, to build and ship mobile apps, to design physical products. The constraint isn’t intelligence or raw capability. It’s the lack of proving grounds where success and failure are visible, measurable, and impossible to hide. That’s what Intelligence is building: the real-world feedback and distribution layer for AI, where models can learn what humans don’t know how to teach.

Behind this vision are two of the most impressive young founders we’ve met. Last summer, in the span of 24 hours, three people reached out independently to say the same thing: you should meet Grace. We met her that morning, and within five minutes it was clear there was something special here. Later that evening, we regrouped with her and Kamryn at our partner Nina’s apartment. The conversation was direct, ambitious, and thoughtful. They challenged us as much as we challenged them. By the time we left, we had conviction that these were founders we wanted to partner with, building at the frontier of what AI can become.

We’re excited to welcome Grace, Kamryn, and the Intelligence team into the Index family as part of their $7.9M seed round. We believe they’re working on a problem that gets at the heart of how AI progress happens, and we’re happy to be on this journey with them.

In this post: Intelligence

Published — Aug. 3, 2026