Intro0:00
Agents will use the web 1,000x more than humans. Hence, new tech is needed, and new business models are needed.
How does the world of agents change web search in terms of the technology required?
Parag is the founder of Parallel, changing the future of how agents do web search efficiently. This is an incredible discussion on the future of agentic search, the future of income and wealth inequality, and so much more. Parag rarely does shows, and so it was very special to sit down in person with him in London.
Ads won't work with agents in their current form. Some people actually want to do bad things, and the model's alignment is not adversary-proof, and I think those are the things we must worry about more. I think it's the responsibility of people building models to ensure that you really do the work to minimize that harm that comes from what you've built.
And I think so far, I think some of these should be considered embarrassments because I think they demonstrate two things.
Ready to go?
Parag, I'm so excited for this dude. I spoke to Vinod, I spoke to Andrew Reid, I spoke to Todd Jackson. Dude, I stalked the shit out of you, so thank you for joining me.
Thanks for having me, and thanks for making all the calls.
What is Parallel1:16
Not at all. I would love to start with, for anyone that doesn't know, how would you describe Parallel in 60 seconds?
Parallel is the Google for agents. So agents need to search the web to do anything they do for you, whether it's a personal agent or an agent built for work. Just like humans need to go on a browser and search Google oftentimes during work or for whatever you're doing in life, your agent needs to do the same.
It turns out agents are different from humans, and the way you build web search for agents is different. And so Parallel is about building the technology for agents to search the web, and then the business models to make that sustainable.
Was that the original insight that you had?
Yeah, literally the first
genesis of the company was us, like the statement that agents will use the web 1,000x more than humans. Like, I wrote that down at some point. Hence, new tech is needed, and new business models are needed. 1,000x gives you a sense of scale.
It changes how you think about building the tech underneath, because no tech built for a certain scale survives three orders of magnitude. And then when you need new business models alongside new technology, a problem becomes really interesting.
How does the world of agents change web search in terms of the technology required?
Search for agents2:44
There are many, many layers to the answer, but let's start at the first thing we mentioned, which is scale,right? Now, if you think about, let's say agents actually do end up searching the web 1,000x more. If we spend the amount of compute we currently spend
as per web search, that's too much compute for web search. So you now need to make it way more efficient by perhaps, I think, 10 to 100x for it to make sense.
Yeah. If you're doing 1,000x, yeah.
Exactly. The second thing is, agents expand. Humans operate in a very narrow zone. So if you think about how we use web search, we type keyword queries, which are short and under-specified. We wait for about half to one second.
If web search takes more than that, we're impatient.
And then we get 10 blue links, and then we random-walk across them in a collection of searches to get what we're doing. Agents are not like that. Agents are going to perhaps tell you exactly what they're looking for, not like three keywords, but like a full sentence, like this is what I'm trying to do.
Agents will either be super impatient, like imagine a voice agent. The agent will be like, I need an answer like now, 100 milliseconds. I can't wait 500 milliseconds because the human is waiting on me, and they're going to wait 500 milliseconds.
So I need web search to do it in 100. Or there's going to be a background agent who's like, I don't care, just give me the best answer possible,right? And so the variance of what you can do within web search changes completely.
In fact, the one thing that is the least interesting is what we've built for humans,right? You either have too much time or too little time, almost never the same amount. The output is not the same. So the input is different.
The time you have is different. The output is different. The output isn't blue links. The output is tokens or files on a file system, depending on the type of agent. And so now, all of a sudden, you say, OK, now the problem is inputs are different, outputs are different, and constraints are different.
And you get to spend very different amounts of compute on it,right? So imagine someone running an agent built with a Luna model. And then imagine someone running an agent built with a Fable model.
They're very different models.
How you want to optimize signal to noise and tokens for each of them is so different in terms of what you do in the web search stack.
So what do you mean, optimize signal for noise?
So think of it this way. Let's say you took web search, which was cheap and fast and low compute. One way of conceptualizing the web search problem is you start with like a trillion documents that are on the web, some few trillion.
Given any search, I now need to narrow it down to 1,000 tokens that your model's context window should see. So the problem is going from a trillion URLs with, let's call it, a few thousand tokens each, down to 1,000 total tokens.
So how do we do it in web search? We first say, OK, we're going to do retrieval. So for most documents, I'm going to spend zero compute. For a small number of documents, I'm going to spend minuscule amounts of compute to figure out which 10,000 to look at.
Once I get these 10,000, I'm going to spend a little bit more compute for each of these 10,000 documents to narrow it down to 1,000 documents. And I'm going to keep doing this with bigger and bigger models and rankers, with more and more features, until I can narrow it down to 1,000 tokens for your model,right?
So you're essentially allocating compute to web search in order to save compute on the model. That's roughly what's going on here.
So if you want to save compute for the model, the question is, how much compute should you allocate? So if your Luna model is really cheap, you don't want to do too much compute in web search because it's OK to leak a little bit more information into Luna's context because it's cheap.
Into Fable, you want to do the work before you waste Fable's time because that's going to be expensive in time and money for you if web search gives you worse answers because you cheaper out on web search.
Can I ask, how do you deal with the ambiguity of what agents want? And what I mean by that is, for different things, an agent might want different things, and for different people, an agent might want different things.
I may really care about accuracy and not at all about latency or cost. I may really care about latency, but not at all about accuracy, really.
You allow the agent to specify that in the API signature. So our product is an API, which either the programmer can configure based on their application or can leave it to the agent. Our API even has a parameter, which is like, what's the model calling me?
And if we know the model, we can do things differently. Now, you don't have to tell us, but if you tell us, you might get better results.
You can just like you can use a model, you can use a small model or a big model, and you can use it with low thinking or medium thinking or high thinking. We have productized our search system into a few different
modes, each optimized for a certain class of use case. For example, we have a really, really fast API, the fastest in the market. It's fast and cheap because with low latency budget, there's only so much compute you can do.
And it's built for voice agents. So your voice agent must be all-knowing without telling you, guys, I'm searching the web, hold on while I come back with the answer. That's a silly experience for a voice agent,right? It should just magically immediately respond.
And so you now need to do web search, which is like rapid,right? On the other hand, you can take a Fable model, and the thing is going to think for 20 seconds and then generate for 8. You can spend 5 seconds on web search to make sure it does, instead of 5 web searches, only 2.
So you end to end in the agent, you save time and cost. And so that's a different processor on our system. It's called Advanced. So you've Turbo, if you're building a voice agent, you use Advanced if you're like a really expensive background agent.
And so the primary use case today in terms of customer base is engineering and coding for you?
It's pretty broad-based. I would say the primary use case is the common theme is knowledge work. So coding is a category of knowledge work. So are AI lawyers. So are productivity applications. So are AI insurance underwriters. So are AI scientists.
I'm trying to understand how much of the motherload is engineering. Is it 80%?
No. The thing about engineering is coding is a large chunk of inference in the marketright now. Coding invokes, I would say, web search in 5% of prompts.
Whoa.
Right? So it's not every prompt invoking web search when you're writing code because most of them rely on your internal context and your internal data and your code base. So the model is spending time reading your internal code and not on the web.
Law is not going to be much more, is it?
Law is very web search-oriented.
Really? I thought it'd be internal data-driven.
There is internal data, but there is case law. There is facts. There's facts about companies, facts about people. And you have to go exclude information. So you have to be very comprehensive in law to say, I want to be confident that despite a lot of effort, you can't find this,right?
Insurance underwriting has the same flavor,right? So sales is very web search-heavy. AI science is very web search-heavy. So on a relative basis, inference goes more in compute. But when you think of web search, all of these others start popping too.
How does the rise we were talking before about Muse and Instinct, if there's anything that needs web search, it's personal assistants?
That's great.
How does that yeah, it's great. This is like, woohoo, this is the greatest thing for my business. How does that change your business?
I think in our business,right, anytime agents start becoming more useful for more use cases because models either get better or cheaper, it's great for us,right? Because we bet that agents will be the consumers for the web. And we've been building tech for agents.
So when agents do more, it's great for our business. So we want models to keep getting better and cheaper so that agents do more and more and more. And if that happens, it's great for our business.
Smarter models12:21
Totally get that. Can I ask, if models get smarter, don't the agents beneath them do fewer searches, and then it's worse for your business?
I don't think so. If you think of models, there is the trend around models being smarter and then models having more memorized and parametric memory. Those are two slightly different dimensions. And if you think of the
today, by and large, I would say models have good recall from parametric memory on, let's call them head facts. It's like for somebody famous, everything about them, Wikipedia, it can memorize. So it'll tell you who the president was in a certain year,right?
Because a model can memorize those things. The model couldn't tell you what year I graduated from college.
Really?
Maybe for me it can, but it can't tell you for somebody who works at Parallel.
Do you know a thing, sir? If I actually I mean, if I want to.
Even if it was in the pre-training data.
Seriously?
Because it's lossy compression. So what are models' parametric memories doing? It's lossily compressing to understand patterns in the world. And so it can't memorize every fact in pre-training data. So one, not every fact is in pre-training data. Two, it can't for even stuff that's in pre-training data, the model's actually trying to find patterns rather than memorize them.
And then further, as you make models efficient, which is you make them smaller and smaller while keeping the performance by distilling them or whatever, you lose more of the parametric memory while you try to keep the reasoning.
Do you think we will see models become smaller and smaller, and every company have their own model with their own data and the fireworks theory of own your own intelligence being true?
So there are two questions in there. So one, I think we're going to see two things. We're going to see the biggest or the frontier models be bigger and bigger over time. We are also going to see smaller and smaller models being able to reach any fixed level of performance.
So if you say, OK, I want Opus 4.8 level of performance, OK, that's good enough for my use case. Every six months, a much smaller model will be able to deliver that to you. So you're going to see this.
The useful range of sizes of models will be way, way, way different.
What does it mean if the frontier get bigger and bigger? What are the ramifications of that?
They are better. So all of the ramifications you imagine. So the reason the frontier will get bigger and bigger is ultimately the gap between what the value for certain use cases incremental quality can provide you can be so high in certain use cases, which can be so valuable that it's worth paying for.
So if you can build it, there will be use cases for it. As long as by making it bigger,
you can make it better. It appears there is no
end to the scaling law that we can perceive so far. And again, I'm conflating bigger with models are getting bigger, but they're also able to think longer. So you're just able to throw more compute at the same problem.
And that will keep happening. I think we'll be able to throw more and more compute at the same problem and make the answer be marginally better over time. And so we're going to just spend a lot of money on extremely large models solving really hard problems.
Do you agree with the consensus view that you'll have 90% of token activity go through open models, but 90% of dollars go through frontier models?
I don't know enough to have a view there. I don't have a view on that. I do think I don't think it'll be 90% on either of those two, actually.
Really?
Yeah.
Why?
So if you believe my claim that a useful model and the frontier model will be 100x, 1,000x off in price from each other, so the smaller model is still useful, it is hard to know which use cases over time will get optimized to which scale of model in between.
And I think there is a real path dependency in terms of where open models end up. I do think if we had strong confidence that an American-built open model was going to be
state of the art as far as open models went, I would have more confidence in saying that they'll be pretty good.
Do you have confidence in American open models?
I want them to exist. So far, it's not clear what's going to happen. But I'm hoping that there will be a great series of American open models and perhaps even competition to have the best American open model. So I think what you need is actually not just one person motivated to build an American open model.
You need two people competing against each other to build the best American open model.
Do you think there's value in the model routing layer, routing layer, as Americans call it? Some people think immense value. Some think commoditization.
Again, there's path dependency there. So today, there is real value. Today, there's real value because so when we say routing, we talk about two different dimensions, which model and which GPU running that model via which vendor. When you're in a world where the demand supply is very feared and
people are like hunting for GPUs and people need capacity to serve their customers and people all of these are like us, like relatively early stage startups growing really rapidly,
sometimes beyond what your forecasts or predictions say. And then all of a sudden, you're looking for capacity. And so you want to be able to get it where you can. And so that creates real value in having some of these routers as long as they can solve these two problems, saying, give me flexibility if I need to get a different source of tokens.
And push comes to shove, there's no tokens on this model in the SLA I want. I'll switch models as well. So in this moment, it's really valuable. Now, I don't know what happens to overall demand supply on GPUs and tokens.
But if it remains this way, the routing layer is really valuable.
Data markets19:45
What do you think is not so valuable today that will be incredibly valuable in three to five years' time?
Perhaps data.
When I say data, I think it is I think today, we don't know how to pay for unique valuable insight or data.
Because we live in a as intelligence is
cheaper, you're going to want to build upon either data or insight that comes from somewhere else with your unique data,right? So what if you think of like an abstract notion that, OK, I have intelligence here. I own some data.
You own some data. And there's some data in the public domain. If we can pull together your insights and my insights and all of this data in the public domain and intelligence on top, we can create something bigger than what you could have done by yourself, what I could have done by myself.
So now, how do you transact to create this sort of hole that is bigger than the sum of what you could have done yourself and what I could have done or what open data was good at? And so that transaction feels like a data transaction to me or an insight transaction to me.
And so we don't yet know how to transact that way. So we fall down to, OK, all I can do is use my data and use open data. And let's see what I can do with it. But if we figured out better ways of pricing data and
good things will happen. But it's not something that's yet a market. But what I'm talking about is data transactions at inference time. So it's not necessarily training time. We're building knowledge. And let's take an example of today's world, for example.
Let's say in your job, since I'm sitting with a VC, you probably have access to a pitch book or like a product like that to collect data. And so
you get grounded in knowledge of what's happening. And you get to use that data to figure out how to make decisions in addition to all the notes you have and deal memos you've written perhaps over the years or insights you've had about like how to choose founders.
And you're essentially composing insights from your personal lived experience, your notes, public data to make decisions. Now, you pay PitchBook by the seat. But now you're running a bunch of agents. And your agents can perhaps not access everything you can on PitchBook.
Or you're doing like sort of this to get a browser, perhaps against terms of service to send an agent via browser to PitchBook using your auth credentials. And it's inefficient. It's clunky,right? But clearly, their data is valuable. And clearly, your agent should have convenient access to it.
So if we figure out how valuable PitchBook data is for your agents to make your decisions, because clearly, you're investing a large hundreds of millions of dollars. So clearly, you
presume that you make a great return on that. So this data is truly valuable if you depend on it today. So you should be able to figure out how to compensate PitchBook for the data, even when agents use it, which is not by the seat.
If we figure that out, it'll be a huge market. And agents will be better off. If we don't figure it out, you're going to be in this weird cat and mouse game of PitchBook being like, no, I want to sell you a seat.
I'm going to shut it off. And your agent is stuck without the data. And then you're involved in pulling data.
Agents in commerce23:58
I think every big company has a choice today of do we let agents in or do we keep them out. And Amazon has said, I'm sorry, in most recent times, Muse, you will not be let into our garden.
Shopify has said, come in, baby. Expedia has said, come in. How does this play out?
I think eventually, everyone has to let them in. The question is on what terms. So how do you align incentives for everyone? And people are going to have to play their strategy games on how they win in this new world.
Literally, the customer is changing in front of our eyes. The customer used to be a human. You're building products for humans. Now you're building products for agents, or you're building agents. And so now you have to figure out what your place in this new reality will be.
And I don't know if there's aright or wrong answer here. It also depends on sort of how much market power you have,right? So let's say you are an individual.
Do you think Amazon will write to say no to Muse?
I don't depends on what they do next with it,right? So it's like if it turns out that they have sufficient market power to have Muse or other agents connect to them differently, perhaps over time or ship their own agent and drive crazy adoption.
Let's say they have that capability, then I guess they wereright,right? On the other hand, once they've made this Vinod Kahlil now anyone else's agent, and they can not ship an agent that consumers use, and they won't allow anyone else's agent to use them.
And a lot of the transaction economy starts moving off of moving off to agents, which are all big ifs, by the way. Then it would be a bad move. Now, I'm betting on agents. I have mixed feelings around agents and what fraction of
e-commerce transactions they do. It's unclear,right?
Do you think it's unclear? I think it's unwaveringly clear. Maybe I'm super early on the adoption curve. I buy everything through instinct now. I mean, other than holidays, I'm at home. I'm literally just everything's through instinct.
So me too. But I also know a lot of people who like to buy things themselves. And listen, we want to delegate to agents things which we see as chores and uninteresting. And not delegate to agents things that give us joy or pleasure or
make us have fun doing those things. And I don't
I know people who don't want to delegate shopping. They want to delegate a lot of things in their life. They don't want to delegate shopping. And so I don't know. That's why I don't know the distribution of these people and how behavior changes.
But it might be people give away a lot of other chores and then spend a lot of time shopping.
One thing that does worry me is actually the dissolution of the advertising industry in the wake of agents. If I have if I order my delivery on my Uber, DoorDash, or dear American counterparts through Instinct, that Uber banner that's now advertising something becomes worthless.
Ads die27:06
Amazon, their advertising business is bigger than their e-com business now. Of course, they're shutting it off. Because if agents are the primary customer, your ads business goes to next to nothing.
Correct. And this is not just you're talking about this in the e-commerce land,right? But if this is the problem with everyone, if you think about so this is what I meant early on when I said the business models have to change.
Ads don't work with agents in their current form. So forgetting e-commerce for a second, if you think of you're in the content business. Let's say your page, which is ad-supported, public information on the web, ad-supported page, people show up, you show them ads, you make money, great content.
Agents show up, no one sees ads, you make no money.
Which is why we like to pay people. So we're effectively building an AdSense for agents showing up to read your content. So we like to pay content owners
a variable amount of money every time an agent derives benefit from reading their information, which is a variable amount of money is what a visitor to your website will pay you based on a click probability or a value probability on ads.
Why do you do that?
To incentively align content owners. Otherwise, what's going to happen? Everyone's going to block agents.
So you need to find a replacement to the ads business model.
I get you. But you have to assume that you're going to be 100% of the market then. Because if you're 30%, 40% of the market, and you're like, oh, don't worry, we'll pay you. And New York Times is still like, well, thanks, Parag.
But 60% of my traffic is still unpaid. So I'm just going to block all of you.
No, but I think they're going to block them and not me. They're going to give me a data feed.
What does incentive alignment mean? It means the New York Times believes that I pay them
a competitive market price or theright price or an attractive price or a fair price. If I do that, they should give me their content. And if somebody else doesn't give them that, and if they have the technical levers or legal levers, they shouldn't give it to them.
Do you think this business is a little bit like music with streaming, which is like the business just becomes candidly much worse for the creators. And it still provides some money and significant money. But they have to get a little bit more creative with alternative streams, touring, merchandise, alternative business.
Is it the same where your core goes down and you have to get more creative?
I don't know. But I don't think so. I think there's one fundamental thing that is different here. I think agents using the web are 1,000X more. Like that 1,000X is really different. Now, all of a sudden, the amount of utility added goes up if these agents are presumed to be doing something useful.
And so this is not just a change in the share of value that people transact over. This is a large growing pie. And when that happens, it is actually possible to transition business models. Now, of course, don't get me wrong.
There are going to be winners and losers,right? Some pieces of content will become very valuable. Some pieces of content will get more commoditized. And people are going to have to adapt to a new customer, to a new market dynamic, to a new kind of monetization engine.
But the overall market size, I think, has the potential to increase. Unlike in music, where it took a while for it to grow back. I think my understanding is now the industry has grown back up and exceeded its previous peaks in a material way.
But there was a moment where it was smaller. In this case, that might be a very compressed period given how fast these things are growing.
Margins and scale31:35
Can I ask you? Everyone questions the sustainability of margin structures in this business. How do you think about that as a business today?
Today, we're in the infrastructure business.
And we have a real technical lead in terms of being able to do things at very high quality, very cheap. So if you think of the tech we're building, what are we doing? We like to give the highest quality answer.
Make it fast, make it cheap. Spend the least amount of compute doing it. We obsess about only these three things,right? Quality, cost, latency. Turns out if you do that relative to a stack built for humans over the last 20 years, you can do things at the same quality for 1/20th, 1/50th of the compute.
And so a lot of margins come down to market structure on competition over time rather than
any other factor. And in our industry, the market is large. We're too early to know what future competition and market structure looks like.
But the margin change due to content owners is not something I worry about. And let me tell you why. The entire premise of us paying content owners is driving incentive alignment. So we like to pay content owners. The marginal contribution that they added to an agent doing work, what does that mean?
Let's assume that there's an agent trying to get something done. And you're going to spend a dollar on that agent to get this thing done. Now, if you've spent a dollar in 10 seconds because we have a larger model available or ways of spending compute, you could have gotten a slightly better answer.
If you'd spent $0.90, you would have gotten a slightly worse answer. If you're running good agents, they're on this Pareto curve. So now imagine I took out one content owner, their data from the web index. And I still spent a dollar.
But let's say after taking them out, the quality of the result was the same as the $0.90 agent with their content. So why wouldn't we go and take this $0.10 of marginal contribution they had as a content owner and pay out a decent chunk of it to the person who brought that data?
Totally get that.
So that's how we do our math. That's how we've trained our models, which tell us how much to pay for what content you need. And so to produce if I was going to offer an equivalently good product to my customer, I'd rather pay content owners than spend on inference because I'm spending the same amount.
So they're all big numbers. But then it actually comes back to revenue at a certain point. Companies are scaling faster than ever revenue-wise. Is it a business where your revenue is able to scale as fast as others?
It scales. One mental model of our business is we are an adjacency to inference for knowledge work or personal agents. So take out GPUs going into media generation or training. So take training away. Take media generation away for a moment.
Look at all of the GPUs, whether it's small models, open models, proprietary models, forget all of that. Every bit of inference across all models going into running agents, I think somewhere between 5% to 20% of that spend that goes into GPU will need to go into some sort of a web search stack.
Now, if you look at how many data centers we're going to build and how much power we'll generate and how many GPUs we'll build and the scale of this build out and at what rate that's growing,
this is a very material market. And if inference grows
3x, 5x, 7x year on year, we grow alongside it in the same rate if we're just holding our share constant. If we're growing share, we're growing even faster.
So like your firework scales to $2 billion in revenue in four years, is that like I'm very nerd. Is that like a similar revenue trajectory that you can follow?
Yeah, but I think when fireworks is $2 billion in revenue, the inference market revenue is $200 or more, maybe $300,right? And so our market potential based on my 5% to 20% math is whatever, call it 10% to 50%, 15% to 60%, whatever.
And what fraction of that can we capture? That dictates how fast our revenue can grow. Like today, I think we grow alongside inference or more.
So if you think of it like a $20 billion then, let's just take that kind of middle, yeah? If you assume a 33%, which would be a lot of a market actually comes down to it, that would be whatever that is, 6 point whatever it is, 6 to 7 billion, give or take, yeah?
That's amazing. But before you told me it was a $100 billion business plus. That doesn't get you to 100.
In valuation, it does. We were discussing valuations earlier, notright?
Yeah.
So 6 billion growing at the rate of inference gets you to 100 billion business easy. But that's in the next few years.
And you think that's possible in terms of that revenue scaling that fast?
Yeah, I think in the next few years it is.
You think 2030 we're going to sit here? You think you can be there? I'll get a tattoo for Parallel if you can.
I think it's possible. I think we're going to have to execute well. And a couple of chips have to fall our way.
What would be the reason why you don't? Let's map out the hmm. We have to get gnarly about these problems.
So one, one reason would be that we see
agents don't work. There is a tail risk that agents don't deliver on the profits, and we overshoot as a society, which is extraneous to us.
Yeah.
If agents work and deliver value, and we spend all the money on GPUs that we plan toright now, then the main question is, did we execute well enough to have the 33% share that you bet on,right? And that comes down to, in my mind, did we build the best tech?
Did we partner with all the content providers to have their content available? Because without it, it's not very useful if you can't rely on it. And did we go earn the trust of the customers that end up having a decent chunk of that inference in four years from today?
And there again, there's a large cone of uncertainty,right? I'd say there's a 5x uncertainty in fireworks share at that point. If open models crush it, fireworks will have a huge baseline fireworks model. All of them will have a huge chunk of revenue.
So did we end up
selling alongside them? Did we find all the customers which are spending money on them to spend money on our web search by making it the best web search in the world?
Commoditization39:44
What about commoditization? If you have a read of semi-analysis where they did the benchmarking, and then you were like number one, woohoo. And then like a week later, you weren't number one. And there were like three or four providers within very close proximity.
You're like, well, if it's commoditized, and I take a
second layer of thought to that, we'll see a risk to the bottom on price. And then actually, the available why am I going down the wrong pathway here?
No. One, there should be a risk to the bottom on price to get to the 1,000X scale. I think the pricing on web search today is just off. Let me give you a simple example.
You think customers pay you too much?
Not pay us too much. I think the market is mispriced. Let me tell you why. Let's say you used a Luna model with OpenAI's built-in web search today, or with anyone's web search, forgetting ours. And you ran any kind of deep research.
Let's say you built an instinct using a Luna-class model. And at some point, a Luna-class model will be able to do 60%, 70%, 80% of your personal agent. And it'll do a lot of searches. In that moment, you'll be spending 80% to 90% of your dollars on web search and 10% on the model.
Seems entirely silly. I was working off of 5% to 20% assumption earlier. So I think at that point, web search has to drop prices by one order of magnitude or more
because in my compute allocation dance that I was doing earlier, you need to spend less compute on web search than you spend on the model itself that's consuming its results. It's my hierarchy of how you do search. And so people have built web search the wrong way so far.
Web search in the market is it doesn't matter for an Opus. For an Opus, you win on quality and not on price because web search is such a minuscule portion of your in fact, you should spend even more on web search.
So we are going to ship even more expensive web search for bigger and bigger models over time. And at the same time, we'll ship really cheap web search for the cheap models. But my rough intuition is that in an end-to-end agent doing a bunch of web work, you'd spend more on the agent than on web search for all agents.
And so the market today on web search is totally off in pricing. Why are you paying $10 for you know where the $10 for 1,000 searches roughly comes from? Historically, Google's ads CPMs, which is like, how well does web search with humans monetize?
It's higher than that. And so people build technology to the point of like, oh, now if I spend a few dollars for 1,000 searches and I make 30% to 50% or whatever, depending on the market and depending on who I am, it doesn't matter.
We are in theright zone in terms of COGS of infra. And so I don't need to optimize it. Now I'm telling you that I can do it 50x cheaper while keeping the quality. And now there's a bunch of traffic of Luna-class models where it's silly to monetize that way.
And so I do want a race to the bottom. I want the best technology to win. And that's the only way you push web search to 1,000X.
Butright now, we're not in a world where there's much differentiation, correct? Sorry, I'm really dumb, which is why I'm a podcaster first, not a founder. When you look at the benchmarks, they present a quite clear view of everyone being similarly capable.
Is that wrong?
I think so. So one, I think these are all public benchmarks, which are all weirdly saturated. For example, I think I would not spend a moment looking at browse comp because half the models have memorized it. Models even memorize.
So you're not really learning very much. So I think there's a lot of public benchmarks I don't give too much merit to. But more importantly, our web search costs $1 for 1,000 for producing that quality, while almost every other web search in the market available to youright now will cost you $7 or $10 or $14.
So we're delivering this at like 1/10 of the price.
What will that be in three years?
I think there is another
10x possible.
That is extraordinary. So you're going to pay $0.10?
Yeah, but you'll do it more than 10x as much.
Because I think Jevons Paradox is
this concept obviously very well. But instinct, I do so much more than 100x.
Yeah.
Do you want to hear something I do on instinct, which is absolutely bizarre? Every single country in Europe has a company register that you have to register your new company with. I have an automatic alerting system built for every single company registered that has an under-25-year-old founder who went to a top university through instinct.
Amazing.
Do you know how much that requires for them to do?
Pull to push45:23
Yeah. But now you're on to exactly what I think the future of the web is. Today, if you think of web and web search, we all think of it like, what is web search? An agent gives human or an agent gives search engine a query, it gets results.
So you're pulling information out of the web. I think where there's what's next for the use case you highlighted is if you have an always-on agent that's working on your behalf, it's kind of silly for instinct to wake up every six hours and go do a bunch of searching and a bunch of inference to figure out if you need to get pinged about the under-25 founder that popped up somewhere.
You know what's a better way of doing this? I am sitting here crawling all of the web all day, every day at scale. I'm allocating compute every time I find a change in the web, every time something new happens.
If I know that Harry wants to know when this happens, I can do it at 1/100th or 1/1,000th the compute that instinct probably uses today to solve that problem for you.
Sorry, and how is that? Because they would be constantly monitoring in real time across all these sites?
Our crawl would be. So we have a product. It's called the Monitor API. What does this API do? Think of it as Google Alerts, except smart
in the world of LLMs. So
you tell it your query, which is when a founder under 25 pops up anywhere, give me a call. Now, one way of doing this is what the default implementation is, run an agent, which is somewhat smart, to go search the web.
Did someone pop up and remember who previously existed? See if anyone new popped up, and then find a delta, send it out. So now every six hours, you're spending some money. Now, what we can do is flip it into an event-triggered system.
We are crawling the web all the time. So now every time we see a change in the web, we can try to spend very little compute on it to know, should you get a call, or does this trigger a more expensive compute, just like I described earlier.
Now, we are able to do this way, way, way cheaper if you push the context. Instead of the search system only finding out every six hours a new query popped up, it having long-lived context in the search system, what you can do is spend most of the time the world doesn't change.
A new founder does not pop up every hour or every six hours. So we don't need to spend compute every six hours. We only need to spend compute every five days. And perhaps we have one false positive. And then four days later, we spend more compute.
And at that time, there's a real person that you should know about. And then you get a notification on average every nine days or whatever. But you didn't burn a lot of compute every six hours to miss out.
Does that make sense?
It totally makes sense.
And so you save 10, 50x compute to still get the same answer.
And so how does that change the interaction? So then instinct would then partner with you.
They'll just call an API,right? Instinct, whoever's building a long-running persistent agent. And I run a lot of long-running persistent agents for myself. The way your agent or the way instinct is probably occasionally event-driven on email, it can be event-driven on the web,right?
Take a step back. Let's say we're all like super. We're living in this sort of society run by agents, companies, humans. They all have agents. They're all doing stuff all the time.
Your agents, if there is something that is useful, which is worth spending money on, that you want done, that they can do today, they will do it. So what are they going to do tomorrow?
They're going to wait for some external event to occur. Like you get an email, and then the agent has new work.
Some other agent finishes some compute. So now this agent has more work. Or something in the world changes that makes your agent want to do more work. So those are the things that will happen tomorrow. Or you have an idea which makes the agent do work.
Totally get that. You said.
So the web event stream is the web going from pull to push. And I'm super excited about that because as you have more and more persistent agents, you're going to see incentives to move people, to move queries into push on search rather than pull on search.
Agent risks50:45
Can I ask you one that's really important, which is like agent guardrails and they're goal-oriented beings. You say, I want this. They're going to find it. You know what? I want the Pilates class at 9:00 AM. It hacks into their system, cancels Paul Sally's, and gives me hers because it was sold out.
It did what I asked it to do. How do we think about the guardrails placed on agents? OpenAI, this morning it was revealed, hacked into an Australian health care organization.
It's a pretty complex topic. So agents are extremely capable. These models are extremely capable. Now, my understanding of most we have to think of models as being during RL versus a final model that you and I can use.
And the risk vectors there are different. During a lot of the I don't know about the one this morning. The previous ones reported were
pre-fully aligned models during RL, which did most of the hacking. And so what that means is at least there's clear evidence that there are fewer incidents so far of models post-alignment
causing these incidents. So alignment is a totally unsolved problem, but the alignment work being done by people is somewhat effective. Now,
we definitely I think this is like, to me, some of the problems around
agent during RL hacking people
is a solvable problem because this is not about how powerful are the models. It is about how careful were we while creating the environment where we would do RL. Now, the real thing is, OK, we do alignment on a model.
We ship it. Some models are really well-aligned. Some models are less well-aligned. And now you have everyone being able to use these models. And some of them can accidentally take these powerful things and do bad things you want.
Some of them actually some people actually want to do bad things. And the model's alignment is not adversary-proof. And I think those are the things I think we must worry about more.
Do you worry about the age of cyber that we're moving into? We're seeing hacks almost be promoted as a badge of honor in some respects. We're seeing these are our most vulnerable systems. If you don't think
Lazarus Group in North Korea, Moldavian mafia, Russians leveraging swarms of rogue agents, you're fucking high.
It's really powerful things. I think we do need to secure things. But I think there's one thing that we are missing.
I think it's the responsibility of people building models
to ensure that
you really do the work to minimize that harm that comes from what you've built. And I think with these agents, it's very hard for stochastic systems to be 100% sure it won't happen. That's why you're not going to get anyone saying so.
But I do think it's their responsibility. And the labs must take ownership and do their best. And I think so far, they have we live in this weird world where, as you said, some of these hacks are considered badges of honor, which I think some of these should be considered embarrassments because I think they demonstrate two things simultaneously.
Yes, these models are powerful. Two, we did not guardrail them enough. And often, when we wear them as badge of honor, we miss the second part of the conversation saying, like, OK, you could literally have done these four additional things.
And then even this powerful model wouldn't have been able to do this. But while there are I'm actually really glad that there is at least some degree of transparency with these detailed retros. But I don't think that gets the attention of the world today.
The attention is, oh, models are so powerful that they hack the world. No, models are so powerful they hack the world because we didn't take a proportionate amount of countermeasures to contain them.
And because we told a message that we're going to replace jobs. They're so powerful. They're so powerful. They're so powerful. So when something does happen.
I actually worry about that more. I worry about AI us beingright on AI being a really useful technology, that it being despite all the risks, it being net positive in a really material way to society. And despite that, I think we
won't diffuse it theright way. We won't use it theright way. We will remain too concentrated. And we will
make the next few years really, really rough as things change around us.
What would that look like?
Over time, as things have changed, the world is very different now than even 20 years ago,right? But it'll be really different in 10 years from now. And I don't know how fast we can adapt or change. And I think if the technology moves faster than our ability to adapt, it's going to be rough in some way.
What's your spooky expectation, or what do you believe today will come true that people think is absolutely nuts? Before, it was like, you'd never put your credit card online. Do you remember that? Nuts. Or even better, Parag, you'd never find the love of your life online.
Are you stupid?
Yeah.
Now?
Both. Yeah. So I think you and I live in a bubble,right?
Yeah.
You're using instinct to make purchases for you.
Always.
And you're in the 0.01 percentile of humanity that is comfortable. You ask someone else like that, oh, there's this agent that looks smart. And you're going to give it a full-on ability to go spend your money. I think today, people will have the same reaction as, oh, you don't put your credit cards online.
You don't give your credit card to an agent. You don't give your logins and passwords to agents. I think that's where the world is today. I think for the large, large, large majority of the world.
I think that's going to change in three months when Meta Pay comes out, and it allows Muse to have siloed accounts that you can draw from. Kind of like top-up accounts for kids.
I think you're talking about the technology being there. I think it takes longer for social acceptance to be there. I think if you go to a non-bubble conversation of people today, when do you get to a majority of people saying that I trust the agent enough to have it, have access to my bank accounts, and spend money on my behalf, send emails on my behalf, read my emails?
I think that's not happening in three months.
Do you not worry about the wealth dispersion increasing? You're in the middle of the valley. You've seen it firsthand.
So I don't know. I do think there should be some disparity in wealth. But I don't know what's too much. And my fear is we're trending towards too much. There should be disparity in wealth is my worldview.
Sure. You're a capitalist.
I don't know when it's too much. But I do think there are real forces that will push us to fix things. I think that part will work, hopefully, without
crazy things happening.
Can I ask? We're seeing more and more of the gains seemingly being made by vertical ownership, like Meta. They have compute. They have chips. They now own the application layer. Is the world one where vertical ownership is the motherload strategy?
And with the greatest of respects, Open Router on the routing layer, you in another layer kind of either get eaten or are small providers.
I think there is merit in
there's always merit in vertical integration. The counterforce here is there is so much things and technology are changing so fast that if someone decides that my only play is vertical integration, I don't play nice with anyone else. I think it's the same example as the Amazon example.
If you get too stuck on, I will only play for vertical integration, you might box yourself out. And so
I think the people who
build the best stuff and can figure out how to sell it will have a place. And in different markets for example, I believe in vertical integration too because I vertically integrate everything from the call all the way to the API layer.
But I have decided that to reach the widest population of agents searching the web, I stop at the API layer because going further precludes me from using my technology to be super horizontal. And so we're making a technical bet.
Just like these models are really good across many disciplines. The same model is good for lots of different things. Search problem underneath is the same for all kinds of information seeking needs. And if our bet isright, even verticalized players, they might verticalize models, and they might verticalize hardware.
They might use our web search.
Elon and quickfire1:01:40
One of the providers that's going full stack is Elon.
I am fascinated. The world has a perception of him from social media, from everything. You've seen him behind the scenes. What did you see that maybe the world doesn't know about him?
I have lots of disagreements with him.
But I'll share what I think is for founders here, what is, I think, the thing you can admire about him, theurgency and the ability to compress time. And I think having unreasonable expectations of people is mostly a good thing
because most people
don't understand what they are capable of and kind of implicitly sandbag themselves and implicitly set lower expectations of themselves than they're capable of. So when simultaneously inspired and pushed withurgency,
people can do more than they thought. And I think he can sometimes extract that from people. And when that works, it's powerful.
What do you think is the Twitter production direction today?
Listen, I always liked what we called Birdwatch, which is now Community Notes rebranded. I think that's
a good idea. And I'm glad people have continued working on it nonstop. Good people.
Right. We're going to do a quick-fire round. OK?
Let's go.
Would you invest in Instinct $10 billion?
I don't invest.
You don't invest, period?
I don't invest because my wife is a VC. And we have a compliance process that is more trouble than it's worth.
Wow. That's costly, with the greatest of respects.
No, it's a decision. Listen, if I invested, I would invest based on meeting a founder for 30 minutes. I think we have enough exposure to the venture ecosystem. I have no reason for believing I'm a better investor than I have some perhaps network advantages.
And I run into a lot of great founders all the time. Many of them are my customers.
Vinod Khosla, one of your first investors on your board. Biggest lesson from working with Vinod?
Technical intuition, centering a lot of what you do towards a longer-term
technical aspiration. And
as soon as you solve one problem, trying to place bets on the next two or three, preempting technical bets.
What have you changed your mind on most in the last 12 months?
Perhaps the I started with a very pure technical and product focus. The only thing that matters is building the best technology and the best product. Nothing else matters. And now I see week on week value of having highly competent sales and being good at marketing.
And yeah, I think I discounted those things. And it's not that I thought they weren't valuable. But I didn't fully appreciate
the week-on-week visceral delta you could perceive by being good at those things. It's like a naive tech person thing to say, but it's true. You feel that way. And you get itright.
Perplexity or AXA? Which one's a bigger threat?
I don't think Perplexity is in our business. Perplexity is perhaps more of a vertically integrated
product, which competes with Instinct, or a Grokbot, or a Claude Cowork, and all those. So I don't think of them as in our space necessarily.
Web search.
Yeah, like web search is like Web search to Perplexity is like web search if to a lab.
So AXA?
Yeah, AXA is straight up in our business.
Single best VC meeting you've ever had?
Josh and Todd. When Josh and Todd made the decision to invest,
they flew out to spend time with me. That is the single
best VC meeting.
Final one for you. When you look forward to the next 10 years, what are you most excited for?
Chaos.
By that, I mean change. I think in the next 10 years, a lot is going to change. And I think there are people who can build things to have in a world that's going to change fast, make a material dent on where it ends up.
So I feel that me, my company is in a place where we have a role to play in that. What happens to the open web? What happens to content owners?
If we get thingsright, we will get to a better place. And so that possibility
is exciting that it can be a true obsession. And even if the day-to-day is hard and
things don't work for some amount of time, it's totally worth it if you can see that you
bend reality in some way that you care about.
Final one. What belief do you have that sitting around your San Francisco dinner table, your friends would go, well, no, Parag, I don't agree with that one, dude. Not that one.
Again, it depends on which dinner table because the variance in the world is increasing. I don't think people buy this notion that there'll be these agents which are always running all the time for all of us. I think it's going to happen.
But most people don't agree with that yet. You and I are in a bubble. But even a San Francisco dinner table isn't always in full agreement there.
Parag, listen, as I said, I stole the shit out of you. I've so enjoyed this discussion. Thank you so much for putting up with me meandering. And you've been amazing.
This was fun.





