# How to Build Your Own Data Center & Why Every Startup Should Do It

20VC with Harry Stebbings · 2026-09-05

<https://20vc.podhood.com/431bf603-185d-467d-a2cc-b561e5264c99>

Cliff Weitzman, co-founder and CEO of Speechify, tells Harry Stebbings that owning NVIDIA GPUs beats renting because renting costs 1.5x the purchase price yearly, so Speechify pays $100,000 per GPU to get Rubins and B300s early. He covers inference versus training workloads, liquid-cooling and insurance logistics, and NVIDIA's 25% buyback underwriting with Blackstone, BlackRock, Apollo and Goldman Sachs. Weitzman admits his biggest mistake was staying B2C as ElevenLabs leapfrogged Speechify through B2B APIs, and defends entering B2B against ElevenLabs and Bret Taylor's Sierra. He also argues founders should hire math olympiads over coders, details agent orchestration with Claude Code and Cursor, and explains how GPU clusters help him sequence genomes to cure his brother's orphan disease.

## Questions this episode answers

### Why does Speechify buy its own NVIDIA GPUs instead of renting them from hyperscalers?

Cliff Weitzman explains that renting an H100 costs $3.5 to $5 per hour, or $35,000 to $50,000 per year, versus $30,000 to buy one, so renting is 1.5x the cost of owning. Owning also lets him co-locate memory with large training clusters and run open-source models for a fraction of a cent per token.

[2:10](https://20vc.podhood.com/431bf603-185d-467d-a2cc-b561e5264c99?t=130000)

### Does Cliff Weitzman think ElevenLabs leapfrogging Speechify was his fault?

Yes, he calls it the biggest strategic mistake in Speechify's history. He met Piotrek and Mati in 2022, thought a text-to-speech API would become commoditized, and skipped B2B. He now realizes an AI lab's first product is a wedge, and the way to win is to constantly innovate and sell users more products over time.

[23:13](https://20vc.podhood.com/431bf603-185d-467d-a2cc-b561e5264c99?t=1393000)

### Is this the hardest time ever for startups to hire great talent?

Cliff Weitzman pushes back partially: growth-stage companies face brutal competition, since Anthropic and OpenAI pay people $15 million a year minimum. But he argues it is the easiest time ever for seed companies, because raw technical intelligence can be taught fast, so founders can hire math olympiads and Kaggle winners who have barely coded.

[31:20](https://20vc.podhood.com/431bf603-185d-467d-a2cc-b561e5264c99?t=1880000)

### What will be a bigger company in five years, Sierra or ElevenLabs?

Cliff Weitzman says Bret Taylor has the best resume in the world, from Google Maps to Meta CTO to Salesforce co-CEO, and he would never fight him. He believes the AI agents space is enormous, like LLMs in 2019, so both companies will be massive, though they are playing different games: Sierra does tool calling broadly while ElevenLabs stays voice-centric.

[53:38](https://20vc.podhood.com/431bf603-185d-467d-a2cc-b561e5264c99?t=3218000)

## Key moments

- **[0:00] Intro**
  - [0:00] Cliff Weitzman: ignoring B2B was "the biggest strategic mistake I made in the history of Speechify"
- **[0:56] Buying GPUs**
  - [1:10] Cliff Weitzman's Michael Jordan analogy: why Speechify bought its own GPUs instead of renting court time
  - [3:03] Buying an H100 at $30,000 beats renting it for $35,000-$50,000 a year, says Cliff Weitzman
  - [4:55] Speechify still runs K80s: old GPUs are fine for inference at 100 milliseconds, says Cliff Weitzman
  - [6:59] Cliff Weitzman: a GPU is a better investment than a 5% bond because renting costs 1.5x owning
  - [7:15] Inside Speechify's data center: racks of Rubin and B300 GPUs, rented space, and energy as the biggest constraint
  - [11:24] NVIDIA will buy back GPUs for up to 25% of value, creating a liquid secondary market like Elon's SolarCity financing
- **[12:29] Chip buying**
  - [12:29] Cliff Weitzman on NVIDIA circular-economy fears: unlike Bitcoin, a GPU has real intrinsic value in teraflops
  - [13:50] Buying AI chips means Dell vendors, French delivery delays, and paying $100,000 extra to skip the queue
  - [15:54] A truck carrying house-priced GPUs needs insurance, and Rubin racks need sidecar liquid cooling most data centers lack
  - [17:01] Harry Stebbings vs Cliff Weitzman: is owning GPUs a mistake when renting offers flexibility and no logistics?
  - [18:37] Cliff Weitzman: the winning AI stack is team, data, and compute — and owning GPUs unblocks all three
- **[20:56] Data market**
  - [22:25] Speechify's Simba 3.2 API costs $10 per million characters vs ElevenLabs' $100 and OpenAI's $196
- **[22:59] B2B mistake**
  - [23:00] Cliff Weitzman explains how ElevenLabs leapfrogged Speechify: he wrongly assumed text-to-speech APIs would commoditize
  - [25:24] Harry Stebbings vs Cliff Weitzman: is going B2B a mistake when ElevenLabs and Bret Taylor's Sierra own the market?
  - [27:41] Cliff Weitzman: Anthropic and Facebook both came in second and still won — the AI voice space is an oligopoly, not a monopoly
- **[30:18] Compound startups**
- **[31:10] Hiring wars**
  - [31:10] Harry Stebbings vs Cliff Weitzman: is this the hardest hiring market ever with $15 million AI comp packages?
  - [33:45] Cliff Weitzman hires math olympiads who may never have coded: raw intelligence beats handcrafted code in the AI era
- **[37:27] Agent orchestration**
  - [38:34] Cliff Weitzman's milk delivery rule: engineers get no credit until code ships to production with users on it
  - [41:02] A 19-year-old Speechify engineer fixed Cliff Weitzman's entire duplex-model note list while waiting on training runs
  - [42:32] Cliff Weitzman will fire engineers who burn 50,000 tokens on a tiny feature — and the 25,000-node scraping cautionary tale
- **[45:30] AI-era hiring**
  - [45:30] How founders should hire in the AI era: functional interviews, live codebase tests, and hiring for slope over intercept
- **[47:03] Commoditization**
  - [47:03] Cliff Weitzman on Whisperflow and Willow: announcing everything built their brand but invited competition — "only losers compete"
- **[50:18] Customer support**
  - [50:18] Harry Stebbings vs Cliff Weitzman: why the AI customer support market is doomed when Klarna and Navan build their own
  - [53:35] Sierra vs ElevenLabs in five years: Cliff Weitzman says Bret Taylor's tool-calling platform and voice AI will both be massive
- **[55:19] Quick fire**
  - [55:43] Cliff Weitzman predicts voice replaces the screen as our primary computer interface within five years
  - [56:33] SpaceX or Meta: Cliff Weitzman picks Meta because Zuck has 20 extra years, but argues Elon's energy and manufacturing edge is unmatched
  - [1:00:36] Cliff Weitzman is sequencing an orphan-disease community's genomes on his GPU cluster to cure his brother — after tech solved his dyslexia

## Speakers

- **Harry Stebbings** (host)
- **Cliff Weitzman** (guest)

## Topics

Tech Market Trends

## Mentioned

Anthropic (company), ElevenLabs (company), Mercor (company), Meta (company), Nvidia (company), OpenAI (company), Sierra (company), SpaceX (company), Speechify (company), Tesla (company), Whisperflow (company), Claude Code (product), Cursor (product), Simba (product)

## Transcript

### Intro

**Harry Stebbings** [0:00]
You said ElevenLabs kind of leapfrogged you; is that on you?

**Cliff Weitzman** [0:04]
Yeah, 100% it's on me. It was the biggest strategic mistake I made in the history of Speechify.

**Harry Stebbings** [0:08]
How do you reflect on that?

**Cliff Weitzman** [0:09]
So.

**Harry Stebbings** [0:10]
Today's a real freaking discussion. Cliff Weitzman, founder and CEO of Speechify, one of the fastest-growing text-to-speech startups in the world, on the show.

**Cliff Weitzman** [0:19]
The best way to lose is not to be in the race. Be in the race. You don't want to be a fat manager who's like a general sitting in the back saying, "Take that hill." You want to be the warrior who runs up with their sword and engages the enemy first.

**Harry Stebbings** [0:30]
Ready to go?

**Cliff Weitzman** [0:42]
Cliff, it is so good to have you back in the studio, dude. I was looking forward to this one, because when I was writing it up, it's a very different thread of conversation to how I'd normally go. And so thank you so much for joining me again today, dude.

**Harry Stebbings** [0:54]
My pleasure. Glad to be here as always.

**Cliff Weitzman** [0:56]
Now, I wanted to start with you're spending tens of millions of dollars on NVIDIA GPUs, and you're paying an additional $100,000 per GPU to receive them four months early. Why? Like, what do you know that the market doesn't know?

### Buying GPUs

**Harry Stebbings** [1:10]
So in 2022, we bought a huge rack of GPUs from NVIDIA. And the reason we bought them is for training,right? We have a bunch of models. The newest Speechify Simba 3.2 model is ranked number one in the world for quality, above all the frontier labs, 10x more affordable than stuff like ElevenLabs.

And we used to rent GPUs. And we found that engineers at Speechify would be parsimonious with how they used the GPUs, because they were like, "Oh my God, I'm costing the company tens of thousands of dollars. Like, I don't want to do that."

And the analogy my brother and I came up with is, imagine you're Michael Jordan, and you want to be in the NBA. It's the only thing you care about. And you need to pay $20 an hour just to train in a basketball center.

Well, that sucks. You want one that you can go to whenever you want to. In fact, you want a hoop in your house. And so our initial idea was, we want a hoop in our house. And so we bought a bunch of our own GPUs.

And that deal ended up being really good for us. And we ended up training really good models. So with time, we invested more and more and more and more. So that's the first part. The second part is actually how the economics work out.

So if you look at it, the Transformer was invented inside of Google in 2017. NVIDIA came out with A100 GPUs in 2019. Shortly after, they came out with H100 GPUs,right? The original ChatGPT was trained on A100s. And then they came out with Blackwells, so then B200s, B300s.

And now they came out with Rubins, which is the GPUs that Elon is sending to space. And they're like liquid-cooled. They're very, very cool. And we're like, OK. Huh. One, every class of GPU is more affordable per 1 trillion flops,right?

So a flop is addition, subtraction, multiplication, any mathematical operation. And you measure them in how many trillion of operations happen per second in a GPU. And so the more affordable as it relates to this. If I was to buy an H100 for, let's say, $30,000, that's how much kind of a single card would cost.

If I wanted to rent an H100 for one hour, spot instance, from GCP, it could cost me $5. If I rented it from, like, you know, Azure or AWS, maybe it'll cost me $3.5 per hour. So if I multiply that times 24 hours, and then times 365 days in a year, I'm actually going to end up paying $35,000 to $50,000 to rent that GPU for one year, but I could buy it for $30,000.

So it's 1.5x the cost of owning the hardware to rent the hardware for a year. Now, the hardware is typically warrantied for three years to work properly. But it'll keep working after the warranty for, I imagine, I don't know, 10 years.

So the math just maths where it makes way more sense to buy them. The other big part is, if you want to do large-scale training like we do, you need the memory to be co-located with a large cluster of GPUs.

I can't just rent from Google or Microsoft or even Base10 and run the size of training that I want, because I need a gigantic memory card next to it with all of my data that all the GPUs are accessing.

So that's why we first started buying them. The next thing that we found is, actually, if you run open-source models for coding, you could pay Anthropic. And then, you know, you're paying for all the tokens and the fact that you're doing the branded,right, Fable 1.

Or you can run an open-source model, and instead of running it on a spot instance from Azure or anyone else, you run it on your own hardware. And then you're paying a fraction of a fraction of a cent per token.

And so for all those reasons, it made a ton of sense. But we can go into all the depth that you want. I just want to dig in. The first thought that I have is, I completely understand the rationale.

But chips depreciate. You have chip cycles, and they are accelerating. We are seeing newer and newer chips being created. We're seeing specialization within chips by buying your locking yourself in, so to speak, to one chip architecture. How do you think about that?

**Cliff Weitzman** [4:55]
At Speechify, we still use K80s for a lot of specific operations for inference. And we use older models of GPUs constantly. And there's essentially a difference between when you do inference and when you do training. For training, I'm like, OK, I have this hypothesis.

I want to know the answer to this hypothesis as soon as possible. Like, every minute that it doesn't come out, I'm in competition with everybody else. And so having a GPU architecture that is much faster by orders of magnitude is a huge advantage.

But if you go speech to text or text to speech with Speechify, I can afford to give you a lower quality GPU, and it'll give you what you need still in, you know, 100 milliseconds. So it's like totally good.

And so I can always use these older GPU models for inference. That's number one. Number two, we have so many experiments that we're running at every single point in time. Not all of them need to run on, like, the newest hardware.

So the analogy I always give, let's say you bought an iPhone back in 2011, and it's an iPhone 3G. And then you bought another iPhone and another iPhone and another iPhone. You could have a drawer in your house with, like, five iPhones that are collecting dust, because you can only use one iPhone at a time.

But if I own 100,000 GPUs, I'm still going to use all of them at the same time. And so I'm not losing anything by having more GPUs, because not only do I own a bunch, I still rent from the hyperscalers all the time.

And I rent both dedicated instances that I prepaid for, and I rent spot instances. For example, more people use Speechify in September, because everybody goes back to school. So I need to, like, level out the load. And so the parts of that load that I know for sure I'm always going to use, whether it be training or it be inference, I might as well just own it.

And then on top of that, it's also the case that I have so many other friends who are running training and running inference. I can always rent it out to other people if I have excess capacity, which I don't expect to have.

But, like, every once in a while, you have an interesting situation. So for all those reasons, it just makes mathematical financial sense. Lastly, if you have excess capital, really, you either stick it in the bank or you buy a bond,right?

Like, the best long-year the best bond you can buy long tail, I don't know, will yield you like 5%. Or you can buy a GPU. And because renting it would cost me 1.5x, buying it for the year, the return is, like, way higher.

**Harry Stebbings** [7:12]
So how many GPUs do you buy, then?

**Cliff Weitzman** [7:15]
So let's talk about Rubins, for example. So Rubins come in the form of 72 cards in one rack. So we'll buy multiple racks of Rubins. And then on top of that, we'll buy B300s, which are, like, the newest form of Blackwells, because we can get them earlier.

And then the same thing, like, you know, when we bought our first instances of DGX H100 GPUs, we bought just, like, a bunch of racks of those. And then those get delivered in a truck to the data center.

We rent the data center space. And so the data center provides the networking capability. It provides the energy, which is actually the largest constraint now. And it provides, like, physical engineers that take it off the truck. They install it.

If it has an issue, they fix it. And then it just runs.

**Harry Stebbings** [7:57]
Does ElevenLabs do this?

**Cliff Weitzman** [7:59]
Yeah. ElevenLabs is amazing at this. ElevenLabs, I think Piotrek at ElevenLabs literally bought a bunch of GPUs early, early on and set them up in his house. And then they just kept building bigger and bigger and bigger clusters.

They do the same thing that we do.

**Harry Stebbings** [8:11]
How do you think about forecasting chip buying? It's incredibly difficult to know, A, demand, but also, B, supply of chips. How do you think about forecasting chip purchasing?

**Cliff Weitzman** [8:21]
Yeah. So number one, I want to explain again, it's very different than buying an iPhone or buying a MacBook. I can only use one MacBook at a time, one iPhone at a time. But I can use all the chips I have at any given point in time, and I still will have more demand, especially when I have multiple teammates and 60 million users who are using inference on my Speechify software that's providing text-to-speech and helping them, you know, read their work and dictate their work and, you know, use Speechify work, which is our newest product.

It's a gen tech, kind of like Jarvis from Iron Man. And so I go, OK, let's imagine I have 100% capacity that is the average usage per month that I need for GPUs for training of my AI models and for inference on my AI models.

Inference is when you actually make a call to Speechify, and you give me text, and I give you back audio. Like, there's math that happens in the background. That's inference. Training is I take a gigantic amount of data.

I take all the architecture and software engineering that we're doing, and I go, I think that this will give me a better model. I'm kind of baking that model in the oven, and I'm going to come out with a new black box.

And then when I give you text, that black box is what calculates it and gives you back the audio. So those are the two usages. Let's say I have 100%, which is what I would have in, let's say, a month like November.

In October, I'll have 140%, because it's, like, a big month for us. In December, you know, everyone's at home. You know, they're not necessarily studying or working. So I might have 80% utilization. OK. So I go, cool. Well, I can take 20% of the usage that is normal and let me buy it, because it's the best deal.

I'll take another 25% of the usage, and I do long-term contracts with hyperscalers. The rest, I'll rent what's called spot instance from the hyperscalers. And then I'm still not even close to over-committing myself. And so that's kind of how we think about the math.

And then we go, OK, well, also, we have 45 engineers. But we want the team to be 150 engineers. And even inside of my 45 engineering person team, like, there's a couple of people who are rock stars. They have dedicated, like, DGX racks just for that one person.

And 25% of my team are almost, like, waiting. And I want to double the size of the team. They just need, like, it's like you have a football team, and you just need another field, because they don't have enough field to practice on.

And so that's how I think about how to allocate. And then in terms of depreciation of the asset over time, I go, OK, well, these are still amazing GPUs. Like, even A100s, you can run amazing experiments on. So it's completely valid to use that as long as it's hooked up and as long as it's not stopping to work.

And so think about the mileage of a car,right? If a car gets, like, 250 miles, you know, it's kind of going to break at this point. That's not necessarily true for a GPU, because it doesn't have as much wear and tear.

Yes, it's moving, and yes, all these things. But, like, it's in a very clean environment. It's very much cooled. It has constant maintenance, because it's not moving around. It's very expensive. And NVIDIA just does a really good job.

And so that asset is going to stay for a very long time. And let's say it got so not good, so outdated, that I no longer can run training on it. Cool. Now I'll use it for inference. There's one more thing that's very interesting that just happened.

So I believe earlier this month, NVIDIA did a huge deal with Blackstone, BlackRock, Apollo, and Goldman Sachs. And they said, listen, we want more people to buy more GPUs. We're going to underwrite for you up to 25% of the value of a GPU that if you lend money to someone who buys a GPU, let's say Google

or a startup, CoreWeave, and that startup goes out of business, and you have that GPU as collateral against that investment, we'll buy back the GPU for up to 25% of the value of the GPU. And so they're succeeding in creating a liquid secondary market for GPUs that they're underwriting.

So now the large banks have an incentive to loan money at much better interest rates. This is actually exactly what Elon did in the beginning of Solar City. He went to Morgan Stanley and Merrill Lynch and got them to amortize the price of a solar panel over 30 years.

So the whole invention behind Solar City was the fact that you could take a loan against the collateral of your solar panel. So NVIDIA has done an amazing job now in creating a clear floor for the value of the GPU over time.

**Harry Stebbings** [12:29]
Do you think the circular economy fears that people often cast against NVIDIA are justified or not? We saw their CFO push back on them and say, enough. Enough of this bullshit. Do you think they're justified or not?

### Chip buying

**Cliff Weitzman** [12:41]
I think that a lot of the things that about a year ago, like, were going on between Oracle and OpenAI, like, that was way too much. Like, that was ridiculous. I think the NVIDIA stuff is not, because you're talking about a real asset.

So if you think, for example, about the logic behind the value of Bitcoin,right? Bitcoin, what is the intrinsic value of Bitcoin? I can't really tell you,right? What's the intrinsic value of gold? Well, gold, you can use it for some medical stuff, because it's a really amazing metal.

And, you know, it's jewelry, whatever. But a GPU, it has intrinsic value. Like, you can actually use that asset for something that's really, really valuable. And it doesn't matter where that GPU is. It could be in Iceland. It's still useful to anybody all over the world as long as it's networked.

And so, actually, it has a pretty good store of value, even if new GPUs come online. Really, my one question, and this is the math for everybody to come back to, is how many teraflops per second can this device do?

And that is, like, essentially a token. That's the value. And so, like, there is intrinsic value. So yes, you can have all these, like, circular things. But at the end of the day, NVIDIA is making a project, a product that's real.

It's not complete tool-of-mania. Like, there's a real, real intrinsic value here.

**Harry Stebbings** [13:50]
What does no one know about buying chips that they should know? What's the, like, oh my god, people are so naive about this?

**Cliff Weitzman** [13:58]
I mean, it's not that people are naive. It's just they haven't been in the space. So I'll give you an example. Imagine you're buying a GPU. Well, you're going to buy it, you know, as an NVIDIA-produced product. But NVIDIA is not going to waste the time talking to Cliff Weitzman.

So who do I buy it from? Well, one of the best rated vendors is Dell. So everybody thinks Dell is a personal computer company. No. Dell is a GPU rack supplier at this point. And then, OK, I want to buy it from Dell.

Well, Dell has a constraint, because there's not a lot of, like, you know, Blackwells out there. Well, it happens to be that they have some in France. Allright, well, I'm going to order mine from France. OK.

Shoot, it was supposed to come a month ago, and it's still not here,right? Because of whatever. Like, you know, there's demand. So then you have to negotiate to make sure that you get it, which is why we're very willing to pay 100K per month extra to get them earlier.

**Harry Stebbings** [14:48]
So you'll call up Pierre in France and say, you know, hey, we'll give you an extra 100K kicker if you get them here in a month?

**Cliff Weitzman** [14:55]
Even more than that. So in that France situation, which is something that happened to me, I was like, Pierre, what the heck? We have a contract. You're not delivering on time. And so it is the case that we had a contract with another company beforehand.

And they were, I don't know, a few weeks late. And I called them, and I was like, listen, I've got a better deal. I'm canceling our contract, because you're you're out. Like, you didn't deliver. So I'm going to go with this other contract.

But if you have a better price, like, we'll go with you. But, like, I just need the GPU now. And remember, I'm paying for the renting space of my data center. So the most expensive part of a delivery of a GPU is if it's late.

I'm still paying rent for that data center space. Now that GPU and so, you know, you put pressure on Pierre to send you the thing when he said he was going to send it to you. And then you go to NVIDIA or Dell or whatever, and you're like, well, it's a market.

Hey, can I pay more to get it earlier? Skip the queue. Yeah, you can. Cool. Now there's a truck somewhere in the United States with a GPU whose value is the value of a house that's coming to my data center,right?

Well, I should have insurance on that,right? Because if that truck gets hit or there's too much humidity or the GPU gets flipped, like, I lost multiple houses' worth of GPUs. So, OK, the insurance is really, really important. And then also, the value of the amortization is really, really important.

And it's like, there's all these nuances of how to do the math through. Then there's the cooling,right? So, like, you're not only paying for the physical space and the networking and the energy. And the energy is the biggest constraint.

We'll talk about it in a second. Well, how do you cool that thing? Because you have a thing that's just, like, moving and moving and moving and moving and moving. Well, the thing that's, like, most new now is liquid cooling, because air is just not enough.

And the thermal load of water is much better. And there's other liquids that are even better than water. And so Rubins are liquid-cooled. But most of these data centers don't have liquid cooling installations already approved. So we had to do a bunch of research, and we find, OK, we could buy what's called a sidecart of liquid cooling that you enter into the data center.

Then you pay someone at the data center to install it for you. Cool. Now you can have, like, the rack that you want. And so there's a big difference between running a purely software company and running a company that includes hardware.

And.

**Harry Stebbings** [17:01]
But when I listen to all of this, I'm now more sure than ever that it is a mistake to price-optimize and to spend the money to buy it versus to rent it, because I get you on the optimization.

But you're not saving 10 times more. It's 0.5x more per year.

**Cliff Weitzman** [17:17]
No, no. Per year. Exactly.

**Harry Stebbings** [17:19]
Yeah, per year. But you have the flexibility to tailor it up and down. You don't have any of the logistical nightmares of insurance, transportation, security, water cooling, logistics. And then you can build your product actually, what matters most, against ElevenLabs, who are fucking running fast.

I don't want to worry about water cooling and insurance for a freight truck.

**Cliff Weitzman** [17:40]
ElevenLabs worries about the same thing, because for them to train excellent models, they need to have co-located GPUs with a lot of memory available.

**Harry Stebbings** [17:47]
You can't do it if you rent it.

**Cliff Weitzman** [17:49]
You can. It just becomes, one, ridiculously expensive. Two, you need to commit for many, many years ahead of time, because you need to build a co-located cluster. And then you don't have as much control, because you don't own it.

So it's, like, difficult to, like, you suddenly need an InfiniBand cable, which allows for the memory to flow from one DGX to the other one. And the answer then is, well, now my ability to train is so much bigger.

I can have a larger AI team. Every person in the AI team is leveraged. And I could just, I could shoot ahead of everybody so much faster. And let me just make one thing clear. If I want a Rubin, which is, like, these much faster GPUs, I'll get it faster if I buy it than if I wait for Google to buy it.

And then there's other people in front of me in line. So I'm going to skip the queue by, like, a lot. And then I'm going to have, like, a year of access to Rubins before everybody else does. The way I think about it is the following.

How do you build an amazing company in a world where there's so much competition today? The team is the most important part. But the team is the most important part, because the team gets you the other resources. And so what are the missing pieces?

The missing pieces are data, compute, and architecture. In a world where intelligence is commodified and no one needs to handwrite code at all anymore,right? Our engineers, really, what I'm looking for is 10 really good decisions per day, which is very tiring, not, like, optimizing random parts of the code.

And each one has, like, you know, 5 to 18 agents running at any point in time, doing long-horizon tasks on these GPUs, coming up with theses, testing them, going back and forth, back and forth, back and forth. If they don't have the capacity to train, the team is limited.

If they don't have the data to train, the team is limited. And by the way, a lot of data is cleaning the data,right? You get this raw data in the beginning. Well, you need to organize it into data sets.

And so you need the GPUs to also organize the data sets, too. Like, I have one of my best engineersright now. He's not even writing models. He's making synthetic data sets to train models. And so, like, it really becomes an indispensable asset.

**Harry Stebbings** [19:44]
Would you ever buy data?

**Cliff Weitzman** [19:46]
We have, but, like, small data sets.

**Harry Stebbings** [19:48]
So I suggest you use Fireworks.

**Cliff Weitzman** [19:50]
OK.

**Harry Stebbings** [19:50]
But, I mean, Fireworks is amazing. Linda and the founders, she's one of the co-founders of PyTorch. But I had the very obvious realization that you'd have every company having their own specialized models of a certain size trained on their own data.

But you would need supplemental data.

**Cliff Weitzman** [20:06]
100%.

**Harry Stebbings** [20:06]
Like this synthetic data or real-world data that you just don't have, and that you would buy that from data providers like McCall, which is why I was saying.

**Cliff Weitzman** [20:14]
Micron One, Surge, all these companies are amazing. And they shorten the cycle, by the way, to get into revenue.

**Harry Stebbings** [20:19]
100%.

**Cliff Weitzman** [20:20]
Because if you are a Mercor, shout out to Brendan Foody, ElevenLabs or OpenAI or Anthropic is going to make money for the next decade or two on the data that they buy from you. And so they're willing to pay a fraction of that 10 years of revenue to you today to supply them the data.

And again, it's all a speed thing. Yes, ElevenLabs or, you know, OpenAI can go and make a team that will get the data, but they don't want to manage it. And so the data is key. All this, like, you need all three things.

You need compute, you need data, and you need team that writes great products. And ideally, you need users that use you a lot and a lot of them to have a feedback loop of whether the stuff is good or not.

So benchmarking.

### Data market

**Harry Stebbings** [20:56]
And so for me, when I was doing the, as a venture investor, we do outcome scenario planning, which is the most bullshit exercise to pretend like you're smart predicting the future. We do it because, you know, it makes us feel important.

But, you know, they've predominantly sell to Frontier Labs today. And that's where 90% of their revenue is from. With the rise of specialized models on a per-company basis with their own data, I believe that you move that customer base from purely Frontier Labs to every large-scale enterprise who needs supplemental data.

If that is the case, how big an outcome is the data marketplace?

**Cliff Weitzman** [21:30]
So the first problem to understand about the data marketplace is it's not ARR,right? It's not annual recurring revenue. It's one-time deals every single time. So the buyer of the data is not required to buy it from you again.

So it's a very risky business. And if you look at early days of companies like Mercor, they didn't raise significant funding off the bat because investors were very skittish about that fact. Let's put that aside. Well, very important that the Ts are crossed and Is are dotted about how you got that data,right?

So, like, you know, like, you need to indemnify the companies who are using you. And that's part of why they buy it from you as opposed to sourcing it themselves. We've seen the lawsuits. But it's a great business.

And if you could do it well, but, like, you need to be an ops monster. Like, you need to be really, really good at operations. You need to be very fast. And really, the key is the company training on your data needs to actually see improvements in their model at the end of the day.

The thing that has always been challenging for Speechify compared to other companies is B2C customers pay a lot less than B2B customers. So ElevenLabs, huge credit to them, leapfrogged us because they sell to B2B. Well, historically, we've only sold to B2C.

And so our big constraint is we needed to do this on a cost basis of it needed to cost us, you know, less than $10 per million characters. Eleven charges $100 per million characters. The OpenAI model on the benchmarks costs $196 per million characters.

So ours, when we sell it to other B2B companies now, we just launched our API, Simba 3.2. It costs $10 per million characters.

### B2B mistake

**Harry Stebbings** [23:00]
Dude, I am I'm too old to not ask the painful questions. And I think the joy is when you asked him and kind of less, you know, worried you were about asking, you said ElevenLabs kind of leapfrogged you.

Is that on you for not doing B2B?

**Cliff Weitzman** [23:13]
Yeah, 100% it's on me. 100% it's on me. It's the biggest strategic mistake I made in the history of Speechify.

**Harry Stebbings** [23:17]
How do you reflect on that?

**Cliff Weitzman** [23:19]
So I met Piotrek and Mati. I was living in London at the time, in my house in London. I think it was 2022. And we were very impressed by them. And we wanted to use their model, by the way.

It was just too expensive for us to use. And I looked at it, and my thought to myself was, they're very smart. They're going to do well. But I don't like their strategy, because I think that an API for text-to-speech will become commoditized with time,right?

You're going to get to the point that you can run that API on your computer and then on your phone. And then, like, what are they selling anymore? So I don't want to go into that business. And I made a critical error.

What I didn't understand is that the point of an AI lab like Speechify or like ElevenLabs is to continuously innovate. And the first product that you release is your wedge that gets other people to then later use your other technology.

So for example, if you're in text-to-speech, you build the best text-to-speech model in the world for one specific voice. Cool. Well, now you can do other voices. Now you can add emotional prosody. Now you can add voice cloning.

Now you can add speech-to-text. Now you build duplex models where it makes the um, ah, laughter, interruption handling, turn-taking. You add a harness for voice conversations. Then you optimize it for sales. And you optimize it for customer support.

And you optimize it for all these things. And so what they did is they first built an amazing API. They were great at launches. They built a really great product for creators. Then they built their best product ever, which was agents.

Agents is amazing because the buyer is no longer a software engineer. The buyer is a CTO, CIO, CEO, executive in the company. Sierra has this concept called outcome-based pricing. Bret Taylor is amazing. And so you can start finding on the outcome.

And having an AI agent is like having an AI coworker. But it was my mistake to think that an API product was a bad strategy, because I thought it was something that would become commoditizable. And I forgot the central thesis about selling and value, which is constantly innovate.

Get the user to start using your product. I don't care if it's free. Then you sell them other things. And so that was my big, big, big, big mistake.

**Harry Stebbings** [25:24]
How possible do you think it is? I think people underestimate the complexity of building out a B2B GTM.

**Cliff Weitzman** [25:29]
Super hard.

**Harry Stebbings** [25:29]
I think it's a strategic mistake for Speechify to go to B2B.

**Cliff Weitzman** [25:34]
Hmm. A lot of people think that. Tell me your position.

**Harry Stebbings** [25:37]
You are now competing against ElevenLabs and Sierra, really. And those two are competing. Whether they like to admit it or not, they absolutely are competing. And they will, I'm sure, if you ask them off camera. That's Bret Taylor.

**Cliff Weitzman** [25:50]
Yeah, you don't want to compete against Bret Taylor.

**Harry Stebbings** [25:51]
Motherfucker, I don't want to compete against Bret Taylor. And that is the tidal wave of ElevenLabs. Now, ElevenLabs is an unstoppable machine at this point, to the point where it has government buy-in across all of the large major Western democracies.

Actually, it's insane the government buy-in they have. And they started three months ago.

**Cliff Weitzman** [26:10]
You just said the key thing. They started three months ago.

**Harry Stebbings** [26:12]
Yeah.

**Cliff Weitzman** [26:13]
And so you know the graph of OpenAI.

**Harry Stebbings** [26:15]
But I think they've reached a tipping point where, actually, they've just taken the market. I think Sierra are running behind them chasing. And they're doing a decent job of it. But they've got Bret and they've got Sequoia and GreenOaks and every royalty of Silicon Valley behind them.

And they're still running behind chasing ElevenLabs, with Sequoia kind of pretending to be neutral, because they're in both of them, which is incredibly challenging. And I just think being third, the postmates effect, is never a good market to be in, when I could be the dominant consumer brand that leads with a really different and compelling story.

**Cliff Weitzman** [26:51]
So here's the two things to consider. The first one is, if you go to the App Store and you search text-to-speech, Speechify has 98% of the installs in text-to-speech for B2C. Speechify has served more than 770 billion words to users over the last few years, which in terms of times of listening, it's like 6,000 years of listening,right?

If you go from today to 0BC and back, you still have, like, thousands of years left. So we've, like, completely dominated that market. And it's still a business that's growing really, really fast. And we're constantly adding more features into that product.

The thing is, we have a pretty big engineering team. And now everybody is capable of doing 10X what they did before. So I have extra staff. I have a huge AI engineering team with the ability to make amazing models.

So, like, where is the highest ROI for that to go? Well, it needs to go both B2C, but it should also go B2B. And one thing that I will never be is a person who doesn't learn. So I might as well just freaking learn B2B.

Now, to your point about competing against giants like Sierra or ElevenLabs, hey, Anthropic came into the market as a second to OpenAI. And they were second for a very long time. And now they're not second. Facebook came as a second to Friendster and MySpace.

And now they're not second. And so the nice part is this space is not a monopolistic space. It's an oligopoly space. And if you look at what happened with ElevenLabs, I'm going to exclude Sierra because the Bret Taylor effect is huge.

It's just amazing to see how good of a business that is. And so it might very well be that for the core offering that they're currently winning on, I will not win. But what did I learn last time?

It's fine if I offer my product essentially for free because I'm an AI research lab. And as long as people start to use me, with time, I'll be embedded in the system. And I'll keep coming out with more and more and more innovations that are useful to them.

And so there's an unbelievable demand from all these companies and governments and everybody else for great tools, whether they be AI agents or APIs or products. I just want to be on your phone if you're a user or in your stack if you're a company and supply you with the best front-to-point engineer experience and AI orchestration experience and API experience to give you an amazing experience.

And there's room for everybody.

**Harry Stebbings** [28:55]
I agree there's room for everybody. I think value accrues to top one player.

**Cliff Weitzman** [29:00]
I agree.

**Harry Stebbings** [29:01]
I think, you know, it's kind of like the inference market where Fireworks will be a multi-hundred billion dollar company. And then, like, I think a GenuineBaseM will be a $100 billion company. And then Together and a load of the others will be $15, which is amazing.

**Cliff Weitzman** [29:12]
It is completely true.

**Harry Stebbings** [29:14]
Hugely amazing, valuable companies.

**Cliff Weitzman** [29:16]
But you would then think that OpenAI would be the place where value accrues for Voice AI,right? That's what you had thought three years ago. And that's not what ended up happening. So you can't not go into the race because there's a big incumbent.

**Harry Stebbings** [29:26]
Well, I think with all candor, that's because of incredibly poor management.

**Cliff Weitzman** [29:30]
I agree. But that's the thing.

**Harry Stebbings** [29:32]
And hiring. And, like, that was theirs to take. And they fumbled the bag across every spectrum.

**Cliff Weitzman** [29:37]
And for every company in the world, no matter how exceptional the leadership team is, niches get fumbled,right? So Voice AI was a niche for OpenAI,right? LLMs were the core. And by the way, they also fumbled AI coding. Now they're trying to cache because it's such a big space.

All respect to Piotrek and Mati. I think they're absolutely amazing. And I love working adjacently to them.

**Harry Stebbings** [29:55]
I just don't think they're going to fumble the bag. That's my trouble.

**Cliff Weitzman** [29:58]
But they have so much in their netright now.

**Harry Stebbings** [30:02]
That's true.

**Cliff Weitzman** [30:02]
And so much is getting added to the net constantly.

**Harry Stebbings** [30:04]
That's true.

**Cliff Weitzman** [30:05]
And so you just you have to, like, go where the football is going.

**Harry Stebbings** [30:08]
Yeah.

**Cliff Weitzman** [30:08]
And so I think that it's too expensive for Speechify not to be playing in B2B as well as playing in B2C. The best way to lose is not to be in the race. Be in the race.

**Harry Stebbings** [30:18]
In terms of the products that we built, we were chatting earlier. And you said that every startup, say, has to be a compound startup. Can you talk to me about that and how you think about that?

### Compound startups

**Cliff Weitzman** [30:27]
It's not that every startup has to be a compound startup. It's at a certain point you can't afford not to be that.

**Harry Stebbings** [30:32]
Do you not think there are a few companies that are just absolutely fucking running rings around everyone else?

**Cliff Weitzman** [30:38]
Yeah, absolutely. Those are the winners,right? ElevenLabs is an example. Anthropic is an example. Ramp is an example. Speechify is an example. All the companies that have absolutely maniacal leadership teams and engineering teams, like, that's why people care about a team more than almost anything else.

Because theright team will iterate fast, get there, and then figure it out. And now, when everything can be turned into a reinforcement learning problem where you can have long-horizon agents and orchestrating agents thinking about the problem for, like, two weeks at a time, if you set that up, of course you're going to win.

### Hiring wars

**Harry Stebbings** [31:10]
I got into a lot of trouble, as I always do with most of my social posts. I used to be quite a sweet little boy, actually. No, really, I used to be like the Harry Potter revenge capital. And now I'm more like.

**Cliff Weitzman** [31:19]
Yeah, you lost the glasses.

**Harry Stebbings** [31:20]
Lost the glasses. And it kind of became more like Piers Morgan, if you know Piers Morgan in the UK. Highly despised figure. Very opinionated. But a question that I have, it's like, I said, if you're a startup, it's never been harder to hire great talent, because OpenAI and Anthropic, candidly, have such a carrot reward mechanism in front of you, especially with impending IPOs, that the best talent just wants to go there.

And talent follows talent. And you're seeing the fucking founder of Monzo, a multi-billion dollar bank in the UK, go there from YC as a partner. Matt Clifford, the founder of EF, which is a multi-billion dollar company. I mean, he should be fucking prime minister.

And he's going to join Anthropic. Am I wrong that this is the hardest time ever for startups to hire, because the prizes of Anthropic and OpenAI are so great?

**Cliff Weitzman** [32:11]
My favorite type of person to hire is a CTO of another company. We have, when we were 21 people at Speechify, 18 of the folks at the company were previously either CEO, CTO, or VP of engineering at their last company.

Anthropic, I have never seen a company like this, hires so many CTOs of publicly traded companies and other successful startups.

**Harry Stebbings** [32:30]
Workday, one of them.

**Cliff Weitzman** [32:31]
The reason is they build the best, most beloved product for engineers in the history of the world. So it's easy to hire CTOs. By the way, they hire much more CTOs than CEOs, because CTOs are the ones who get the most excited about this product.

And, like, you'reright. They're the fastest growing company ever, especially at the scale that they are. So they're going to keep growing. OpenAI is going to keep growing. You had this, like, very condensed period, like fireworks of growth in both of those companies.

Yeah, it's very hard to hire. But remember, they're hiring people that their annual compensation needs to be $15 million a year minimum. What startup is hiring someone and paying them $15 million a year? You're not. Like, your seed founder, that was not someone that you were going to hire.

And so I will push back against it. The competition for growth stage companies hiring exceptional leadership talent is more difficult. For seed companies, I would say it's the easiest time ever, because the impact of even just the founder on their own is bigger, because they can orchestrate agents.

But the same thing for hiring. So one thing that we have changed about our hiring in the last even six months is we really cared that you read a ton of textbooks about software engineering and that your handcrafted code was amazing.

I still care that you read a lot of textbooks about software engineering and you understand it. But the thing I care about the most today is technical aptitude and just, like, raw technical intelligence, because I know that we could teach you everything else.

And in six months, you could be a machine. And so we hire a lot of math olympiads and leetcoders and, like, Kaggle award winners and people who, like, studied physics and math. Like, they might have even not coded before, because I just need the hunger and the work ethic and the intelligence.

And anyone can become so good so fast now. And so the pool for hiring exceptional talent is bigger than ever before. And Duolingo did this really well. They love hiring college grads and then coaching them. And so I wouldn't say that it's harder to hire than ever before for seed companies.

Seed companies now, almost anyone can be someone that you hire if they're smart and hardworking, because you could teach them very fast. What is more challenging to hire is for growth companies, because you're fighting with just absolute juggernauts.

**Harry Stebbings** [34:42]
Are you not a growth company?

**Cliff Weitzman** [34:44]
So it's challenging for us. Why do you think it's hard to hire a really good salesperson?

**Harry Stebbings** [34:48]
I totally get that. And I completely agree. I will see CRO packages in the $50 million plus range, by the way.

**Cliff Weitzman** [34:53]
Yeah, exactly.

**Harry Stebbings** [34:54]
15 is, like, kids' play.

**Cliff Weitzman** [34:56]
Hire.

**Harry Stebbings** [34:56]
By the way, by the way, with the greatest of respects, I will even see $15 million on the table for comp packages for seed companies today. That is the dislocation that I think, with the greatest of respects.

**Cliff Weitzman** [35:07]
Sorry, sorry. But so this is a seed company that has raised how much money at what evaluation?

**Harry Stebbings** [35:12]
Well, I mean, you've got to understand that CDR and SA will be $150, $200 million.

**Cliff Weitzman** [35:17]
And.

**Harry Stebbings** [35:17]
And there are several of them. I mean, there's 30, 40 companies that at seed have raised $100 to $300 million.

**Cliff Weitzman** [35:23]
And this is a company of, like, a guy who's, like, one year out of university?

**Harry Stebbings** [35:26]
No, no, no. This is a guy who's probably spent four years at OpenAI or spent four years at DeepMind.

**Cliff Weitzman** [35:31]
So then what about the company that's, like, you know, the guy who's been in university for, like, two, three, four years. And now they're starting a company. Or do you think that those people are out of the water now?

**Harry Stebbings** [35:39]
No, I think that that's just a very different world. And so, yeah, they'll raise $10 million seed rounds.

**Cliff Weitzman** [35:43]
Yeah. So for the company that you just described, they raised a seed round at $150 valuation. And they raised, I don't know, $20 million?

**Harry Stebbings** [35:51]
No, I said it was a $150 million raise.

**Cliff Weitzman** [35:53]
Oh, I wouldn't call that a seed round.

**Harry Stebbings** [35:57]
But my point is, and that's my point, though, which is, like, the talent is concentrated. The people who really fucking get AI and systems and have seen the magic inside OpenAI, Anthropic.

**Cliff Weitzman** [36:09]
I agree with you that if you have a company that's raised $150 million at $500 to $2 billion valuation, definitely that company should give $15 million comp packages. No question.

**Harry Stebbings** [36:18]
And there's a lot of them.

**Cliff Weitzman** [36:19]
Yeah, that makes perfect sense.

**Harry Stebbings** [36:21]
And but there's a lot of them. There's 30. And those 30 take 30 people. And there is 1,000 people now. That is fucking hard.

**Cliff Weitzman** [36:28]
And so but what you just described is exactly what used to happen with Google and Meta, let's call it, six years ago, which is if you were really cracked, there was essentially a maximum amount that you can get paid at a company like Google or Meta.

And the best way for you to make a life-changing amount of money is to go to a company that is small and ride from the beginning all the way through and be a really solid founding engineer at that company.

**Harry Stebbings** [36:50]
I think people want more certainty of cash today than upside, which sounds more.

**Cliff Weitzman** [36:55]
No, I think that the equation is the same as always, which is each person has their own equation in their head of how much certainty and how much risk they're willing to take. It hasn't changed. It's the same.

Humans are still humans.

**Harry Stebbings** [37:04]
But I think people would rather know that the certainty of a $10 million from Anthropic versus a $60 from that quirky startup they could make.

**Cliff Weitzman** [37:14]
This is the reason why companies IPO,right? There's two reasons. Either you want a ton of money, or you want a lot of credibility in B2B, like Zoom did, or you're hiring. And the value of the package that you offer is so much better when your stock is liquid.

**Harry Stebbings** [37:27]
When we look at that dev team feed today, you said, hey, I wanted to go in. I want to see how we're orchestrating agents. What did you find? What did you learn in that discovery process around agent orchestration internally?

### Agent orchestration

**Cliff Weitzman** [37:39]
So inside of our AI research team, everybody's orchestrating agents. It's when you go lower not lower. If you go then into the product-facing things that we build, for example, the platform team or the iOS team or the Mac team or the Chrome team or the web team or the Android team, these are super smart folks who have been working in those domains for, like, 10 years.

And they know iOS like the back of their hand. They know Kotlin, JetBrains, like the back of their hand. And so it's very easy for them to hand code things, because you're not dealing with something that's, like, super, super new.

So why change? People, you know, it's hard to change,right? And so you just need to force them to change. So one, the best thing is to inspire. So you do a Zoom screen share, and you show them how the best engineer in the team is orchestrating agents.

And they're like, oh, wow. I didn't know you could even do that. And then you go, yeah, like, please do it. You recommend blog posts for them to read, books for them to read, Twitter threads for them to read.

**Harry Stebbings** [38:31]
What's the team using? Claude Code, Cursor, Codex?

**Cliff Weitzman** [38:34]
Cursor and Claude Code. Those are the two most popular. Yeah. It's a little bit of Codex usage. It's not that big. I would say Claude Code is number one, then Cursor, then Codex. We want you to use as many tokens as possible in whatever harness way is the best for you.

You mentioned Linear. Linear is amazing. Like, automatically cutting tickets from Linear is fantastic. And just, like, being able to go into your agents and be like, OK, I have these, like, six Linear tickets. Start on them. And then really, a good engineer today is just an exceptional QA,right?

The AI will make them a feature. You will test the feature, see if it's good. You'll figure out where the edge cases are. You'll prompt it to fix it. And then you try to make it as efficient as possible, which is hard to do.

And then you need to make essentially, like, you know, roughly 10 really good product and engineering architecture decisions a day.

**Harry Stebbings** [39:19]
How do you think about token allocation internally? You know, we've seen leaderboard speeds, which is, I think, the most fucked up form of incentive kind of playing. You don't want to, like, prevent people.

**Cliff Weitzman** [39:31]
There's a lot of people who are a lot of talk. And I'll ask for examples. And I'll read the examples. And, like, you know, I'm doing this. I'm doing this. I'm doing this. I'm doing this. And then you look.

And I'm like, eh. And so I think about it in terms of demos. Can we hop on a Zoom call? And you'll show me what you built. And then I use it myself. And I'm like, wow, that's amazing.

Or you send me a screen recording of a feature or technology that you built. And I'm like, wow, that's so good. And so we give credit when things get shipped to production to users. So even inside of the AI team, if you build a really amazing and this is part of why Speechify ended up winning.

You asked, how did you build BiggerLabs? The answer is we ship to production all the time. That's how we won. We are not in the theory space. We are an applied AI company. That's why we win. And so if you're an engineer at Speechify, the analogy I always give people is imagine that you are in the milk delivery business.

And you make me a beautiful bottle of milk. And you leave it down the road. The milk will spoil. You have to get it to my door. Knock. If you didn't do that, you get no credit. If you carry the football all the way to the line, but you don't cross over to the end zone, if you don't kick it into the goal, you get no credit.

If you bring the ball just to the rim, you don't put it in the rim, you get no credit. And in the rim means push to production with no bugs. And users are actually using it. And then we get feedback.

How many companies do that iteration cycle fast? Almost no one. Definitely not with the user base number that Speechify has. And so in the AI team at Speechify, you make some amazing discovery. We're like, great. Push it to production.

And then you go, oh, wait. There's this QA problem. And this QA problem. And if you have this many people use it on the AI serving layer, then you have this other issue. Cool. You get no credit from me.

It's not in production. I can't use it on my phone. When I can use it on my phone, I will give you credit. And so this morning actually, not yesterday. Yesterday, I had a call with our AI engineering team.

And I said, listen, the project that we have running for duplex models and for AI conversational harnesses is something I'm really excited about. And it's been moving fast. I want it to move faster. Here's, like, 14 notes that I want.

And then what I do always is I'm on a Zoom call. I flip my computer around to face my phone. And I use the product in front of them. And we record it. And so then they see all the bugs.

And then I send the recording in the chat. Someone on our team, he's 19 years old, sent me a demo this morning off of that conversation that solved all of my problems. And he was like, hey, I was waiting for, like, three training runs to finish.

So I had a little bit of time while I was waiting. So I implemented everything that you asked. And it blew my mind. It was so good. That's using AI correctly. So it's not a token leaderboard. It's what did you show in production that was good?

**Harry Stebbings** [41:55]
How many companies do you think are actually as token-pilled, AI-centric as we think in terms of devs?

**Cliff Weitzman** [42:05]
There's a guy, Jason Yeager, who used to work at Speechify. And now he has MyTech CEO on Instagram. He's super funny. And so he makes a lot of videos about, like, you know, crazy CEOs who all use tokens, use tokens.

I think all founders in some way have that animal inside of them, because you know that it's theright path. But there is a difference between reality and theory. And you need to make sure that you don't overdo it.

**Harry Stebbings** [42:30]
Do you have any price sensitivity on tokens?

**Cliff Weitzman** [42:32]
Yeah, of course. Absolutely. I mean, I'll lose my mind if to implement a tiny feature, you use 50,000 tokens. Like, why did you do that? And, like, we will let people go if they just go bananas with something for no reason.

**Harry Stebbings** [42:46]
Are you able to accurately budget tokens on a?

**Cliff Weitzman** [42:49]
Not accurately, but within bounds.

**Harry Stebbings** [42:51]
Yeah.

**Cliff Weitzman** [42:52]
The other thing is, like, a lot of engineers are look, you go into engineering because you like optimization. Most engineers are not blind. And it physically hurts them to overspend tokens. And, again, I always think that the best way to interact with AI is you are chatting in the chat or actually doing it verbally.

And you're essentially pseudo-coding with your words constantly. And you're explaining architecture. And a great example would be

I know someone who has no engineering background. And they wanted to build an app. And they built exactly what they wanted. It took them two hours. But they needed an API call. And they needed to scrape this website.

And they basically scraped every single page of the website, every single part of the website. And so the bill that they got for the scraping was gigantic. And then I was like, why are you doing like that? Why aren't you going into the database to this exact URL and then scraping that from the URL?

So the amount of nodes they needed to hit became, like, 20 instead of 25,000. And so an engineer will spend their time making sure that the thing is optimized like that. So that's how you build, like, a good database or a good architecture system, whatever.

You do the same thing when you're interfacing with the agent. You want the agent to take the path of least resistance, not the path of most resistance.

**Harry Stebbings** [44:06]
I think one of the biggest problems is that agents are goal-seeking. And so they are, like.

**Cliff Weitzman** [44:10]
It's all about the target. You need to be good at picking theright target. And I think Anthropic published this paper when Fable 1 came out about long-horizon tasks with Fable. So the first thing is it was much better at, like, running a two-week task.

And it could burn $12,500 worth of tokens in two weeks and basically make a better model with that. That's a perfect, amazing way of using tokens. That's exactly what you want. And what you don't want is burning 12,000 tokens in the span of five hours doing something that's, like, just totally unnecessary and doesn't make any sense.

You need the loops to happen. And then you need to check the result. So what you want to build and Boris, who's the inventor of Claude Code, talks about this all the time. It's all about the loop. You say, here is the target.

Here's how you measure the target. Now iterate against the target over and over and over again until you get it.

**Harry Stebbings** [45:01]
What did you not know about building an AI-centric dev team that you wish you had known?

**Cliff Weitzman** [45:07]
How useful is it to own your own GPUs?

**Harry Stebbings** [45:10]
Was that a realization moment? Just did you see a bill on there?

**Cliff Weitzman** [45:14]
The realization moment was when we realized that we had really talented engineers who were essentially moving at 1/7 of the speed they could have if they had the compute one-to-one with their creativity and ideas.

**Harry Stebbings** [45:30]
How if you're a founder listening to this, how should I change my hiring process in a new AI world?

### AI-era hiring

**Cliff Weitzman** [45:37]
Number one, functional interviews. Build this. And then you see if they can build the thing. And then you run it through unit tests. The second one is give them a large code base, even an open-source repository, and have them understand the code base, make changes, and then check what they broke.

And then, yeah, like, they have to be able to orchestrate agents well. And if they're not doing that, it's kind of not worth to have the person. And then the next thing I'll say is it is more fun to have a smaller team.

Like, having a big team is great as long as everyone's carrying their weight. But the way I kind of think about it is, yes, I can have multiple agents running on my computer. Or I can have several Slack chats with really smart people who are bigger domain experts than I am.

And basically, that human being is the outcome owner for that task. And they have the agents. And so I can run as a founder multiple projects at the same time to a really amazing level of granularity. And so I think about moments earlier in the year when my brother Tyler would literally have an alarm to wake up at 3:00 in the morning, because he needed to check what the agent was doing at 3:00 in the morning.

And then he'd wake up, make sure it's good, go back to sleep. Like, you want to babysit your agent basically every three hours. And the beautiful thing now is you can go work out. And the agent will tell you the answer.

And then you, like, voice note back with Speechify what you want to do next. And, like, that'll happen. And so you want people who are essentially that level of addicted. Obviously, that creates massive AI fatigue. So make sure your teams don't burn out.

But you want someone who is that level of excited. And so I think hiring for slope more than intercept is more important today than ever before. Said another way, I look for the potential the person has more than I look for where they are today.

### Commoditization

**Harry Stebbings** [47:03]
When I look at Whisperflow and Willow and I did this tweet. And I deleted it, because I don't ever want to be soggy and miserable. And it's an amazing thing to build a company. And you should be incredibly credited for doing so as an entrepreneur.

But I found Whisperflow's product was just getting worse. And I said it on Twitter just because I honestly just wanted alternatives. I really need this product. And I wanted alternatives. I got 500 different alternatives. And I was like, motherfucker.

We'll talk about the commoditization of a market. That is not one that I want to be in. Can you help me understand? Have we seen the complete commoditization of that Whisperflow, Willow, speech-to-text for productivity?

**Cliff Weitzman** [47:45]
What they came out to the market with first was not necessarily their own model. Part of the reason they got worse is they switched to their own model, because it's a lot more affordable. And so they had a harness that ties together a bunch of other things.

Probably it was DeepL under the or DeepGram under the hood with a bunch of optimizations.

**Harry Stebbings** [48:01]
When they're trying to do notes and they're trying to move more into productivity.

**Cliff Weitzman** [48:03]
And I think they're being successful with it. Yeah.

**Harry Stebbings** [48:05]
100%.

**Cliff Weitzman** [48:06]
Yeah. So that, like, that's to your point of the compound startup. One of my biggest mentors.

**Harry Stebbings** [48:12]
When you look at them, do you not reflect on your we said before, not announcing fundraisers, not announcing anything?

**Cliff Weitzman** [48:18]
They're the opposite.

**Harry Stebbings** [48:19]
They've announced everything. They announced going to the bathroom.

**Cliff Weitzman** [48:21]
Correct.

**Harry Stebbings** [48:22]
And hence they have, I would say, a bigger brand.

**Cliff Weitzman** [48:26]
Not in terms of users. Like, if you walk down the street in New York City, way more people will know Speechify than know Whisperflow just by virtue of the fact we have way more users. But in the tech world, way bigger brand,right?

Investors know who Whisperflow is because they announce. We intentionally don't announce. But we don't have any competitors. Who are you going to use instead of Speechify to do text-to-speech for your models? Like, the closest thing is ElevenLabs. And we're, like, so much bigger than ElevenLabs or B2C.

**Harry Stebbings** [48:49]
What about Gradium?

**Cliff Weitzman** [48:50]
What's Gradium?

**Harry Stebbings** [48:52]
Brazilian company that does text-to-speech, competitors to ElevenLabs.

**Cliff Weitzman** [48:55]
OK, but they do B2C?

**Harry Stebbings** [48:57]
Yeah, no.

**Cliff Weitzman** [48:57]
Yeah. So there's unlimited numbers of companies doing B2B text-to-speech. But, like, we are unique in our market. So because Whisperflow was so public about it, they now have a lot of competition. And so Peter Thiel,right, only losers compete.

Try to not compete. And so, yes, they.

**Harry Stebbings** [49:15]
What happens to that market? Whisperflow take majority and then there's thousands of angle biters?

**Cliff Weitzman** [49:20]
I don't know. I mean, I want to be in that market,right? I think that market becomes oligopoly as well. Obviously, and this, by the way, I think it's another mistake that I made. I built my own speech-to-text experience that I've been using on my computer for the last, like, seven years, side-loaded on my iPhone and on my computer.

But I figured it's a commoditized product,right? Apple's going to release it instead of the button. It'll be great. And, like, there you go. But Apple keeps not doing it. If you remember, two years ago, Apple announced a partnership with ChatGPT that will improve Siri.

Nothing happened. And so that's also the reason why I never went after Siri. And so now we've launched a product to compete with Siri. And we've launched a product to compete with Whisperflow. And we've launched a product to compete with ElevenLabs, because I learned a lesson that I should have learned before, which is the same lesson from ElevenLabs.

The way that you win is you offer an excellent product for free. And then you have a wedge. And then you add more and more and more things. And so I don't know what happens with the Whisperflow space.

I just know that if you're a founder, you should also always try.

### Customer support

**Harry Stebbings** [50:18]
Final one. Another one that I get in trouble for, but I stand by strongly, is I just think the customer support market's a challenging market to really get behind.

**Cliff Weitzman** [50:27]
Yeah.

**Harry Stebbings** [50:27]
You have Sierra and DAC on out in front with the majority of funding and attention. But to say that there are 18 companies that have now raised over $100 million in the last 18 months. There is kind of what I would call, like, the mid-tier, which is that your intercoms and your talk desks and your crescendos, all these ones, where it's like, you know, they're not old, but they're old enough.

**Cliff Weitzman** [50:48]
Yeah.

**Harry Stebbings** [50:49]
8 to 10 years old. And they're pretty good.

**Cliff Weitzman** [50:50]
Yeah.

**Harry Stebbings** [50:50]
And then you've got Salesforce, Atlassian, and the much older ones. And then the worst thing about this market is that for any sophisticated buyer, an Airwallex, a Klarna, a Navan, a technology-facing company, everyone has built their own.

**Cliff Weitzman** [51:07]
Yes.

**Harry Stebbings** [51:07]
Because they need a sophisticated.

**Cliff Weitzman** [51:08]
But why would you pay a tax for it?

**Harry Stebbings** [51:11]
So what am I missing?

**Cliff Weitzman** [51:13]
Yeah. So the first thing you're missing is the core product that we're offering B2B is the API, not the agents,right? So Sierra doesn't have their own model team. They use other people's models,right? Because the value of Sierra is the go-to-market.

It's Bret Taylor. And so that's why if you talk to Matti and Piotrek, they'll tell you we're not competitive with Sierra, because their main business historically has been the API. So that's the first thing. In the API business, you have Speechify, ElevenLabs, Gemini, Grok, so SpaceX is now in the race, and Kartheeja.

And that's kind of it. And so that's not that competitive a space compared to, yeah, the B2B customer support thing. Everybody's in that space. Finn, everybody. And so I'm not building that product. What can I offer you that's 10x better than the next person?

Not much. And so in the core API side, I can offer you better quality, faster speed, and 10x cheaper. Good offering. But then I have to also offer agents, because there are so many pockets of value that have not been unlocked.

And unless I am again, I have this model for leadership. You don't want to be a fat manager who's, like, a general sitting in the back saying, "Take that hill." You want to be the warrior who runs up with their sword and engages the enemy first.

You need to be the same thing with your product. You need to be the number one user of your B2C product. And you need to help your customers use your product better. And if you do that, you will learn their problems.

And then you will figure out what the next product is that you need to offer them. So unless I have front-of-flight engineers working with my B2B customers, building agents for them using our technology, I will not figure out what the really amazing next innovation across the hill is.

And so you mentioned theright thing, which is ElevenLabs now has all these partnerships with governments. Governments is not exactly customer support. They would have never gotten to governments had they not done a great job on the private sector first.

I agree. ElevenLabs is, in addition to OpenAI, the most integrated companyright now, AI company with governments. That means they figured something out. But you got to start in something like customer support. Now, we support the models. So if you're a person building a customer support product, and your CFO is frustrated with the size of your ElevenLabs bill, and you want to cut it by 10x, go to speechify.ai and or hit me up, cliff@speechify.com.

I'll give you some discounts. But again, if you're a founder, you need to try. You cannot not try. You cannot give up before you're even in the race.

**Harry Stebbings** [53:35]
What will be a bigger company in five years, Sierra or ElevenLabs?

**Cliff Weitzman** [53:38]
Bret Taylor has the best resume, I think, of anyone in the world,right? I think he started Google Maps. Then he was CTO of Meta. Then he was co-CEO of Salesforce. He's on the board of OpenAI. And now he founded Sierra.

I would never try to fight Bret Taylor. And I think the field is so large. Like, no, like, we don't understand how big the space for AI agents is. Like, not even like, AI voice agents, not even close.

In the same way that people didn't understand how big the field was for LLMs in 2019. And the same way people didn't understand how big the space was for AI coding agents in 2021. Like, this is the next huge space.

And so both those companies are going to be massive.

**Harry Stebbings** [54:16]
I think they're playing very different games.

**Cliff Weitzman** [54:17]
Allright.

**Harry Stebbings** [54:18]
I think Bret Taylor's actually trying to recreate a next generation of Salesforce. He is absolutely not playing the customer support game. He's moving into presales.

**Cliff Weitzman** [54:28]
Everything.

**Harry Stebbings** [54:28]
He's moving post-sales.

**Cliff Weitzman** [54:29]
But neither is ElevenLabs. ElevenLabs has a product that also does customer support. But they do everything else, too. That's why I call it AI agents, not customer support.

**Harry Stebbings** [54:36]
But I think.

**Cliff Weitzman** [54:36]
Like, ElevenLabs is still.

**Harry Stebbings** [54:37]
A very opinionated voice-centric company.

**Cliff Weitzman** [54:41]
Correct.

**Harry Stebbings** [54:41]
Voice and I think Bret Taylor is doing all of it.

**Cliff Weitzman** [54:45]
Put another way, if you use a tool like Sierra, the wedgeright now is voice. But the important part is tool calling. ElevenLabs lets you do some tool calling. But that's not the bread and butter. There was a really good presentation that Bret Taylor did, a screen share of him building a guitar store on Shopify and how he uses Sierra to do customer support and sales and everything else.

It was extremely impressive. If you haven't searched this, you should search this. Bret Taylor is a big guitar guy. That is a very different product than what ElevenLabs is doing. And so they're both going to crush. I agree with you in the Sierra conclusion.

**Harry Stebbings** [55:19]
What crazy thing today will be incredibly and this is a quick fire, my friend, because I could talk to you all day. What crazy thing today will be very common in five years' time? You know, before, it was like, find your partner online.

### Quick fire

**Harry Stebbings** [55:34]
Duh, no, weird. Put your credit card online. Fuck, no. What today is, no. And in five years' time, we'll be like, yeah, of course.

**Cliff Weitzman** [55:43]
Human-computer interface is going to become primarily voice as opposed to a screen. So part of the reason why Google succeeded, it is a very simple interface. There's a text box and a button. That's it. Anyone can learn how to use it.

The reason why ChatGPT worked as opposed to GPT-3 is because it was also a very simple interface, just chat. There's a text box and a button. You get a response. That's it. The simpler version of that is just having a conversation.

I say something. I hear something in response. If you use voice AI from ChatGPTright now, it sucks. It's too slow. The LLM is much dumber than the core LLM. The escalation to the higher quality LLM is pretty weak.

I think what will happen and Meta has theright idea, by the way. So go Chris Cox is people are going to be talking to their computer and phone and some wearable constantly throughout the day and using screens a lot less.

**Harry Stebbings** [56:33]
You can buy one. SpaceX or Matter? Which do you buy?

**Cliff Weitzman** [56:37]
Meta.

**Harry Stebbings** [56:38]
Why?

**Cliff Weitzman** [56:39]
Elon's distracted.

**Harry Stebbings** [56:40]
Is he distracted or is he building full stack? Because actually, I think he's never been more strategically positioned. And he has an outlet for each of the different products that he's built. And each one feeds the next. When you look at Zuck and Meta, bluntly, the compute spend that he's producing, the outlet is increased conversion on an ads business, which is the biggest ads business in the world.

So 7% on $240 billion is a lot of fucking money.

**Cliff Weitzman** [57:05]
Yeah.

**Harry Stebbings** [57:05]
But it's actually not in the same quantum league as doing space data centers.

**Cliff Weitzman** [57:11]
Yeah. So let's take the space data centers out for a second. I think the space data centers is a very interesting idea. And what it does really well is it lets me underwrite a gigantic TAM for my expectations.

**Harry Stebbings** [57:21]
Reverts all estimates.

**Cliff Weitzman** [57:22]
Right? And so that's, like, that was a great rabbit out of the hat by Elon in order to pitch investors really well. Let's take that out for a second. And I'm going to talk to you about SpaceX and Tesla like they're one company, because really, I'm assessing Elon.

I'm not assessing everything. Like, you know, SpaceX is an individual stock. You know, for data centers, the biggest constraintright now,right now, is memory cards. And then very soon, it's going to be energy. And it's energy a lot of the times.

So what do you need for energy? You need energy supply. And you need energy storage. And so the best energy storageright now actually comes from Tesla. Tesla also has a chip manufacturer that they're doing, basically competing with everyone else.

And, like, that's going to do really well. And if you saw that Joe Rogan interview with Elon maybe two years ago, he was explaining that the hard part is not building the product. The hard part is building the manufacturing for the physical product.

So Elon is number one in the world for manufacturing complex items like that. So, like, that's very exciting. And so the TAM for Elon's companies are bigger. However, I think that Meta trades what does Meta trade atright now?

Less than SpaceX.

**Harry Stebbings** [58:27]
It's less than SpaceX.

**Cliff Weitzman** [58:28]
Yeah.

**Harry Stebbings** [58:29]
It's fucking so.

**Cliff Weitzman** [58:29]
And so I think Meta has more data than anybody else in the world. I think Meta is actually super hampered by laws like GDPR. Like, if GDPR didn't exist and the other laws in the US didn't exist, Meta would be ripping.

They just can't train on their data properly. And so they'll figure that out at some point in some way. I don't know how. But I believe in Zuck. And at the end of the day, I'm a huge believer in founder-led companies.

And so both we're talking about two of the best founders in the world. And the last thing I'll say, look at Zuck's age and look at Elon's age. And Zuck's not going to stop. And Elon's not going to stop.

But at a certain point, one of them will expire. And so Zuck has, like, 20 extra years. And so depending on how long you're investing, I'm younger than Zuck. Let's see what happens.

**Harry Stebbings** [59:07]
And if Zuck expired, Meta's stock price would increase.

**Cliff Weitzman** [59:10]
What? I disagree completely.

**Harry Stebbings** [59:12]
No, because you'd have a CEO who comes in and understands that. And this may be a short-term.

**Cliff Weitzman** [59:17]
Yeah.

**Harry Stebbings** [59:17]
But they're saying, we're going to invest more and more and more and more and more and more and more in CapEx.

**Cliff Weitzman** [59:21]
Yeah.

**Harry Stebbings** [59:21]
When we don't have an outlet for it.

**Cliff Weitzman** [59:23]
Yeah.

**Harry Stebbings** [59:23]
You'd actually see a stock price appreciation in the short-term. Every time Zuck steps out on the podium and says, CapEx, CapEx, CapEx, he's like fucking hammered for it.

**Cliff Weitzman** [59:30]
And that's why Meta is a good investmentright now, because what Meta doesn't have is what Palantir has, which Palantir has the Alex Karp effect. Alex is really good at pumping up the P/E ratio of the stock. And Zuck, I agree, is the opposite.

**Harry Stebbings** [59:43]
The same as Elon is the Elon green room.

**Cliff Weitzman** [59:44]
Exactly.

**Harry Stebbings** [59:45]
If Elon were to be removed.

**Cliff Weitzman** [59:47]
The intrinsic value of that.

**Harry Stebbings** [59:48]
70% of that value.

**Cliff Weitzman** [59:49]
Exactly.

**Harry Stebbings** [59:50]
If Zuck is removed, you definitely don't lose 70%. You maybe lose I don't think you lose anything. I think you get an experience exec in who says we're an ads business.

**Cliff Weitzman** [59:59]
Charlie Munger and Warren Buffett actually, no, it's Benjamin Graham has this concept of the cigar butt,right? Like, what's the intrinsic value of a company? And they approach it from an accounting perspective. I think about it as from an underlying technology and business perspective.

So the underlying asset, the intrinsic value of Meta is so large in relation to how it's valued in the market today. And you're correct. What's the P/E ratio of Meta? 32, something like that? And SpaceX is insane,right? Tesla is also multiple hundreds.

I think that there has to be a correction that happens, unless Elon succeeds with a big, big, big, big, big vision, in which case then he wins.

**Harry Stebbings** [1:00:36]
Final one for you. What are you most excited by?

**Cliff Weitzman** [1:00:39]
We talked about Jeff Dean. I'm Jeff Dean?

**Harry Stebbings** [1:00:42]
Yeah.

**Cliff Weitzman** [1:00:44]
I'm most excited by applications of AI to pharmacology and biology. So I have a family member who has very severe autoimmune neuroinflammation. He's had it for six years. I took a blood sample from him every week for 15 weeks, sent it to a lab, sequenced his genome, did proteomics on it to figure out how the proteins are expressing in his body, and run an RNA analysis in each one of those weeks.

And then I compared that to self-reporting data on what the quality of life is and what his mood effect was. Every day, I have, like, six years' worth of data on him and ran it on a GPU cluster.

And, like, I found so many things that no doctor could ever tell me. And he has a very rare disease. It was called an orphan disease, because there's not that many people. There's a Facebook group for this disease.

I'm buying now basically, like, you know, it's like a $5,000 device you can fit in your pocket. But if you put a piece of hair or saliva or blood into it, it can sequence your entire genome. And so I'm organizing meetups with all the people who have this disease to sequence all of their genomes and then compare them all on a gigantic GPU cluster to figure out what epigenetic common thread there is between them.

And I am I know I'm going to solve this disease. I would have never had an edge to do that in the past. And, like, it gets even more beautiful, because I can then take all the conclusions that I have about it and put it into alpha folds from Isomorphic.

And I can design not just the protein that is creating these issues, but I can design the molecule that needs to bind to that protein to either turn it on or off. I can use CRISPR to do the same thing.

I can use a lab like Twist, where I can tell it, I want you to make me this RNA sequence or this DNA sequence. And it can make it for me and ship it to my lab or my house.

And I can create amazing outcomes with it. And I can simulate all of it on my computer that's SSHed into my GPU cluster in Scottsdale, Arizona. And I can cure my brother. And so my experience is I'm a kid who, when I was eight years old, I couldn't learn how to read.

And my dad had to open a book and read Harry Potter to me. And that's how I learned how to read. And then when I was 13, I moved to the United States of America. And I didn't speak English.

And I listened to Harry Potter audiobooks 22 times in a row. And I still have the first chapter memorized. And then I couldn't get into the private high school that my brother went to and that my sister went to.

And I was really bummed. I went to, like, you know, a lower quality high school, whatever. And I didn't get into AP US history because I made a bunch of spelling mistakes in my essay. And I couldn't read the passage in time.

And I needed to train myself to read the SAT English portion. I wouldn't read the passage. I would read the answers. And then I'd go and hunt for the answer. And then when I got to college, somehow, by the grace of God, I ended up going to Brown and, like, starting a major for renewable energy engineering, because I couldn't do literature.

And I built a text-to-speech tool that would read out all my books to me. And that's why I graduated. Technology solved my dyslexia. And it solved my ADHD. And it's going to solve my brother's disease. And it's already solved my dad's prostate cancer, because I figured out with a bunch of help from other people how to use GPUs to identify where in his body the lesion was.

That's what I'm excited for, is there's better quality of life for literally everybody, because you have this magical machine that can run a trillion operations per second on as many GPUs as you want. And it can solve problems that we can't.

**Harry Stebbings** [1:03:55]
I find it staggering still today. We have orphan diseases, which is like, oh, there's too few people to make it economically viable for us to try and solve. And there are hundreds, thousands, low thousands, but low thousands.

**Cliff Weitzman** [1:04:09]
And again, it's the same thing. You just need data. You need compute. And you need to ask good questions. Like I said, 10 good decisions per day, either hypotheses or actual product decisions. And you can solve these problems.

Freaking amazing.

**Harry Stebbings** [1:04:22]
Cliff, it's been so great to have you on the show. I much prefer it when it's a discussion. It's been an amazing discussion. So thank you so much for putting up with me.

**Cliff Weitzman** [1:04:29]
My pleasure.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
