
Why People Are Paying 10x More for AI | Sid Sheth, d-Matrix
August 6, 202650 min · 8,629 words
Show notes
The AI chip market looks monolithic from the outside - NVIDIA dominates, and everyone else is fighting for scraps. But d-Matrix's CEO Sid Sheth argues that the market is quietly splitting into two distinct tiers, and the one that's exploding right now is the one NVIDIA's architecture isn't built for.
Highlighted moments
you can't really build a single chip that serves all of inferencing's needs, right? And so you've got to really kind of break it down.
“you have an architecture that puts compute and memory together in a way where you have order of magnitude more memory bandwidth available to your compute than a GPU would have, right? Because GPUs are based on HBM, and they, you know, have a certain cap on the amount of memory bandwidth.”
“we don't use the bleeding-edge process technology at TSMC. We use kind of N-1, N-2 process nodes. We don't use any HBM technology. We don't use any COWAS technology.”
Transcript
0:00People talk about organizations being entirely run by agents, right? How far are we from that? Eventually, I could run large portions of most companies. Do you see emerging companies, startups that are agent first? The New York Times had a piece recently about a guy and his brother running a $1.6 billion sales company to basically an agent first company. Do you see that segment of the economy growing? It's really an opportunity for everyone to go into something new. I think at the Dmatrix, we are being very proactive about the use of AI
0:31throughout the whole organization for various different use cases. We have an AI team in the company that is entirely focused on building AI across the world. So it is not an option. It is mandated.
0:43Tell us what you're talking about at HumanX. I mean, I read that you've had some acquisitions since the last time we spoke. So we're moving into a larger system, not just providing chips.
0:58And I'm interested in that. You know, my interest is generally in algorithmic research, but I do try and follow hardware. And you guys are a player in that space. I think I told you last time I talked periodically to Andrew Feldman at Cerebrus, to Rodrigo Liang at Semba Nova. I haven't spoken to Brock. If I'm not mistaken, you and those three are the primary players in this new inference-ship space.
1:35Or maybe not. I mean, if you could give me an overview on what you guys are doing and how you're differentiating. The inferencing space, specifically, right, is kind of bifurcating, right, because I think everyone looked at inference as one, you know, kind of, you know, singular entity, right? But I think there is lots of nuance when it comes to inference. It's not a one-size-fits-all, something we've been saying for a long time.
2:05So, you know, depending on where you are doing the inferencing, what your target markets are, what applications you're going after, you can't really build a single chip that serves all of inferencing's needs, right?
2:24And so you've got to really kind of break it down. And not every company is playing in every pocket of the market. It can't, right? Now, GPUs, you know, just because they are so general, the ecosystem is so broad, clearly, GPUs can go into many different applications for inference, right? But they're not going to be very efficient because of that, right? I mean, so, yes, they're very broad in their applicability, their general purpose. But when it comes to certain breakaway applications where applications need a very specific metric fully optimized, then GPUs will not do well in that market.
3:05And I think, so I think the way to look at it is, okay, you have the inferencing market, it's going to be the largest part of AI compute, it's expected to be over a trillion dollar market, you know, call it in the next five years. But it's, you know, then GPUs will play in that market clearly, but there's going to be portions or segments of that trillion dollars that is, you will need, you know, lots of optimization. And a highly optimized silicon, and the one that is truly breaking out right now is the segment that is called low latency compute, and low latency compute, low latency inference is specifically around applications that need high levels of interactivity.
3:47So, you know, a great example would be cloud code, this happened, you know, a few months ago, and you have people who are not programmers, who are now using cloud code and are great programmers, right? I mean, you know, the joke is, right, the new programming language is English, right? And everyone can program, and you can be an ace programmer, and you know how to prompt the model and how to use it. So, but that has led to people wanting to interact.
4:19And the faster the model is, when it responds back to the user, the more likely the user is to stay with the, perhaps, to the application, right? So I think interactivity has become very important, and companies are beginning to charge more for those high levels of interactivity. So you can always say, look, I don't want that level of interactivity, in which case, you know, I'll pay, call it $2 for a million tokens. But if I want that high level of interactivity, I'll pay $20 for a million tokens, and people are willing to pay that $20. They want that interactivity.
4:49So I think, given that this new tier of inferencing is emerging, this new tier of tokens is emerging, right, where we call it a premium token economy. So we'll look at the entire token economy, which is inference, this is the premium token economy.
5:07Not many companies play in that segment, because you need a specific compute architecture that is well-suited for the fast token generation, right? So it's about generating tokens very quickly. And that typically means you have an architecture that puts compute and memory together in a way where you have order of magnitude more memory bandwidth available to your compute than a GPU would have, right? Because GPUs are based on HBM, and they, you know, have a certain cap on the amount of memory bandwidth. With solutions like D-Matrix, and you mentioned, you know, Grok and Cerebris are the other, have architectures that allow, you know, compute to get access to very fast memory.
5:50And it's an order of magnitude faster than HBM. So that is the breakaway category right now. And D-Matrix falls into that breakaway category. And because that category has exploded on us literally in the last few months, we're just trying to keep up. And I think we need to scale up into that opportunity rapidly. Scale up, meaning manufacturing? Manufacturing, just delivering on product. Deploying product faster, supporting product faster, you know, making more of our product.
6:25All of these need to happen. Bringing on more people to help, you know, support more customers. All of this needs to happen simultaneously. And that is why we did this acquisition of a company called Giga.io, which allows us to scale into, you know, faster deployment of servers, faster deployment of racks.
6:45We actually had started our rack scale journey last year, right? We announced the protocol squad rack in October of last year at OCP. You can explain what rack scale means. Yeah, sure. So rack scale is, you know, you know, you have, you have, if you go into a data center. You know, you know, you have these aisles or rows of hardware and the unit of compute deployment in that data center is typically a rack, right?
7:17And the racks are then put together into what they call a cluster. Then the clusters are kind of put together into, you know, so, you know, into, you know, eventually multiple clusters will form a data center. They have, you know, sub, you can actually break it down into pods. I mean, different people use different terms, right? So there's racks and pods and clusters and eventually you kind of scale up to the data center, right? So the unit, but the unit of deployment in a data center is a rack. Yeah. And these days, I think just given how quickly AI computing is, is, you know, the demand is a shooting, you know, through the roof.
7:55So I think you need lots of compute and the unit, everyone is, is looking of compute deployment in terms of racks, right? So I think you need to be in their own racks, in their own racks or even in data centers, right? I mean, like you go talk to a customer, any customer, hyperscale customer or a NeoCloud or, or, or, or a Frontier Lab. And there's like, how many racks of compute can you deploy? How quickly? That will be the first question they ask you. They're like, okay, D matrix, how many racks of compute can you deploy for me and how quickly?
8:26Because then they can do the math in their head. You know, one rack is about a hundred kilowatts, you know, depending on, you know, everybody has different ratings. Sure. Somebody has 50 kilowatt racks, somebody has a hundred kilowatt racks, somebody has 150 kilowatt racks. So depending on the customer, they can make the math. They can run the math very quickly because they're like, yeah, you know what? Okay. So D matrix, you're telling me you can do like a hundred racks in three months or five months or whatever the number is. A hundred racks times, you know, a hundred kilowatts. This is the amount of power. And this is the amount of compute capacity I need, which is measured in, typically measured in terms of power.
8:57You know, so many megawatts of compute or so many kilowatts of compute or whatever. Typically it's always megawatts. Now we're talking gigawatts, right? So they're kind of working backwards from there. They're saying, okay, I need X gigawatts or X megawatts of compute that translates into X number of racks. And so the unit of deployment has become a rack. And so I think rack scale is just, you know, is kind of very prevalent unit in terms of, you know, how quickly you can deploy compute. Yeah. And so that's, that's really why we, yeah, let me ask because a hyperscaler, they have their own racks.
9:32And you were selling previously packages or cards, cards, cards packaged into, as I recall, two cards in a, you know, what do you call the unit? Yeah. Yeah. We had trays and cards, right? Trays and cards. Yeah. Okay. And then the hyperscaler, if they were going to have a D matrix cluster, they'd buy a bunch of those cards or, and they would slot them into their racks.
10:06What's the advantage of you having your own racks? Because the hyperscaler, I mean, how does that work? Yeah. So I think two things, right? One is that model has not changed from our perspective, from the point of view of the hyperscaler. Keep in mind, even when we work with the hyperscalers or the big neoclouds who also tend to buy trays and cards and will deploy their own racks, you still need to understand rack scale. As a company, we need to understand, because you're talking to people at the other end who understand rack scale work.
10:40So you can't be like, look, we don't understand any rack scale, but, you know, we leave it to you. No, they expect you to understand. We, at the end of the day, have to help them deploy our solutions into their racks. If anything goes wrong, we need to have rack scale people at our end who can help them. So I think there's an expertise and a skill set gap that we are addressing with this in our company. So that's one piece. But then there is another type of customer that we're going after, which is more the sovereign customer, longer term will be the enterprise customers who really want racks. They want whole racks to be deployed.
11:11Even there, we are partnering with Supermicro today to deploy those racks, right? So it's not like we are going to build our own rack. We're going to have our partners build the racks. But somebody needs to spec out the rack. Somebody needs to design the rack. And we would have to do that for our hardware, right? And then Supermicro can go build it to spec. But you still need to have the rack scale engineers. Now, that whole process has been accelerated is all I'm saying. It's like the deployment process has just got accelerated because of what we talked about earlier in the conversation, which is the low latency computing opportunity has exploded on us, and we have a solution that is well suited.
11:44So that whole process of taking our solution from where it is today to deploying it in a data center, whether it's a sovereign data center or a hyperscale data center, it doesn't matter. We need to have the rack scale people at our end, the system scale people at our end who can accelerate the process of deploying the product. And that is what the giveaway is. Because when I read about it, it sounded to me, or I misunderstood, that it's a hardware that you're moving from simply producing cards to producing racks. And then, I don't know, having your own clouds with your own racks or providing the racks directly.
12:23But you're not building racks. It's the expertise that you require. Is everybody doing this? Not everyone. Not everyone because I would say, depending on your business model, right? I think in our case, our business model was, and I'm assuming your question is, is every other chip company looking to acquire rack scale expertise? The answer to that is yes. I would say pretty much every company is looking to acquire, but not every company is looking to build a partner the way we build a partner, right?
12:57I think a lot of the other companies that you mentioned earlier, they're building their own racks. I mean, in many cases, they sell a rack. They sell, their unit of sale is a rack. In our case, the unit of sale is still a card or a tray. We partner with, you know, folks like Supermicro who would then sell the rack, right? And we would revenue share with them, or there would be another business model. Or in the case of a hyperscaler, we would not even worry about selling a rack because they build the rack, right? Yeah. It's fascinating, this breakout premium token market.
13:32I haven't heard it described that way. Are there other examples beyond coding that you can give that's driving that market? I mean, OpenClaw? I mean, what happened with OpenClaw at the end of last year? The agentic, we haven't even seen the true fallout of that. But, I mean, you've seen a lot of agentic tools now being unleashed, right?
14:00Claude is, you know, kind of incorporating a lot of agentic capabilities, right? Pretty much a lot of, pretty much every enterprise class software tool is trying to incorporate agentic capabilities. But I think to me, two things, right? One is Claude code and the significant improvement in capabilities there. And then the OpenClaw. So this is agentic communication, machine-to-machine communication, where machines are kind of, you know, you can spawn multiple agents across multiple virtual machines. And a task is, you know, broken down into multiple agents.
14:32And these agents are talking to each other, you know, essentially orchestrating all the different work or the workflows that you have running on your computer. And now that is very latency-sensitive communication because, you know, machines are just waiting for other machines to finish tasks, right? And you don't want that to happen for too long. So latency begins to matter. Yeah. Right? And in something like OpenClaw, it's also writing code on the fly. Right, yeah, yeah. It could spawn an agent to write code. So it can spawn cloud code to go write some code for you on the side, right?
15:04Wow, that's amazing. Yeah. What about edge compute? It seems to me that this would apply the latency question. Are you moving at all into the edge where latency is so important? We could. I mean, you know, the big, absolutely, we could. We are not currently. The platform, the way we have built it is a platform that is built with, you know, like a Lego block approach, right?
15:35We have these chiplets that are the Lego block and then we can scale the solution up. We can, you know, put many chiplets together into a very large solution. So for the cloud market and the data center market, we, you know, essentially take our chiplet approach and we package, you know, close to, you know, eight of them on a single card. And we have 64 cards in a rack. So, you know, we call it rack scale. We have 64 cards, each with eight chiplets, right?
16:08So that gives you an idea of how many chiplets we sprinkle through the rack, right? But you don't have to, I mean, if you went into an edge market application where you're just doing, say, for example, if it's a robot or it's a, you know, physical AI application or autonomous, you know, right, driving, right, where you need AI capabilities. You don't need that many chips. You might need, you know, like a few chips and we can actually do that because of the Lego block approach. We just scale it down, sprinkle fewer Lego blocks into a physical AI application or an autonomous application and the software would not change.
16:43The software was essentially built to scale with more chiplets or less chiplets, right? So the platform has been built today to go into the, you know, edge market.
16:53However, the edge market is very fragmented. So everyone has a slightly different software stack and, you know, each vertical, I mean, like, so if you go to the physical AI vertical, the software stack looks different. Autonomous vertical, the software stack looks different. Or industrial, the software stack looks different. So we just don't have the scale today as a company to go address all these markets while we are trying to address the data center market. So our goal is to, you know, stay focused on the data center market. Then at the right time when the company has scale, we will branch out into other markets. Did you foresee this emerging market with coding and agent-to-agent communication?
17:31And if so, what do you foresee coming? I wish I could sit here and say, you know, I saw it all. I envisioned it all. No, I think some of it we did. Some of it was luck, obviously. And the luck is really about, you know, how quickly a lot of this stuff happened and how useful a lot of this stuff became so quickly, right? So I think that is certainly caught even us by surprise, right? I mean, how quickly this is maturing. And this is because AI is learning fast, right?
18:01So AI is learning fast. And it is, AI is actually learning to help itself, right? I mean, so there's this iterative recursive loop where as AI gets smart, it learns to build tools that it can use in more efficient ways and more capable ways, right? So I think you have this kind of very quick, you know, positive feedback loop that is, you know, a virtuous cycle that is underway. And that's why you're seeing these impact. You know, the impact is just so rapid, right? But, you know, we had a broad idea, right?
18:34We said, okay, the bet was on inference first, you know, when we started the company. And then we said, you know, inference is going to be the big application. But then, you know, what, you know, it's not just going to be about inference. It's going to be about, you know, interactive inference. And then it's not just about interactive inference, but it's about spawning machines that can take all of that interaction and do something with it, right? So I think we kind of had a broader vision of how this could evolve. And then, but, you know, it's even surprised us and pleasantly surprised us at how it's all come together so quickly, right?
19:08And I think what we would be most excited by next is I think we have seen agentic capabilities unleashed. I think we haven't yet truly seen organizational capabilities unleashed, right? So meaning you have human spawning agents and augmenting humans, but can you create teams of agents? And so the abstraction of the task will keep going higher and higher. Right now, the tasks that are getting abstracted are like, okay, I have, you know, I want to make, I want to extract this data from this database,
19:43prepare a PowerPoint presentation based on the data for a board meeting that I have in a few hours, call it. And that can, you know, the agents can go spawn up and do a task like that. Now, imagine you start abstracting to a higher and higher level, right? Where you say, wait, it's not just about, you know, mundane tasks that I want to get done on a day-to-day basis. Imagine if we could describe a task at an extremely high level, like, look, I have a customer that I want to win, right? And you go figure out who the decision makers are, what product, what my product needs to be changed to.
20:15And so this thing can cover the whole strategy, like what an entire sales team and a marketing team would essentially do collectively. These agents could go spawn off and create a whole strategy plan. So I think the tasks will get more and more abstracted as you go higher. And very soon, I think, you know, people talk about organizations being entirely run by agents, right? And so I think where are we and how far are we from that, we don't know yet. But I think we are going to levels of abstraction here that, you know, eventually AI could, you know, run large portions of most companies, right?
20:49Quite reliably. Let me ask first how D-Matrix is employing agents. So we are, we just rolled up, we have an AI team in the company that is entirely focused on rolling AI across the whole company, right? We just launched Cloud Code to the whole organization. We have everyone who has access, we are mandating that everyone uses AI, finds ways of using AI and shares any new tricks or tips that they find with the rest of the organization, right?
21:22So it is not an option. It is mandated. This is across the hardware team, the software team, the operations team, the CTO teams, the marketing teams, the Corp.com teams, everyone, everyone is going to be using AI, right? And not everybody really knows what they will do with it, but they have to start using it. And the moment you start using it, you know, the use cases will emerge, right? Right. So, yeah, no, I think at D-Matrix, we are being very proactive about the use of AI across the whole organization for various, various different use cases.
21:53And Cloud Code is one thing. Open Claw is another. I'm, you know, not a coder, not a technologist. I'm a journalist, but I started using Open Claw through, you know, my Claw is the interface I use. And it's incredible. And I sort of imagine if I'm doing it, you know, I'm 70 years old, I'm a retired journalist.
22:24It must be happening all over the place. Do you have an agentic framework that, I mean, beyond Cloud Code that you have the organization using to spawn agents? Yeah. So Open Claw, I mean, we don't use Open Claw because it's an open source frame. So there's obviously security concerns around the, there could be. I mean, we don't know what the agents are doing behind the scenes yet, right? There is a certain level of transparency that's missing, but we use enough Cloud Cowork.
23:01So, you know, now Anthropic is incorporating all the same agentic capabilities that Open Claw has into, into Cloud Cowork. So we have Cloud Cowork Enterprise, Cloud Code Enterprise, right? So, yeah, absolutely. We have this agentic layer that comes along with the Cloud tools that are available to the whole organization. So, yes, they can, they can unleash agents across multiple different applications.
23:28And the good thing is there is a zero data retention policy with all the enterprise tools. So anything that the employees enter into, into Cloud is not retained by Cloud, right? It's only used for the accession. And, yeah, so I think, I think, yes, we are, we are, we are already got an agentic overlay framework on top of all the Cloud tools that we're using. I've interviewed companies. There's one, Aerotechnologies, which builds essentially a C-suite co-pilot.
24:02And the idea is that a CEO will have this co-pilot in his office and we'll be talking to it throughout the day or using it in sessions to brainstorm strategy, business strategy, or to change business processes. Do you use anything like that when you're looking at the market, looking at your organization, trying to decide, you know, should I buy a rack?
24:44Rack scale company. A scale company. Oh, yeah, absolutely. I mean, again, a lot of this is, you know, happened in the last three to four months, keep in mind. I mean, but even before that, I always used ChatGPT at the time. Really? For that kind of thing? Absolutely. And it's surprisingly, now, of course, we use Cloud. We use Cloud Code. So Cloud Code is, you know, phenomenal. So for both, you know, I'm doing, by the way, we did one acquisition. We are looking at more, right? And so now our M&A strategy, we actually use Cloud as a sounding board for our M&A strategy, right?
25:18And it's amazing what it can do. It comes up with, you know, beautiful reports on what, how the companies could be integrated. So this used to be entire teams, remember? You know, banking teams and advisory teams that used to come in and help companies integrate and come up with a whole strategy on how to put teams together. Guess what? Cloud does it in 15 minutes, right? So, and it presents a beautiful report on how the companies can be integrated and put together. Yeah. Right. Now, you know, it's not perfect because sometimes the data is, you know, not stale. But I mean, once you have a data room from the other company, I mean, you can throw Cloud at it
25:51and it will help you create a beautiful plan on how to put, what an integration plan looks like, right? Now, till this was not available, I mean, till about six months ago, I didn't, we didn't have access to this. So, you know, I used to use ChatGPT just to run strategy questions. Yeah. You know, we're thinking about a product like this. I mean, what do you think, right?
26:09I would say sometimes, you know, ChatGPT tended to be very polite, right? I mean, it never disagreed with anything that I asked. I mean, so I think what you're seeing more of now is the models are just more capable and more well-reasoned and well-thought-out. And they do tend to disagree. I mean, they will tell you stuff that you're missing, right? Of course, you know, the personality that each of these chatbots has is, you know, I think they've been curated to be very likable and amenable.
26:41But I think I've seen more debate in the conversation with these chatbots than in the past, right? And that's probably because I think as they acquired the step function improvement in reasoning capability, I think they were able to kind of reason through situations. You could say, you could argue that they became more confident, right? The models have now understood that they have better capability. It's almost like a human being, right? Once you realize that you have acquired a new capability, you become more confident.
27:14The same thing with these models. They tend to become more confident with disagreeing in a polite way, but highlighting, you know, things that you might have not seen, right? So I see more of that in the conversation. There's a level of trust. I mean, with the sycophancy that that's addressing, were you ever concerned that the feedback or the ideation that it's helping you with is just reinforcing your own biases?
27:48A little bit of that. Initially, I did. So I wasn't sure how much to use it. I mean, now I'm feeling a lot more comfortable because it feels more like a human conversation. Yeah. I think six to eight months ago, it felt more like a echo chamber conversation, right? There's somebody out there just echoing my sentiments and maybe adding a little more color to, you know, what, you know, I was already saying and presenting it in a slightly different way. And so if you are looking where those tools were more useful for me personally was if I wanted to make a point or if I had an idea that I already wanted to convey, a presentation I wanted to do, like, it kind of augmented the way I did it.
28:26Yeah. It didn't change anything. It was augmented. Like, look, you know, you can present this, you know, maybe talk about this this way. This is another way you can say the same thing. So it really was an augmentation tool. Now it has become more of a, it's a copilot. I mean, I don't know if copilot is the right word, but it's almost like, it's like, you know, another version of me that is willing to debate with me, right? It's not just an echo chamber, right? Yeah. I mean, then on the, on the sycophant, see, that's also a problem with employee sympathy. Right. CEOs tend to surround themselves with people who are eager to agree with the CEO.
29:03Correct. Exactly. Yeah. That's fascinating. Is Anthropic a customer? Could be someday. Yeah. Yeah, yeah, absolutely. But this kind of inference that you're talking about is, who are the customers that are buying the chips for that low latency? Oh, it's all the Frontier Labs. So Anthropic would be one of them. OpenAI would be one of them. XAI would be one of them. Meta Super Intelligence would be one of them. Google DeepMind would.
29:33I mean, all of them need this. Everybody needs it. Yeah. There's no other way, right? I mean, they're going to have to find a way to use a different type of low latency computing to augment the throughput-based computing that they've been, you know, kind of doing with HBM-based solutions. So the GPUs or other accelerators that all use HBM just don't have the same levels of interactivity that, you know, the solutions we build are having, right? I was asking about, you mentioned, you know, sovereign AI as opposed to models like from the foundation labs that you're hitting with an API.
30:17Can you talk about that market? Is that because there was this huge transition? You got to feel for the IT guys. There was this huge transition, get out of your data center into the cloud. Now, a lot of people have gone through that. And now it seems like things are flowing back into private data centers. Yeah. Can you talk about that, how you see that evolving?
30:47You know, I know there was a big outflow we saw over the last, call it 15 years, where, you know, cloud became more and more popular. And, but I think it was, it was, you know, a lot of the outflow that happened to the cloud was smaller companies, you know, companies looking to get started. I mean, I'll talk about Dmatrix. When we got started, we didn't want to deal with an on-prem, on-prem cloud, right? I mean, we wanted to deal with, you know, we could just offload our entire IT infrastructure to somebody like an AWS or a Microsoft Azure.
31:20And we were often running in like days, you know, right? As opposed to setting up our own data center, setting up our own hardware, even if it's in a colo, right? I mean, doing all, I mean, that was a traditional model, even for small companies, once upon a time, where you set up your own, you know, small data center or, right? And you, you, you, you kind of ran your own hardware, right? And we didn't have to do any of that. We didn't have to worry about security. We didn't have to worry about scaling that hardware, deploying the tools, none of that, right? I mean, so you see, I mean, that, that trend has been very strong and will stay strong, even to this day, all the AI native startups, right?
31:57Companies that are building these new age applications for AI, they're all running in the cloud. They're all still running in the cloud, right? Now, what you're talking about is maybe companies that were, you know, kind of further along in their journey, much larger, mid-size to large-size software companies, Fortune 500 companies that carry a lot of data. See, they never fully migrated to the cloud, right? There were, there were applications that they migrated to the cloud, but there were certain applications that they didn't migrate into the cloud. There were certain things that were data sensitive. Their data never migrated into the cloud. Now, the only thing that is different now is you still have the same structure.
32:31It's a hybrid approach for some of the larger organizations where some of their applications run in the cloud, but some of their applications that tap into their native or, you know, domain-specific data, they don't want to, they don't want to have those applications running in the cloud. So they're still running on, so they're maintaining both, right? They're maintaining an on-prem data center. They're maintaining, you know, they have a cloud data center or partner also, right? Now, AI doesn't change any of that, right? AI just is an overlay on top of that. So now you use AI to access the applications running in the cloud, depending on the type of applications you're running.
33:03Those applications will get infused with AI. And you have an AI overlay on top of those applications. The same thing goes for what they're doing inside their enterprises, right? The only thing is that the cloud, AI overlay that is happening in the cloud is happening much faster, right? Because the hyperscalers, Google Cloud, Microsoft Azure, Amazon, they are very AI savvy. So they are able to bring AI capabilities into all their cloud offerings a lot faster than an enterprise that has its own data and its own tools internally can do.
33:35So bringing AI capabilities natively to a Fortune 500 company, for instance, right, will take much longer than their applications running in the cloud at, say, Microsoft Azure will become AI-friendly a lot faster, right? So I think that's the only thing that you're seeing. It's not like they're looking to pull back stuff that was running in the cloud back into their organizations. Because they left the organization in the first place because they were not concerned about the data that they were accessing.
34:08Yeah. Although in regulated industries, there is concern about sending data to the cloud. Do you deal with a lot of... That concern has always been there. It's not like a new concern with AI. It's concern has been there like 10 years ago. So I don't think anything changed. I think it's all about where the data is resident, right? And that concern never changed, right? Because this data has been around for decades in many companies. And they were never quite comfortable sending that data into the cloud, right?
34:39I think the one thing that you might be referring to is most of these companies have access to an open AI API or an anthropic API, right? And now there is new data. I mean, and what people are doing is they are kind of conducting searches. And that tool has to get in and access databases and has access to that data, right? And, you know, is that something you want to do, right? I mean, I think so there is a... We have the same issue. We have all our proprietary data sitting in a data silo somewhere.
35:11And now Anthropic can access all of that data, right, to do things for us, right? And that's where they have the zero data retention, right? None of that data ever goes to Anthropic, right? And at least the enterprise class too. So that's a guarantee that they're giving you that all the data that they touch within your organization is not going to them, right? So it's only... And I think maybe that was your question originally. It's like, hey, look, there is an API that you guys are accessing through Anthropic or OpenAI. And that API gets access to some of your internal data because you are taking some of your internal databases and asking Anthropic or OpenAI to do things for you.
35:52Yes. So that, you know, is a concern, I think, has been addressed by these companies, right? And in some ways, it's probably not very different from you running, you know, for example, you know, Microsoft tools. Microsoft Open, you know, Office 365 or, you know, M365 has a lot of, you know, your information, a lot of emails, you know, but it's sitting, it is, you know, Microsoft tools have access to it, right? So in many ways, it's not very different from that.
36:23This market, as you said, it's expanding maybe faster than even you anticipated. Are you primarily focused on the U.S. market? Pretty, yes, yeah. Yeah, I travel a lot and, you know, every country. I ran into Rodrigo Liang in Saudi Arabia a year or so ago. Are you looking at these other centers that are trying to build out compute and AI infrastructure for either for their own country or for, you know, there's a lot of talk about who's going to be the leader of the global south of the AI movement?
37:08Right, right, right. We are spending just the right amount of time, not too much. I think we are focused a lot on the U.S. market because that is where the opportunity will emerge and then eventually it will diffuse into other countries in the global south, right? So I think we want to start at the source of the problem, not go to the destination directly. So I think we're starting at the source where a lot of the innovation is happening. People are trying to figure out how to deploy these solutions, what makes the most sense, which applications really need it, which applications don't, where does it work well, where does it not work well?
37:43Once those questions have been answered, then it will diffuse into a lot of the sovereign applications, into a lot of the enterprise applications. We'll ride that diffusion process, but it has to start from the source, you know. Yeah. Do you pay attention to what's happening in China? Not so much. We, of course, pay attention to all the work that's happening there and, you know. Yeah, I just wonder, are they moving in the same direction on, you know, this premium token market and all of that?
38:15They have to be. They have to be. I mean, I just don't see why. I mean, that's a very intuitive thing, right? Like, I want to create, I want people to access applications. I mean, everybody wants to access an application in a highly interactive and a real-time way. Why would that be any different for a segment of the population? Yeah. Right, I mean, if you need it in the U.S., you need it in China, right? Yeah. Are you allowed to sell to China under the... Under export controls, we would be allowed to sell to China, yeah. I recall most of your chips, if not all, are manufactured in Taiwan at TSMC.
38:49Or... Yeah. Silicone is etched in Taiwan. That's right. That's right. You know, TSMC is building a fab here. I don't know if it's online yet, but... Yeah, it is not online. It is online. They're building multiple fabs in Arizona. Yeah. And at least one of them is online since 2024. Have you seen any of your manufacturing or fabrication moving back from Asia?
39:22Not yet. Not yet. Still in Taiwan, yeah. Yeah. I would imagine, and I asked you this last time, and I get the same answer from everybody. Oh, we have relationships. But is there any sort of a bottleneck? I mean, can TSMC serve the world as the demand for inference explodes? Well, you know, the chips are going to be in short supply for the next, you know, call it three to four years, at least, if not more.
39:58So we have memory, you know, chip shortage. We have compute shortage, right? There's a reason why Elon wants to build his own terafab, right? Because he feels TSMC cannot meet all the demand for compute and memory. So we'll see how that goes, and Intel is helping him out with that, so we'll have to see how it all works out. But I think for the demand that we have, D-Matrix, if I would look at the D-Matrix demand picture for the next, call it five years, I think TSMC has got plenty of supply to build our demand.
40:36I'm, you know, we are a small company. We're looking to grow into this opportunity. This opportunity is also very nascent. It's just exploded on the scene literally in the last few months. It's got a long, long way to go. So I think we are growing into an opportunity that's pretty new, and the choices we have made on how we build our product, like we don't use the bleeding-edge process technology at TSMC. We use kind of N-1, N-2 process nodes. We don't use any HBM technology. We don't use any COWAS technology.
41:06So all these decisions that we made early on help us. So, yeah, if you look at a supply chain that uses those technologies, like the ones I just outlined, I think TSMC should be able to make, you know, will it be tough? Yes, for everyone. Tough for us, tough for everyone else. But I think we have a better shot at making supply because of the choices we made. Is that a constraint on your growth at all? Not yet. Not yet. But it would be a good problem to have, right? I mean, you know, if I get to the point where my growth is constrained by how much TSMC can supply to me, I think I'm in a good place, right?
41:41I think then I have to just kind of work through the problem, which we will work through the problem with them, right? And to me, I think it comes down to the relevance of the problem, right? If your problem that you're solving is very relevant and there's lots of big customers who care about solving their problem, yes, you have relationships and that's all great. But at the end of the day, you know, follow the money, right? I think if there's a very relevant large problem that a very relevant large customer is trying to solve and we think we are helping, you know, solve that type of problem, then the supply chain will make, you know, a place for you, right?
42:16Because the world needs this, right? So I think they'll find a way to make place for you. Yeah, you talked about the agentic explosion and about the auto, whatever you call it, AI coding explosion. Which of those is the larger market for you? And what kinds of companies are you selling to?
42:44I mean, you were saying you're selling to the foundation labs, not necessarily anthropic. But whose coding co-pilots are you providing chips for and whose agentic systems are you providing chips for? So agentic, I think, is not a separate agentic is a capability that I think pretty much every company that is building a coding tool is embracing
43:16or is introducing into their product offering, right? So I don't think you should look upon it as, okay, there is a segment of, you know, section of companies that's only working on agentic. It's agentic is like, you know, kind of a capability that is now being, OpenClaw demonstrated what's possible. But now you take that capability and introduce it into your products, right? So everyone is doing it. OpenAI is doing it. XAI is doing it. You know, obviously Anthropic, because they sell into the enterprise, they do it, they're doing it faster than anyone else. But it's a capability, right? And that capability adds more influencing needs on top of the coding generation tools that are already, already need, you know, need inference, right?
43:58So you have, you have, it's kind of putting it together. We, we said, we, we are working with everyone, right? I mean, we're working with all the Frontier Labs that build coding tools, agentic coding tools. We're working with Frontier Labs that are building video agents, voice agents. We are working with AI Native Startups building voice agents.
44:19We just have to be careful about not going too broad, because I think the company is still not at the stage where we can support all these customers simultaneously. So we are to be, you know, kind of gradually grow into this opportunity with, you know, at the right time with the right company. But we're doing, we're doing it with everyone, because it's relevant. Again, going back, you know, I mean, what we are building is relevant for everyone. Everybody wants it. It's not like there's somebody who says, like, you know, I don't care about, you know, high levels of interactivity. I don't care that my users are able to interact with my application in real time. Nobody's telling us that, right?
44:50Everyone wants fast tokens. Is there a segment of the market where you think, you talked about an agentic enterprise with where it's, you know, agents are doing much of the work? Do you see emerging companies, startups that are agent first? And do you think that segment of the economy is going to grow? I mean, everyone's sort of watching, you know, there was a New York Times had a piece recently about a guy and his brother running a $1.6 billion in sales company.
45:34I mean, that's basically an agentic first company. Do you see that segment of the economy growing? I mean, that's the future. That's the future. Is that going to challenge legacy companies in different? I think it depends. I think it depends on what kind of product offering or products or what kind of service you're building. I think there is a place. I mean, you have those companies today.
46:07I mean, even pre-AI, right? I mean, there was companies, you know, you had people running a consulting practice, right? I mean, you are a one man. Are you a one man? There you go. So you are right. You're already there. Now you're going to start using AI, right? I mean, so you could be one of those people like, hey, look, I've been far ahead of everyone. I've been a one-person company for a long time, right? And now I just use AI as an augmentation tool. So it really depends on the product offering and service, right? I mean, I think there is a lot of companies that had a product offering or a service that can be taken over by AI completely, right?
46:45And you don't really need humans for the kind of work that they were doing. So there will be kind of the low-hanging fruit, I would say, that will immediately get consumed. And either they embrace it or then you have what you just said is a new breed of companies just come and say, wait a minute. Why are those companies, you know, why do they have humans in those companies? You know, we can just do it with AI and then they will just become a lot more efficient. Or those companies embrace it themselves and, you know, re-skill the humans into something else, right?
47:17Or they find that as an opportunity for those companies to grow into something different. Because humans are good at something which AI is not good at. So why don't you re-skill the humans to do stuff and expand the opportunities you can go after as opposed to sticking with what you have. It's really an opportunity for everyone to grow into something new. But absolutely, there will be many more companies that will be tooled only with AI to address those products and those services, right? But then there are certain products and services you really just cannot, you know, build, you know, with AI only.
47:49I mean, you need manufacturing lines and you need, you know, good product design. And there's emotion that goes into how the product is sold and, you know, how the messaging is done and how you appeal to consumers. I mean, all this is something that you need humans. Because at the end of the day, you're making a sale to a human. At the other end is a human who's buying the product and you want to appeal to their sensibility and their emotions. So I think not every section of the economy is going to get consumed by AI only companies, but there will be a lot of them, right?
48:22A sizable portion. Yeah. Healthcare is an industry that's going under through tremendous change. Do you see much demand for agentic AI in the healthcare space? Are you serving that space or are you focused on a few verticals? Yeah, we are not very active in the healthcare space. So I can't meaningfully comment on what the latest and greatest is there.
48:56Because we are serving the tools. The tools get used by various verticals, right? Healthcare is certainly one of them.
49:06But I think to me, the most exciting thing in healthcare is really drug discovery, you know, disease, you know, curing, you know, really complex diseases, you know, producing new forms of treatments that could accelerate the development of drugs or find cures, right? I think, I think, to me, to me, that segment is a lot more exciting. And, you know, that is one of the goals of AI, should be one of the goals of AI is, yes, you know, more equitable wealth distribution, sure.
49:43But really, can we make, you know, benefits, you know, broadly available, can be easily accessible to parts of the population that don't even have access to it, right? And so, so can we get humanity to a point where, okay, everybody has enough, they can live long lives, long, healthy lives, right? And they can live peacefully. I mean, that's maybe the third quest, right? We can, now moving maybe into a bit of a utopian vision here, but, you know, that would be the vision, right?
50:17It's like, you know, there's enough wealth for everyone, enough, you know, health for everyone, enough, you know, peace for everyone. And then AI has really helped humanity, right? Yeah, that's right. Well, we're hoping. It also requires political leaders. That's right. That's right. We'll take it in that direction. Right. Okay.
More from Eye on AI

AI Agents Fixing Your IT Before You Even Know Something Broke | Erhan Giral & Ryan Manning, BMC Helix
Aug 3, 202659 min

Real AI Transformation Costs HALF of Everyone's Salary for 2 Years | Chris Blackburn, Liatrio
Jul 30, 20261h 5m

"According to NASA's Definition of Life, I'm Not Alive" - Why Nobody Can Define Life | Dr. Kate Adamala
Jul 29, 202646 min

Video Is About to Stop Being One-Way (and That Changes Everything) | Victor Riparbelli, Synthesia
Jul 28, 202639 min

Video Is About to Stop Being One-Way (and That Changes Everything) | Victor Riparbelli, Synthesia
Jul 28, 202639 min