Investor Event Transcript
Nvidia Corp (NVDA)
Conference Transcript - NVDA 2026-05-28
Shauna Laughlin, Analyst
Thank you, everybody. Welcome to day two of our 54th annual TMT conference. Really pleased to be joined on stage by Shauna Laughlin, who heads up our networking coverage, and Galad Shanir of NVIDIA. How'd I do?
Gilad Shanir, Analyst — Other
Almost, almost. Close enough.
Shauna Laughlin, Analyst
All right. I think my bosses are in the room, so I am obligated to ask you for an Extel vote if you think we've earned it this year, and if the Wi-Fi password wasn't subtle enough, we'd really appreciate it. Gilad, maybe just to start with that out of the way, you guys reported earnings last week. The networking numbers you gave, I think, were $14.9 billion, up 199% year over year. A lot of that is obviously captive in your NBL racks, but maybe you could walk through what are the key components that are driving all the momentum you're seeing on the networking side.
Gilad Shanir, Analyst — Other
yeah um so just just to tell a secret you know i got i got to pick on the questions beforehand um and the original question was uh 199 percent growth and nearly 15 billion of revenue and i couldn't sleep at night yesterday because i tried to figure out who wrote the question you know 199 percent and nearly 15 billion right it could have been engineer going engineer would say 199 and 14.8 and it could be marketing person because it would say nearly 215 billion right so i could try i couldn't slip it nicer for that i tried to figure out but Correct it now. So when you look on what we built, on what we design, we design a single unit of computing. We design an AI factory, which is a single unit of computing. And when you design a full data center, full AI factory that needs to behave like a single unit of computing, There is a lot of infrastructures, a lot of networking infrastructures that you need to bring into that AI factory to make it work like one. There is scale up with Envilink. There is scale out, and scale out we have InfiniBand as one option, and we have Spectrum XEterent as another option. We have scale across that we're using with Spectrum XGS. and then we have introduced a new storage infrastructure with Bluefield as a storage processor and we also have an access network that we're using Bluefield as a device to enable access into the AI factory and provide all the security capabilities and so forth. All of those networks, all of those areas, infrastructures are growing. So we see growth in Envilink as a scale-up domain. We see growth on InfiniBand and Spectrum X-Ethernet as a scale-up domain. And we see growth in Bluefield as a storage processor, as also a DPU to enable access. So there's growth on all those infrastructures, all those elements, and that contributes to the numbers that you mentioned.
Shauna Laughlin, Analyst
Okay, thank you. I'm going to go back in time all the way to 2020. NVIDIA made the acquisition of Mellanox that brought you and your team over. We've referred to this on our team as perhaps the most important and successful technology M&A has ever happened. Can you talk about how that deal came together? What did NVIDIA see and why they felt they needed to bolster that networking asset so early? And how is it paying dividends now? And what are your expectations going forward as well?
Gilad Shanir, Analyst — Other
There's another thing that I saw in the questions, by the way. Those are very long questions. very long questions yeah i'm an engineer you know so if you have more than four words and a question i completely lose you know so you know it's it's i need to recap what you ask um how the acquisition happened i think it was simple you know jensen came we talked and he put a deal and we signed and that's it, you know, simple as it is. I think that Jensen saw that the world needs computing data centers or accelerated data centers, AI factories. He saw that NVIDIA needs to become a computing company, not a device company, not an AC company, but a computing company. And the way that you connect computing ASICs will determine what those compute ASICs can do. If you connect it in one way, you just got a server farm. If you connect it in a different way, you actually can build a supercomputer. So in order to go to a direction to enable the company to become a computing company, you need to bring the right networking infrastructure that enables all of that magic and I think this is what he saw in Mellanox and that's the reason that he came we talked, you know, there was a live in first sight, put a bid and we agreed and we joined NVIDIA Joining NVIDIA, you know, Mellanox was kind of one team there was no there are no different business units in a sense. Melnox was one thing and we were focusing on building networking infrastructure for distributed computing workloads. We built a great technology that used in high performance computing and AI is another example of distributed computing workload and that's why Melnox was a great fit to NVIDIA. When we joined NVIDIA, when I joined NVIDIA. It was a great experience because NVIDIA actually behave and work the same as Melonox. It's one unit. It's actually one unit. There's a group discussions, groups meetings, networking, and compute, and infrastructure, all work as one team. It's the same as Melonox. So it was actually felt like home. A larger home. you know there's more people in that house more rooms in that space but it was felt like we didn't leave Mellanox it was a great experience and it still
Shauna Laughlin, Analyst
is okay per your direct feedback I'm going to ask two questions at once
Gilad Shanir, Analyst — Other
I'm not going to remember the second
Shauna Laughlin, Analyst
we'll get through it together and then I'll pass it to Sean to ask about scaling up out across and diagonally so you know I think there's You guys have shifted from selling GPUs to selling fully integrated racks, and I think there has been some pushback from ecosystem partners that don't like being captive into one, not having optionality of which components to pick and choose. Can you talk about the pros and cons of that go to market, and then also how NVLink Fusion came about? Was that a reaction to this trend, and what that offers your customers?
Gilad Shanir, Analyst — Other
Yeah. Well, you did combine the questions. um so so when when you when you build when you build a supercomputer when you build an air factory you need to build it as one unit because that's actually the compute unit and and when you build one compute unit that has a lot a lot of components inside you need to have a extreme co-design that combines the software and the hardware and the compute asic and the networking ASICs and storage element and so forth because you build one unit. So we design it vertically. Everything needs to work as a balanced system. If one element does not give what the rest of the elements are required, then that system will not work. When we deal with distributed computing workloads, I'll give you one example. When you deal with distributed computing workloads, you need all the compute ASICs to work like one. if one of those ASICs let's say I have hundreds of thousands of GPUs in my factory, in my data center if one of that GPU ASIC gets data a little bit late versus all others, all others will wait okay that's how serious it is and therefore you need to design it vertically but after we design it vertically and we bring all the co-design elements and making sure that everything works as a single unit We actually sell it horizontally. You can take pieces. You can take pieces of it. You can take the GPU. You can take the CPU. You can take the networking. You can take NVLink separately. And then you can mix and match with your own designs if you want to. So what we do, it's actually vertically, but everything can be used as a different separate unit. And nothing is kind of closed. Everything is very open. All the interfaces are given, are known. You can actually put your own software and all modifications and your own enhancement on top of what we do. And if we can choose what you want to take. Envilink Fusion, you mentioned Envilink Fusion, and that's actually an answer that it's not a black box. Because everything we design, we are so proud of them, then we're happy if you take any piece of it. So Envilink Fusion, because I think it's the only scale-up network that is proven from performance and from production perspective and if we build something that great why don't we want our customers and partner to enjoy that as well even if they have their own cpu or even if their own gpu that they have built and they want to use it and therefore nothing is a black box all the components are available you can choose you can mix and match and fusion actually enable our customers to also take Envilink as a separate element if they want to do that. And we're also working with an ecosystem. So we have already made announcements on our partners and customers that are part of Envilink Fusion ecosystem or using Envilink Fusion for their own AI factories.
Shauna Laughlin, Analyst
I wanted to pivot a little bit to some more geeky and more fun questions about tech rather than these lame business questions. Maybe if I could just ask an open-ended question about Spectrum and its approach to Ethernet in a system way where there's both intelligence on the NIC side and within the switch as opposed to a more purely switch-centric architecture. What are the benefits on the Spectrum side and how does that translate both in a training environment and in a more distributed inference environment?
Gilad Shanir, Analyst — Other
Yeah, well, we can take an hour to answer this question. So if you have time, when we start working on Spectrum XEternet, well, the reason that we start working on Spectrum XEternet, first, we had InfiniBand and we'll still have, and it's growing and it's one of the best technologies ever created for distributed computing workloads. That's why Melnox did so great in high performance computing and if you look on high performance computing supercomputers you're going to see a lot of infiniment there. It was built for low latency, it was built to eliminate jitter which is a key element and so forth. But as AI is growing and every data center become accelerated, every data center becomes an AI factory, we knew that we also need to bring an option for Ethernet because we have customers that invested in Ethernet they know how to run Ethernet they build their software management on top of Ethernet and it's gonna be hard for them to go and do something else so we have InfiniBand for people to use InfiniBand and we also wanted to design an Ethernet version that can also be used for scale-out that can also be used for AI workloads and distribute computing workloads now when when people refers to Ethernet it's important to know that there is no one Eternet out there. There are different kinds of Eternet. And different kinds of Eternet that were developed for different kind of workloads. There is Eternet kind that was developed for high virtualized, small reddix infrastructure. There is another kind of Eternet that was developed for single server workloads, large cloud infrastructure. There is another kind of Eternet that was developed for telco and DCI and kind of long distances and based on debuffers approach and so forth. The issue that we had is that none of those were built for distributed computing workloads. None of those were designed to eliminate GTR. GTR was fine. If I built Ethernet for single server workloads, I don't care if there is a skew in time between one server to another server. because there is no communications between them. If I'm building something for log distance or DCI and I base it on debuffers, I actually based it on creating GTAR. So none of them were dealing with GTAR, and GTAR is the biggest problem when you deal with distributed computing workloads or AI training and inferencing, which is our example for distributed computing workloads. And that's the reason that we actually created Spectrum X. And Spectrum X is the only Ethernet that is purposely built for AI. Now, something that we learned from InfiniBand is that there is no way to build a network that's going to eliminate Jeter and do that on a single device. No way. And it's simple to explain it, okay? Data that comes out from the GPU goes out in an order. Same as we speak, right? and there is an order of the words. Data that's gonna be written to a remote GPU needs to get to that remote GPU memory in order. And if that data is gonna go through a switch, and that switch needs to maintain that order, then that switch will introduce GTR. And the reason is that every switch has a lot of ports that you can use. There is a lot of paths in the network. If the switch will start doing a distribution of every packet can go to a different route, to a different road, because there is less busy roads that I want to use, then that will create, by definition, out of orderness in the delivery of data. And that means that I cannot use it on the other side. And if you look on all the designs of the off-the-shelf of switches that exist today, they're actually based on not creating out-of-orderness. So they're using approaches like flowlets, which means if there is a flow, I'm going to keep that flow even though there is an empty road that I can get it faster. No, I'm going to keep it the same path because the data must get in order to the other side. That's your enemy, okay? That's how you create GTR and we didn't want that to happen. We actually wanted to make sure that there is no GTR. And in SpectrumX Ethernet, the switch needs to unconditionally distribute traffic across the entire infrastructure that exists. The switch will choose for every packet a different port. What is the fastest path? What is the less busiest path I'm going to use? And by definition, I'm creating out of order of data delivery. And in order to put the data back in order, I need a super nick on the other side. So I'm using RDMA, because RDMA enables me to put the data directly in the GPU memory, no buffer copies, no delay on the other side, but I need a smart element that sits next to the GPU on the server that will take data that's going to come completely out of order, but place it in the right order in the GPU memory. And that's the purpose of the SuperNIC. And that's why when you build an infrastructure for distributed computing workloads, you need to have a switch element that does the distribution unconditionally, and then you need the SuperNIC that will put the data back in order. That's why it's an infrastructure, and it's not a single device.
Shauna Laughlin, Analyst
I think that's a perfect segue to kind of expand this conversation about out-of-order and packet spraying-type concepts and talk about, maybe if you could just talk about MRC and the recent announcement that you made with your consortium partners, as well as maybe contrast that with some of the goals that the Ultra Ethernet Consortium is going for. Because it sounds, to a layman like myself and I would assume most in the room, a lot of what Ultra Ethernet Consortium is attempting to do is solve for that problem.
Gilad Shanir, Analyst — Other
Yeah. You know, there is more and more focus on AI workloads. Every data center is going to be accelerated and AI is going everywhere. So obviously there is a good attention on it. what we did in Spectrum Max Ethernet is two things one of them is we brought a lot of learning from infinimental Ethernet lossless the reason that we prefer lossless is because we don't want to draw packets because of congestion because once you draw packets you need to retransmit it and that means g tier that means extra delay so we don't want to draw packets and we're focusing on lossless focusing on RDMA and by the way, the other protocols that you mentioned are also based on RDMA. ROGKEY is just RDMA over Ethernet. If you say RDMA and ROGKEY, actually you said the same thing, okay, twice.
Shauna Laughlin, Analyst
ATM machine.
Gilad Shanir, Analyst — Other
MRC, it's also RDMA or ROGKEY, for example, based and so forth. We also brought adaptive routing into the infrastructure that has been done in hardware because you actually want the decisions on the different paths to be done very, very quickly, immediately. So we brought all those things into SpectrumX. We also enable in SpectrumX a flexibility to support other routing protocols on top of that. MRC is an example for that. MRC is another way or another algorithm to how to distribute the traffic across the network. And SpectrumX does not support just one protocol. SpectrumX actually supports multiple protocols on that infrastructure. It supports the adaptive RDMA protocol. It supports MRC protocol on top of that. And I can tell you it supports other customized protocol that our other customers or large customers have developed and are using. So there is a variety of routing protocols that can run on top of SpectrumX. and they are optimized as an entire infrastructure on the end-to-end side because, again, for any protocol, you need two elements at least. There is an element on the supernick and there is the element on the switch. A lot of things, by the way, that we built into Spectrum Max, those are the things that were discussed later on in other consortium, like you mentioned, and there is also other groups of companies that work on more algorithms and so forth. But as we have customers that build very large infrastructures, very large AI factories, those are expensive AI factories, they would like to optimize their infrastructure to the way they run their own workloads. So that's why we brought the ability in Spectrum X to support different kind of routing protocols, to do it in a zero-GTA approach, and it could be the adaptive RDMA, MRC, and several others.
Shauna Laughlin, Analyst
Just to briefly clarify when we talk. So would it be fair to compare MRC to, for example, BGP? That is another routing protocol that could be built on top of Spectrum.
Gilad Shanir, Analyst — Other
Yeah, exactly. And in Spectrum Max, first in Spectrum Max, it was important for us to use all the standard protocols that exist in Ethernet. The way that we implement that, that was done differently in order to eliminate jitter and so forth. MRC is another way to route packets, for example. You mentioned other protocols. Yes, there is multiple protocols that you can use. There is ways to implement that in a way that you eliminate jitter, have zero jitter, which is the key element. And all of those options are supported on Spectrum Max Ethernet.
Shauna Laughlin, Analyst
I'll maybe steal one more one more geeky question and that is if you could just talk about how the networking problem changes moving from large scale pre-training to maybe a multi-tenant inference workload type and maybe how are your customers thinking about provisioning fungibility across those two deployments Is there maybe an over-provision of a back-end network in the eventual inference because it gives you flexibility to scale up and down? Not scale up in the networking sense, but scale up and down.
Gilad Shanir, Analyst — Other
So there is a lot of commonality between actually pre-training and inferencing. Both are distributed computing workloads. Both require a zero-G tier. Now, when we say zero-G tier, Zero jitter means that if you're running a single workload, that workload will not impose different delays on different communications to different GPUs, because that's going to be a nightmare, okay, from a performance perspective. But it's the same thing when you run multiple workloads on the same infrastructure, like a cloud, like AI cloud. One of the key problems in traditional clouds and off-the-shelf Ethernet that was used is as Jitter was not a thing that was a focus, what happened is that one workload could impose performance issues on another workload. One workload can create delays in the network that will impact another workload that share the same infrastructure. And one of the common best practices in traditional cloud was never have two different users run on the same switch because one will impact the other and then your SLA is gone out of the window. So there was a heavy focus on how do I schedule different jobs in a traditional cloud that one job will not be on the same switch as another because that will negatively impact the performance on another job. Once you deal with Jeter, once you eliminate GTR it means that there is no traffic that will create congestion in the infrastructure and if there is no traffic that will create congestion in the infrastructure there is no way from one workload impact another workload so it doesn't really matter if those are two training workloads running on the same infrastructure or it's a hundred inferencing workload that runs on the same infrastructure you need the same solution for both so what we brought in Spectrum X for training, that was the first workload that we're running, it's so amazing now when you have inferencing. You can actually see the difference in that. Now, inferencing does enable or burden it to create more infrastructures. And recently, we announced a new storage infrastructure for context, for memory context, for inferencing. Because now as we move to the world of agenting AI, there's AI talks with AI, There is much more data that you need to hold. There is larger sizes of KVCache. Not everything can be stored in a local server, in a GPU server. And you need to go to an outside storage. And the outside storage that exists is network storage. And network storage is great for a variety of workload, but it's not really optimized for inferencing. Because network storage was built to make sure that the data is not going to get lost. So I'm going to invest in replicas of the data and making sure that if an SSD went down, I still have still replicas in others. It's too expensive if you look on inferencing because in inferencing, for the rare cases that something's going to happen to an SSD, I can actually recalculate the data. So instead of investing in replicas and so forth, I can build something that is going to be much more effective and optimized for inferencing, And that's what we did with Bluefield and STX and CMX and creating a new storage infrastructure for inferencing for KVCache. So what we built for training works greatly for inferencing, actually. But inferencing created or drove the creations of more infrastructures as part of the big factor.
Shauna Laughlin, Analyst
All right. I'm going to ask one that Sean's going to have to deal with the answer to. so it seems like you know the debate on CPO has shifted from scale out to scale up more recently you know what's your view of you know what CPO can bring to both of these domains and what sort of a reasonable time frame at which you know we should expect CPO adoption more broadly in you know your computer
Gilad Shanir, Analyst — Other
ecosystem yeah and I'll combine two answers if it's okay you combine two questions, I'll come by two answers. There is also the, you know, I heard that there is a debate between, you know, copper versus CPO, copper versus optics. It's actually a funny debate. It's like you're going to ask, you know, how do I look on an airplane versus a car? If I need to drive to the next city, if I need to drive to New Jersey, I'm going to take a car. If I need to fly to Taiwan, which I have a flight tonight, I need to take an airplane, right? There is no way can I use a car? The same thing goes to copper versus optics. If I can use copper, which means is that the distance that I need to cover is applicable for copper, I'm going to use copper because optics will be too expensive for that. If I need to go to New Jersey, I'm not going to take a airplane. I can, you know, I can fly from New York to JFK, right? For example, I can do that with an an airplane, so copper consumes your power. It's very cost-effective, it's very reliable. The problem is short distance. But if that distance is okay for where I'm designing, I'm going to use copper. If the distance is not applicable and copper cannot cover the needed distance, I'm going to use optics. That's simple. Now, in the optical world, in optical connectivity, there is different ways to connect optics. There is different kind of transceivers and so forth. But optics, in order to cover distances, require to use active devices. They require us to use different kind of light sources and DSPs and optical engines and so forth. All of those consume energy. and we live in a world today that power is the number one limit of air factories of the compute capacity I can build in air factory right that's my limiting factor and of course I want to try and optimize power consumptions I want to reduce power consumptions wherever I can in order to be able to bring more compute because this is how I'm limited and since optical connectivity is more and more used scale out requires optical connections because of distance and it consumes more and more power scale up domain scale up domain if that scale up domain is within a rack I'm gonna use a car I'm gonna use copper if that scale up domain is start to have multiple racks I need to use optics so when we talked about for example connecting 1152 GPUs with the fireman platform we also So mention, hey, that will also use co-package optics or optics in order to run the distance. Now, if I'm using optics, an optics on scale-out infrastructure today can get close to almost 10% of the compute capacity on power perspective. That's a big number. Co-package optics is a technology that enables to minimize the power consumption that is going to be done or run or used on the optical network. And that's why we went to co-package optics. That's why I invested in co-package optics, because if I need to go to distances, I need optics. If I'm using optics, I want to have the best technology that consumes the least amount of power. And that's called co-package optics, regardless if it's scale-out, scale-up, scale-across. It all depends on the distance.
Shauna Laughlin, Analyst
All right. Well, unfortunately, we're out of time. I think we could have sat up here for another hour. But Gilad, we really appreciate you joining us and providing all of your insight. I mean, it's a privilege to get a front row seat to see what the innovation you and your team is driving, and good luck.
Gilad Shanir, Analyst — Other
Thank you very much.