Historical Opportunity for Domestic Chips: CUDA Moat and TileLang
AI and new technologies like TileLang are dismantling NVIDIA's CUDA moat. The ecosystem issues for domestic Chinese chips are expected to turn around within a year.
There is a historical opportunity for domestic AI chip replacement now.
Another point: because CUDA evolved from gaming cards, in many aspects, the design of gaming cards and their settings are consistent. CUDA is compatible with gaming cards. In the past, AI computing was a small field, smaller than the gaming card market, so this was reasonable. But now the market for compute cards is larger than gaming cards, so there is no reason for the two to remain coupled. The current trend is that they will no longer be coupled. So dedicated chips, whether from Huawei or NVIDIA itself, will be dedicated chips, not the previous ones.
In this context, the role of the original NVIDIA ecosystem is greatly reduced. Because for a dedicated chip, it has little to do with CUDA; it is not bound to CUDA. Or rather, when this chip was designed, it already considered how to build this ecosystem.
It's a bit complicated, but anyway, there is a historical opportunity for domestic AI chip replacement. We believe that within the next year, we will see something verified: the ecosystem of domestic chips has no problems at all. Previously, it was thought to be problematic, unusable or hard to use, but in the future, I think within a year, we can turn around this perception, or use facts to change it.
Domestic AI chips' hardware and ecosystem have no issues; the only problem is insufficient production capacity. There are no obstacles in adapting domestic cards, and NVIDIA can't stop it. If it were a normal business environment where I could buy NVIDIA cards, then domestic replacement would be difficult. But under the circumstances where NVIDIA cards are unavailable, everyone is forced to adopt domestic chips.
In this context, there are no obstacles to adapting domestic cards. I believe there are no obstacles to building an ecosystem for domestic cards that is as good as or even better than NVIDIA's, but it still needs time.
We are mainly cooperating with Huawei now. Huawei does the adaptation themselves, but we participate in the ecosystem and will deeply engage in Huawei's. Huawei's problem is still insufficient production capacity. For example, Huawei supplies us with about 16,000 cards, while internet giants might get hundreds of thousands. We have over ten thousand, and I think the ratio is relatively... but this might be all the capacity Huawei has.
So we can't count on training the next larger model on Huawei, or training a model with hundreds of billions of activated parameters, which some say might happen this year. But perhaps next year or the year after, there will be opportunities.
For Huawei card adaptation, our main work is to improve its high-level language compiler and develop TileLang. Once TileLang is ready, the problem may be solved. This is complicated to explain, but we are working on it, and it will be clearer after completion.
You can understand that when training V3, it still used NVIDIA cards, but no longer relied on NVIDIA's ecosystem. V3 used NVIDIA cards but not NVIDIA's ecosystem. Instead, we first wrote a high-level compiler called TileLang, and then based on TileLang's ecosystem to complete everything else, almost completely independent of NVIDIA's ecosystem.
As long as I redo this whole process on Huawei cards, it's done. I think this may be a historic mission: to completely reverse the previous perception that the domestic card ecosystem is bad. Now we have the right timing and conditions, just lacking time. I think within a year, many people will have or will understand that this problem is solved, and the only remaining issue is production capacity.
I am quite optimistic about domestic compute power. I think on this point, NVIDIA is digging its own grave. Huawei's supernode, the Huawei 950 supernode, can fully replace NVIDIA's GB200 and GB300 in terms of performance and price. The price will definitely be higher, but within a limited range. Fifty percent or one hundred percent more expensive is fine; even two hundred percent more expensive is acceptable.
For example, at 100% higher price, I think it can already be considered a price equivalent. In terms of tasks, it's also equivalent. All tasks that GB300 can do, Huawei supernode can do, with the same latency. The only cost is that four Huawei cards equal one NVIDIA card, and it's two years behind.
Four to one is understandable. Two years behind means four Huawei 950s can match one GB300. The two-year gap is in timeline. The Huawei 950 supernode is available in Q3 or Q4 this year, while NVIDIA's GB200 was available in Q3 two years ago, a two-year difference. NVIDIA may already have a new generation by Q3 this year. So our gap with the US in chips, I believe there will be no gap in ecosystem in the future, but in chips it's four times plus two years.
Another question: will we vertically integrate upstream? I hope not. So the question is: will we vertically integrate into upstream applications? We hope not. I hope others do that. I don't want to take everything. I just want to take one piece, the piece I am best at, or the piece we consider the core, and then with users directly related to us, I think many of our industry partners care more about it than we do... that should be their meaning, not for me to interpret. Will we build large-scale clusters ourselves in the future? I think we definitely will. We've been doing that all along-all our clusters are self-built. But whether we will develop our own chips depends on how much benefit there is.