CONFIDENTIAL · CONFIDENTIALCONFIDENTIAL · CONFIDENTIAL
Recording timestamp 01:15:12

Only the AGI Main Line: Boundaries of Cooperation and the Starting Point of the Resource Gap

Externally, we only focus on the AGI main line; video generation, world models, and other non-mainline work are left to others. Leads to discussion of the resource gap with the US.

When dealing with the outside world, our attitude is: we only do the main line of AGI.

When dealing with the outside world, our attitude is: we only do the main line of AGI. That is what I just mentioned: GPT, CoT, Agent, etc. Only the main line. The AI field is vast, and we consider many things not on this main line. For example, 3D, video generation-I think they may not have much to do with the main line of intelligence, so we won't do them.

There are also things like world models. I think they have little to do with the ceiling of intelligence right now, so we won't do them either. But if others do them, we're happy to help. Whether we have time is another matter, but there's no conflict of interest. We also hope these AI technologies can be applied in various production environments, improving social productivity and helping various industries. We are very motivated to do this.

Whether I have the time, whether we have the manpower, or whether our team is interested is a separate issue. But in terms of interests, there's no conflict. We hope to achieve this goal. And we believe this doesn't conflict with business at all; the benefits we deserve are still there.

I think that by holding this attitude, we haven't actually lost anything. We haven't gotten less because we open-sourced, or because of our goodwill, or because we helped others. For example, last year's consumer-side users-we still have a lot of consumer-side users, and they're quite stable. And our enterprise side this year, I think, looks positive.

Our goodwill hasn't affected our commercial interests at all. No impact whatsoever. In fact, it might have been a plus. This seems counterintuitive, but it's true. Or think of it the other way: if we violated that, would we necessarily get more? No.

One question is: how to understand that world models have nothing to do with raising the ceiling of AI intelligence? I'm talking about at this stage, that's our judgment. We have our own AI roadmap; it's not the only one.

From our understanding and judgment, the most important thing right now is to do AI training well. Doing AI training well doesn't require world models, or even multimodality. Because if you narrow the scope of AI training a bit, without multimodality, there are just some tasks you can't do, but it doesn't affect the validity of the algorithm.

Multimodality will eventually be necessary. For now, training is important. The next step is to solve continuous learning, and after that, having it ask questions on its own. But world models and video generation are not on this roadmap. When video generation first came out, it was very hot, as if you had to do it or you weren't an AI company. I found that strange. If you think about it carefully, it has nothing to do with the intelligence roadmap.

In fact, you also noticed that after Sora came out, everyone did video generation-big companies and small companies alike. But later, small companies all dropped it. It has nothing to do with the ceiling of intelligence. Commercially, it's a good business, yes. But it has nothing to do with intelligence. We won't do it just because it's a good business; we only do things that are on the intelligence roadmap.

Video generation is relatively clear, so I use it as an example. World models are less clear because many things can be called world models. In our judgment, world models and intelligence are not the most important things at this stage. The most important is AI training, and then how to solve continuous learning after training. That's our company's judgment; of course, each company has different judgments.

As I just said, for our company, the most important issue is personnel stability. From another dimension, what are we lacking? What is the gap with the US? Actually, there's only one point: resources. We don't have that many GPUs; our number is still relatively small. We now have about 20,000 H-equivalent compute, most of which arrived recently, in the past month or two. Many machines may not have arrived yet.

Our total compute last year was relatively small. This year we are aggressively expanding it. We now have about 20,000 H-equivalent, and in the next few months, we will buy a large number of machines, mostly NVIDIA. How many GPUs do we need? Definitely more is better. Within what we can afford, more GPUs is definitely better, no question. So our current strategy is to buy as many GPUs as we can at a reasonable price.

If we use up this financing, I will buy as many GPUs as I can. The speed of spending is not planned; as long as the price is reasonable, I'll buy however many I can. If I spend all the money within half a year, I think that's a good thing. If I can spend it all in half a year, that would be too fortunate, too ideal. In reality, spending that much money is very difficult. You can't buy that many GPUs; it's hard to get them, and the price is high. You also can't pay an extremely high price-you have to ensure the price is reasonable.

Open this chapter in the guided reader →
This transcript was produced by automatic speech recognition and edited with AI. Speakers are not separately labelled; chapter titles, summaries and the argument map are editorial aids. Names and figures may contain recognition errors — refer to the original recording. Liang Wenfeng noted during the meeting that some figures are sensitive; please do not redistribute.