Simulated Data: Can It Surpass Human Experience?
Investors ask whether simulated data can allow models to surpass the upper limit of human experience. Liang Wenfeng cites AlphaGo as an example, saying it can exceed but has limits.
I think there are two points where it can surpass. For example, in Go, AlphaGo played a move that humans had never seen before.
We don't need that many companies. Right now, people might think the profit margins on this are very high, so they insist on doing it themselves. But when they realize it might not be that profitable, they might stop. Recently, it certainly isn't that profitable. I don't believe it's that profitable because that would go against objective laws. That means we are at a stage where if there is an extremely high profit margin, it must be against objective laws. We should have reasonable profits. So this is the current state of the industry, and I think it will definitely converge.
When it comes to building large models, no single company should say, 'I want to take all the profits.' That certainly won't work. If every company only takes reasonable profits, then we don't need so many people building large models. In China, if three or four companies end up competing, that's already enough competition, and prices will definitely be low enough for a price war. For large models, maybe two big companies and two small ones would be enough.
As for the ecosystem, I don't have too many thoughts. We hope to support more people, but we don't have that much energy. We have the willingness and there's no conflict of interest, but whether we actually do it is another matter. At least there's no conflict of interest here; we want win-win cooperation.
First, I absolutely do not think that large model companies can take most of the profits; that's impossible. Because there are so many large model companies, the gaps don't need to be that big. There are only two differences: time and cost. So no company will have excessive profits; I think it's unlikely. Those who control costs well will earn a bit more, and those who don't will earn a bit less. That's all. Does that answer the first question?
Can everyone... do you believe that in the future many people can... actually in the future... everyone's data usage and iteration cycles... can you hear me now? Thank you.
Yes, thank you for your answer. I deeply understand and respect your strategic positioning in the ecosystem. For example, regarding data: now public data, I believe model companies already have channels to obtain it. The method shouldn't be a problem, just a matter of time and cost.
Then, when we truly reach AGI, one possible scenario or ideal state is that the model can iterate and learn by itself-training itself. Regarding this, do you think simulated data can be used, or is real data of the highest quality?
If it still relies on real data, wouldn't that limit AI's intelligence to the scope of human past... because it depends on real data that humans have actually generated? Or can it break this upper limit through simulated data, synthetic data, created data, etc., so that the model's capabilities surpass all of humanity's past real...
I think it can surpass. I think there are two points where it can surpass. For example, in Go, AlphaGo played a move that humans had never seen before. That is, it certainly surpasses humans within a certain range.
But it may also have an upper limit; it may have limitations. But we cannot see those limitations now. So broadly speaking, we believe it can surpass based on the knowledge that humans already have and can articulate.
So in that regard, does it rely on real data or simulated data? Will it work?
It's impossible to have just one method; there are many ways.
Okay, thank you. I've taken up your time, and I'd like to continue asking about AI Infra. What was the second question? I'm a bit... let me quickly repeat it. I wanted to ask about AI Infra: in the future, compute power is being invested in at the scale of hundreds of billions of dollars. In this regard, we believe the Chinese people will achieve the hardware mission, and one day we may have efficient compute power, but in practice, it is still a bottleneck.
In the future, will two trends work against each other? On one hand, as model capabilities improve, they will shift from brute-force compute to clever compute, so the compute requirement per model or task will gradually decrease marginally. On the other hand, hardware like GPU capabilities will increase, making iteration faster and improving single-card efficiency.
What will this phenomenon look like in the process? For example, if we build a 10,000-card cluster and buy H200s or 960s now, will they become relatively less optimal compute after two or three years, becoming obsolete?
NVIDIA's cards can basically be depreciated over five years. Huawei's cards can be depreciated over three years at most. Huawei's 950 will be fine this year, still usable next year, but the year after that, it might be too power-hungry. The lifecycle of Huawei cards is definitely shorter because they are already two years behind NVIDIA. But I think the gap isn't that big.
If B200s were available now at any price, I think it would be worthwhile. For Tencent, if Alibaba could buy them, depending on the quantity; if the price is reasonable, it would be worth it. But this is not the time to calculate costs; I believe they won't be available.
Understood.
Our compute power is behind; that's a fact. We offset this fact in three ways. First, we accept that our models lag behind. We can only use smaller models than theirs. The question is how much smaller for training. We have to accept a certain degree of model lag and smaller model size. This lag has a benefit: it means we have more time and some technology, so we can use clever methods.