CONFIDENTIAL · CONFIDENTIALCONFIDENTIAL · CONFIDENTIAL
Recording timestamp 02:08:18

Don't Go Upstream: Stick to What You Do Best

Don't make chips, don't do vertical integration, focus on one piece; multimodal is a component, not the main line; scaling is far from its limit.

I only need to take one piece, the one I'm best at.

Tesla, cultivation... that kind of thing. If you run a power plant, you don't necessarily have to build the generators yourself, right? Power equipment can be made by others; as long as the price is reasonable, why would you build it yourself?

So, I hope we don't have to make chips. I hope we can buy chips at a reasonable price, so we don't need to make them. I think that's likely the case-recently, the profit margins on NVIDIA chips are not necessarily sustainable... although...

We want to focus on just one piece. I think AI is huge, but we don't need to cover everything. We'll just do one piece. If we focus, and I believe the business opportunities here are large enough-if the AI era will produce many trillion-dollar companies, I think we will be one of them.

We've always been doing a small piece of it. It's fine for us to be one of those companies. I'm not ambitious to dominate everything. I think it's hard to change focus, and if you really want to do something else, you'll have more ideas. That's our attitude.

For example, at least in B2B and B2C businesses, as far as we can see, those who truly want to close the loop in B2C end up doing well, and those who want to close the loop in B2B might not do as well as we do. The more you want...

Also, I can say more about B2B. The ceiling for B2B business should still be demand. In the context of current-generation AGI and AI technology, B2B demand is limited. It will grow quickly, but it's not infinite; in the end, it's constrained by demand, not compute.

Under current technology, revenue ultimately depends on demand. Think about it: if I can recoup costs in ten months, I would certainly buy more if there were places to use it. But there aren't that many use cases. As technology continues to break through, demand will grow larger.

We've been working on multimodal capabilities. It's important for products, especially consumer-facing products. But for the ceiling of intelligence, it's a component, not the main line itself. That said, as a component, we will definitely do multimodal, and we are already working on it. We should release relevant models-our V4 and subsequent versions will support native multimodal. But we see multimodal as a component of intelligence, not intelligence itself. Its main line is like search.

Search is also a component. Multimodal is the same-we see it as a component. As for scaling, we believe in it. Larger scale definitely yields better results and unlocks more capabilities. What prevents us from scaling is simply compute. It's not that we don't want to scale; we don't have enough compute to do so.

We haven't reached the ceiling. We train models of this size not because we think that's enough, but because that's the amount of resources we have. We calculate based on our resources how large a model we can afford to train. It's not that this model is enough.

Currently, the benefits of scaling are still very clear. We haven't had the chance to hit the scaling wall-it's still far away. When Silicon Valley talks about scaling hitting a limit, that's for them. For us in China, we are still far from that; we haven't even scaled to that level.

Scaling includes scaling data, model size, and training cost. We're still far from exploring the upper limit of scaling, simply because we don't have enough compute. But we will spare no effort to push that limit.

After we finish fundraising, we'll have more compute and can train larger models. The sooner we can buy the compute, the better; if we could use up all our resources in half a year, that would be ideal, but it's not possible. The core capability of the next-generation model, I think, must be continuous learning. Before that, what we can do is reduce cost, improve effectiveness, and increase speed. But for a big breakthrough, it should have continuous learning.

Another question is why we seem to care so much about model compute efficiency. I've noticed that not everyone cares about it. For a commercial company, there's no incentive to pursue model efficiency, because if efficiency is higher... So they don't have a strong drive to make models highly efficient. Some startups don't even mention low cost as their goal, because that doesn't align with their interests. If costs are low, how can you make money? If costs are low, you can't charge much.

On the contrary, this is part of our vision. Our team members care about cost because they are ordinary people; they know that using these models costs money. They have empathy: since others have to pay to use it, if it's cheaper, they'll be more willing to accept it. So many of our team members hope we can lower the cost even further. But from a business perspective, you wouldn't think that way. From a business standpoint... From a cost or business perspective, this is not the top priority. For both startups and large companies, the cost of this service is not high at all. But I think we want it to be lightweight, I want it to be affordable, especially in the context of China's tight supply of computing power-affordable and usable on domestic Chinese chips.

Open this chapter in the guided reader →
This transcript was produced by automatic speech recognition and edited with AI. Speakers are not separately labelled; chapter titles, summaries and the argument map are editorial aids. Names and figures may contain recognition errors — refer to the original recording. Liang Wenfeng noted during the meeting that some figures are sensitive; please do not redistribute.