CONFIDENTIAL · CONFIDENTIALCONFIDENTIAL · CONFIDENTIAL
Recording timestamp 02:53:59

Trailing by One to Two Years, but Using Only One-Twentieth the Compute

An investor inquires about the pace of technological catch-up and Huawei 950 production capacity. Liang Wenfeng proposes a narrative of trailing but using only one-twentieth the compute.

Trailing by one to two years, but using only one-twentieth of its compute.
Investor question

Thank you.

Liang Wenfeng responds

So our gap with the US might be trailing by 12 months, maybe 12 to 18 months, or 6 to 12 months. Anyway, to put it simply, we're trailing the US by two years, but we accomplished this using only one-twentieth of their compute.

That narrative is trailing by one to two years, but using only one-twentieth of its compute. In the future, we want to rewrite that narrative: using only a fraction of its compute, but shrinking the time even further, to six months, three months. I think that's a goal.

And we might even surpass them in some areas. But given the overall order-of-magnitude gap in compute, comprehensive surpassing is unrealistic; however, in certain key areas with trade-offs, we might be able to exceed them in some places.

Investor question

Understood, thank you, Mr. Liang. Full of confidence, let's work together. I'll leave time for other colleagues, thank you.

Thank you for the sharing just now. I have two quick technical questions. In the discussion of the technical roadmap, you mentioned that the core problem we need to solve at this stage is continuous learning, which is also a hot topic in overseas research, called Recursive Improvement. I'd like to ask, from a technical perspective, what is the biggest difficulty right now? From your viewpoint, when can this be solved? That's the first question.

Second question: you also mentioned just now to first solve continuous learning, then tackle intelligence. I'd like to understand the technical reasoning behind that. Does it mean that after solving the continuous learning problem, DeepSeek will subsequently work on general intelligence? Please advise on these two questions.

Liang Wenfeng responds

Technical questions are actually a bit tricky to explain... The difficulty is that we haven't yet found a method that really works. No one in the world has found a good method; everyone is still exploring. So we're still in the exploration phase-we don't know who will stumble upon the next solution to this problem. We have many ideas, many promising-looking approaches, but none have fully materialized yet. Yes, that's the first point.

Second, internally we emphasize this narrative: training our next model, we hope it helps our own development. It can improve DeepSeek's efficiency. Our model's primary goal is to increase DeepSeek's own productivity, so that when we develop the next model, it can provide more assistance.

Or to put it simply, the primary goal of our model is not for others to find it useful, but for ourselves to find it useful. It needs to be useful for us first. Once it's useful for us, we can develop the next model faster.

Many people internally think this way: it must be useful for us first, used by us first. And this is the fastest path to AGI. When it's useful for us, it might also be useful for others, but we must ensure it's useful for us first.

This narrative is a bit odd, but many people really think this way. It's not about users finding it useful; I hope it helps us more, so we can achieve AGI much faster. So the logic of this narrative is to help us achieve AGI. But first, help us achieve it. We need this help to achieve AGI. Now it's very certain that we indeed need AI to help us achieve AGI. Although it's not yet working autonomously, it still works with humans, but it's already very useful.

Investor question

Second question: you just mentioned solving the continuous learning problem first, then moving into general intelligence. Is that your expectation for the future? I'd like to understand the technical reasoning behind this. Why solve continuous learning first before entering general intelligence? What's your subsequent understanding on this?

Liang Wenfeng responds

Because solving the continuous learning problem can greatly accelerate our R&D progress. If we solve continuous learning first, then the general intelligence problem will be a piece of cake. With AI assistance, if AI can continuously learn, its capabilities should be very strong.

Current agent capabilities are limited because they cannot continuously learn, they cannot do effective continuous learning. If we can first get continuous learning done, AI's capabilities will be very strong, and it can vastly improve our own research efficiency.

If we develop continuous learning first, general intelligence might become very easy-easy to achieve with it. So I say this is the outcome we'd prefer to see, it saves us effort, makes things easier. Otherwise, if we have to manually work on general intelligence now, it's a laborious, painful, data-intensive, and human-intensive task, and not cost-effective.

Investor question

Thank you for sharing.

A quick question-please check the online questions in the chat group. How long do you think it will take to achieve AGI? Can domestic hardware catch up by then? It's in the Zoom meeting chat window.

Liang Wenfeng responds

Okay, I saw that. Huawei 950-currently Huawei is providing us with 16,000 cards, which should be okay to disclose publicly. That's about an order of magnitude less than what internet giants have. Huawei can only give us this many because it's not cheap.

Internet giants have bigger ambitions, they need more. For us, we can buy some non-compliant cards. So our purpose in buying Huawei 950 is to help Huawei build a good ecosystem.

16,000 cards of Huawei 950 are only equivalent to 4,000 cards of the B-series. So it's not a large quantity, not very significant. It's not enough to train a next-generation model; it can only train our current generation, not the next. But it allows Huawei to get it right first. That's about Huawei 950.

Then, how long until AGI? Can domestic hardware catch up by then? I think in the AI field, in AI matters, domestically we should be able to reach a similar level to abroad in one to two years, or maybe even this year we can achieve parity in replacing foreign models. In AI, under the current approach and paradigm, it's not too difficult, so we should be able to do it this year. But it's still not AGI.

Open this chapter in the guided reader →
This transcript was produced by automatic speech recognition and edited with AI. Speakers are not separately labelled; chapter titles, summaries and the argument map are editorial aids. Names and figures may contain recognition errors — refer to the original recording. Liang Wenfeng noted during the meeting that some figures are sensitive; please do not redistribute.