Organizational Structure, Release Cadence, and "Nonlinear" AGI
Covers organizational restructuring, a release cadence of every two to three months, and a discussion of AGI as gradual yet nonlinear, ultimately leading to embodied intelligence.
It has no tipping point, but it is nonlinear. We currently believe in a narrative: AI can accelerate AI research.
I at least think it needs to be capable of continuous learning. Can domestic hardware catch up at this point? I think domestically it might take a few years. First, we need to solve the ecosystem problem domestically, because the ecosystem is a matter of confidence. After solving the ecosystem problem, then we solve the capacity problem. I think it should be solvable step by step.
I don't really believe that in five years we'll still be stuck on the capacity issue. Right now we're certainly stuck on capacity-this year, next year, the year after, I think we'll probably still be stuck on capacity-but in five years, maybe not. I'm still relatively optimistic.
Then the second question is about thinking about future organizational structure and the scale of personnel planning.
First, our previous organizational structure was very fragmented, because there was no structure. But as we further expand our headcount, this will certainly need some changes.
For now, I can only say that many changes will be needed here, but it's hard to articulate all at once. In the end, we should split into different departments. Some departments we need to set up with a relatively strict hierarchy. Other departments may maintain a looser, flatter structure.
As headcount increases, we will make this adjustment. We should make it immediately, because I'm already making it. If we don't make this adjustment, many things can't move forward. Indeed, many departments should have an organizational structure.
Then everyone also had a question: which version did CV precede, right?
I think with the GCV4 version we've released online, it's still quite rough, and its capabilities still need a lot of time.
Generally, for me, a comfortable release cadence is about one version every two to three months. Last release was probably end of April, then next might be end of June, roughly like that. If nothing unexpected happens, each version should be better than the last.
At the 50B active parameter scale, I feel that ultimately we won't be too different from the current wave of open source. In terms of inference speed and performance, I think there won't be a huge difference.
But compared to their larger model-the one they haven't disclosed-the gap should still be significant. That gap means our active parameters probably can't achieve it; it would definitely require a larger model, maybe 150B.
As for 150B, I think with our current training progress, optimistically we can start training by the end of this year, or at least by next April... the gap is still large. Yes, that's the gap with OCE.
Hi, I actually have a question. You often say that the process of achieving AGI is gradual rather than sudden. So can I understand it as a process without a tipping point?
It has no tipping point, but it is nonlinear. We currently believe in a narrative: AI can accelerate AI research-AI can accelerate AI research. That is, it's not linear, because you can use AI to accelerate your own research, so it may become nonlinear later on.
I see. So currently, from what I understand from your earlier conclusion, continuing to scale language models is sufficient, enough to reach this state.
I can only say that for language model scaling, I don't see an upper limit yet. Our current intelligence level, or the level in the US, I don't see an upper limit.
I see. Because I'm very curious: you said that for the US, the 800B active parameter model-you said they can train it but can't really use it, so they can only train it but it's hard to deploy for everyone because it's too expensive. I was actually curious before: humans have had language ability for only about a hundred thousand years, but before that, evolution took 3.7 billion years. But in training AI, perhaps the order can be reversed. But eventually we may still enter what is called, not necessarily the world model, but the physics model or embodied part, right? That is, after this upper limit.
Yes, I think embodied intelligence definitely needs to be entered, eventually embodied. So for our company, naturally, the end point might be embodied. Because for a normal person, their needs are not computers, right? Because normal people, in terms of eating, drinking, entertainment, clothing, housing, transportation, they don't need computers.
What they need is, so they still need embodied intelligence to solve specific labor needs. If the goal is to reduce labor needs, then embodied intelligence is unavoidable, I think.
I see. So in terms of stages, supposing we reach something like-not necessarily a tipping point-but a point where it can self-evolve, improve on its own, or close to that point of ascent with SV, I'm curious what the first deployment of AI in this state would be...
It might be different from now. We hope that it can-if there's no embodiment, then our definition of AGI, or what we want AGI to do, is: it can help us iterate the next version of the model, just like that. Then if we have embodiment, we also want it to iterate the next version of the embodiment, to make the next version of the robot.
I'm still curious about one thing: from previous interviews with DeepSeek and so on, it seems that in choosing important directions and research, taste and intuition are very important, not just simple engineering optimization. So if AI can self-evolve later, will these things like taste, taste, intuition still be important, or what will be important?
AI currently doesn't lack taste and intuition; what it lacks is the ability to continuously learn. AI's taste and intuition are fine. If you ask it to write an article, its taste and intuition, I think, are fine.
There are a few more questions from earlier. Let me look at that question. I see several questions on the screen, but I can't see them on my end.
Mr. Liang, I have a question left on the screen, let me read it to you. Actually, I wanted to ask about continuous learning-which you mentioned, and many researchers also mentioned it's an unsolved research problem. Then the coding agent, especially catching up to MILES and scaling to the Office level and MIS level, is a relatively certain goal. For a research problem that hasn't been solved yet and a relatively certain scaling goal, how do you think research resources-especially the talent pool of researchers-should be allocated to achieve the best balance and results?